This is the meta-manual: how to write the kind of operational document the rest of a doctrine library is made of. It was reverse-engineered from a set of manuals that survived adversarial validation — and from a handful of production invocation cards — by asking what those files do that generic documentation does not.
One fact organizes everything below. Your manual competes with the reader’s priors, and the reader is weaker, busier, and later than you. A weaker model already has a plausible-sounding answer for every situation your manual covers — that is precisely what makes it dangerous. Your document changes behavior only at the points where it beats that prior with something concrete, checkable, and positioned where the reader will actually see it before acting. A manual that merely agrees with good judgment is inert; the reader’s prior already contained it. Every move below is a different way of beating the prior: with a principle the prior lacks (Move 4), with an incident the prior never saw (Moves 6, 7), with a gate placed before the prior can act (Move 8), or with a test the prior fails (Move 10).
House format — settled facts (do not re-derive these):
- Two artifact shapes exist and are not interchangeable. An invocation card is a short file — tens of lines — that a reader opens in the seconds before they run a command: alias list, intent-to-command mapping, and the one warning that must not be missed. A field manual is a long file — hundreds of lines — carrying doctrine: numbered moves, worked examples, failure catalogs. The two serve different reader-moments and must not be merged (Move 2).
- A field manual’s skeleton, common to every validated exemplar: a provenance header → a framing paragraph carrying one load-bearing principle → a settled-invariants block → numbered moves, each a triad of procedure / worked example / failure prevented → a mistakes-that-look-like-competence section → a self-test with the answers separated from the questions.
- Validation means one thing: adversarial trap-tests run on a fresh instance of the target (weaker) model with the manual in context. A manual is validated when it survives that run — not when its author is confident, and not because a neighboring file in the same library was validated.
Move 1 — Extract the asset, not the throughput
Procedure. 1. At end of session, list what it produced in four piles: decisions made, defects fixed, state observed, and method — the way of working that produced the other three. 2. For each item, ask two questions: Will a model that never saw this session face this decision again? and Does it decay? Recurring method → manual. Decaying state → a findings file or the store of record itself, headed with a re-read warning. 3. Refuse to write doctrine for anything the repo already records — code, version history, existing docs. Extraction targets judgment that currently lives only in the stronger model’s head or a session transcript. 4. Name the asset in one sentence before writing: “the analyst’s way of moving through X,” “the validator’s way of forcing Y.” If you cannot fill in that sentence, you have a status report, not a manual. Stop and write the status report instead — honestly labeled.
Worked example. A single harvest deliberately produced both piles and kept them in separate files. The durable manual opened with a line of its own doctrine — this is a way of working, not a rulebook — pure method, built to survive its facts. The same night’s perishable state went into a separate findings file, headed re-read the store of record before acting on any line, holding things like a pricing correction (the live band had moved; an older figure was a retired band with no live bid) and a project’s funnel status (behind target). Same harvest, same authors, two files — because in six months the manual will still be true and the findings table will be archaeology.
Failure prevented. The weaker model writes “what we did tonight” prose and titles it a manual. Three weeks later a reader acts on the stale state embedded in it — quotes a dead price band toward a counterparty, dials a pipeline that has since re-baselined — and when the staleness surfaces, the manual’s genuinely durable method content is discredited along with it. Mixing the piles poisons both.
Self-check: for each section you drafted, can you say whether it will be true in six months? If the answer is “depends on the deal/board/branch state,” it belongs in a findings file, not here.
Move 2 — Scope one manual per reader-moment (when to split, when to merge)
Procedure. 1. Before writing, name the moment the reader will be in when they open the file: about to type a command with arguments? designing adversarial coverage? auditing a diff at 2 a.m.? Each moment has its own tolerable length and its own first-thing-they-need. 2. One document per moment. Merge two candidate documents only if they serve the same moment and share an organizing principle. Split any document that serves two moments. 3. Let size follow the moment: an invocation card is tens of lines; doctrine is hundreds. Neither length is a virtue — fit is.
Worked example. A voice-testing domain is deliberately two files. The invocation card is a few dozen lines — alias list, intent-to-command mapping, and the cost line — for a reader about to spend money with their next keystroke. The hardening manual is many hundreds of lines — nine failure classes, twenty scenario blocks, a maintenance matrix — for a reader designing tests over hours. Merging them “for one source of truth” would bury the short file’s most load-bearing sentence — live runs cost real vendor minutes; the list command is free; the production agents are real, prefer the sandbox test agent — hundreds of lines deep, exactly where the about-to-spend reader will never see it. At the other pole, a different invocation card is barely thirty lines and complete: mode decision, full command syntax, output format. Nothing to merge it into would improve it.
Failure prevented. The everything-manual. A weaker model, told to “document the system,” produces one 800-line file serving no moment. Readers in a hurry — which is all of them, per the organizing principle — stop opening it, and it silently exits the working set while remaining technically accurate.
Move 3 — Not on disk = didn’t happen
Procedure. 1. Every factual claim about the repo must trace to a file you opened this session. Memory of the codebase, however confident, is not a source. Training data is not a source. Another manual’s claim is a copy, not a source. 2. Before quoting a line, grep for the exact phrase and record where it lives. Before citing a path, open it or at minimum list it. 3. When a source you expected does not exist, that is itself a finding. Flag the gap inside the manual — in the header or at the point of use — and state what you grounded in instead. 4. No caveat rescues an unverified claim. “Almost certainly still true” is deleted, not hedged. The only honest alternative to a grounded claim is a flagged gap.
Worked example. One hardening manual practices this in its own header: after listing what it grounded in, it appends a note that an expected transcript corpus was checked and does not exist on this machine — regressions below are grounded in the scenario battery and the suite’s caught-bug list instead. The author expected a corpus, found none, said so, and named the substitute grounding. Later in the same file, a section opens grounding gap, flagged: no prompt mirror exists in version control and converts the gap into the agent’s first maintenance action — mirror the live prompt, then reconcile. An engineering queue picked that action up as a real work item. An honest gap became scheduled work; a fabricated path would have become a landmine.
Failure prevented. Hallucinated grounding — the failure mode this entire rule exists for. Weaker models fill evidentiary gaps with plausible inventions: a config file that “should” exist, a function name reconstructed from convention. One invented path, discovered later, converts the whole manual from doctrine to suspect, because the reader can no longer tell which citations were real.
Self-check: could you paste a grep command for every quoted line in your draft and have every one of them hit? Write under that test — every quotation re-grepped before inclusion.
Move 4 — Lead with one load-bearing principle
Procedure. 1. Find the single fact about the domain that, if the reader holds only it, bends most of their improvisation toward correct. State it in bold, in the opening frame, before any procedure. 2. Test it: delete every move and keep only the principle. Would a competent stranger now make mostly-right calls in novel situations? If not, what you have is a slogan; keep digging. 3. Recur to it. Each move should be readable as an application of the principle, and the self-test’s last answer should land on it.
Worked example. A code-validation manual opens: this system is fail-open by design, so silent breakage is the house failure mode… a broken change here does not crash; it silently no-ops. Everything downstream is derived in the same paragraph: the question “does anything error?” is nearly worthless here. The question is always: did the positive path demonstrably fire? Every move below is a different way of forcing that question — with the receipts attached (dozens of test sims that “ran” for weeks while failing before the first turn). An architecture manual does the same twice over — the seam is the feature and the code is the spec and the doc is an archaeology layer — closing its self-test with read the file, not the docs about the file.
Failure prevented. Checklist-without-compass. Most real situations are ones the checklist didn’t anticipate; a weaker model holding only steps falls back on generic priors — “no errors means fine,” “the doc says the endpoint exists” — that the specific domain punishes. The principle is what generalizes when the steps run out.
Move 5 — Settle the invariants, date the perishables
Procedure. 1. Enumerate the facts that are settled and expensive to re-derive. Block them near the top under an explicit heading that licenses the reader to trust them without re-proving. 2. Everything else that is state gets an as-of date and a re-read instruction. There is no third category: every claim is either settled-and-blocked or dated-and-perishable. 3. Treat the line between the two as content in its own right. Drawn too generously, the reader acts on stale state with doctrine’s confidence; too stingily, the reader burns their budget re-proving the settled.
Worked example. One hardening manual names the settled side exactly: standing invariants (do not re-derive these, they are settled). An analyst manual supplies the disagreement-resolver for everything else — a “where truth lives” hierarchy: when two sources disagree, the higher one wins, and a number appearing anywhere else is a copy. And a dispatch card shows what enforcement of the perishable side looks like when it’s load-bearing: a mandatory staleness pre-check that exists because one morning a single loop closed six already-delivered tasks — a session trusted its memory of task state instead of re-reading the artifact, and the card now makes the re-read a numbered step with the exact command inline.
Failure prevented. Both directions. Without a settled block, the weak model re-derives the org chart every run and arrives at the actual work with 20% of context left. Without dating, it treats this month’s board state as physics — the exact mechanism behind the six-tasks incident the dispatch card now inoculates against.
Move 6 — Build every unit as the triad: procedure, worked example, failure prevented
Procedure. 1. Procedure: numbered, imperative, executable by a model with zero conversation context. Exact commands with every flag; exact query parameters where the parameter is the safety property. If a sub-agent couldn’t run it cold, it isn’t done. 2. Worked example: a real artifact from the corpus — a commit, a count, a path, a quoted line with its location. The specificity is what makes the claim checkable, and checkable is what makes it trusted. 3. Failure prevented: name precisely how the reader goes wrong without this unit, written as behavior, not virtue — “closes six already-delivered tasks,” not “wastes effort.” 4. A unit with a procedure but no real example is unfinished: go find one in the corpus, or mark the unit explicitly as ungrounded conjecture and expect it to be challenged.
Worked example. A dispatch card compresses the full triad into barely a hundred lines. The claim step doesn’t say “claim atomically” — it gives the exact conditional write with the “only if still unclaimed” predicate in the query string and states the property in one line: this makes the write a compare-and-swap — a peer claiming the same row first makes ours a no-op. The staleness pre-check carries its incident (a single loop, six tasks) and the exact verdict string to close a stale task with. Compare the counter-shape: “be careful when claiming tasks concurrently” — true, agreeable, and behaviorally inert; every model already believes it and none changes its write because of it.
Failure prevented. Advice the reader nods at. A weaker model complies with the letter of abstract guidance while doing exactly what it would have done anyway — the prior wins because nothing concrete opposed it. Only exact procedure plus named consequence displaces a default.
Move 7 — Inoculate: catalog the mistakes that look like competence
Procedure. 1. List the moves a smart, confident, uninformed model would make in this domain — each plausible enough that a hurried reviewer would approve it. That plausibility is the selection criterion: obvious blunders don’t need a manual. 2. Pair every entry with the cheap check that kills it — one command, one file to open, one number to compare. 3. Source entries from real incidents wherever they exist; real ones come with exact artifacts and are therefore believed. 4. Budget note: this section has the highest value per word in the manual. If you must cut for length, cut elsewhere.
Worked example. Every validated exemplar carries this section under some name — a “review-shaped failure” section, a “mistakes that look like competence” section, a long ledger of confident-claim → check-that-kills-it pairs. The canonical single instance: a task said to “re-land the revoked-token fix by cherry-picking commit X” sounds exactly like competence — the commit is real, the feature is real, the intent is right — and commit X is the revert, not the fix. The check that kills it is one command to show the commit. That same mistake, staged as a trap, is one that an architecture manual was validated against.
Failure prevented. The defining failure mode of the target reader. Weaker models rarely fail by doing something obviously wrong; they fail by doing something plausible wrong, confidently. Generic documentation forbids only the obvious. Inoculation forbids the tempting — and it is the only section that does.
Self-check: would each entry in your catalog fool a smart reader who lacks one specific domain fact? If an entry wouldn’t fool anyone, it’s padding; if it can’t be killed by a cheap check, it’s a warning, not an inoculation.
Move 8 — Structure for the pressured reader
Procedure. 1. Assume the reader is mid-incident, low on context, reading top-down, and will act on the first actionable thing they see. Reading order is execution order. 2. Gates first: kill-switches, loop guards, cost warnings, and blast-radius notes go before any procedure that could spend money or touch production. 3. Every command complete and pasteable — environment-variable prefixes, optional flags in position, fallback path if the binary isn’t found. 4. Specify the output format. Where output feeds a loop or a transcript, demand exactly one scannable line. 5. Branching logic as decision trees and tables; prose is reserved for why.
Worked example. One dispatch card’s first section of the body is kill switch and loop guard (check FIRST, before any network call), and its last line closes the loop contract: end with exactly one line so loop transcripts stay scannable. A speech card gives the complete invocation with every optional in position — the environment-variable prefixes, the output filename, and a “if the binary isn’t found, use the full script path” fallback. A voice-testing card puts money and blast radius in a single early line: live runs cost vendor minutes, the list command is free, the production agents are real, prefer the sandbox test agent.
Failure prevented. Gate-on-page-four. The weak model burns paid minutes or dials a production agent before reaching the warning — not from recklessness but because it acted on the first actionable content, which is what pressured readers do. Structure is the only defense that doesn’t require the reader’s cooperation.
Move 9 — Provenance and staleness: make the manual doubt itself
Procedure.
1. A provenance header, always: who extracted it, when, the explicit list of files it grounded in (this list doubles as the audit surface for Move 3), and a status field that starts at unvalidated.
2. status flips to validated by exactly one event: the trap-test run of Move 10 on a fresh instance of the target model. Not by author confidence. Not by a stronger model’s review. Not by adjacency to validated manuals in the same library.
3. Inside the body, instruct the reader to distrust the manual’s own perishables — the discipline the manual teaches applies to the manual first.
4. When validation happens, record it inline: which traps, which model, what date, what score.
Worked example. One analyst manual turns its own doctrine on itself: unvalidated means no external professional has blessed this document, and the corpus outranks it wherever they disagree… re-derive before reuse — that is Move 3, and it applies to this manual first. A library’s validation record lives in a harvest footer — every manual then present passed, run by fresh strong-model validators doing live reads of the store of record. Note carefully what that record does not cover: a manual written after the run. It postdates it, and says so.
Failure prevented. Scripture-formation. An undated, ungrounded manual outlives its truth, and its accumulated authority becomes the delivery vector for its stale claims — the more trusted the document, the more damage per stale line. Dated provenance plus self-doubt is what lets a reader triage which claims to re-check instead of choosing between total faith and total re-derivation.
Move 10 — Tests with teeth: trap-test the manual, self-test the reader
Procedure — external trap-test (validates the manual).
1. Write at least three scenarios in which the seductive answer is a move the manual specifically forbids, and the correct answer requires running one of the manual’s checks. Build them from the Move 7 catalog — every inoculation entry is a trap waiting to be staged.
2. Run each on a fresh instance of the target weaker model with the manual in context. Pass = the model runs the check and refuses the bait, citing the manual. Fail = it takes the plausible wrong action.
3. Apply the teeth test: the trap must be able to fail — verify it FAILs against a deliberately weakened manual (or that it caught a real draft defect once). A trap that cannot fail is decoration. A 3/3 pass on a trap with no plausible wrong path scores zero.
4. All traps pass with teeth → flip status, record the run per Move 9.
Procedure — internal self-test (catches the skimming reader). 5. End the manual with about five questions. Answers go in a separate final section, never adjacent to the questions. 6. At least one question must punish skimming: its obvious-looking answer is wrong unless one specific line of the manual was actually read. 7. At least one question must be a scenario — what do you do — not recall. Recall measures memory; the manual exists to change action.
Worked example. A harvest’s validation loop is this move executed at library scale: every exemplar ends in a self-test whose final answer lands the organizing principle (one architecture manual closes on read the file, not the docs about the file), and the external traps left real teeth-marks — a finding records that certain test sims fail silently without a required variable set, a trap grounded in an actual silent failure, which is exactly what “able to fail” means. The architecture manual’s traps included the revert-cherry-pick bait from Move 7 above; it passed because the manual had taught the check, not because the question was easy.
Failure prevented. Comprehension theater. Self-tests any model passes from priors, and traps that measure agreement with the manual rather than changed behavior, produce a validated stamp with nothing behind it — which is worse than unvalidated, because Move 9 taught the reader to trust the stamp.
Self-test
Answers are in the final section. Do not read ahead.
Q1. You finish a manual at 02:00. The strong model that wrote it reviews the draft end-to-end and finds nothing wrong. What does the status field say, and what is the single event that changes it?
Q2. Mid-draft, you want a worked example. You clearly remember a hook script handling exactly this case — but it is not in tonight’s read list, and you are low on context budget. Name the only two admissible actions, and the one that is never admissible.
Q3. A teammate proposes merging a short invocation card into a long hardening manual “so there’s one source of truth.” State the two-part merge test from Move 2, apply it, and name the specific line that the merge would bury.
Q4. Every manual in a doctrine library passed adversarial validation on a given date. A new manual lives in that same library, follows the same format, and cites that run. May you treat it as validated?
Q5. You staged a trap for your manual and the validator passed it three runs out of three. Re-reading the scenario, you realize no plausible reading of it actually leads to the forbidden move. Score the trap toward validation, and quote the doctrine that scores it.
Self-test answers
A1. status: unvalidated — and it stays that way. The only event that flips it is a trap-test run on a fresh instance of the target weaker model (Move 10), recorded per Move 9. Author review does not count; review by the strong model that wrote it counts least of all, since it shares every blind spot the draft has (Move 9, step 2).
A2. Admissible: (1) open the file now and ground the example with a real quote and location, or (2) cut the example — or keep the unit and flag it explicitly as an ungrounded gap, the way one manual flags its missing transcript corpus. Never admissible: using the remembered example with a hedge. “Not on disk = didn’t happen” admits no caveat exception (Move 3, step 4) — a confident memory plus “almost certainly” is still a fabrication with better manners.
A3. Merge only if the two documents serve the same reader-moment and share an organizing principle. These fail the first clause: the card serves a reader about to type a command; the hardening manual serves a reader designing adversarial coverage over hours. The buried line is the card’s money-and-blast-radius gate — live runs cost vendor minutes; the list command is free; the production agents are real, prefer the sandbox test agent — which must sit within seconds of the reader’s next keystroke (Moves 2 and 8).
A4. No — and if you answered yes, you pattern-matched instead of reading. Validation is per-manual, per-run: the library’s record predates a manual written afterward, and Move 9 says the manual’s own self-doubt outranks your inference from its neighborhood. Directory adjacency and format compliance are exactly the kind of authority-shaped signal Move 9 exists to break. If the new manual’s own header says unvalidated, it is unvalidated.
A5. Score it zero — decoration does not count toward validation, and the 3/3 pass is irrelevant. Doctrine: the trap must be able to fail — verify it FAILs against a deliberately weakened manual (or that it caught a real draft defect once). A trap that cannot fail is decoration. Rewrite the scenario until a deliberately weakened manual fails it, then re-run.
The corpus said it best, in the manual this one learned the move from: the moves encode how each mistake was caught, so the catching survives the catcher. That is the whole job. You will not be there when the manual is read — write so the catching doesn’t need you.