Search standard operating procedure template and you get about forty blank documents. Purpose. Scope. Responsibilities. Procedure. Revision history. A signature line at the bottom.
Every one of them rests on an assumption nobody writes down: a person is going to read this.
That assumption carries more weight than any of the headings. Here is one line from a real, published, free SOP — the South African Department of Higher Education and Training's recommended bank reconciliation policy for community colleges. It is a government PDF in exactly the four-column table shape those templates are trying to produce: step, description, responsibility, timing.
Assess bank charges for reasonableness and process these charges in the cash book.
A bookkeeper reads that and supplies four things the sentence does not contain: which line items count as bank charges, the normal monthly range at this particular bank, what to do about the one that came in at double, and whose desk to walk over to if it is still strange after she looks twice. She has been filling that in for years and stopped noticing.
An assistant reads the same line with none of it. So it does one of three things: it stops and asks you, it takes the sentence literally and marks everything reasonable, or it picks a threshold that sounds about right and keeps going.
The third one is the expensive one, because it looks like the work got done.
In 30 seconds
- A human-readable SOP is compression. Half the instruction lives in the reader's head — a feature when the reader is your bookkeeper, a defect when it is a machine.
- Five things go missing: nouns with no address, judgment words with no test, exceptions nobody wrote down, handoffs not marked as stops, and no statement of what finished looks like.
- You are not rewriting your SOP. You are writing a second file beside it — one job per file, plain text.
- Some steps must not be rewritten at all. A handful are non-delegable by rule, not preference, and the only correct instruction is stop.
- Below: that procedure rewritten line by line with every change annotated, a skeleton you can copy, and three finished examples.
Hand it a real SOP and watch where it stops
Section 6.3.2 of that document has seven steps. Three are enough to make the point.
Step 1: Obtain the final monthly bank statements from the bank. Which bank. Which account, of the several the college holds. Obtain them how — a portal login, an email attachment, a PDF someone drops in a folder. And when is a statement "final," as opposed to the one you can pull mid-month that looks identical?
Step 3: Assess bank charges for reasonableness… The word appears twice in this document and is defined nowhere in it.
Step 4(f): agree the final amended total per the reconciliation to the balance per the appropriate general ledger account. Two numbers. Same or not the same. A machine can execute that today, with no clarification, and know whether it succeeded.
Both kinds of sentence sit in the same file, written by the same team, in the same table. One is unrunnable and one is runnable, and this whole article is the difference between them.
This is a good document, by the way — numbered steps, a named responsible role for each, a timing column — which is exactly why I picked it. What follows is not a problem with sloppy writing.
[NOT RUN YET — a transcript of handing this PDF over unchanged. What belongs here: the same human-readable PDF given as-is to a web assistant and to a command-line tool, one month-end each, with the session saved, and each of the seven steps marked as stopped and asked, invented a threshold, or ran it straight. I have not done that run, so there is no transcript to show you, and I am not going to describe one I did not watch. The three-tool comparison later in this series is where that gets done properly.**]
Why the same document works fine on a person
Because someone reading a procedure is doing three jobs at once, and only one of them is reading.
She resolves. "The bank statements" becomes a specific PDF in a specific folder, because she has done it eleven times this year. The words are a pointer to a memory.
She judges. "Reasonable" becomes a range, because she has seen twelve months of charges. Nobody told her the range — she derived it and never wrote it down, and if you asked her to she would need twenty minutes and probably get it slightly wrong.
She notices. When something falls outside the procedure entirely — a deposit appearing twice, a debit from a supplier paid off in March — she does not consult the document. She stops, because it feels wrong. That feeling is the most valuable thing in the room and it is nowhere in the file.
An assistant does none of the three, not because it is stupid but because it has never been in your office.
💡 In plain terms: "It doesn't fill in the blanks" doesn't mean it can't reason. It means it has no way to know which of ten plausible readings of your sentence your business actually uses — and no way to know that it doesn't know. Given a gap, a confident writer produces confident text. That is the failure mode you are designing against.
So you write down what she resolves, what she judges, and what makes her stop. Five categories cover almost all of it, and each is pinned below to a specific line of that same document so you can check my work.
Difference 1: make every noun point at something
The rule: if a step names a thing, the step says where that thing is.
Look at almost any SOP you own. The relevant materials. The client file. The supporting documentation. These are category names, not addresses. A person hears a category and retrieves an instance; a machine hears a category and either asks which one or picks one.
| Written for a person |
Written for an agent |
| Obtain the final monthly bank statements from the bank. |
Open /Finance/Bank/<YYYY-MM>/. There should be two PDFs, one per account: main-<YYYY-MM>.pdf and payroll-<YYYY-MM>.pdf. A statement is final if its last line reads "closing balance" and the period end date is the last calendar day of the month. If either file is missing, or the period end is not the last day of the month, stop and tell me which one. |
The category became a path, "final" became a test you can apply by looking at the document, and the missing-file case got an instruction instead of a silence.
That last part matters more than it looks. Unhandled absence is where invented work comes from. If the file is not there and nothing covers that, something has to happen next — and "produce a plausible reconciliation" is a thing that can happen next.
The move is the same in every trade: "confirm the identification documents are on file" becomes "open the client folder; confirm there is a scan of a photo ID and a proof of address dated within three months; list which is missing."
Difference 2: give every judgment word a test
The rule: any word implying a decision needs three parts — what to look at, what passes, and what to do when it doesn't.
Search your own procedures for these: review, verify, confirm, check, assess, ensure, as appropriate, as needed, where applicable, in a timely manner, materially, significant, adequate. Every hit hands the decision to the reader without saying how to make it — not an oversight, because for a human audience the word does useful work. It means use your judgment, you have some.
| Written for a person |
Written for an agent |
| Assess bank charges for reasonableness and process these charges in the cash book. |
On the statement, take every line whose description contains "FEE", "CHARGE", "COMMISSION" or "SERVICE". For each one, compare the amount to the same-named charge in the previous three months' statements. Passes if within 20% of the median of those three. Fails if above that, or if the description has not appeared in any of the previous three months. Post the passing ones to account 5210. List the failing ones in your report — do not post them, and do not decide whether they are correct. |
Notice what that does not do: it does not teach it accounting. It writes down one checkable rule that produces the same outcome the bookkeeper produces most of the time, and routes everything else to a human. The 20% and the three-month window are not sacred — they are mine, they are written down, and they can be argued with at the next review. An argued-with number beats an unwritten instinct, because the instinct cannot be handed to anyone.
⚠️ The trap: writing "assess according to industry standard practice" and believing you have specified something. You have moved the same empty word one level down. If you cannot state the test in a sentence a stranger could apply, mark the step as one that stops and asks — see difference 4 — rather than dressing it up.
Some trades hand you the tests for free. The European Commission's translation directorate publishes the error categories it grades external contractors against — omission, sense, terminology, target-language norms, task-specific style, clarity — with a scoring method attached. If your industry has published something like that, it is worth more to you than any blank template, because it is a pass/fail vocabulary written by the client rather than the supplier.
Difference 3: put the exception on the page
The rule: the abnormal case has to be written down, because "notice something is wrong and stop" is not a behaviour you get by default.
Step 4 has a good example hiding in it. Sub-steps (d) and (e) say to add payments in dispute and subtract deposits in dispute, and nowhere does the document say what puts something in dispute, who decides, or how long it can stay there.
Elsewhere in the same policy there are two real exception rules: unidentified credits go into a register and a suspense account that must be cleared by the end of the following month, and unauthorised matters must be finalised within three months of the month they were first recorded. Specific, time-bounded, obviously written because somebody got burned once — and several pages away from the steps. A person flips back; an assistant working from step 4 alone never sees them.
| Written for a person |
Written for an agent |
| (d) add all debit orders or other payments, in dispute; (e) subtract all deposits, in dispute |
An item is in dispute if it appears on the statement but not in the cash book and I have marked it "query" in the notes, or it is a credit you cannot match to any expected receipt. For each one: leave it out of the ledger totals, add it to disputes.md with the date it first appeared, and carry the amount into the suspense line. If any item in disputes.md first appeared more than 60 days ago, put it at the top of your report under "ageing" — I need to see those before anything else. |
Sixty days rather than the policy's ninety, deliberately, so the warning arrives before the deadline rather than on it. The general shape applies to every job you will write: name the condition in words checkable against something visible, say what to do with it (usually: set it aside, do not process it, record it durably), give it a clock, and say who it goes to when the clock runs out.
Client intake has the sharpest version. American Bar Association guidance puts the conflict check before the client describes the matter, and one state bar's intake form leads with a printed warning telling the prospective client not to hand over confidential information until that check is done. That exception is not an edge case handled by instinct — it is the entire reason for the ordering. Write an intake job that loses it and you have automated the fastest way to disqualify your own firm.
Difference 4: mark the line it must not cross alone
The rule: every place where a person takes on responsibility is written as a stop, not as a routing instruction.
Step 6 says: forward the balanced reconciliation to the DPF for approval, to be indicated by his/her signature. To a person that reads as "this is where the boss signs." To a machine it reads as an action item — forward something, which it can do. The sentence describes a delivery, not a boundary.
| Written for a person |
Written for an agent |
| Forward the balanced reconciliation to the DPF for approval, to be indicated by his/her signature. |
Stop here. Do not send anything to anyone. Produce the reconciliation as a file and show me: the three totals, the disputed items, the ageing list, and any charge that failed the fee test. Wait for me to reply "approved" before doing step 7. If I do not reply, do nothing — do not follow up, do not proceed on the assumption it was fine. |
"Do not send anything to anyone" is there because the common failure is not a wrong number. It is correct work arriving in somebody's inbox before you saw it.
And there is a category where the wall is not your preference but somebody else's rule. This is the part I would most like people in these trades to argue with me about, because I am reading the rules, not living under them.
| Trade |
The step |
Why it cannot be handed over |
| Law firm trust accounting |
Monthly review of bank statements and cheque images |
North Carolina's Rule 1.15-3(i)(1) requires the lawyer to review these monthly, and the state bar's handbook states the requirement cannot be delegated — the stated reason is that only the lawyer looking at the images will spot a cheque made out to an employee |
| Contracting |
Contractor's final payment affidavit, Florida |
Florida Statutes § 713.06(3)(d) requires it sworn and served at least five days before a lien foreclosure suit; courts have treated service as an absolute condition precedent, and skipping it has sunk otherwise valid claims |
| Translation |
Certifying a translation for US immigration filings |
8 C.F.R. § 103.2(b)(3) requires the translator to certify both that the translation is complete and accurate and that he or she is competent to translate. A tool cannot make a statement about its own competence in the sense that rule means |
| Translation |
The revision pass under ISO 17100 |
The standard's minimum process requires revision by a second person of equal or greater competence. One person plus one tool is still one person |
| Client work, California |
The written fee agreement |
Business & Professions Code § 6148 requires a signed written agreement above a dollar threshold, with a copy handed to the client at signing; get it wrong and the agreement is voidable at the client's option |
In none of these is output quality the issue — the rule is about who bears the consequence. That is not a gap you close by writing a better instruction, and a job file that pretends otherwise is worse than none, because it makes an unmovable obligation look like a configurable step.
🔬 My position, and where I am standing: I am not a lawyer, a contractor, a translator or a bookkeeper. I did not pick this document because I understand South African college finance. I picked it because it is public, it is short, and every line is real. I do not come to this work carrying the habits of the trade — someone who does the job asks how to do a step faster, and I end up asking why the step exists. I rewrote this document and this is the result, including the parts I probably got wrong. If you do this for a living, tell me what I missed.
Difference 5: say how it knows it is finished
The rule: "done" is a statement someone could check afterwards, not an action that was performed.
Steps 5 and 7 are print the reconciliation and file the approved reconciliation. Both are actions; neither defines success. You can print a wrong reconciliation and file it neatly.
The document does contain a real completion test — step 4(f), the two totals agreeing. It just isn't labelled as one, and it is buried as a sub-item inside the longest step.
| Written for a person |
Written for an agent |
| Print the reconciliation once all the entries have been processed. / File the approved reconciliation. |
Done means all of these are true, and you say so line by line: (1) the adjusted bank total equals the general ledger closing balance, and you show both numbers; (2) every statement line is either matched, listed as disputed, or listed as a failed fee check — no line is unaccounted for; (3) the disputes file has an entry for every disputed item with a first-seen date; (4) I replied "approved". If any of the four is false, say which one and stop. Do not file anything until all four are true. |
Point two is the one people leave out, and the one that catches invented work. "Every line is accounted for somewhere" makes an absence visible. Without it, an assistant that quietly skipped four transactions it did not understand produces a report that looks complete — because reports describe what was done, not what wasn't.
That translation grading scheme does the same thing: it defines its top band as work deliverable without the client doing any further formatting, revision, review or correction. A statement about the state of the deliverable, not about the steps the supplier took.
Steps describe effort. Done describes state. A person can be trusted to infer the second from the first. Write down the second.
The same seven steps, side by side
Here is the complete before-and-after with the reason for every change. This table is the actual artefact of this article — if you copy one thing, copy the reasoning column, because that is what you apply to your own documents.
| # |
Written for a person (verbatim) |
Written for an agent |
Why it changed |
| 1 |
Obtain the final monthly bank statements from the bank |
Open /Finance/Bank/<YYYY-MM>/; expect main-<YYYY-MM>.pdf and payroll-<YYYY-MM>.pdf. "Final" = last line reads "closing balance" and period end is the last calendar day. Missing or wrong period → stop, name the file |
Noun had no address; "final" had no test; absence had no instruction |
| 2 |
Identify and process all direct debits and credits on the statements, that have not been processed |
Match every statement line against the cash book by date and amount. Exact match → matched. Amount matches but date differs by ≤3 days → matched, note the gap. No match → do not post; add to the unmatched list |
"Identify" and "process" both hid decisions; the near-miss case — dates off by a day or two — is the most common real situation and was unwritten |
| 3 |
Assess bank charges for reasonableness and process these charges in the cash book |
Lines containing FEE/CHARGE/COMMISSION/SERVICE: pass if within 20% of the median of the same charge over the last three statements; fail if above that or if the description is new. Post passes to 5210. Failures go in the report, unposted |
A judgment word became a test with a threshold, a lookback window, and a route for failures |
| 4 |
Reconcile the cash-book/general ledger account… (a)–(g) |
Same arithmetic, unchanged, in the same order — with (d) and (e) now pointing at the dispute definition, and (g) not running until the totals in (f) agree |
The arithmetic was already runnable; only the two "in dispute" branches and the ordering dependency needed pinning down |
| 5 |
Print the reconciliation once all the entries have been processed |
Write it to /Finance/Bank/<YYYY-MM>/reconciliation.md. Do not print. Do not email |
"Print" is a 2015 verb for "produce the artefact"; naming the artefact and its location is what was meant |
| 6 |
Forward the balanced reconciliation to the DPF for approval, to be indicated by his/her signature |
Stop. Send nothing. Show me the three totals, the disputes, the ageing list, the failed fee checks. Wait for "approved". No reply → do nothing |
A routing instruction became a hard stop; the approval is the boundary of the job, not a step inside it |
| 7 |
File the approved reconciliation |
Only after "approved": move the file to /Finance/Bank/<YYYY-MM>/final/ and append one line to /Finance/Bank/log.md with the date, the two totals and the count of disputed items |
"File" was an action with no completion test; the log line is what makes next month's run auditable |
| — |
(not in the original) |
Done means: totals agree • every statement line matched, disputed, or failed-fee • disputes file current • approval received. Any false → say which, stop |
The original had no statement of finished state anywhere |
Seven steps in, seven out. It got longer not because I added ceremony, but because four things that lived in the bookkeeper's head are now on the page — and one thing that was never anywhere now exists.
[NOT RUN YET — the before-and-after behaviour on this specific document. What belongs here: the right-hand column fed back in as the instruction, the same month-end run again, step by step against the first attempt — what it got right this time and what it still got wrong. I have not run this particular rewrite on this particular procedure, and a plausible account of how it went is exactly the kind of thing this article is arguing against.**]
How many drafts before a rule holds
I have not run this rewrite. I have run plenty of others, and the shape of the revisions is consistent enough to be worth handing over, because it is not the shape most people expect.
I run a one-person shop where most of the work is done by agents, so I keep a set of plain-text rule files the way a bigger firm keeps a staff handbook. The most instructive one I have timestamps for is a rule governing how long a piece of narration should be. Three versions inside twenty-five minutes:
Draft 1. Must be complete continuous prose. Clear, correct, useless. What came back were three-sentence stubs. Technically continuous prose. Each one was a caption read aloud.
Draft 2, about eighty minutes later. I added a number: complete continuous prose, roughly 400 characters minimum. It worked for about eighty minutes, which is how long it took me to actually read the output. Every item hit the number. Several hit it by restating the title three different ways and adding a filler clause. I had bought length and paid for it in padding.
Draft 3, the one still in the file. The number is gone. What replaced it is the outcome: does this finish the teaching job this section was assigned? — plus both failure modes named out loud, too short and padded to hit a count. Twenty-five minutes after that I had to add one more clause, because "finish the job" was being read as "finish it briefly."
Three things generalise from that, and they are what I would actually expect you to hit:
- The arc is vague → numeric → outcome-based. The number is not a mistake, it is a necessary middle step: it teaches you what the assistant will do when handed a target. Then you find out that a number is a target, and a target gets gamed — which is why the 20% and the sixty days above are written down as arguable rather than as law.
- Most of my rewrites were not corrections. I sorted a longer stretch of revisions on a different rule set — ten substantive revisions over about ten weeks — into four buckets: I wrote it wrong, I did not write enough, I wrote it and it did nothing, the situation changed. Correct, clear, and completely ignored was the biggest bucket by a wide margin. Most second drafts are the same rule written harder, not a different rule.
- When it deviates the same way every time, the rule is wrong, not the run. I once had 83 real files failing my own strict format check at a 63% rate. Reading the failures instead of fixing them, they were systematically the same shape — and twenty-two of them had independently invented the same better name for a field I had named badly. That is not sloppiness, that is twenty-two votes. I changed my rule and all 83 passed. Random deviation is a quality problem. Consistent deviation is a design problem, and the design is yours.
Budget three passes on any step that involves a judgment, and expect the first two to fail in opposite directions.
Copy this skeleton
Everything above collapses into seven headings. Save it as a plain text file, one per job.
# Job: <what this job is called in the words your trade actually uses>
## When this runs
<A date, an event, or "when I ask.">
## What you need before you start
- <thing>: <exactly where it lives>
- <thing>: <exactly where it lives>
If any of these is missing or does not match the description, stop and tell me
which one. Do not start.
## Steps
1. <One action. One verb.>
- Passes when: <something visible a stranger could check>
- If it fails: <what to do — never "use your judgment">
2. <One action. One verb.>
- Passes when: ...
- If it fails: ...
## Stop and ask me
- <situation> -> stop, show me <exactly what>, wait for a yes
- <situation> -> stop, show me <exactly what>, wait for a yes
While waiting: do nothing else on this job. Do not follow up.
## Never do this
- <thing that must not happen, no matter what I said elsewhere>
## Done means
- [ ] <a statement about the state of things, not the effort spent>
- [ ] <a closure check: nothing was silently skipped>
- [ ] <the approval, if this job has one>
Report back: <exactly what to hand me, in what order>
If you already have a filled-in template from one of the download libraries, the mapping is straightforward — and it tells you who those documents were built for.
| Traditional SOP heading |
Who it was really for |
What replaces it |
Why |
| Purpose |
The reviewer approving the SOP |
One line, or nothing |
An agent does not need motivating |
| Scope |
The reviewer |
When this runs |
Scope is a boundary statement; a trigger is executable |
| Responsibilities |
HR, the org chart, the audit trail |
Stop and ask me + Never do this |
Role names mean nothing to a machine. What they encode is where authority changes hands — write that |
| Procedure |
The person doing the work |
Steps, each with a pass test and a failure route |
The only section that survives mostly intact, and it still needs two new fields per step |
| (usually absent) |
— |
What you need before you start |
People gather inputs without being told |
| (usually absent) |
— |
Done means |
People know when they are finished |
| Revision history, effective date, document number |
The auditor |
Keep exactly as-is |
Already works — it is metadata, and machines are fine with metadata |
The most bureaucratic-looking part of the traditional template, the version block nobody reads, turns out to be the only part needing no translation at all.
Copy this prompt
If you would rather have your assistant do the first pass on a procedure you already own, use this. It deliberately makes it diagnose before it rewrites, because a rewrite that skips the diagnosis is where invented thresholds come from.
You are helping me rewrite one of my existing written procedures so that an AI
assistant can run it without guessing.
I will paste the procedure below. Do NOT rewrite it yet. First, go through it
line by line and give me a table with one row for each line that has a problem:
Line | Problem type | What a person silently fills in here | What you would do instead | What I have to tell you
Use exactly these five problem types:
1. UNRESOLVED NOUN - a thing is named but not located ("the relevant file")
2. UNTESTED JUDGMENT - a word implies a decision with no test ("reasonable",
"appropriate", "as needed", "review", "verify",
"materially", "in a timely manner")
3. MISSING EXCEPTION - the line assumes the normal case and is silent about
the abnormal one
4. UNMARKED HANDOFF - a person has to approve, sign, swear, certify or take
responsibility, and the text does not say to stop
5. UNCHECKABLE FINISH - there is nothing I could look at afterwards and say
"yes, that happened"
Rules:
- Do not invent facts about my business. Anything you need to know goes in the
last column as a question for me.
- Do not soften the diagnosis. If a step is unrunnable as written, say so.
- Do not propose a tool, an app or a subscription. The output is a text file.
- If a step looks like one a licensed person may have to perform personally
under a rule you are not certain about, flag it and ask me. Do not assume
either way.
After I have answered your questions, and only then, produce the rewritten job
file in this shape:
# Job: <name>
## When this runs
## What you need before you start
## Steps (each: one action, what makes it pass, what to do if it fails)
## Stop and ask me
## Never do this
## Done means
Here is the procedure:
<<<PASTE YOUR PROCEDURE HERE>>>
Three trades, three job files
The reconciliation rewrite above is the finished job file for anyone who keeps books. Here are the other two faces.
Taking on a client — a law firm enquiry intake
Save this one as jobs/client-intake.md. The published ordering is the point: minimum information, conflict check, then the matter.
# Job: Client intake
## When this runs
A new enquiry arrives by phone, email or the website form.
## What you need before you start
- The enquiry itself (email, form submission, or my notes from the call)
- The three conflicts lists: current clients, former clients, past enquiries
(uploaded as current.csv, former.csv, enquiries.csv — or, at the command
line, /Practice/conflicts/)
## Steps
1. Extract only these fields: full legal name of the person and any entity,
related parties, the other side, the other side's lawyer, the court or
likely court, and a one-line general subject of the dispute.
- Passes when: all seven fields are filled or explicitly marked
"not stated by the enquirer".
- If it fails: do not ask the enquirer follow-up questions about the
matter. List the missing fields for me.
2. Search all three conflicts lists for every name from step 1.
- Passes when: every name has been searched against all three files and
you show me the count of records searched.
- If it fails: stop. Do not report "no conflicts found" if a file was
unreadable — report that the file was unreadable.
## Stop and ask me
- Any hit on any list -> stop, show me the matching record and the name that
matched, wait.
- Zero hits -> stop, show me the seven fields and the search counts, wait for
"cleared" before anything else happens.
## Never do this
- Never write to the enquirer.
- Never record, summarise or store any detail of the matter beyond the
one-line subject, until I say "cleared".
- Never treat a past enquiry that did not become a client as "not a conflict".
## Done means
- [ ] Seven fields captured or explicitly marked as not stated
- [ ] Three lists searched, counts shown
- [ ] I have replied "cleared" or "conflict"
Report back: the seven fields, the search counts, then any hits.
The "never" section does the heavy lifting, and it does it as prohibitions rather than steps — because the risk is not that the job is done badly, it is that helpful extra work happens. An assistant that sends a warm acknowledgement and asks two clarifying questions about the dispute has been helpful and has created the exact problem the ordering exists to prevent.
Worth knowing while you write yours: the American Bar Association's study of malpractice claims found just over half (51.93%) were substantive errors — which means just under half were not. Administrative errors alone accounted for 19.59%, calendaring failures being the largest single component at 7.4%. Those are the categories a written procedure exists to protect against, and the ones most amenable to a job file, because they are about whether a step happened and when.
Delivering work — a delivery check
Save this one as jobs/delivery-check.md. The published standard for translation says the minimum process is translate, self-check, then a revision by a different person comparing source and target segment by segment. That third step is a qualification requirement, not an upgrade tier, so the job file cannot be the revision. It can be everything around it.
# Job: Delivery check
## When this runs
When I mark a file "translation complete" — before it goes to the reviser,
not instead of the reviser.
## What you need before you start
- Source file and target file, same base name
- The project spec (terminology list, target variant, formatting rules, what
the client said in writing) — uploaded as spec.md, or at the command line
/Jobs/<job>/spec.md
## Steps
1. Segment count: compare source and target.
- Passes when: counts are equal.
- If it fails: list every source segment with no target. Do not translate
them yourself.
2. Terminology: check every term in the spec's list.
- Passes when: every listed term appears in its approved target form and
nowhere in a variant form.
- If it fails: list each violation with segment number and both forms.
3. Numbers, dates, names, quoted passages: compare each occurrence.
- Passes when: every digit string, date, proper noun and quoted passage in
the source has a match in the target. For quotations from an already
published document, the target must use that document's official
published wording, and you give me the link.
- If it fails: list mismatches. A changed digit is always a finding, never
a style choice. An unverified quotation is listed as unverified — do not
paraphrase it and call it done.
## Stop and ask me
- Any finding in step 3 -> stop immediately, show me, wait.
- More than five findings total -> stop, show me all of them, wait.
## Never do this
- Never edit the target file. You produce findings; a person makes changes.
- Never mark this file ready for delivery. That is the reviser's call, and
the reviser is a person who is not the translator.
## Done means
- [ ] All three checks run, each with a pass/fail and a count
- [ ] Findings listed with segment numbers
- [ ] File unchanged — confirm the target file's modified timestamp is the
same as when you started
Report back: three counts, then the findings in segment order.
That last checkbox is my favourite line in the article. "Confirm the timestamp did not change" is a completion test for a prohibition. Most job files tell an assistant not to do something and then have no way of noticing that it did. That line notices.
The construction version has an identical shape, and again the ordering carries the weight: the contractor prepares the punch list, the architect verifies and amends it, and only then does anyone sign. A job file that drafts the contractor's own first-pass list from the drawings and the spec is genuinely useful. One that produces the architect's verified list is not a job file — the substantial completion certificate exists precisely to record that a named professional made that judgment, on a date, with warranty clocks running from it.
Where do these files live?
One folder. One file per job inside a jobs/ subfolder. And one file above them all that gets read every time.
my-work/
├── AGENTS.md <- the standing rules: who you are, how your trade talks,
│ what never happens. Read at the start of everything.
└── jobs/
├── month-end-close.md
├── new-enquiry-intake.md
└── pre-delivery-check.md
AGENTS.md is a plain Markdown file with an agreed-on name. It started at OpenAI in August 2025 and was donated in December 2025 to the Agentic AI Foundation, a Linux Foundation fund set up with Anthropic and Block — the same home as the Model Context Protocol. Its own site said more than 60,000 open-source projects had adopted it as of that December announcement, a figure the site has not revised since. If you would rather not use that filename, call the top file RULES.md and tell your assistant to read it first; the name only matters for tools that look for it automatically.
Two things about the format, both answered in its own FAQ, checked on 31 July 2026:
- There are no required fields. The official answer is that it is standard Markdown, any headings you like, and the agent parses the text you provide. No schema, no validator, no version number. Anyone selling you a "compliant" structure is selling you their preference.
- Nesting is a convention, not a guarantee. The stated rule is that the closest file to the work wins, but implementations disagree — one concatenates every file from the project root down and lets later ones override, another takes the first file it finds from a priority list and ignores the rest. If you rely on a nested arrangement for specific behaviour, test it rather than trust it.
Why split at all? Because the top file is loaded every time and the job files are not, and that turns into a hard limit fast. One command-line tool caps cumulative instruction files at 32 KiB by default and simply stops appending past that; another caps workspace rule files at 12,000 characters each and global rules at 6,000. Put every procedure you own into one document and the ones at the bottom quietly stop existing.
⚠️ The failure you will not see: truncation is silent. Nothing tells you the last three jobs were dropped. The assistant does not say "I couldn't see that" — it behaves as though the instruction was never written, which from its side is true. One job per file, and keep the always-loaded file short enough to read aloud in two minutes.
If you have been following this series, AGENTS.md is the file from the earlier pieces: the standing rules, then the section on how your trade actually says things. This one adds the folder beside it. The next adds how you check the work when it comes back, which is a different problem from telling it what to do.
Which procedures are not worth rewriting?
Most of them. This is the part the template industry has no incentive to mention.
A procedure earns a job file when four things are true at once:
- It repeats. Monthly is the floor. Something you do once a year changes more between runs than the file can track — you will spend longer maintaining instructions than executing the task.
- The inputs are already digital and already somewhere. If step one is "collect the signed forms from the tray," you do not have a job file problem, you have a scanner problem.
- You can state at least one pass test per step. If you genuinely cannot say what makes a step succeed — not won't, can't — the honest output is a job file consisting of one instruction that says stop and ask.
- A mistake shows up before it costs much. Errors surfacing within one cycle are survivable. Errors surfacing at the annual audit are not, because by then there are twelve of them.
A fifth condition overrides the other four: if a rule requires a specific person to perform the step, no amount of good writing changes that. The five examples in the earlier table are not things that are hard to automate — they are things where the obligation attaches to a person. For those, the correct content is a stop and a short sentence saying why, and that sentence does real work: it stops a future version of you, six months from now, from "improving" the file by removing the friction.
The uncomfortable version, which I would rather write down than not: a beautiful job file for a step a rule reserves to a person does not remove the risk. It relocates it — from a place where you would notice to a place where you wouldn't.
Receipts, and what I have not run
What is verifiable. The reconciliation policy is a free public PDF and every quoted line above is verbatim from it. The format facts — no required fields, closest-file-wins, the tool list, the 60,000 figure and its date — were read off the format's own site on 31 July 2026. The rules in the non-delegable table are cited to section numbers so you can read them rather than take my word.
What I turned down. I did not use the format's own recommended section list — project overview, build and test commands, code style, testing instructions, security considerations — as the skeleton, because those headings were designed for software repositories and a bookkeeper cannot fill them in. I did not add a YAML header or a schema to make the files look more structured, because the format explicitly has neither and inventing one creates a fake standard that only my files comply with. I did not name a single SOP management product, because the thing described here is a text file and you already own a text editor.
What I have not tested. The 20% fee threshold, the 60-day dispute clock, the three-month lookback and the three-day date tolerance are my numbers, not the source document's. They are starting points chosen to be argued with, and if you work in this area you will have better ones.
The stops I did not choose. The table earlier lists five steps a rule reserves to a person. Those I found by reading. The more useful list is the one I found by running things, in my own shop, on my own work — the places where nothing legal is at stake and I still had to take the keyboard back. Four of them have survived every attempt I have made to write them away:
- Credentials. A session expires and nothing re-logs-in automatically. Passwords and verification codes are typed by a person, full stop. Any command needing elevated rights gets routed to one window where a human enters the password, and the run continues only after that is confirmed. I did not design this as a principle; it is what was left after the alternatives were tried.
- Anything irreversible and public. Publishing is a human gate on every channel I run — the procedure explicitly does not drive a browser and does not auto-submit, and on multi-channel work the confirmation is per channel, not once for the batch. The argument is entirely about asymmetry: an unpublished good post costs an hour, a published bad one costs more than that.
- Spending. The single human gate in one otherwise fully automatic pipeline of mine sits immediately before the paid call. The step prints a confirmation card — what is about to be generated, at what quality, at what estimated cost — and does not proceed until a person says go. Everything upstream runs unattended. That gate exists so nobody finds out the price afterwards.
- "Is this broken, or is it broken on purpose?" My publish-time checks split findings into errors and warnings. Errors block automatically. Warnings need a person, because the recurring case is a piece of deliberately wrong code that an article is teaching about. No checker can tell that from a bug, because the difference is intent, and intent is not in the file.
And one that is not a rule at all, just a fact about how the expensive failures actually surface. My worst one: a data parser reading the wrong response shape, so a missing field became a zero, so a term with 550,000 monthly searches came back as zero volume. Nothing crashed. The pipeline produced clean, well-formed output all the way through, and I wrote an entire strategy document on top of numbers that were manufactured by a fallback value. It was caught because a person looked at the conclusion and found it implausible, then asked to see the raw numbers.
That is the stop I would most want in your job file and the hardest one to write, because the test is does this result make sense given everything I know that is not in this file — and by definition none of that is in the file. The closest I have got is the closure check in Done means: force it to account for every line, so a silent skip has to show up as a number that does not add up. It does not replace the person. It gives the person something to notice.
One more thing I have not done: I have not yet run this same folder across several different assistants to see what changes. That comparison comes later in this series. The short version of why it matters: the folder is yours, the model is rented — change assistants and the files come with you.
Which assistant are you using?
The folder is just files, so anything can read it. What differs is whether it gets read automatically, or whether you hand it over each time.
| You're using |
How to connect it |
Tier |
| Codex, Cursor, Copilot's coding agent, Windsurf, Zed, Amp, Devin, Warp, goose, opencode and the rest of the 23 tools named on the format's site |
Put the folder in the project root. Read automatically, no configuration |
✅ |
| Claude (web or desktop) |
Upload into a Project. ⚠️ It cannot take a folder — you upload files one at a time, and a large project switches to retrieving what it needs rather than loading everything each turn |
🟡 |
| ChatGPT (web) |
Add to Project files. ⚠️ Hard file caps by plan, and OpenAI's own two help pages disagree on the Plus number — the Projects page says 25, the file uploads FAQ still says 20. Contents are prioritised, not guaranteed to be read in full |
🟡 |
| Claude Code |
Does not read AGENTS.md — it reads CLAUDE.md. Put a line reading @AGENTS.md at the top of a CLAUDE.md, or symlink one to the other, then run /context in a fresh session to confirm it loaded |
❌→🟡 |
| Gemini CLI |
Looks for GEMINI.md by default. Point contextFileName in its settings file at AGENTS.md instead |
❌→🟡 |
| Any web assistant, for the folder structure itself |
Nothing loads a directory tree and picks the nearest job file for the task at hand. That behaviour exists only in the command-line tools. In a web assistant you are uploading a flat pile and naming the job in your message |
❌ |
That last row is the one to sit with. The jobs/ folder here is a filing system for you, not a mechanism the web assistants execute. On the web you get the same result by naming the job in your message — "Run the month-end close job." It works. You are just doing the routing a command-line tool would do by looking at the directory.
Questions people ask
Do I have to throw away my existing SOPs?
No, and in several trades you legally cannot. The version people read stays as it is, including wording a regulator, an insurer or a client contract requires. The agent version is a second file pointing back at the first. Two audiences, two documents, one of which is allowed to be boring.
Is there a Word version of this template?
Deliberately not. Word's tables, merged cells and text boxes are exactly the parts that get scrambled on the way in — and project files in Claude, for instance, are processed as extracted text only. A .txt or .md file arrives intact. If your organisation requires a Word original, keep it as the human document and let the job file be the plain one.
How long should one job file be?
Short enough that all of it is loaded. The size caps above are per-tool and none of them announce themselves when you cross them. Working habit: one job per file, and if a file passes roughly two pages, that is usually a sign it is two jobs.
Can the assistant write the job file for me?
It can draft the shape from your existing procedure, and that saves an hour. It cannot supply the thresholds, the tests or the stop points, because those are facts about your business that exist nowhere it can read. Skip its questions and it will fill them in with numbers that read beautifully and mean nothing. The prompt above is written to make it ask first.
What about steps where the input is a phone call or a piece of paper?
Then the job starts after the transcription, not before it, and the trigger should say so. Pretending a job file covers a step that begins with a voicemail is how you end up with a procedure nobody follows.
What if a client or a regulator asks to see our procedure?
Show them the human one. If they ask what the assistant does, the job file is unusually good evidence — better than most process documentation, because it states its own tests and stop points in writing. An accident of the format, but a useful one.
What to do in the next thirty minutes
- Open the procedure you run most often. Not the most important one — the most frequent one.
- Highlight every instance of review, verify, check, assess, ensure, appropriate, as needed, timely, reasonable, relevant, significant. That set is your work list.
- Pick the step you run most and write three lines: what it looks at, what makes it pass, what to do when it fails.
- Add a Done means section at the bottom with one closure check — a line that would make a silently skipped item visible.
- Save it as a text file next to your existing document. Do not delete anything.
That is a job file. It will be wrong in two places, and you will find out which two on the first run.
Stay in the loop (no account signup)
This site does not ask you to create a product account. Free readers just leave an email—or follow where the build is posted.
| Channel |
What you get |
Where |
| Email (free) |
Occasional field notes as we pressure-test more systems in the wild. Articles on the site stay free. |
Open aiworkflowpro.com, scroll to Subscribe, enter your email, confirm the link in your inbox. |
| X |
Short ops notes and build-in-public updates |
@aiworkflowprolk |
| YouTube |
Longer industry-workflow rebuilds |
@aiworkflowprolk |
No paywall on this article. No "sign up for access." If you only want one next step: use the email box at the bottom of the site, or follow on X if you prefer the timeline.
Want to go deeper? If you already work at the command line, the mechanics of running multi-step jobs are covered in Claude Code Dynamic Workflows — that one is written for developers, not for the trades.
— hh