The Best AI for Small Business Is the One You Can Check

Everyone ranks the tools. Nobody tells you how to check the work. Four kinds of check, three copy-paste checklists, and the part a checklist can't catch.

A short written check sheet lying beside finished work, with the best AI for small business waiting to be judged against it

I searched best ai for small business on the day I started writing this and read the whole first page. Eight text results, a video block, and Google's own summary panel above all of it — a panel that is itself a list of tools with monthly prices beside each name.

Of the eight: three are ranked lists of tools, two of those published by companies selling one of the tools on the list. Two more are straight product pages. One is a US government explainer, one a Reddit thread asking strangers for recommendations, one a YouTube video.

Not one answers the question you actually have. You are not confused about which chat assistant exists — you have already used one. Your question arrives at 6pm on a Friday, after the thing has produced four pages of something that looks correct: how do I know this is right?

Short answer, and the rest of this article is the long version. Quality does not come from watching it every time. It comes from four lines you wrote down once, that get read before the work leaves your hands. Those four lines are a file. It sits next to your work, takes twenty minutes to write, and does not care which assistant you are using this year.


Choosing the best AI for small business is the wrong first question

The list posts all compare the wrong number. Twenty dollars a month, thirty, forty. None of that is what this costs you.

What it costs you is the reading.

Think about how the first month goes. You hand over a job — a client summary, a draft scope of work, a month of transactions categorised. It comes back fast and it looks good, so you read all of it carefully, because you do not yet know where it goes wrong. That takes thirty minutes. The job would have taken fifty. You have saved twenty minutes and acquired a new task called proofreading, which most people in the professions find slower and more tiring than doing the work, because you are hunting for an error you cannot predict.

Month two is where it gets expensive. The output has been fine four times in a row. So on the fifth you skim. On the eighth you skim harder. Nothing has gone wrong yet, and every day nothing goes wrong is a small argument for reading less tomorrow. None of that is a character flaw. It is what happens to any human doing repeated inspection of mostly-correct material, and it is why the professions stopped relying on attention a very long time ago.

And month three is the one nobody writes a list post about, because it involves an error that got out.

[NEEDS REAL RUN: this is where the wreck goes — the one job we handed over before we had a written check, what came back, when we noticed, and what it cost. Requires actually running a job pack unchecked for a few weeks and keeping the log. We have not done this, and I am not going to invent it.]

I can tell you what the industry-wide version of that error looks like, because someone counted — but this is somebody else's data, not my receipt. The American Bar Association's committee on lawyers' professional liability profiles malpractice claims. In the 2016–2019 profile, substantive errors, meaning actually getting the law wrong, account for 51.93% of claims. Administrative errors account for 19.59% — broken open: failure to calendar properly 7.4%, clerical error 4.08%, procrastination in follow-up 3.45%, failure to react to the calendar 2.54%. Client relation errors add 16.7%; conflict of interest is 4.9% on its own.

Nearly half the claims against lawyers have nothing to do with whether the lawyer knew the law. They are a date not entered, a document not sent, a question not asked, a step done in the wrong order. Exactly the failures a written check catches and a careful read-through does not — because a read-through checks what is on the page, and these are failures of what is missing from it.

So the bill is not the subscription. The bill is the reading you eventually stop doing.


Why watching does not work

Three reasons, and the third one is the interesting one.

Watching costs about what it saves. If you bill by the hour the arithmetic is immediate: an hour supervising is an hour not billed, worth exactly as much as the hour you saved. You moved the work, you did not remove it. Bill by the project and it is less visible, same trade.

Watching decays, and it decays fastest when things are going well. There is no version of this where you stay as alert in week nine as you were in week one. The check that exists only in your head is not a check; it is an intention, and intentions have a half-life.

Your profession already solved this, before anyone had an assistant to supervise. This is the part I did not expect when I went looking, and it is the strongest thing in this article. Look at what three unrelated trades independently decided to do about exactly this problem.

In construction, the AIA's general conditions put the burden somewhere specific. Section 9.8.2 requires the contractor to prepare and submit the list of everything still to be completed or corrected, before the architect walks the site. The architect verifies and amends it. Then owner, architect and contractor each sign. Nobody's judgement is trusted alone, and the person who did the work writes the first list rather than waiting to be caught.

In translation, ISO 17100 defines a minimum production process whose third step is a revision performed by a second person — someone with equal or greater competence and relevant subject experience. A translation cannot be revised by the person who wrote it. Not a premium tier, not an upsell: it is the condition of calling the job compliant at all. And revision here specifically means comparing the translation against the source segment by segment, not reading the translation on its own — because reading the translation on its own cannot catch a sentence that was never translated.

In trust accounting, North Carolina's Rule 1.15-3 does something even more specific: it defines the check as an equation. Three numbers — the control ledger balance, the sum of every individual client sub-ledger, and the adjusted bank balance — must be identical. And Rule 1.15-3(i)(1) requires the lawyer to review the bank statements and check images every month, personally. The state bar's handbook says plainly that this review cannot be delegated, and gives the reason: the lawyer looking at the actual check images is how you notice a cheque made out to an employee.

Three trades, three centuries of accumulated grief, one shared conclusion:

Write the criteria down. Check against the written thing, not against your memory of what good looks like. And separate the person doing the work from the person confirming it.

None of that was designed for AI. All of it applies unchanged. Your assistant is just a new occupant of a chair your profession already built rules for.


What a check actually is

A check is not "review the output." That instruction is worthless to a machine and nearly worthless to a tired human. A check is a small number of lines, each of which states what good looks like precisely enough that two different people would reach the same verdict.

The move that made my own checklists stop being vague: checks come in four kinds, and the four behave completely differently when you hand them over. Sort your lines into these buckets and you see immediately which parts can go and which cannot.

Kind 1 — Present: is everything that must be there, there?

The cheapest and most valuable kind. You are not judging quality. You are counting.

Bad line: Make sure the intake form is complete.
Good line: Does the record contain all six conflict fields — full legal name of the person or entity, related parties, adverse party, opposing counsel, court or venue, general subject of the dispute? Fail if any field is blank or says "unknown".

The good version works because it names the six things and defines failure. "Unknown" counts as blank, which sounds pedantic until the third time you find a form where someone typed "TBC" and everyone downstream treated it as answered.

Presence checks catch absence errors, and absence is what humans are worst at — you cannot notice a paragraph that was never written by reading the paragraphs that were. The European Commission's translation directorate gives omission its own error code, OM, alongside meaning, terminology, target-language norms, task-specific style and clarity. Omission earns its own category because it is invisible in the deliverable.

Kind 2 — True: does it match something you can look up?

Facts that have a source. A figure that came from a bank statement. A date that came from a court notice. A clause reference that either exists in the contract or does not.

Bad line: Check the numbers are correct.
Good line: For every figure, date and document reference quoted, state which source document you checked it against. Fail any that were checked against an earlier draft rather than the source.

That last clause matters more than it looks. The commonest way a wrong number survives four rounds of review is that everybody after round one checked it against round one. DGT's specifications for external translators require that quotes from already-published documents have been checked — against the published document, not the previous version of your file.

Kind 3 — Shaped: does it look and sound like ours?

Format, house terms, the words your profession requires and the ones it forbids. Everyone writes these first because they are easy, and they matter least of the four. Write them anyway — nearly free, and they stop the small embarrassments.

Bad line: Use professional language.
Good line: Does every ledger entry state who, what matter, what for, and what authorised it? Flag any entry whose description is a single word.

A ledger line that says "Deposit" is not wrong. It is untraceable, which is a different failure and a slower one to discover.

Kind 4 — Allowed: did anything cross a line?

The one people leave out, and the one that matters most.

Bad line: Follow the rules.
Good line: Does the intake record contain any facts of the matter? Fail if it does. Facts recorded before the conflict result is a stop, not a note.

That line exists because of a documented failure mode. Model Rule 1.18 makes a lawyer's duty to a prospective client unconditional — any information, not just harmful information — and disqualification under 1.18(c) reaches the entire firm. Which is why Kentucky's Ethics Opinion E-455 attaches a conflict consultation form headed by a warning telling the person not to hand over anything confidential until the conflict check is done. The order of operations is the control. And the ABA has confirmed twice recently that you cannot delegate your way out of it: Formal Opinion 506 for non-lawyer staff doing intake, Formal Opinion 512 for intake done through an AI tool.

Boundary checks are the only kind of check whose failure mode is not "we did it badly" but "we should not have done it."

Four kinds. Present, True, Shaped, Allowed. Every good check I have written fits one of them, and every vague check I have written failed to fit any.

[NEEDS REAL RUN: how many lines it took before the leak stopped. The honest version of this paragraph reads "my first check for job X had 3 lines, it took 4 more before the failure stopped recurring, and here is which 4." I do not have that number. It needs one job run repeatedly with a check in place and every escape logged.]

[NEEDS REAL RUN: the before-and-after. Same job, same kind of error, counted over the same number of runs without a written check and then with one. Without that comparison, everything above this line is an argument rather than a result, and I would rather say so than dress it up.]


Where the file lives, and the three lines that make it run

You do not need new software for this. You need a folder and about twenty minutes.

If you have been following this series, you already have a folder with a rules file in it, and the previous article covered how to write the steps for a job so that something other than a human can follow them. If you are arriving cold, the short version is: a plain folder, one plain text file at the top called AGENTS.md that holds your standing instructions, a jobs/ folder holding the steps for each recurring piece of work, and now a checks/ folder holding what "done properly" means for each of those jobs. (There is a longer walkthrough of that folder here: Keep rules in one folder.)

my-work/
├── AGENTS.md              your standing instructions
├── jobs/
│   ├── new-client-intake.md
│   ├── before-delivery.md
│   └── month-end-close.md
└── checks/                ← this article
    ├── new-client-intake.md
    ├── before-delivery.md
    └── month-end-close.md

One check file per job, same filename as the job. That naming is not cosmetic — it is how the instruction below stays one sentence instead of a lookup table.

Then add this to AGENTS.md:

## Before you hand anything back

Before you tell me a job is done, open the file in `checks/` with the same name
as the job and go through it line by line.

Answer every line with PASS, FAIL or CAN'T TELL, and quote the exact text you
used to decide. If any line is FAIL or CAN'T TELL, tell me that first, before
you show me the work.

Never edit a file in `checks/` to make something pass.

Four sentences. Whole mechanism.

Why AGENTS.md and not some product's settings panel. AGENTS.md is a plain-Markdown convention that OpenAI published in August 2025 and handed to the Agentic AI Foundation under the Linux Foundation in December 2025 — a body co-founded by OpenAI, Anthropic and Block. As of my check on 31 July 2026, the project site lists 23 named tools that read it, and there are no required fields, no template to fill in and no version number. It is a plain text file with whatever headings you choose.

And there is one line in that project's own FAQ that is the reason this article exists at all. Asked whether the agent will run testing commands found in the file automatically, the answer is: yes, if you list them — it will attempt the relevant checks and fix failures before finishing the task.

That sentence was written about test suites. It works exactly the same way when the checks are a lawyer's conflict sequence or a bookkeeper's three-number equation. Nobody has to build anything for this to work. You just have to write the checks down where it looks.

One honest caveat while we are here: the same documentation that describes this behaviour is careful to say that instruction files are context, not enforced configuration. A rules file makes something much more likely. It does not make it certain. Anything you need to be certain about belongs in the last section of this article.


Three checks you can copy today

Three of them, one for each shape small professional work tends to take. Each is short on purpose, and each line names its kind so you can see the pattern. Read the one closest to your work, then rewrite it in your own words — the wording matters less than the file existing at all. Everything they are built from is public and free to download, and the sources follow each one.

1. Taking on a new client

For anyone whose risk lives at the front door: solicitors, agents, consultants, anyone who has to know who they are dealing with before they can help.

# Check: new client intake

Run this before anyone here reads the facts of a matter.
Answer every line PASS / FAIL / CAN'T TELL. Quote the text you used.

1. PRESENT — Does the record contain all six conflict fields: full legal name of
   the person or entity, related parties, adverse party, opposing counsel, court
   or venue, general subject of the dispute?
   FAIL if any field is blank or says "unknown", "TBC" or similar.

2. PRESENT — Did our first message carry the non-confidential warning in full,
   before we asked for any facts?
   FAIL if the warning came after the first request for facts, or is absent.

3. TRUE — Was every name from lines 1 searched against all three lists: current
   clients, former clients, and people who consulted us and did not hire us?
   Report the three counts separately. FAIL if any list was not searched.

4. ALLOWED — Does the record contain facts of the matter?
   FAIL if it does. Facts before the conflict result is a stop, not a note.

5. SHAPED — Is the engagement letter written, and does it state scope, fee basis,
   and each side's responsibilities, with a countersigned copy ready to hand over?
   FAIL if any of the three parts is missing or the copy is not countersigned.

STOP RULE: if line 3 or line 4 fails, stop and tell me. Draft nothing.

Built from: ABA Model Rule 1.18 and Comment 4 (the three lists, and limiting the initial consultation); Kentucky Bar Association Ethics Opinion E-455 and its conflict consultation form (the warning notice and its position); ABA Formal Opinions 506 and 512 (delegating intake to staff or to a tool does not move the duty); California Business & Professions Code § 6148 (written fee agreement above $1,000, countersigned copy at signing, three required contents — and the agreement is voidable at the client's option if you skip it); New York 22 NYCRR § 1215.1 (written engagement letter before the representation starts).

What this does not cover: whether you should take the matter at all. Not a checklist question.

2. Before it leaves the door

For anyone who delivers a finished thing: builders, designers, translators, anyone with a client waiting at the other end and a sign-off in between.

# Check: before it leaves the door

Run this on the package before it goes out.
Answer every line PASS / FAIL / CAN'T TELL. Quote the text you used.

1. PRESENT — Is there an itemised list of everything still to be completed or
   corrected, written by us, before anyone outside inspects it?
   FAIL if the list is empty, or if the only list came from the other side.

2. PRESENT — Does the list carry the sentence that an incomplete list does not
   reduce our obligation to finish everything the contract requires?
   FAIL if that sentence is missing.

3. TRUE — Does every item on the list have a date it must be done by and a named
   person responsible? Report any item missing either. FAIL if that count is
   above zero.

4. TRUE — For every figure, date and document reference quoted, name the source
   document you checked it against. FAIL any that were checked against an earlier
   draft instead of the source.

5. ALLOWED — Has a second person compared this against the original, and are the
   maker and the reviewer different people?
   FAIL if the same name appears in both roles.

6. SHAPED — Does the sign-off sheet carry a signature line for every party that
   has to sign?
   FAIL if any party's line is missing.

STOP RULE: you can never answer line 5 with PASS. Report it as
"needs a human" every time, and name who it should be.

Built from: AIA A201-2017 § 9.8.2 (the contractor prepares the list first) and the G704 instructions (the architect verifies and amends it; owner, architect and contractor each sign); § 9.8.4 (certifying substantial completion does not waive warranty obligations) and § 12.2.2 (non-conforming work found within one year is corrected at the contractor's cost); the punch-list disclaimer sentence in line 2 is standard industry wording — you can read a real signed certificate with its real punch list in the City of Hammond, Louisiana council records from March 2020, which is free to download while blank AIA forms are not; ISO 17100's minimum process, in which revision is bilingual, segment-by-segment, and performed by a second person; and DGT's tender specifications, which make full revision before delivery a contractual obligation whose breach can terminate the contract.

If you deliver documents rather than buildings, swap lines 1–3 for the six error categories DGT scores against — omission, meaning, terminology, target-language norms, task-specific style, clarity — and keep lines 4–6 exactly as they are. Line 5 in particular becomes the whole game: a translation cannot be self-revised, and no amount of thoroughness by the person who wrote it substitutes for a second pair of eyes.

What this does not cover: anything requiring an oath. In Florida, a contractor's final payment affidavit has to be sworn and served on the owner at least five days before suing to enforce a lien, and the courts treat that as an absolute condition precedent — miss it and the lien claim fails on its own, no matter how good the rest of your paperwork was.

3. Closing the month

For anyone who has to say a set of numbers is right: bookkeepers, one-person companies, anyone holding money that is not theirs.

# Check: month end close

Run this after the ledger is final and before I sign anything.
Answer every line PASS / FAIL / CAN'T TELL, and always show the three numbers.

1. TRUE — Do these three match to the cent?
   (a) the control ledger balance
   (b) the sum of every client or job sub-ledger, listed individually
   (c) the adjusted bank balance = statement closing balance
       + deposits in transit − outstanding items
   FAIL if any two differ. Show all three even when they match.

2. PRESENT — Does any sub-ledger show a negative balance?
   List every one. FAIL if any exists without a written explanation attached.

3. SHAPED — Does every entry this period say who, what matter, what for, and what
   authorised it? Flag any entry whose description is a single word.

4. TRUE — Was payroll tax withheld this period actually remitted, and is there a
   confirmation reference? FAIL if there is no confirmation reference.

5. PRESENT — Are there open reconciling items older than our cutoff?
   List them with their age in days.

6. ALLOWED — Have the statement images for this month been reviewed, and by whom?
   This line is mine to answer, not yours. Report it as "needs a human".

STOP RULE: if line 1 fails, stop. Do not post adjusting entries to make it
balance. Tell me what the three numbers are.

Built from: North Carolina Rule 1.15-3(d)(2), which defines the three-way reconciliation as an equality between exactly those three figures, and 1.15-3(d)(3) and (h), which require the lawyer to sign, date and keep it for six years; 1.15-3(i)(1) and the state bar's trust account handbook for line 6 — the monthly review of statements and check images cannot be delegated, and the reason given is catching a cheque written to an employee; California's rule 1.15(d)(3) and (e), which require the three-way reconciliation monthly with a written record, and the state bar's handbook, which says in as many words that hiring a competent bookkeeper is allowed but you remain personally responsible; Tennessee State University's bank reconciliation policy for the no-entries-after-close discipline; the University of Rochester's policy, which forces reconciling items older than 120 days into a designated account rather than letting them sit; and IRC § 6672 for why line 4 is in there at all.

Why line 4 is the one I would not remove: under IRC § 6672, if payroll taxes were withheld from employees and not remitted, the IRS can assess 100% of the unpaid trust fund portion against any individual who was both a responsible person and wilful about it. That liability is personal, it can be assessed against several people at once for the full amount, and it is not dischargeable in bankruptcy. For a one-person company, the most dangerous line in the month-end close is not whether the numbers balance. It is whether the withheld money actually left the account.

What this does not cover: whether the categorisation was right. A perfectly reconciled set of books can be reconciled to the wrong accounts.


Making it grade its own work — and where that stops working

Once the check file exists, the obvious next step is to have the assistant run it before handing anything over. That works, with sharp limits worth knowing before you lean on it.

Write the file so a verdict is forced. Three rules make the difference between a self-check that catches things and one that flatters you:

  1. Three verdicts, not two. PASS, FAIL, and CAN'T TELL. If the only options are pass and fail, everything ambiguous becomes a pass. CAN'T TELL is the most useful verdict in the set, because it tells you where the check itself is badly written.
  2. Demand the evidence, not the conclusion. Every line should require it to quote the text it used. A verdict with no quote is not a verdict. This single rule kills most of the confident-and-wrong problem, because it is much harder to fabricate a quotation from a document in front of it than to fabricate a judgement about that document.
  3. Report failures before the work. If the failures come after four pages of nicely formatted output, you will read the output first and the failures with whatever attention you have left.

Now the limits, sorted by the four kinds, because they are not equally reliable:

Kind Self-check reliability Why
Present Good It is counting against a named list. This is the thing it is best at, and the thing you are worst at.
Shaped Good Pattern matching against a stated format. Same reason.
True Unreliable It can only compare what it can actually see. If the source document is not in front of it, "checked against the source" means "checked against my impression of the source".
Allowed Do not rely on it It has no stake in the answer and no consequence for getting it wrong. The person with the licence has both.

That table is why every checklist above puts its boundary line under a stop rule and marks it as needing a human. Not because the model is stupid — because grading yourself on whether you crossed a line is not a task anyone should be given, machine or otherwise. It is the entire reason ISO 17100 puts revision in somebody else's hands.

There is a fourth limit that is easy to miss, and it is structural rather than about capability. Instruction files are treated as context, not as enforced rules. The tools themselves say so. A check file makes the right behaviour much more likely and much more repeatable. It does not make it guaranteed. If you need a guarantee, the guarantee is you.

[NEEDS REAL RUN: the count that would make this section land — out of N self-checks, how many lines came back PASS that were actually wrong, and what kind were they. My guess is that they cluster in "True", but a guess is not a receipt and I am not presenting it as one.]


The counter-intuitive part: keep it short

Every instinct says a longer checklist is a safer checklist. It is not, for two separate reasons, and one of them is measurable.

The measurable one. These files get truncated, and they get truncated silently.

Tool Limit on instruction files What happens past it
Codex CLI project_doc_max_bytes, default 32 KiB Stops appending. No error.
Windsurf, global rules 6,000 characters Hard cap.
Windsurf, workspace rules 12,000 characters per file Hard cap.
Claude Code (CLAUDE.md) Loaded in full at any length But its own documentation says shorter files produce better adherence, and that specific, concise instructions are followed more consistently.
Claude Code auto memory first 200 lines or 25 KB Truncated.
Chat products with uploaded files File count caps, and retrieval instead of full reading Consulted, not guaranteed to be read each turn.

(Checked 31 July 2026. These change; check the current figure for whatever you use before you write a long file and assume it is read.)

Nothing warns you when a check falls off the end. The check simply stops being run, the output keeps looking fine, and you have less protection than you think you have — which is worse than having none, because none at least keeps you reading.

The unmeasurable one, which is bigger. You will not run a forty-line check either. You will run it twice, skim it the third time, skip it the fourth — exactly how you skimmed the output before you wrote anything down. And forty lines of instruction produce forty lines of verdicts, mostly PASS, in a wall nobody reads to the bottom of.

Four to six lines. More than six means one of two things: you have more than one job and it wants more than one file, or you have included things that never actually went wrong. Cut those. They are spending the attention the real ones need.

The way to grow a check is not to add lines up front. It is to add exactly one line the next time something escapes, and to write that line so specifically that the same escape cannot happen twice.


What a checklist cannot catch

This is the boundary of everything above, and it is the part I would keep if you only kept one section.

Judgement. A checklist can confirm the intake form is complete. It cannot tell you the client is lying. It can confirm the three numbers match. It cannot tell you the run of small round-number transfers is worth a second look. Every profession's competence lives in the judgements that resist being written down — that is roughly what makes it a profession.

Things that require having done this before. Some criteria you can only write after you have been burned in a specific way. Which is why the "add one line per escape" rule matters more than any template, including the three above. My three checklists are reconstructions from published rules. Yours will eventually be better than mine, because yours will have lines in it that only make sense to someone who has had that particular Tuesday.

Anything with your name on the consequence. This is the hard one, and it is not about capability at all. There is a category of task where the law has already decided that a person must do it, personally, and no quality of work product changes that:

  • North Carolina requires the lawyer to review trust account statements and check images every month, and the state bar handbook says explicitly that this review cannot be delegated.
  • A Florida contractor's final payment affidavit must be sworn — an oath is an act performed by a person, and skipping it makes the lien unenforceable regardless of the merits.
  • A translation submitted to USCIS must carry the translator's signed certification that the translation is complete and accurate and that they are competent to translate. No notarisation is required and no credential is required — but a statement of one's own competence is not a statement a tool can make.
  • ISO 17100's reviser must be a different person from the translator. Not a different pass. A different person.
  • California's own trust accounting handbook says you may hire a bookkeeper and remain personally responsible to your clients and the state bar for the money.

Notice what those five have in common. None of them is about whether the work is good. They are about who is answerable when it is not. A checklist can tell you whether the work meets the standard. It cannot take on the consequence of the work, and the tasks where the consequence is the whole point are tasks that do not get handed over at all — no matter how well the check performs.

Which is a different question from the one this article answers, and it is the next one worth asking: not how do I verify this, but should this have been handed over in the first place. Some work can go over entirely, some can only be assisted, and some has to stay with you. That sorting is the subject of the next article in this series, and it is the one I would read before you scale any of this up.


I'm not a lawyer, a contractor, or a bookkeeper

And I am not going to pretend otherwise, because the pretending is where this kind of writing usually goes wrong.

What I bring is not knowledge of your trade — the opposite. I do not carry your profession's assumptions about why a step exists. Someone inside the trade asks how to do step four faster; I ask why step four is there at all, and often enough the answer is that it is a control. Controls are exactly what translates into a written check. So I read what your professions already published — bar rules, standard contract conditions, an ISO process definition, a European Commission tender document, three university treasury policies — and cut them down to the shortest file that still catches what they were built to catch.

What I have not done is run these three checks against live matters and come back with results. That receipt does not exist yet, and inventing it would poison everything else in this article. So the ask is narrower than usual, and more useful to me:

Read the checklist for your trade and tell me which line is wrong. Which one is unenforceable in your jurisdiction, which one is missing the thing that actually goes wrong, which one would fail every time for a stupid reason. None of that comes out of a PDF.


Receipts

What I actually did, and what I did not.

Checked on 31 July 2026, first-hand:

  • Searched the phrase this article is named after and read the entire first page. Eight text results, one video block, one AI summary panel, four "people also ask" entries, and a reported total of 182 results for the query. Zero of them address how you verify the output. What each result was is described in the opening.
  • Pulled the search data for that phrase: 480 searches a month in the US, competition index 3 out of 100, cost-per-click $25.00. The twelve-month series runs 210 / 170 / 140 / 140 / 170 / 140 / 210 / 210 / 260 / 260 / 390 / 2900 — a slow climb with one large jump in the most recent month, which I would not call a trend on one data point.
  • Read the AGENTS.md project site and counted the compatibility list: 23 named tools. Read the FAQ, including the answer about running listed checks before finishing a task, which the section above restates.
  • Read Anthropic's memory documentation for Claude Code, which states that it reads CLAUDE.md and not AGENTS.md, that instruction files are context rather than enforced configuration, and that shorter files produce better adherence.

Solutions I looked at and rejected:

  • A single big checklist for everything. Fails on both counts in the "keep it short" section: it gets truncated by at least two tools at real-world lengths, and nobody finishes reading it.
  • Putting the checks inside the job steps. Tempting — in real professional documents the procedure and the criteria usually do share a file, and the European Commission's quality evaluation pack is exactly that. I split them because a check living inside the steps gets read while you are working, when you already believe you are doing it right. The value is in reading it afterwards.
  • Scoring instead of pass/fail. DGT scores translations against a credit pool, which works at their volume with trained evaluators. At small-firm scale it turns a decision into a number and then you have to decide what the number means. Pass, fail, can't tell.
  • Asking the assistant to write its own checklist. You get something plausible and generic — "verify accuracy, ensure completeness" — the exact failure this article is about. Useful for the format, useless for the content, because the content has to come from what has actually gone wrong.

Not tested:

  • Whether these three checks run the same way in a chat product with uploaded files as they do in a tool that reads the folder directly. I know the mechanisms differ. I have not measured the difference.
  • Whether self-grading holds up over dozens of runs, or drifts.
  • Any of this against a live matter, in any of the three trades.

Missing on purpose, marked in place above: the wreck that predates the checklist, the before-and-after on repeat failures, how many lines it took to stop the leak, and the self-check false-pass rate. Same kind of gap in every case — they need months of running rather than an afternoon of reading — and each is marked where it belongs instead of filled with something plausible.


How this connects to whatever you are using

The checklist is a plain text file. That part does not change. What changes is how your particular tool gets hold of it.

What you use How to connect it Grade
Codex, Cursor, Copilot's coding agent, Windsurf, Zed, Amp, Devin, Warp, goose, Junie and the rest of the 23 on the project list Put AGENTS.md and checks/ in the folder. Read automatically, no configuration.
Claude, web or desktop Upload the files into a Project. Note: folders are not supported — upload each file individually. Not guaranteed to be read every turn. 🟡
ChatGPT, web Add the files to a Project. There is a hard cap on file count, and the files are prioritised rather than guaranteed to be read. 🟡
Claude Code Does not read AGENTS.md. Create a CLAUDE.md whose first line is @AGENTS.md, and it will pull the other file in. (There is also a one-line command-line shortcut, ln -s AGENTS.md CLAUDE.md, but on Windows it needs administrator rights, so the one-line import is the safer route everywhere.) ❌→🟡
Gemini CLI Reads GEMINI.md by default. Point contextFileName at AGENTS.md in settings, or rename. ❌→🟡
Anything else Paste the check into the conversation before you ask for the work. 🟡 every single time
Whatever the vendor remembered about you There is no export format any other vendor can read. Not for any of them.

(Grades as of 31 July 2026. ✅ reads it as-is · 🟡 works with a detour · ❌ this one does not have it.)

Everything above that last row is a detour — annoying, five minutes, solvable. The last row is not a detour; it is a thing that does not exist, at any vendor, in any tier. Which is the whole argument for keeping your standards in files you can read with your own eyes: the four lines survive a switch, the assistant's private notes about you do not. That gets its own article later in this series, with the vendor-by-vendor table.


Common questions

How long should a check file be?
Four to six lines. Not for elegance — because the tools truncate. Codex stops appending past a 32 KiB default, Windsurf caps global rules at 6,000 characters, and Anthropic's documentation says shorter instruction files are followed more consistently. A forty-line check is a five-line check that nobody finishes.

Can I trust it to grade its own work?
For presence and format, yes, and it is better at those than you are. For facts requiring an outside source, only if the source is genuinely in front of it. For whether a line was crossed, no — mark those lines "needs a human" and mean it.

Does this work in the chat products or only in developer tools?
Both, by different routes. Coding tools read the folder. Chat products need the files uploaded into a project, with the caveats in the table above. The file itself is the same file, which is the whole point.

What happens to my checks if I switch next year?
Nothing. They are plain text files in a folder you own.

Where do I start with twenty minutes?
Pick the job that most recently went wrong. Write four lines, one of each kind. Add the four-sentence section to your AGENTS.md. Then run it once on a job you already finished and already know is correct — that tells you whether the check works before you rely on it.


Want to go deeper

These are for readers already working at a command line. If you are not, skip them; nothing above depends on them.


Stay in the loop (no account signup)

This site does not ask you to create a product account. Free readers just leave an email—or follow where the build is posted.

Channel What you get Where
Email (free) Occasional field notes as we pressure-test more systems in the wild. Articles on the site stay free. Open aiworkflowpro.com, scroll to Subscribe, enter your email, confirm the link in your inbox.
X Short ops notes and build-in-public updates @aiworkflowprolk
YouTube Longer industry-workflow rebuilds @aiworkflowprolk

No paywall on this article. No "sign up for access." If you only want one next step: use the email box at the bottom of the site, or follow on X if you prefer the timeline.

— hh, AI Workflow Pro

Successfully subscribed! Check your inbox for confirmation.

Successfully subscribed! Check your inbox for confirmation.

Successfully subscribed! Check your inbox for confirmation.

Successfully subscribed! Check your inbox for confirmation.

Done.

Cancelled.