AI Virtual Assistant for Small Business: Who Signs?

An AI virtual assistant for small business can take some jobs whole, help with others, and must never touch a third kind. Sort them by who takes the blame.

A three-tier sorting table for an AI virtual assistant for small business, showing which jobs stay with the person whose signature is on them

There are now 1,814 court decisions worldwide, 1,252 of them in the United States, where somebody filed AI-generated material that turned out to be invented — fabricated cases, fabricated quotes, authorities that said the opposite of what was claimed. That count comes from the database Damien Charlotin maintains at HEC Paris, updated 30 July 2026.

Here is the part worth sitting with: not one of those people was punished for using AI. They were punished for putting their name on what it produced.

The first famous one cost $5,000. In Mata v. Avianca, Judge P. Kevin Castel sanctioned two attorneys and their firm on 22 June 2023 after a brief cited six decisions that did not exist (opinion and order). Three years later the ceiling sits somewhere else entirely. On 15 April 2026 the Nebraska Supreme Court suspended an Omaha attorney from practice after a divorce-appeal brief in which opposing counsel counted 57 defective citations out of 63 — and after he told a justice, on the record, that he had not used AI, before admitting that he had (WOWT, Nebraska Public Media).

So the useful question about an AI virtual assistant for small business is not "what can it do." It can do a startling amount. The useful question is the one nobody in the search results answers: which jobs can you stop watching, which ones can you never stop watching, and how do you tell the difference before it matters?

What follows is a sorting rule, three filled-in tables for three kinds of trade, a 40-minute method for sorting your own work, and a file you can hand to your assistant so it stops offering to do things it must not do.

[NEEDS REAL RUN: this opening should also carry one job of ours that nearly went wrong — what it was, what the assistant produced, who caught it, and what it would have cost. That requires actually running a job with real consequences attached and logging the near-miss. Right now we have other people's court records and none of our own.]


Sort by who takes the blame, not by how hard the job is

Everyone sorts the wrong way. Open any list of tasks to hand to an assistant and the implied axis is difficulty: simple and repetitive at the top, complicated and creative at the bottom. Hand over the boring stuff, keep the clever stuff.

That axis fails in one specific direction — it makes the dangerous jobs look safe. Reviewing a bank statement is not hard. Confirming a payment is not hard. Certifying that a translation is complete is not hard. All three are trivial, all three are fast, and in three different trades each one is a job the rules say a named person has to do personally. Meanwhile plenty of genuinely difficult work is safe to hand over whole, because a wrong result shows up within the hour and the cost of being wrong is doing it again.

The axis that actually predicts trouble: when this comes out wrong, whose problem is it? Sort on that and everything lands in one of three tiers.

Tier Name The rule What the assistant is for
1 Hand it over whole There is a clear test for done, you would see an error, and being wrong is cheap It finishes the job. You look at the result, not the work
2 It works, you decide It can produce the material but the decision has consequences you own Drafts, candidates, reminders, first-pass checks. You make the call
3 You do it A licence, a signature, an oath, or a person is on the other end of it Nothing. It may prepare inputs. It may not perform the act

💡 In plain terms: this is how you already treat a new hire. Day one, they order the supplies without checking (tier 1), they draft the client email and you send it (tier 2), and they do not sign the cheque (tier 3) — not because signing is difficult, but because your name is on the cheque. None of that changes when the new hire is software. The only difference is that software never asks "are you sure you want me to do this?"


Tier 1 — hand it over whole: three tests, all three required

Two out of three is tier 2.

There is a test for done. Not a feeling of done — a test. "Every file in this folder has been renamed to the client-matter format" is a test. "The intake summary is good" is not. If you cannot write down what a finished result looks like in a sentence someone else could check, the job is not tier 1 yet. It might become tier 1 once you write the test.

An error is visible to you. The failure has to surface where you will trip over it. A misfiled document surfaces the next time you look for it. A wrong figure buried mid-way through a 30-page summary surfaces when a client asks.

Being wrong is cheap. Cheap means you can redo it. Ten minutes of rework is cheap. A missed filing deadline is not, no matter how simple the job looked.

The same three tests across three trades:

Trade A tier-1 job Test for done How you would see the error Cost of being wrong
Law firm intake Pull conflict-check fields out of an enquiry into a standard form Every field filled or explicitly marked "not given" Blanks are visible on the form Ask the enquirer again
Contracting Turn a walk-through voice note into a numbered punch list grouped by trade Every spoken item appears once, numbered, under a trade heading You were there; you read the list Add the missed item
Bookkeeping Flag every transaction this month with no matching receipt Flagged plus matched equals the full transaction list The arithmetic does not balance Re-run it

All three produce material for a decision. None of them make the decision. The conflict check is not run. The punch list is not certified. The unmatched transactions are not written off.

The most common error here is promoting a job to tier 1 because it feels mechanical. Extracting fields from an enquiry is tier 1. Concluding from those fields that there is no conflict is not — a missed conflict does not announce itself, and under the American Bar Association's Model Rule 1.18 the disqualification that follows is imputed to the whole firm, not just the person who took the call (ABA Business Law Today on Formal Opinion 510). Same enquiry, same five minutes, two different tiers.


Tier 2 — it works, you decide: four shapes and one trap

Tier 2 is the biggest tier and will stay the biggest tier. It arrives in four recognisable shapes.

It drafts, you finalise. The engagement letter, the client update, the scope note. It writes; you read it as if a stranger wrote it, because a stranger did.

It finds candidates, you pick. "Here are eleven transactions that look like duplicates." Narrowing is most of the work and none of the risk — as long as the narrowing is not silently removing something you needed to see.

It reminds you, you act. "The completion certificate was signed 41 days ago; the one-year correction period starts from that date." A reminder stays tier 2 even when it is dead simple, because acting on it is yours.

It checks its own work against your written test. Before handing anything back, it runs your list — which only works if the list already exists.

The trap in tier 2 is not over-delegation. It is that a good draft feels finished, so the decision step quietly degrades into scrolling to the bottom and hitting send. The degradation is invisible from outside; your process diagram still shows a review step.

Two things stop it. Make the check a separate action with its own output — a filled-in list, not a feeling. And make the check look for something specific. "Does this look right" catches typos, not a quotation that reads perfectly and does not exist in the case it is attributed to. Norton Rose Fulbright's review of this year's sanctions describes exactly that shift: fabricated quotations and mischaracterised holdings rather than obviously invented case names, "making errors far more difficult to detect" (AI in litigation: update on Gen AI sanctions in 2026). For the long version of why a claim has to be checkable against a source rather than merely fluent, we took that apart when we fed 8.5 million words of law to an assistant and watched what it did with citations.


Tier 3 — you do it: four tests, and every one is about a person

A job is tier 3 if any one of these is true. Not all four. Any one.

A licence stands behind it. Someone can take your ability to work away over this. The North Carolina State Bar is unusually blunt: Rule 1.15-3(i)(1) requires lawyers to review bank statements and cheque images for all trust and fiduciary accounts monthly, and the handbook states that "this review requirement cannot be delegated" (

📄

Lawyer's Trust Account Handbook

PDF document

Download PDF

). Read the reason they give, because it is the most instructive sentence in this article: the monthly cheque-image review exists so the lawyer might see a cheque made out to an improper payee such as an employee. The rule is not there because the task is hard. It is there because the person with something to lose has to look.

Someone signs or swears. Florida requires a contractor in direct contract with the owner to serve a sworn, notarised final payment affidavit at least five days before suing to enforce a lien. Courts treat service of that affidavit as an absolute condition precedent — miss it and the lien claim fails even if everything else was perfect (Fla. Stat. § 713.06(3)(d); Delta Painting, Inc. v. Baumann, 710 So. 2d 663 (Fla. 3d DCA 1998); statute text). An oath is not a document format. It is a person accepting a consequence.

Translation has the same shape. US immigration rules require a full English translation accompanied by the translator's certification that it is complete and accurate and that he or she is competent to translate — 8 C.F.R. § 103.2(b)(3), text at eCFR. A tool cannot make a competency statement about itself. Somebody has to, and that somebody is on the hook for it.

A wrong answer cannot be undone. Filing deadlines. Sent client communications. Money that has left the account. Anywhere "we'll fix it next round" is not an available option.

A person is owed the duty. Not an output — a person, with a claim against you. Telling a client something. Deciding not to tell a client something. Once you are discharging a duty to a human being rather than producing an artefact, delegating the artefact does not delegate the duty. The California State Bar says the quiet part out loud: hiring a properly trained and supervised bookkeeper is allowed, "however, you are still personally responsible for accounting to your clients and the State Bar for the money" (

📄

handbook

PDF document

Download PDF

).

One more, for the solo operators, because it is the most expensive line in most one-person businesses. Under IRC § 6672, when a business withholds income tax and the employee share of FICA and fails to remit it, the IRS can assess a penalty equal to 100% of the unpaid trust-fund amount against any individual who was a responsible person and wilfully failed to pay. Personal, assessable against several people in full simultaneously, and not dischargeable in bankruptcy (IRS Internal Revenue Manual 5.7.3). Your assistant can prepare the numbers. It cannot be a responsible person, which leaves you as one.

⚠️ The mistake almost everyone makes: "it's fine, I check it" is not a tier. It is an intention. Tiers are properties of the job — of what the rules require and what the failure costs — and whether you personally are diligent this month is not an input. Write the tier down while you are calm, so the version of you at 6pm on a Friday does not get to re-decide it.


Why tier 3 does not shrink when the model gets better

This is the section that will still be true in five years, and the one the search results never touch — because the people writing them sell the tools.

Everyone's intuition is that tier 3 is a temporary consequence of models being unreliable, and that as they improve, jobs migrate downward until the list is empty. That intuition is wrong for three separate reasons, and the reasons stack.

The rules name a person, not a quality level

Read the actual texts and you notice something. They do not set an accuracy bar and then permit delegation above it. They name who must act.

EU AI Act Article 14 is the cleanest example, written recently and specifically with capable systems in mind. Paragraph 1 requires high-risk systems to be designed so they "can be effectively overseen by natural persons during the period in which they are in use." Paragraph 5 goes further for certain high-risk uses: no action or decision may be taken on the basis of the system's identification unless it has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority (European Commission AI Act Service Desk, Article 14).

Two natural persons. Not "one natural person, or zero if the system scores above X." The requirement is stated in units of people.

The construction is everywhere once you look. ISO 17100, the translation services standard, requires that revision — the bilingual check of the translation against the source — be performed by someone other than the translator, with equal or greater competence, with no clause removing the second person for high-quality first drafts. A project either has two people on it or it is not a compliant project. North Carolina's cheque-image rule cannot be delegated to a bookkeeper today and will not become delegable to a better bookkeeper.

The load-bearing point: delegability is a property of the rule, not of the worker. Improve the worker as much as you like. The rule does not read the release notes.

The reason for the rule is often detection, not competence

When a rule exists because the task is hard, better tools genuinely relax it. When a rule exists because someone needs to be watching for a specific kind of betrayal, better tools do nothing.

Look again at why North Carolina makes the lawyer personally examine cheque images every month: so the lawyer might notice a cheque written to an employee. An adversarial control, assuming that somewhere in the process a person may be acting against the firm's interest, and placing a human with skin in the game at the checkpoint. An assistant that is 100% accurate at reading cheque images does not satisfy that rule, because accuracy was never what the rule was buying.

Most tier-3 jobs are like this. Conflict checks exist because two clients may have opposed interests, not because matching names is difficult. Sworn affidavits exist because someone must be exposed to perjury, not because the form is complicated. The American Institute of Architects' G704 asks the owner, the architect and the contractor each to sign at their own line, so that three parties simultaneously own a date that starts warranty clocks, occupancy rights and retainage release (AIA instructions for G704-2017). You cannot automate exposure.

Better output makes the review harder, not easier

Here is the uncomfortable one.

In 2023 the failure mode was cartoonish: six citations, none of which existed, in a brief a first-year associate would have caught by typing a case name into any research service. You find that failure by looking.

By 2026 it has moved. Norton Rose Fulbright describes quotations that do not appear in the cited opinions, real cases cited for propositions they never discussed, and affirmative misrepresentations of what a lower court held — far more difficult to detect. The Sixth Circuit's March 2026 decision in Whiting dealt with over two dozen fake citations and misrepresentations of fact across three consolidated appeals.

The output got better and the review got harder. A more capable assistant produces failures that cost more to find — because the tell that used to give it away, obvious wrongness, is gone. Capability shifts work into tier 3, not out of it: checks that used to be a glance now have to be real verification against a source.

🔬 What I actually think: better models change the second tier enormously and the third tier not at all. In tier 2 the gains are real and large — better drafts, tighter candidate lists, fewer round trips. In tier 3 the only thing that changes is the cost of failure, and it moves the wrong way. Which makes this list the most durable file you own. Everything else in your folder describes what the assistant can do. This one describes what you are accountable for, and nobody has shipped a model that takes over your accountability.

A fourth reason needs no argument. Read the terms of any of these products: the vendor does not accept liability for your professional outcomes, and your regulator's rulebook does not care which product you used. The risk was never distributed. It has been sitting in one place the whole time.


The table: three trades, one row per real job

This is the thing to take away. Steal the format; your rows will differ.

Face 1 — dealing with clients (law firm, agency, consultancy, estate agency)

The job Tier Why What the assistant may do
Turn an enquiry into structured intake fields 1 Test for done exists; blanks are visible; cheap to redo All of it
Send the "not yet confidential" warning before intake 2 Wording is standard; the moment of sending is a judgement about the relationship Draft it, prompt you to send it
Run the conflict check across current, former and prospective clients 3 Model Rule 1.18; a missed conflict is imputed to the whole firm and stays invisible until it is fatal Assemble the search terms and party list. Not run the decision
Decide whether to take the matter 3 Duty owed to a person; not reversible once you have heard the facts Nothing
Draft the engagement letter 2 Content is patterned; California requires written fee agreements above $1,000 in non-contingency matters and makes a non-compliant one voidable at the client's option (Bus. & Prof. Code § 6148) Draft against your template
Sign the engagement letter and hand the client a copy 3 Signature, with a statutory consequence for getting it wrong Nothing
Sign a written buyer agreement before the first showing 3 Required before touring a home under the NAR settlement practice changes effective 17 August 2024 Prepare it
Diary the deadlines that come out of the matter 2 Calendaring failures are 7.4% of legal malpractice claims in the ABA's 2016–2019 profile — worth automating and worth double-entry Create entries; you confirm them

Face 2 — delivering work (contracting, translation, design, any trade with a handover)

The job Tier Why What the assistant may do
Turn walk-through notes into a numbered punch list 1 Complete, visible, cheap to fix All of it
Group the list by trade and flag items with no responsible party 1 Structural, checkable All of it
Decide the work is substantially complete 3 Professional judgement that three parties sign; starts warranty, occupancy and retainage clocks Prepare the certificate draft
Sign the substantial completion certificate 3 Signature — three separate signatures, in fact Nothing
Cross-check lien waivers against the subcontractor list 2 It finds the gaps; you decide whether a gap blocks payment Find gaps, list them
Swear the final payment affidavit 3 Sworn and notarised; absolute condition precedent to enforcing the lien in Florida Nothing — you want to read every line yourself
Translate a document 1 or 2 Depends entirely on what happens to the translation afterwards Draft it
Revise the translation against the source 3 where the job is claimed as ISO 17100 compliant Revision must be a different person from the translator Pre-flag omissions and number mismatches for the reviser
Certify a translation as complete and accurate for an immigration filing 3 8 C.F.R. § 103.2(b)(3) requires a competency statement from a person Nothing

Face 3 — keeping the books (bookkeeping, accounting, one-person company)

The job Tier Why What the assistant may do
Match transactions to receipts and list the unmatched 1 The two lists must sum to the whole; arithmetic catches errors All of it
Categorise recurring, previously-seen transactions 1 Precedent exists in your own history; errors show in the category totals All of it
Categorise a transaction it has not seen before 2 Novel judgement with tax consequences Propose with reasoning; you confirm
Prepare the reconciliation working paper 2 It assembles; you check the three balances agree Assemble it
Review bank statements and cheque images monthly 3 Explicitly non-delegable under NC Rule 1.15-3(i)(1) for trust accounts, and a sound habit where no rule applies Nothing
Sign and date the monthly and quarterly reconciliations 3 Signature; North Carolina requires the record be kept six years Nothing
Investigate a client-ledger negative balance 2 → 3 It can find and explain; you must fund the shortfall or attach a written explanation Find and explain
Remit payroll withholding 3 IRC § 6672: personal, joint and several, not dischargeable in bankruptcy Prepare the figures and the reminder
Decide an unreconciled item is immaterial and write it off 3 Irreversible, and exactly the decision a policy time limit exists to force Surface the ageing

[NEEDS REAL RUN: these tables are derived from published rules, not from jobs we handed over ourselves. What is missing is row-level history — which jobs we put in tier 1, ran for a few weeks, and pulled back to tier 2, and what nearly happened first. That requires running a real workload against these tables for at least a month and logging every reclassification with its trigger.]


Sort your own work in about 40 minutes

Do not start from a list of things AI can do. Start from your own week.

Step 1 — list the jobs, not the tools (10 minutes). Open the last two weeks of your calendar and sent items and write down what you actually did, one line each, in your own words. Not "admin" — "chased three outstanding invoices", "wrote the scope change for the Hendricks job". Aim for 20 to 40 lines.

Step 2 — ask four questions, in this order (20 minutes). The order is the whole trick, because whichever question comes first sets the answer.

  1. If this comes out wrong, whose name is on it? Yours personally, the business's, or nobody's?
  2. Is there a licence, signature, oath, certification or filing deadline attached?
  3. If it came out wrong, how would you find out — and how long would that take?
  4. What is the worst realistic outcome, and can it be undone?

Question 2 is a hard stop. If the answer is yes for anything beyond the business's name, the job is tier 3 and you do not continue. Do not negotiate with yourself here; avoiding that negotiation is what the written list is for.

If questions 1 and 2 come back clean, question 3 decides between tier 1 and tier 2. If you cannot describe a check that would catch a wrong result, the job is tier 2 at best — no matter how mechanical it looks. And if question 4 says "cannot be undone", it is tier 3 regardless of everything above.

Notice what is not on the list: "can the assistant do this well?" That one comes last, and it only ever moves a job between tier 1 and tier 2. It never touches tier 3. Ask it first and you end up sorting by capability, which is where everybody starts and why everybody's list is wrong.

Step 3 — handle the ties (5 minutes). Three rules for the jobs you genuinely cannot place:

  • Unsure means tier 2. Never round up to tier 1. Being over-cautious costs a few minutes a week; being under-cautious is why this article opens with a body count.
  • A job that is sometimes tier 3 is always tier 3. "Tier 1 unless it's client-facing" is a rule you will misapply at 6pm.
  • Split jobs that straddle the line. Most tier-3 jobs contain a tier-1 job wearing a coat. "Run the conflict check" splits cleanly into "assemble the party list" (tier 1) and "decide there is no conflict" (tier 3). Splitting is how you get the speed without giving away the signature.

Step 4 — write it down (5 minutes). Not in your head. In a file.

⚠️ The most common failure of this exercise is doing it once, feeling clear-headed, and never writing it down — so six weeks later you are making the same judgement again, differently, under time pressure, with an assistant that is now much more persuasive than it was. The file is not documentation. It is a decision you made once so you never have to make it again badly.

[NEEDS REAL RUN: "about 40 minutes" is a design target for these four steps, not a measured time. It needs timing against a real work list — ideally three people in different trades — before it stays in the article as a number.]


Write it into `boundaries.md` so it gets read every time

The point of writing it down is not tidiness. Your assistant reads it before it starts, so it stops offering to do things it must not do.

The file sits next to the rest of your rules — the same folder holding what your business does, how you say things, and the steps for each job. If you have not built that folder yet, we walked through the minimum version in keeping your rules in one folder.

Here is a complete boundaries.md. Copy it, replace the rows with yours, delete the rest.

# What I do not hand over

I sort work by who is accountable when it goes wrong, not by how hard it is.
Three tiers. If you are unsure which tier something is in, treat it as Tier 2
and ask me. Never treat an unlisted job as Tier 1.

## Tier 3 — I do these myself. Do not perform them. Do not offer.

You may prepare inputs for these. You may not carry them out, and you may not
produce a finished artefact that looks like the act itself.

| Job | Why it is mine |
|-----|----------------|
| Deciding whether a conflict exists | A missed conflict is imputed to the whole firm |
| Signing anything | My signature is the thing being relied on |
| Swearing or certifying anything | An oath needs a person who can be prosecuted |
| Monthly review of bank statements and cheque images | Non-delegable under our bar rules |
| Telling a client something material, or deciding not to | Duty owed to a person |
| Anything with a filing deadline attached | Cannot be undone |
| Remitting payroll withholding | Personal liability, survives bankruptcy |

If I ask you to do one of these, remind me it is Tier 3 and ask what input I
actually need instead. Do not comply on the second ask either.

## Tier 2 — you do the work, I make the call

Draft, shortlist, remind, and check against `checks/`. Then stop and hand it
back. When you hand something back, tell me in one line what you were unsure
about. If you were not unsure about anything, say that too — I will read it
more carefully.

| Job | Your part | My part |
|-----|-----------|---------|
| Engagement letters | Draft from template | Read, amend, sign |
| New-category transactions | Propose with reasoning | Confirm |
| Deadline calendar | Create entries | Verify against the source document |
| Lien waiver gaps | Find and list gaps | Decide whether payment is blocked |

## Tier 1 — finish these, tell me the result

| Job | Done means |
|-----|------------|
| Intake fields from an enquiry | Every field filled or marked "not given" |
| Punch list from walk-through notes | Every spoken item appears once, numbered, by trade |
| Unmatched transactions for the month | Matched + unmatched = full transaction list |

## Rules about this file

- If a job is not listed, it is Tier 2. Ask me.
- Do not move a job to a looser tier because you have got better at it.
  Only a change in the underlying rule moves a job, and only I can make
  that change.
- Quote the row you relied on when you decline something. I want to know
  which line stopped you.

Two lines in there are doing real work. "Do not comply on the second ask either" exists because the failure mode is not the assistant volunteering — it is you, tired, asking twice. And "quote the row you relied on" makes the file debuggable: when it declines the wrong thing, you can see exactly which line is badly written.

The file lives wherever the rest of your rules live. If you keep them in a chat workspace, upload it alongside the others. If you keep them in a folder on disk, AGENTS.md is the conventional entry point that a long list of assistant tools read without being told to — put a pointer to boundaries.md in it, and put the tier-3 table itself in it too, since that table is short and important enough to be worth repeating. AGENTS.md is plain Markdown with no required fields and no schema; its own FAQ says "Use any headings you like; the agent simply parses the text you provide" (agents.md, checked 31 July 2026).

Copy this prompt to build your own table

Paste this into whichever assistant you use — Claude, ChatGPT, Copilot, or anything else that holds a conversation — and answer honestly. It will not let you skip step 2.

I want to sort my own work into three tiers by who takes the blame when it
goes wrong — not by how hard the work is.

Ask me for the list of jobs I actually did in the last two weeks. One line
each, in my own words. Do not suggest jobs. Do not group them yet.

Then, for each job, ask me these four questions in this exact order, one at
a time, and wait for my answer before moving on:

1. If this comes out wrong, whose name is on it? Mine personally, my
   company's, or nobody's?
2. Is there a licence, a signature, an oath, a certification, or a filing
   deadline attached to it?
3. If it comes out wrong, how would I find out — and how long would that
   take?
4. If it comes out wrong and I don't catch it, what is the worst realistic
   outcome, and can it be undone?

Rules you must follow:
- If the answer to question 2 is yes for anything except "my company's
  name", the job is Tier 3. Stop asking. Do not argue with me about it.
- If I cannot describe a check that would catch a wrong result, the job is
  at best Tier 2, no matter how simple it sounds.
- If I am unsure, put it in Tier 2 and mark it "unsure". Never round up to
  Tier 1.
- Do not tell me you can handle a Tier 3 job. Do not offer.
- Do not give me legal, tax, or professional advice about my obligations.
  If a question turns on what my rules require, say so and tell me to check
  the rule, then move on.

When we're done, output a table with four columns: Job | Tier | The test it
passed or failed | What you are allowed to do on it. Then list separately
every job I marked "unsure".

And one for three months from now, when you will not remember why a row says what it says:

Read boundaries.md. For each job listed there, ask me one question: has
anything changed in the last three months about who is accountable for it —
a new rule, a new client type, a new filing requirement?

Only ask about jobs where something might plausibly have changed. Do not
walk the whole list.

If I say a job should move to a less strict tier, ask me what specifically
changed in the rule — not what changed in the tool. If my answer is about
the tool getting better, tell me that is not a reason to move it, and leave
it where it is.

Output the diff only: what moved, in which direction, and why.

What I got wrong, what I refuse to do, and what I have not tested

Four approaches I tried and threw away.

Sorting by risk score. A 1–5 rating per job collapsed within an hour. Risk scores are continuous; the decision is not. A job is either one you may stop watching or one you may not. Numbers invited me to negotiate — "it's a 3, so I'll mostly check it" — which is the exact behaviour the list exists to prevent. Tiers cannot be averaged; that is their whole advantage.

A blanket "human reviews everything" rule. Standard advice, nearly useless alone. A review with no written test degrades into a glance in about two weeks, and a glance is precisely the check that fails against fluent-but-false output.

Letting the assistant self-assess the tier. Tempting, wrong for a structural reason: the thing being asked has no exposure to the consequence. The tier depends on facts about your licence, your jurisdiction and your client, and those facts live with you.

Conditional tiers. "Tier 1 for internal work, tier 3 for client work" reads fine and fails in practice, because the classification then happens at the moment of use, under time pressure, by the person the rule was meant to protect against. Every conditional row I wrote eventually split into two unconditional rows.

What I refuse to do here. I have not recommended a product and will not. Not neutrality theatre — a recommendation would undercut the argument. Anything I named today would change its pricing, its limits or its ownership inside a year, and the tier-3 list would still say exactly what it says now. The list has to outlive the tools, so it cannot be written in terms of them.

What I have not tested.

  • Whether an assistant actually respects a declined tier-3 request on the second and third ask, across different products. The instruction is written. The adversarial test that would prove it holds has not been run.
  • Whether the tier-3 table survives being read out of a chat workspace where the assistant is not guaranteed to load every file on every turn. See the next section — I suspect it does not, reliably.
  • Anything about non-US rules beyond the EU AI Act article quoted above. The three tables lean on US bar rules and US statutes because that is where the public, citable material is densest.

[NEEDS REAL RUN: this section is where the strongest receipts belong and where the biggest gap sits. Missing: the specific jobs we handed over whole and then took back, why we took them back, and what nearly happened before we did. Also missing: the jobs we classified wrongly on the first pass and what corrected us. Both require running this against a real workload over weeks.]


Where this breaks: the file only helps if it actually gets read

Two honest limits, both limits of the products rather than the method.

Neither big chat workspace guarantees your file is read on every turn. Claude Projects has no hard file-count limit, but once the material approaches the context limit it switches to retrieving on demand rather than loading everything — documented behaviour, not a bug. ChatGPT Projects has hard file caps by plan and describes its behaviour as prioritising and using your files, which is reference language rather than guaranteed-load language. Neither tells you which turn was the one where your boundaries file did not make it in.

The practical consequence: do not rely on the file alone for the tier-3 rows. Repeat the tier-3 table at the top of any step file where a tier-3 decision could arise, and put it in your standing instructions as well as in the folder. It is duplication, it will drift slightly, and it is still the right call, because the failure you are protecting against is silent.

The second limit is that nothing you write survives in the assistant's own memory. What you write in a file travels with you to any tool. What a product remembers about you on its own does not, and no export format exists that another product reads. That is a whole subject of its own, and it is next in this series: what your assistant remembers, and what it loses the day you switch.


How to wire this into what you already use

The boundaries file is plain Markdown, which is the entire reason it works everywhere. The tools differ only in how they pick it up.

What you use How to connect it Grade
Codex, Cursor, Copilot coding agent, Windsurf, Zed, Amp, Devin, Warp, goose, Junie, opencode and the rest of the 23 tools listed on agents.md Put AGENTS.md in the folder root; they read it with no configuration
Claude, on the web or desktop Upload into the project's knowledge. You cannot upload a folder — files go up one at a time, and full loading is not guaranteed every turn 🟡
ChatGPT, on the web Upload into the project's files. There is a hard file limit by plan, and your files are "prioritised", not guaranteed to be read in full 🟡
Claude Code Does not read AGENTS.md. The official docs say so in one line: "Claude Code reads CLAUDE.md, not AGENTS.md." Fix it by writing @AGENTS.md inside CLAUDE.md, or by making CLAUDE.md a symlink pointing at AGENTS.md ❌ → 🟡
Gemini CLI Reads GEMINI.md by default. You have to set contextFileName before it looks at AGENTS.md ❌ → 🟡
Anything else Paste the tier-3 table into the top of the conversation 🟡 every time
Whatever the tool remembers by itself No cross-vendor standard exists. It does not come with you

Checked 31 July 2026: agents.md lists 23 named tools, and Claude Code is not among them. The site still describes adoption as "more than 60,000 open-source projects" — a figure published with the December 2025 donation of the format to the Linux Foundation's Agentic AI Foundation and not revised since, so read it as a floor rather than a current count.

That last row is the one worth staring at. Everything above it is a wiring problem with a known fix. The last row has no fix, which is why the important things belong in a file you own rather than in a feature you rent.


Frequently asked questions

Am I allowed to use an AI virtual assistant on client work at all?
In most trades, yes. US legal ethics guidance treats generative AI as assistance you supervise rather than something forbidden — ABA Formal Opinion 512, issued 29 July 2024, brings it under the supervision rule that already covers non-lawyer assistants. Supervision is the operative word: the live question is never permission, it is which jobs you may stop watching.

Does "human in the loop" satisfy the requirement?
Only when the loop does something. EU AI Act Article 14 requires oversight measures with a real capacity to intervene, and two people for certain uses. A loop where someone clicks approve with no defined thing to look for exists on the org chart, not in the process.

My trade has no written rules. Does any of this apply?
Remove the licence test and the other three still work: signatures, irreversibility, and duties owed to a person. A freelancer with no regulator still has clients who can sue and deadlines that do not move. The difference is that you write your own tier-3 rows instead of copying them out of a rulebook.

Should the tiers differ for a human assistant versus an AI one?
The definitions are identical, because the rules are written about accountability rather than about who does the typing. Two practical differences: a person tells you when something feels wrong, and a person cannot generate fifty plausible pages in ten seconds. Both push work toward the middle tier.

What about jobs that are tier 3 only because of one small step?
Split them — the highest-value move in this method and the most commonly skipped. Almost every tier-3 job contains tier-1 work wearing a coat, and separating them is where the actual time saving lives.

How do I stop my own list from quietly loosening over time?
Make the direction of change asymmetric. Moving a job to a stricter tier can happen any day, for any reason. Moving one to a looser tier requires a specific change in the underlying rule, written next to the row. Improvement in the tool is explicitly not a qualifying reason — put that sentence in the file, so future-you has to argue with it.

What happens to my list when I switch tools?
It comes with you, because it is a plain text file rather than a setting inside one product. The one thing that does not come with you is anything a product learned about you on its own.

What is the very first row I should write?
The one job you already feel uneasy about. Everyone has one — the thing you keep half-checking. Write that row first, put it in tier 3, and notice whether you feel relieved or annoyed. Relieved means it belonged there. Annoyed usually means it needs splitting.


Before you close this tab

  • [ ] Write down 20–40 jobs from your last two weeks, in your own words
  • [ ] Run the four questions in order, question 2 as a hard stop
  • [ ] Split at least one tier-3 job into its tier-1 assembly half and its tier-3 decision half
  • [ ] Put the "unsure" pile in tier 2 and leave it there
  • [ ] Save it as boundaries.md next to the rest of your rules
  • [ ] Repeat the tier-3 table inside the step files where those decisions come up

One sitting. Afterwards the assistant stops being something you supervise by vibes.

This list is the most valuable file you own. Everything else in your folder describes what the assistant can do, and all of it will be obsolete within a couple of model releases. This one describes what you are accountable for — true before any of these tools existed, and true after the current ones are gone. The folder is yours. The model is rented, and this file is the part of the folder that does not depreciate.

If you want a running start, the minimum working skeleton of the whole folder — the rules file, the step files, the checks, and a boundaries.md stub with the four tests already written in — is what goes out to the email list.

[NEEDS REAL RUN: the starter folder described in the sentence above does not exist yet. Do not ship this line until it does and until it is actually attached to the email list. If it is not ready at publication time, cut the sentence rather than soften it.]

Next in this series: what your assistant remembers, where that memory actually lives, and what disappears the day you switch products.

To go deeper on structuring the folder itself — and if you are already comfortable on a command line — we took that apart in how to build a set of files an assistant can actually navigate. If you are running the whole business alone, the wider picture is in the one-person company playbook.

One request. I am not a lawyer, a contractor, or a bookkeeper. I built these tables by reading the rules those trades are held to, and every row traces to a source you can open and check. If you work in one of them and a row is wrong — or if there is a row missing that would have saved you — tell me. That is the only feedback on this piece I actually want.

Stay in the loop (no account signup)

This site does not ask you to create a product account. Free readers just leave an email—or follow where the build is posted.

Channel What you get Where
Email (free) Occasional field notes as we pressure-test more systems in the wild. Articles on the site stay free. Open aiworkflowpro.com, scroll to Subscribe, enter your email, confirm the link in your inbox.
X Short ops notes and build-in-public updates @aiworkflowprolk
YouTube Longer industry-workflow rebuilds @aiworkflowprolk

No paywall on this article. No “sign up for access.” If you only want one next step: use the email box at the bottom of the site, or follow on X if you prefer the timeline.

— hh, AI Workflow Pro

Sources checked 31 July 2026: AI Hallucination Cases database · Mata v. Avianca opinion and order · Nebraska suspension coverage · EU AI Act Article 14 ·

📄

NC Lawyer's Trust Account Handbook

PDF document

Download PDF

· California Handbook on Client Trust Accounting · Fla. Stat. § 713.06 · 8 C.F.R. § 103.2 · IRS IRM 5.7.3 · AIA G704-2017 instructions · agents.md · Claude Code memory docs

Successfully subscribed! Check your inbox for confirmation.

Successfully subscribed! Check your inbox for confirmation.

Successfully subscribed! Check your inbox for confirmation.

Successfully subscribed! Check your inbox for confirmation.

Done.

Cancelled.