Autonomy Level 1, Assisted: when your team checks every run
Human in the loop AI, in a business, means something practical: a person sits at two points in the workflow. Someone on your team starts the AI employee, and someone on your team checks every result before it counts. For us, that's the first of four autonomy levels: Autonomy Level 1, Assisted. Every process starts here. This is where the trust grows that your team needs before it lets anything out of its hands. And this is where it becomes clear whether a process is even described well enough for an AI employee to take it over. This article covers what your team does at this level, how training works and how you can tell a process is ready for the next level.
What human in the loop means in your business day to day
On our home page, Level 1 is one sentence: the AI employee does its task when someone starts it, and your team checks every run.
That sentence holds three answers:
- Who starts the run? A person. The AI employee waits until someone gives it a nudge. It doesn't have a fixed schedule yet and doesn't react on its own to an incoming email.
- Who checks? Your team, on every run. Not as a spot check, and not only when something looks off.
- What do you see? Everything. Every result passes through someone on your team before it's used further.
That sounds like little relief, and at first it is. The gain at this level is a different one: your team sees every case the AI employee handles and learns where it's confident and where it isn't. The AI employee gets to know every variant of the process. And for the first time, you see in black and white how your process really runs.
The four levels apply per process, not to the whole business. One process can be at Level 1 while another is already at Level 2. How all four levels fit together is covered in the main article, The four autonomy levels: how a process in your business runs on its own.
One distinction belongs right at the start. The levels apply to AI employees that take over processes. AI employees you develop ideas and concepts with keep working with you in conversation.
With a thinking partner, the human is always in the loop, because the conversation is the work. This article isn't about those AI employees. It's about work that simply has to get done.
Why Level 1 shouldn't be a permanent state
At Level 1, your team shifts from doing the work to checking it. While the AI employee is still learning, that's a good trade. As a permanent state, it gets you little.
The reason is simple. If someone has to start and check every run, the process still depends on a person. If that person is sick or on vacation, it stops. And if you are the person doing the checking, you're still the bottleneck, just in a different spot.
That's why, for us, Level 1 is a stage with a clear end. The goal is Level 2, Autopilot: the process starts on its own on a schedule or trigger, runs through to the final report and hands anything uncertain to a person.
What happens at Level 2 is described in Autonomy Level 2, Autopilot: the process starts on its own and reports back at the end.
The human stays in the workflow after that too, just at different points. They get a report, they get the uncertain cases, and they sign off at the points where something is at stake. We decide beforehand which points those are.
When we keep a human involved
For some steps, a house rule applies, no matter which level the process is at: we keep a human involved when decisions are about people, about money or about anything with legal consequences.
In concrete terms:
- About people: who gets hired, who gets let go, who gets a contract and who doesn't. The AI employee may prepare and compile. A person makes the decision.
- About money: payments, discounts, credit notes. Here too, the AI employee prepares and a person signs off.
- With legal consequences: anything that triggers a deadline, makes a commitment or lands with a public authority.
On top of that, there's a second group we look for in every process: steps that can't be undone. A payment, a submission to a public authority, an email to a client. Once it's out, it's out. Steps like these get their own approval point: the AI employee stops until a person explicitly says yes, every single time, for every individual case. If no answer comes, nothing happens.
This is a house rule for our work, not legal advice. Where decisions affect people legally or in a similarly significant way, the law also sets limits on purely automated decisions (Art. 22 GDPR). For filing, the inbox and reports, though, our rule is an operational one, not a legal one. What applies in your case is something to settle with someone who is liable for the answer.
One distinction helps here. The escalation path kicks in when the AI employee is unsure: then it doesn't handle the case, writes to a person and lists the case as open in its report. The approval point kicks in at a fixed spot, always, even when the AI employee is sure. The escalation path catches what it doesn't know. The approval point catches what shouldn't happen without a person.
Training in four steps
Level 1 is the stage in which your AI employee is trained. The approach is the same one you know from a new person on your team. It has four steps: watch, sit alongside, remember and ask, sign off.
1. Watch
Before the AI employee touches anything, we watch your team at work with you, to see how it handles the process today. Ideally the person who does it most often, on a normal day, with real incoming work. They work and say out loud what they're doing and why. We take notes and mark every spot where they pause for a moment. That's where a decision is hiding, and every decision later becomes a rule or a variant.
How those notes become a process profile is covered in Documenting processes so an AI employee can take them over.
2. Sit alongside
Then the AI employee joins in, one variant at a time. You take a real case. It looks at it and says what it sees: which variant this is, how it recognizes it, how it would fill in the fields. The person on your team confirms or corrects. Only then does it carry out the step, and even then only up to just before saving. At the end, it writes down the variant and sums up what it has learned.
The rule here: every variant once. Not just the most common one. It's the rare ones that cause trouble later.
3. Remember and ask
The AI employee saves every variant it has learned. Anything it doesn't know goes to a person on your team via the escalation path. After that comes a short feedback session: the person explains the case, the AI employee saves the new variant as a rule, and next time it knows how.
At Level 1, this is the most important part. This is where your team learns that the AI employee asks instead of guessing. Without that experience, nobody lets a process out of their hands.
4. Sign off
Once all variants have been learned and the escalation path has been triggered and checked once, the runs that count begin. That's what the next section is about.
What your team checks on every run
"Check every run" only stays useful if it's clear what to look at. Otherwise it turns into quick clicking through, and that's worse than no check at all, because it fakes safety. We recommend that the person checking looks at five things:
- Is the result right? Is the document in the right place, and is every field filled in correctly?
- Did it recognize the variant correctly? A misread variant is a typical reason for a mistake that doesn't show up right away.
- Did it ask when it was unsure? Every unclear case belongs in an email to a person, not in a guess.
- Did it stay within its limits? Nothing deleted, nothing sent outside, nothing decided that a person decides.
- Is everything logged? Every run, every intervention and every escalation email are in the training log. Whatever isn't written down, nobody remembers three weeks later.
If someone had to step in because a result was wrong, that points to a gap. The case gets trained in, and the count of clean runs starts over.
Four clean runs: the threshold to Level 2
When a process may leave Level 1 isn't something we decide by gut feeling. The threshold has three steps.
First: four clean runs in a row. A run goes from start to finish; with daily incoming work, that's one day's batch. Clean means: no error and no intervention. An error or an intervention resets the count to zero. We recommend not counting an escalation email as an error, because it's the intended path when the AI employee doesn't know something. It still gets logged, because a feedback session follows. After the fourth clean run, the AI employee is signed off.
Second: watch alongside. The AI employee now works in real operations, and the person who used to do the process checks every result. You decide beforehand for how long, in days or in runs. If you don't decide that beforehand, you'll be watching forever.
Third: on its own, with a report. The AI employee works without a nudge and sends a completion email after every run. From here on, the process is at Level 2.
The sign-off is a sentence with content, not a "looks fine". It names the process, from when it runs on its own and who gets which report. It's recorded with name and date on an acceptance sheet. If someone asks three months from now since when the AI employee has worked on its own, the answer is there.
The four runs are Kevin's criterion from the tax advisory firm, and they work together with the side-by-side phase. The four runs show that the AI employee can do the process. The side-by-side phase shows that it can do it in everyday operations too, with everything everyday operations bring in.
What this looks like for us
At a tax advisory firm, the first process was filing digital tax assessment notices. Up to 30 to 40 come in there each day, and the office team used to need around 2 minutes per notice.
First, Kevin watched the office team: which types of assessment there are and how each one gets filed. Behind it are around 45 rules for the entire notice process, digital and on paper: many types of assessment, different layouts and types of tax.
Then an AI employee for filing was trained, one type of assessment at a time, with someone sitting next to it.
Anything it can't match with certainty goes to the office team by email, and the AI employee is trained on it afterwards. It was signed off after four clean runs, and after that, the team watched alongside it for a while. Today it runs on its own and reports back with a completion email. The notices that arrive on paper were added as the second process. For them, a second AI employee checks the inbox every 30 minutes and hands the emails with paper notices over to filing. Both processes are at Autonomy Level 2, Autopilot (as of September 2026).
In Kevin's own business, an even stricter rule applies: anything that leaves the business only happens after his explicit sign-off. He runs it with 14 AI employees.
"Get it done" doesn't count as a sign-off there, only a clear "send it" or "put it online", noted with wording and time. And before anything leaves the business, a separate role checks it, one that didn't create it.
That's the human at exactly the points where something is at stake. For all the other work steps, nobody asks for interim sign-offs.
Frequently asked questions
How long does a process stay at Level 1?
Until all variants have been learned, four runs have gone through cleanly and the side-by-side phase you set beforehand is over without errors. How long that takes depends mostly on the number of variants and on how often the process comes up. A process that runs every day gets to four runs faster than one that runs once a month. We don't give a fixed time commitment for this.
Who on the team should do the checking?
The person who does the process today. They know the cases, they spot a mistake fastest, and in the end they have to live with the process running on its own. As the owner, you're usually the wrong person at Level 1. Often, the person who does the process every day knows more about the special cases than you do.
Does every run really have to be checked, even when everything looks fine?
At Level 1, yes. That's exactly what sets it apart from Level 2. It's the runs that look fine that show you whether the AI employee also recognizes the rare variants correctly. Spot checks come later, once the process runs on its own and you only read the report.
What happens when the AI employee makes a mistake?
The person checking corrects the result and writes down what was wrong. Then there's a feedback session, the case gets trained in, and the count of clean runs starts over. At Level 1, a mistake shows up in the check before the result is used any further. That's what the level is for.
Who decides, and who only establishes facts?
Behind human in the loop is a question that goes beyond Level 1: what may an AI decide on its own at all, and what stays with the human? Kevin has written down a split on his blog that separates establishing facts from deciding: What Can an AI Lead Role Decide? There, it's about a single role and its authority. Here, it's about how far an entire process runs on its own. The two fit together: the more clearly it's written down what an AI employee may do, the sooner your team can stop checking every run.
Your next step
Level 1 is the stage in which your team gets to know the AI employee and trusts it with more, step by step, until the team only signs off. If you want to know which process in your business should go through this level first, we'll look at it together on a call.