Applied AI
If You Can't Score the Job, Autonomy Is Just an Expensive Guess
Before you let an agent into the live kitchen, decide whether the outcome can be checked the way a folded shirt can — not the way a plate of eggs can.
Businesses keep getting offered the same shortcut: skip the messy middle and let the system just run. The pitch now arrives as world models, physical AI, and agents that emit actions instead of paragraphs. The operating question is older than those labels.
Can you score the job?
If you cannot tell, in a way another competent person would accept, whether the work is done well, you do not have an autonomous job. You have an expensive guess with a better story. That distinction is not about model size. It is about the domain.
Laundry is a job. Eggs are a taste test.
Some work can be checked without a debate. A stack of laundry is folded or it is not. The shirts are in the drawer or they are on the chair. A second competent person can look at the result and agree.
Other work only looks finished. Scrambled eggs can be plated, photographed, and described as perfect. Taste is still a human call. Timing and texture do not live in a checkbox.
Most companies mix the two and then act surprised when “just let it run” produces theater. They put the agent on a job whose success condition is a vibe: a better email, a nicer proposal, a smarter strategy. Those can be useful. They are not, by themselves, a scorable domain.
The missing first page is a Verifiable Domain Card. Before you talk about models, connectors, or unattended runs, write down four things:
- What does done look like in the real world?
- Who can check it without you in the room?
- What evidence would make a skeptic agree?
- What happens if the agent is wrong?
If you cannot answer those four, you are still in the live kitchen arguing about eggs.
Reach is not the same problem. Neither is “should this work exist?”
A lot of failed AI work is misdiagnosed as a model problem.
Sometimes the agent cannot reach the work. The files sit behind a login. The system of record lives on someone else’s laptop. The job lives in a tool the agent is not allowed to touch. That is an access problem. Training more will not open a locked door.
Sometimes the work should not exist. You are automating a report nobody reads, a status meeting that exists because the pickup counter is missing, a process that should have been deleted. That is a cost-to-serve problem. A better agent only makes the waste faster.
Scoreability is a third card. The agent can reach the work. The work is worth doing. And you still cannot tell whether the output is good except by watching it yourself. That is the laundry-versus-eggs test. Autonomy without a scorer is not operations. It is a live kitchen with no recipe and no tasting spoon except you.
This is why “we will add a judge later” fails in practice. A judge that cannot be named, sampled, and disagreed with is just you, delayed. If the only person who can score the job is the owner, the company did not buy labor. It bought a new way to stay in the loop.
Actions are not paragraphs
A paragraph that says “fold the laundry” is not a policy. A policy is a trajectory: observe, act, check, halt. It has a named stop. It can fail a real test.
That is why so many agent demos feel impressive and then die in the business. They emit language about the job. They do not emit a checkable path through the job. The company then staffs a human to interpret the language, which means the agent never left the chat window.
If your current loop ends with “looks good to me,” you have a writer, not an operator. Writers are useful. Do not confuse them with unattended labor. The moment the output needs a taste test, keep a human in the scoring seat and stop calling the loop autonomous.
Named jobs make this easier. “Close yesterday’s invoices.” “File the inbound receipts.” “Move completed tickets to done.” “Draft the weekly exception list from these three sources.” Vague jobs make it harder. “Make it better.” “Sound more like us.” “Figure out what they meant.” Those may still be worth a model. They are not worth an unattended run.
Simulate the kitchen before you cook in it
The practical move is not to ban ambitious work. It is to refuse to train on the live kitchen first.
Spin up a disposable copy. Use a worktree, a held-out sample, a sandbox profile, a fake customer, a replay of last week’s tickets. Run the trajectory a hundred times where a miss costs a log line, not a real invoice, a real email, or a real customer record.
Simulation is not a science-fair extra. It is how you discover whether the domain is laundry or eggs before the customer is in the room. If the sandbox cannot score the job, the production system will not magically grow a scorer.
Specialization belongs here too. A job that only needs a named trajectory should not carry a general-purpose reasoner that wants to discuss the history of breakfast. Extra intelligence that cannot be scored is not insurance. It is more surface area for an expensive guess.
The same rule applies to “physical” work and to ordinary office work. Folding shirts and posting invoices look different. Operationally they rhyme: either a second person can check the result, or they cannot. If they cannot, you do not have a robot problem or a model problem. You have a scoring problem, and you should not let the agent cook unattended.
What to do on Monday
Write one Verifiable Domain Card for the next job you are tempted to just let run.
Name the job in buyer language, not model language. Then fill the four questions. If the answers are crisp, you have laundry. Put a halt owner on it, run it in a sandbox, then promote a narrow path. Keep the live kitchen behind a gate until the scorer is boring.
If the answers collapse into taste, you have eggs. Keep a human in the scoring seat. Use the model as a prep cook. Do not call that autonomy, and do not spend frontier-model money pretending a vibe is a workflow.
The companies that will waste the next year are the ones buying world-model stories for jobs they cannot grade. The companies that will compound are the ones that get ruthless about scoreable work first.
Fold the laundry. Score the fold. Then, and only then, talk about letting the agent cook.