Back to Insights TOC

Operations

Max Effort Is Not a Quality Strategy

If the top of the slider is your default, you are paying for seriousness theater instead of matching effort to the job.

Thursday, September 10, 2026 AgentC Foundry

Most shops do not blow the budget on the wrong model first. They blow it on the highest effort setting they can find.

The move looks responsible. The job is real. The system is capable. The work matters, so the slider goes to the top. The run takes longer. It burns more tokens. It produces a prettier page. The operator feels serious. The business still does not have proof.

That is not a model problem. It is a control problem.

Effort is a job-phase setting, not a status symbol.

Serious work used to have a visible cost. You could see overtime, extra staff, extra meetings. AI hid that cost behind named tiers: low, medium, high, max. The names sound like a quality ladder. They are not. They are consumption knobs. They change how much searching, spawning, and polishing the system is allowed to do. They do not rewrite the definition of done.

When you hold the assignment still and only change the knob, a pattern shows up fast:

  • The lowest setting is often slow and messy, not cheap in any useful sense.
  • A middle setting can finish faster, with fewer tokens, and a cleaner plan.
  • The highest settings spend more clock and more compute on extra sources, extra agents, and extra chrome.
  • A cheaper older route can still complete the research and the plan.

If the highest setting does not improve the Done Contract, it did not buy quality. It bought latency.

Pretty is not proof.

This is where operators get fooled. The expensive run often looks better. The site is cleaner. The canvas has more boxes. The write-up sounds more complete. None of that is the job if the job was: find a real buyer problem, name the workflow, and show why someone would pay.

A cheap baseline with an ugly page that still named the market evidence can beat a polished page that only rearranged the same canvas. If you cannot inspect the claim, the extra effort did not help. It decorated.

AgentC Foundry’s older rule still applies: redesign the work before shopping for tools. The same rule applies before shopping for effort. Write the assignment. Write what would count as done. Write what a human has to see before the output can leave the shop. Then pick a default effort. Do not start at max and hope the spend becomes strategy.

Keep a cheap baseline in the bake-off.

If you only run the expensive route, you have no comparison. You will always conclude that the expensive route “did the work,” because it is the only thing that ran. That is how shops lock themselves into a default they cannot defend.

A useful bake-off is boring on purpose:

  1. Same prompt.
  2. Same tools.
  3. Same definition of done.
  4. One variable: effort, or one cheaper model on the same job.
  5. A human scores the artifacts against the contract, not against how impressive the page looks.

You do not need a laboratory. You need comparable threads and a scorecard. If a middle setting plus one follow-up beats a max-first run, max is not the default. A tighter second pass at a middle setting is often cheaper than paying swarm prices for a first draft.

Put the knob on the route card.

Operators do not need another model religion. They need a Model Route Card with an effort field:

  • Default effort for this job class.
  • Escalate only when uncertainty is high or inspectability failed.
  • Never-max unless the job truly needs wide search or parallel investigation.
  • Cheap baseline that must still be able to finish the plan.
  • Proof artifacts the human will actually check.

That is the operating object. Not “always use the newest high setting.” Not “always use the cheapest setting.” Match the knob to the job.

A first-pass research brief does not need a swarm. A production change with irreversible side effects does not belong on autopilot, no matter how high the slider goes. A client-facing prototype can look finished and still fail the contract if nobody can trace the claims.

The historical analog is not mysterious. Factories did not run every machine at redline because the order was important. They matched capacity to the station. Over-speeding one station created inventory, heat, and delay. AI shops are doing the same thing with tokens.

Stop treating token spend as a proxy for intelligence.

Token burn feels like work. It is not. A low setting can burn more than a middle setting and still miss the assignment. A max setting can burn twice as much and return the same shape of plan. If your reporting only shows spend, you will reward the run that looked busiest.

What belongs on the receipt:

  • What job ran.
  • Which route and effort.
  • What proof was produced.
  • Whether a cheaper baseline could have finished.
  • Whether a human accepted it.

If you cannot name those five, you do not have an AI operation. You have a meter.

Extra questions are not automatically laziness. Some systems stop instead of assuming. If the job wants checkpoints, that is a collaboration policy. Write it into the assignment. Do not try to buy silence with a higher tier.

The practical move this week

Pick one recurring job. Write the Done Contract in plain language. Run it once at your current default. Run it once at a middle setting. Run it once on the cheaper baseline you already pay for. Do not change the prompt. Do not “help” the expensive run. Score the outputs against the contract.

Then lock a default. Most shops will discover they have been paying max-effort prices for first-draft work. Keep the high setting for the cases where the middle run fails inspectability, not as a badge that the work mattered.

The businesses that will get hurt are not the ones without access to the highest tier. They are the ones who confuse access with judgment. The slider will keep getting more names. The names will keep sounding like quality. Quality will still live in the assignment, the proof, and the human gate.

Match the knob to the job. Keep a cheap baseline. Do not buy seriousness from the top of the slider.