Back to Insights TOC

Applied AI

The Model Got Smarter. Your Harness Got Fatter. Delivery Got Worse.

The invisible layer of custom instructions, skills, context packs, and corrective rules you control is accumulating debt faster than the underlying models are advancing — and more “helpful” additions can make the whole system less reliable on actual delivery.

Tuesday, July 21, 2026 AgentC Foundry

Your AI just got a major capability upgrade. You responded the way any serious operator would: you fed it more instructions to take advantage of the new power.

Then delivery got worse.

This is not a story about a bad model. It is a story about the harness — the owned layer of custom instructions, project files, saved prompts, memory, skills, tools, permissions, examples, and checks that shape what the model actually does before the next prompt ever runs.

One recent experiment made the problem unmistakable. After accumulating thousands of corrective words over time, the operator added roughly 5,000 extra words of “good” instructions to a capable model (Fable 5 in the report). Intermediate reasoning improved. The model thought harder and scored better on analysis. End-to-end delivery success dropped from three successful runs out of three to one out of three. The compact, older brief passed all three runs cleanly.

The engine got smarter. The harness did not. In fact, the harness got in the way.

This pattern is familiar once you name it. Every time the AI misses something, you add another rule. “Be more specific here.” “Always check that output.” “Include this format.” “Never forget the stakeholder list.” Each rule feels like a small, rational improvement. Over months they compound into 10k–18k+ words of invisible context loaded before the actual task even begins. The bloat is not obvious until you run a controlled test or an audit. By then the degradation is already baked into your results.

The deeper issue is loss of visibility and governance. The harness is no longer a deliberate design. It is sediment. Rules added to fix yesterday’s miss sit alongside rules for today’s different context. There is no retirement path, no classification of what still matters, and no measurement of net effect on delivery versus reasoning. The system slowly shifts from crisp execution to elaborate internal debate that never quite ships.

This is especially costly because the base models continue to improve. New releases add better autonomous execution, longer context windows, stronger goal-directed behavior. If your harness has not been maintained, the upgraded model can feel less reliable than a fresh chat on the old model. The very capabilities you paid for get diluted by the accumulated instructions that were written for a weaker version.

The practical response is not “use fewer words” as a vague slogan. It is to treat the harness itself as a governed artifact that requires periodic audit and maintenance.

An audit makes the invisible visible. You inventory the surfaces: custom instructions in your primary tools, project-level files, skills and context packs in your operating system, MemoryForge artifacts, cron prompts, daily brief templates, and any persistent state the agents load. You then classify each piece with a simple action lens:

  • Keep (still delivers clear value on the current model and current work)
  • Retire (no longer relevant, superseded, or net negative)
  • Combine (overlapping rules that can be collapsed)
  • Delay or gate (move from always-on to triggered only when evidence requires it)
  • Test in isolation (measure impact with and without)

The goal is not minimalism for its own sake. It is to restore visibility so you can make deliberate trade-offs. The experiment above is the proof: more context is not automatically better. Curation and compression matter more as the model gets stronger.

This principle already lives inside AgentC operating patterns, even if we have not always called it “harness audit.” The 4-bucket processing framework (Keep, Compress, Upgrade, Add) is itself an anti-bloat discipline applied to every signal and lesson. It forces compression of repetitive noise and upgrade into higher-order insight instead of endless accumulation. Durable learning files and skills are designed to be queryable, revisable artifacts rather than growing scrapbooks. The standing directive to “redesign the work before shopping for tools” is exactly the harness mindset: maintain and clarify the packaging layer you own before chasing the next model release or new automation.

Treating the harness as the trainable artifact also aligns with emerging self-improving frameworks. The document (or skill, or context pack) becomes the thing that evolves through rollout, reflection, scored improvement, and bounded edits — with a locked validation step that prevents regression. You do not let the AI endlessly append rules. You run a controlled loop that proposes changes, measures real delivery impact, keeps only honest improvements, and retires the rest.

For most operators the next step is not another tool. It is a recurring practice:

  1. Pick one high-volume surface this week (for example, your primary project instructions or your most-used skill files).
  2. Run a targeted inventory. Search for accumulated rules, examples, and “always” statements that have stacked up.
  3. Score a sample of recent outputs with and without suspect sections. Measure delivery success, not just how smart the intermediate thinking sounded.
  4. Make explicit retirement or refactoring decisions and document them.
  5. Add a lightweight “harness audit” item to your weekly synthesis or Kanban so it does not become another forgotten rule.

The models will keep getting better at reasoning and autonomy. That is not in question. The question is whether the harness you control will let those improvements reach the customer, the deliverable, or the decision. Right now, for many teams, the answer is no — because the harness has been allowed to grow without maintenance.

Audit it. Classify it. Retire what no longer earns its place. The difference between an AI that thinks harder and one that actually ships is almost entirely in the layer you can still see and govern.