---
title: "Your Agent Didn't Get Dumber. Your Harness Did."
description: "A controlled study across 35 releases of five coding agents found no quality improvement tied to harness updates, and nearly double the token cost. The model was the same the whole time."
author: Subhadip Saha
published: 2026-10-02
updated: 2026-10-02
canonical: https://thatdevguy.in/blogs/your-agent-didnt-get-dumber-harness-did
tags: ["AI", "Agents", "Developer Tools", "Backend"]
---

# Your Agent Didn't Get Dumber. Your Harness Did.

Two weeks ago an agent that used to nail a specific class of refactor started botching it. Same task, same codebase, same model I'd pinned explicitly in the config. My first instinct, and I'd bet yours too, was "the model got worse." Weights drift, a quiet distillation pass, maybe they quantized something to save on inference cost. It's the easiest story to reach for, because it's the thing with a name and a version number everyone talks about.

Except the model hadn't changed. I checked. What had changed, three releases back without me noticing, was the CLI wrapping it, the thing that decides how much of the conversation history gets kept, how tool results get truncated before they're fed back in, when to retry a failed call and when to give up. The harness.

Nobody blames the harness. There's a study out of Queen's University that says this isn't just my anecdote, it's the pattern, and it's bigger than most people building on top of these tools realize.

## What a Harness Actually Does, and Why You Don't Think About It

A model takes text in and produces text out. Everything that turns that into "an agent that edits your codebase" is a layer of software sitting between you and the model, deciding what goes into each prompt, parsing what comes back, running the tools the model asks for, and feeding the results in for the next turn. That layer is the harness. Claude Code has one. So does Codex, Gemini CLI, Cursor, every coding agent you've used. The model is the engine. The harness is everything else, the fuel injection, the transmission, the thing that decides how the engine's output actually gets translated into motion.

Here's the part that's easy to miss: two harnesses running the identical model can produce meaningfully different output quality on the identical task, because of decisions that have nothing to do with the model's weights. How aggressively does the harness compact old context when the conversation gets long, and does that compaction summarize faithfully or quietly drop details that turn out to matter three turns later? When a tool call returns a huge response, does the harness truncate it sensibly or just chop it at a token limit mid-sentence? When something fails, does the harness retry with a clear explanation of what went wrong, or does it just resubmit the same call and hope? None of that is the model's decision. All of it shapes whether the model has what it needs to actually do the job well.

I'd always treated the harness as plumbing, invisible by design, the kind of thing you don't think about until it leaks. Turns out it leaks more than I assumed, and the leak gets misdiagnosed as something else almost every time.

Picture the same failing test, fed through two different harnesses wrapping the identical model. Harness A hits a test failure, captures the full stack trace, and on retry includes a short note: "previous attempt failed with AssertionError on line 42, expected value was X, got Y." Harness B hits the same failure, truncates the tool output at a fixed character limit because that's simpler to implement, loses the actual assertion details in the cut, and retries with the bare instruction "fix the failing test." Both harnesses are running the same model. Only one of them gave it something to work with. The model isn't smarter in the first scenario, it's just not being asked to debug blind. From the outside, as a user, both of these look like "the agent," and only one of them looks competent.

## When Compaction Quietly Rewrites What Happened

The compaction case is subtler and, in my experience, the one that costs the most without ever announcing itself. A long session accumulates context: files read, decisions made, constraints mentioned once in passing twenty turns ago that still matter. Eventually the harness has to compress that history to stay under the context limit. A careful compaction step preserves the constraint, "don't modify the legacy auth module, it's being deprecated next sprint," even if it summarizes away the surrounding conversation. A less careful one optimizes for brevity and drops exactly the kind of one-off caveat that doesn't look important in isolation, because nothing in the summarization step knows it'll matter later.

You don't see this happen. There's no error message for "I just summarized away something you needed." You see the downstream symptom, four turns later, when the agent confidently edits the legacy auth module because as far as its current context is concerned, nothing ever said not to. It looks exactly like the model making a bad call. It's actually the harness having already made the call, silently, several turns earlier, by deciding what was safe to forget.

## The Study That Actually Isolated the Variable

A team at Queen's University, Oussama Ben Sghaier, Hao Li, Bram Adams, and Ahmed E. Hassan, ran the study I wish someone had run before I spent an afternoon convinced my model had gotten dumber. Their setup was deliberately boring in the right way: hold the model fixed, vary only the harness, and measure what happens. They tracked 35 sequential releases across five major open-source coding harnesses, Codex, Qwen Code, Gemini CLI, OpenCode, and OpenHands, and then did a deep dive on one of them, Qwen Code CLI, running every release against the same 50 stratified tasks from SWE-bench Verified, a benchmark built from real, resolved GitHub issues.

The scale of what these tools are already doing made the question worth asking in the first place. A separate piece of prior work they cite documented more than 456,000 pull requests authored by five leading AI coding agents, across more than 61,000 repositories, over a six-month window. This isn't a niche workflow anymore. It's hundreds of thousands of real changes landing in real codebases, shipped by software most of those repositories' maintainers have never audited for how it actually makes its decisions.

What they found, holding the model constant across 35 releases: no statistically significant improvement in SWE-bench resolve rate. Continuous development, growing codebase complexity in the harnesses themselves, release after release shipping new features, and the actual bug-fixing quality didn't move in a way that survived statistical scrutiny. Worse, later harness versions consumed nearly double the computational tokens and tool calls of earlier ones, for the same benchmark, without a corresponding gain in how many tasks actually got solved correctly. You were paying almost twice as much per task and getting, on average, the same result.

That's not "harness updates don't matter." It's closer to the opposite: harness updates mattered enormously, just not reliably in the direction anyone shipping them intended. Some changes helped specific cases. Others quietly cost more for the same or worse outcome, and from the outside, wrapped in a version bump that reads like "new features, improvements, bug fixes," there's no way to tell which you got until you've already paid for it.

## This Isn't Hypothetical, It's Sitting in Public Issue Trackers Right Now

What makes this land for me isn't just the benchmark number, it's that the paper's own footnotes point at the exact kind of complaint I'd dismissed as model drift, except now attributed correctly. A Cursor forum thread titled, more or less, "it's getting worse and worse." A GitHub issue on Claude Code specifically about a token consumption spike. Another about Opus quota burning faster than expected, filed against the harness repo, not against anything to do with model weights. A Gemini CLI discussion about quality degradation. A Codex issue describing a harness change that caused what the reporter called a quality cliff. A Qwen Code issue about excessive token consumption.

Every one of those is a practitioner noticing the exact symptom the study measured, in production, and every one of those reports lives in the harness's repository, not the model provider's. The infrastructure for reporting "something got worse" already routes the complaint to the harness maintainers. The instinct for diagnosing it still routes to "the model." Those two things are pointed in different directions, and that mismatch is why this keeps happening quietly instead of getting fixed loudly.

<Callout type="warn">
If an agent's output quality drops after an update and you haven't changed which model you're pinned to, the harness is the first thing to suspect, not the last. The update that shipped between "it worked" and "it doesn't" is very often the thing you weren't watching.
</Callout>

## Why the Blame Lands on the Wrong Layer

This isn't really a technical mystery, it's an attention mystery. Model releases get announced. They get benchmarked publicly, compared across providers, written up in posts exactly like this one. A harness update ships as a changelog entry in a CLI tool that auto-updates in the background, with none of the scrutiny a model release gets, even though it's just as capable of changing what you experience as a user. You notice the thing that has a name and a launch post. You don't notice the thing that silently patched itself between yesterday and today.

There's also a structural reason the harness side resists scrutiny: it's genuinely harder to evaluate. A model's quality is, imperfectly but meaningfully, something you can benchmark in isolation. A harness's quality only shows up in how it behaves across dozens of edge cases, long conversations that need compaction, tool calls that return more data than expected, failures that need sensible retries, exactly the kind of thing that doesn't show up in a quick smoke test before shipping an update. The Queen's University team had to build a whole controlled longitudinal methodology to even measure this properly. Nobody's doing that before every Tuesday's CLI release.

It's the same asymmetry as the one between a database engine and the ORM sitting on top of it. Postgres gets benchmarked, versioned carefully, and upgraded deliberately because everyone understands it's load-bearing. The ORM gets bumped in a routine dependency update because it's "just a convenience layer," right up until a query that used to generate one efficient join starts generating an N+1 pattern after a minor version bump nobody read the changelog for. The database didn't get slower. The layer translating your intent into what the database actually executes changed, quietly, and the slowdown got blamed on "the database must be under more load lately." Harnesses sit in exactly that position relative to models: the unglamorous translation layer that everyone assumes is stable because the flashy thing underneath it gets all the attention, and all the scrutiny.

## What Actually Separates a Good Harness From a Bad One

If you're building on top of any of these tools, or building your own agent loop, the parts worth actually paying attention to aren't exotic:

**Context compaction that preserves what matters, not just what's recent.** A long-running agent session eventually has to summarize or drop old context to stay within budget. The difference between a harness that does this well and one that doesn't shows up three or four turns after the compaction happens, when the agent confidently acts on information that got summarized into something subtly wrong. You usually can't tell this happened until the output is already wrong.

**Tool-result handling that truncates intelligently.** A tool call that returns ten thousand lines of log output shouldn't just get chopped at a token limit mid-line. A harness that truncates with awareness of structure, keeping the error at the end, summarizing the noise in the middle, produces a model that actually understands what happened. A harness that just slices the string produces a model reasoning over half a stack trace.

**Retries that explain themselves instead of blindly repeating.** A failed tool call retried with the same parameters and no added context about why it failed is a coin flip. A harness that feeds the failure reason back into the next attempt gives the model an actual chance to adjust. The difference in practice is the gap between this:

```
retry(toolCall)
```

and this:

```
retry(toolCall, {
  reason: lastError.message,
  attempt: attemptCount,
  hint: summarizeFailureForModel(lastError)
})
```

The first one is cheaper to write and ships just as easily in a changelog that says "improved reliability." It isn't improved reliability. It's the same blind guess, a second time, burning another round of tokens and another full inference pass on a coin flip the model had no better information to win.

**Verification loops that check work before declaring it done.** The harnesses that perform best aren't the ones with the fanciest prompt, they're the ones that build in a step where the agent's own output gets checked, tests run, a diff reviewed, before the task is marked complete. That's a harness decision, not a model capability.

<Callout type="tip">
If you can, pin your harness version the same deliberate way you pin your model version. "Always update to latest" sounds like free improvements. The study suggests it's closer to a coin flip with a token-cost tax attached.
</Callout>

## Stop Waiting for the Next Model to Fix This

The instinct to blame the model is understandable. It's the visible, nameable, heavily marketed part of the stack. But the thing orchestrating how that model actually gets used, what it sees, how its mistakes get caught, how its failures get retried, is doing at least as much work in determining whether your agent is good at its job, and right now almost nobody's holding it to the same standard.

If your agent got worse recently and you're waiting for the next model release to fix it, you might be waiting for the wrong update. Check what actually changed first.

**Sources:** Ben Sghaier, Li, Adams & Hassan, *Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality*, Queen's University, ACM Trans. Softw. Eng. Methodol. (July 2026) · [arXiv:2607.03691](https://arxiv.org/abs/2607.03691)
