Harness Engineering: What 11 Coding Agents Look Like Inside
In January I wrote that the harness is what separates an agent demo from something you can ship. In July, a paper opened up eleven coding agents and read their source code to see what that layer really contains.
The paper is "Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents" by Paul Barbaste, Tristan Darrigol, Germain Vu and Tom Wiltberger, from Wavestone AI Lab and Inclusive Brains. The authors read the source of each system, about four million lines of Python, TypeScript and Rust, and compare how they are built. There are no benchmark runs in the paper.
What they looked at
Eleven production coding harnesses. Four come from model vendors: Claude Code, Codex CLI, Gemini CLI and Mistral Vibe. Seven are open source: OpenHands, Aider, Mini-SWE-Agent, Hermes, Pi, OpenCode and OpenClaw. They also look at Omnigent from Databricks, a layer that runs other harnesses. More on that one below.
The definition they use fits in one line: an agent is a model plus a harness. The harness is everything except the model, which means the loop, the tools, the context, the safety controls, the orchestration and the ways to extend it.
Seven parts every harness has
Every system takes a position on the same seven parts, from a 100-line script to a million-line CLI. Sometimes the position is "we don't have one."

The range inside each part is huge. Mini-SWE-Agent has exactly one tool, bash. Claude Code has 43 typed tools and loads most of them only when they are needed. Mini-SWE-Agent keeps its whole history in one list, while Codex runs a sandboxed sub-agent that maintains memories across sessions in a git-tracked store. For safety, Mini-SWE-Agent has a cost limit and a step limit. Codex has policy rules, an LLM that reviews approvals, and an OS sandbox on three platforms.
Five things that stood out to me
1. The loop is the cheap part
Mini-SWE-Agent is about 100 lines of Python. Its authors report 74%+ on SWE-bench Verified, the same range as systems a thousand times bigger. A more sophisticated loop does not predict a better score.
So where do the other million lines go? Into safety, user experience, extensibility and, more and more, clients. In OpenCode, about three fifths of the non-test code is the TUI, web, desktop and SDK clients.
2. Nobody uses agent frameworks, and nobody does RAG over code
None of the eleven runtimes imports a general-purpose agent framework. That includes LangChain, LangGraph, AutoGen, CrewAI and about a dozen others they checked. Gemini CLI does not even use Google's own. Every loop is hand-written async code.
None of them retrieves code with vector embeddings either. They use ripgrep, glob, tree-sitter, and Markdown files like
AGENTS.md and CLAUDE.md that the harness picks up on its own.A footnote in the paper points out the irony: the term "harness engineering" was defined inside LangChain, whose libraries show up in none of these runtimes. LangChain responded by building a harness of its own, Deep Agents.
3. Skills passed MCP
Nine of the eleven systems support
SKILL.md skills, and eight support MCP. Skills now come with registries, trust tiers, and the first skills written by agents themselves.4. In 90 days, the vendors went from converging to copying
Eight of these systems were also in the paper's April edition, so the authors could diff the same code three months apart. In that window, Codex took Claude Code's hook event names word for word and shipped an importer for Claude Code sessions and settings. OpenHands adopted Claude Code's plugin format. Codex itself grew from 621K to about 1.12M lines of Rust.
The authors' summary: "The half-life of a competitive distinctive in this field is currently measurable in weeks."
Rules are also moving out of the prompt. Codex's newest prompts dropped the "don't commit" and "don't gold-plate" instructions in favor of feature flags, and Mistral Vibe deleted its "Never Commit" rule. Policy is moving from the prompt, where the model reads it, to configuration, where the platform enforces it.
5. The harness turned into a platform
Omnigent runs eleven vendor harnesses behind one API. A Claude Code orchestrator can hand work to Codex and ask Cursor to review it. One policy layer applies to all of them through each vendor's own hooks, and every adapter gets tested for streaming, tool calls, interrupts and policy denials. In the paper's words: "Harnesses are tested like hardware."
That is the paper's main thesis. What vendors compete on now is the ecosystem around the agent loop: plugins, marketplaces, importers and governance.
What they recommend
The paper ends with 18 recommendations, each tied to code you can go read. These are the ones I'd put on a sticky note:
- Start with a while loop and one bash tool. Add a tool only when you see a failure that needs it.
- Past about 15 tools, load them lazily. Claude Code's deferred loading makes the first prompt about 40% smaller.
- For edits with frontier models, use exact unique-string replacement. Never edit by line number.
- Pick up Markdown context files at the project and user level, and read your neighbors' file names too.
- Compact at a fixed buffer below the context limit. Keep the latest turns word for word, and merge summaries instead of starting over.
- Skip RAG over code. ripgrep and tree-sitter are enough.
- Write safety rules as data. If you offer a YOLO mode, keep a floor under it.
- Stay single-agent until you have a real breadth-first phase. Multi-agent runs use around 15 times the tokens of a chat.
- Use skills for workflows and know-how, and MCP for external systems.
They also include a 90-line Python harness that implements ten of the 18. Here is its core loop, simplified:
What I'm taking away
In January I quoted Anthropic's advice: start with the simplest workflow that works, and add autonomy only when the problem demands it. This paper gives the same advice with four million lines of production code behind it. The authors even note that Anthropic's Effective Agents articles line up closely with what all eleven systems do, though only one of the eleven comes from Anthropic.
What changes for me is where I put my attention. I have read a lot about loops and orchestration patterns. The code says the expensive work is in permissions, sandboxes, compaction, edit formats, and the glue that lets other tools plug in.
Two caveats come from the paper itself. The Claude Code analysis is based on a source snapshot that circulated in March 2026, not an official release, so it is the weakest part on reproducibility. And the benchmark numbers come from each project's own reports.
The authors also make a point I want to remember. Inventory claims, like tool counts and versions, go stale in weeks. Structural claims, like the seven parts and the two things nobody uses, have held up. That is a good way to read any post about agents, this one included.
Paper: Paul Barbaste et al., Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents, arXiv 2609.00006, July 2026. CC BY 4.0.
Cover: Eadweard Muybridge, The Horse in Motion, 1878. "Abe Edgington," owned by Leland Stanford, driven by C. Marvin, on the Palo Alto track. A horse, a harness, and a driver. Library of Congress, public domain.