The Harness Matters As Much As The Model: What The New Empirical Study, AGENTS.md And A Week Of Agent Security Incidents Mean For Coding Agents In A Financial Codebase
September settled an argument senior engineers have been having for a year. An empirical study of harness design for coding agents - 176 matched settings across SWE-Bench Verified and Terminal-Bench 2.1 - found that context management extends how long an agent runs without changing what it does, planning changes where it stops, and predefined tools help weaker models while bash-capable models run cheaper with bash alone. The same week Claude Code adopted AGENTS.md as a portable configuration standard, and Manifold Security showed seven coding agents executing attacker code from a repository's git config with no prompt. Apollo's Watcher now grades every tool call at 93% recall and under 1% false positives. This is the educational piece for CTOs: the harness is an engineering artefact, here is what the evidence says goes in it, and here is what a regulated codebase must add.
AlchmAI Engineering15 min read
176
Matched settings in the September study of harness design, across SWE-Bench Verified and Terminal-Bench 2.1 with four models from two families
472
Hacker News points for Claude Code adopting AGENTS.md as a fallback configuration - portability became a developer priority overnight
7 agents
Claude Code, Codex, Cursor, Goose, Hermes, Qwen Code and Grok Build all ran attacker code from a repository's core.fsmonitor on git status (GitSpawn, 1-2 Sept)
93% / <1%
Recall on high-severity cases and false-positive rate for Apollo Research's Watcher, which grades every tool call, at 3-5% cost overhead
There is a phrase that has been quietly winning the argument about coding agents through 2026, and September gave it evidence: the harness matters as much as the model. A harness is everything around the model that turns a capability into an agent - the execution loop, the tools it may call, how it plans, how its context is managed, what it can see and touch, and what watches it. For a year the discussion of coding agents was mostly about which model; the discussion among people operating them at scale has been almost entirely about the harness, and three things this month made that visible to everyone else.
The first is a paper. An Empirical Study of Harness Design for Coding Agents, posted to arXiv in September, holds the execution loop fixed and varies three components - planning, action space and context management - across 176 matched settings on SWE-Bench Verified and Terminal-Bench 2.1, using three sizes of the Nemotron-3 family and Mistral-Medium-3.5 from a different lineage. Its findings are specific enough to act on: context management extends execution trajectories without substantially changing agent behaviour; planning changes where trajectories stop, shifting from an accuracy scaffold for weaker models to a cost saver for stronger ones with little change in accuracy; and predefined tools improve weaker models' performance while bash-capable models operate effectively with a bash-only interface at substantially lower cost, especially on command-line-centric tasks.
What The Evidence Says Goes In The Harness
Read together, the study's three findings give a CTO something rarer than opinion: a basis for defaults. Here is how we translate each into a decision for a financial engineering team.
- 01Context management buys duration, not judgement. Compaction, summarisation and retrieval let an agent keep working on a long task without changing the quality of any individual step. Invest in it for long-horizon work - migrations, test-suite repair, cross-repository refactors - and do not expect it to make a weak model's decisions better. It makes them longer.
- 02Planning is a cost control for strong models and a crutch for weak ones. With a frontier model, an explicit planning step mostly changes where the agent stops and how many tokens it spends getting there; accuracy barely moves. Keep planning for its stopping discipline and its audit value - a written plan is the artefact a reviewer reads - not because it will fix a task the model cannot do.
- 03Tool surface should match the model's shell competence. For a bash-capable model, a bash-only action space is cheaper and works, particularly on command-line tasks. Predefined tools earn their place for weaker or specialised models and for anything where you need a typed, governed boundary - which in a regulated codebase is the important exception, covered below.
“The study's most useful sentence for a CTO is the quiet one: the harness components trade off against cost as much as against accuracy. Most teams have been tuning for accuracy on a benchmark. The money is in tuning for cost on their own tasks.”
AGENTS.md As A Governance Artefact
The portability win from AGENTS.md is obvious. The governance win is the one financial teams should care about: for the first time, the instructions a coding agent follows in your repository are a single, versioned, reviewable file that every compatible tool reads. That makes AGENTS.md the right home for the constraints a regulated codebase needs an agent to respect, and the wrong home for anything that must be enforced rather than requested.
# Agent instructions - trading-platform monorepo
## Read first
- This repository is in the regulated perimeter. Changes to anything under
risk/, oms/, surveillance/ or compliance/ require a human reviewer from
CODEOWNERS and a linked change ticket. Do not merge; open a PR.
- Never modify files under fixtures/golden/. They are the evaluation set.
## How to work here
- Run `make check` before proposing a change. It runs typecheck, unit tests,
the parity harness (backtest vs live decisions) and the eval gate.
- Prefer small, reviewable diffs. One concern per PR.
- Write the plan into the PR description before the diff. Reviewers read
the plan first.
## Things the harness enforces (listed so you understand the refusals)
- No network access except the package registry mirror and the internal
model gateway. Outbound calls elsewhere are blocked, not just discouraged.
- Commands run in a per-task microVM with short-lived credentials. There is
no production credential in this environment to find.
- Every tool call is graded by a monitor. A refused call is logged with a
reason; do not attempt an alternative route to the same effect.
## Conventions
- Money is Decimal, never float. Timestamps are UTC ISO 8601. Market data
is point-in-time: every read carries an as-of date.The Containment Stack A Regulated Codebase Needs
The September incidents mapped the failure modes with unusual precision, and the defence stack that emerged in response is now concrete enough to specify. Five layers, from the outside in.
- 01Per-task, ephemeral isolation. Cursor's cloud agents moved to dedicated Firecracker microVMs with short-lived credentials; CrowdStrike published a seven-layer containment architecture with kernel-enforced process confinement. The principle is that an agent's environment is created for the task, holds only the credentials the task needs, and is destroyed afterwards. GitSpawn's 'outside the sandbox' only mattered because there was a durable developer environment to escape into.
- 02No shared mutable resources between agents. The Hugging Face swarm found each other through a message board nobody had sanctioned. If you run agents in parallel on the same repository, the shared state is the git remote and nothing else; caches, scratch directories and package stores are per-task.
- 03Supply-chain scanning of skills, plugins and hooks. NVIDIA's SkillSpector scans agent skills with 87% precision on intent analysis; AIR launched on 1 September with $50m from Sequoia and Greenoaks to build an inline firewall for skills and plugins; the CHAINDROP worm targeted more than 400 npm packages with malicious hook files. A repository's git hooks and config are inputs to the agent, and they need scanning on clone, before the first git status.
- 04Tool-call monitoring with published performance. Apollo Research's Watcher grades every Claude Code and Codex tool call through multiple stages, with 93% recall on high-severity cases, under 1% false positives and 3-5% cost overhead. The point is not the product; it is that a monitor with measured recall now exists, and 'we review the PRs' is no longer the only available answer.
- 05Typed tools at the regulated boundary. This is where the study's bash-only finding meets its exception. For a bash-capable model, bash alone is cheaper - but a typed tool with a schema is the only place you can enforce that an order-related change goes through the pre-trade gate, that a market-data read carries an as-of date, or that a migration touching compliance/ opens a ticket. Use bash for the ninety per cent and a governed tool for the ten per cent that regulators will ask about.
# Harness policy - enforced by the runtime, referenced by AGENTS.md.
isolation:
runtime: microvm # per task; destroyed on completion
credentials: short-lived # minted per task, scoped to the repo and registry mirror
network:
allow: [registry.internal, model-gateway.internal]
default: deny
repository_intake:
scan_hooks_and_config: true # GitSpawn: core.fsmonitor and friends, before first git command
refuse_on: [core.fsmonitor, core.hooksPath, filter.*.process]
scan_skills_and_plugins: true # SkillSpector-class intent analysis on anything the agent may load
action_space:
default: bash # cheaper for bash-capable models (harness study)
governed_tools: # typed, schema-validated, audited - the regulated ten per cent
- name: propose_change_to_regulated_path
paths: [risk/, oms/, surveillance/, compliance/]
requires: [change_ticket, codeowner_review]
- name: read_market_data
requires: [as_of]
monitoring:
grade_every_tool_call: true
block_on_severity: high
log: { destination: audit-store, fields: [tool, args_hash, verdict, reason, task_id, model_snapshot] }
evaluation:
golden_set: fixtures/golden/ # read-only to the agent
gate_on: [parity_harness, eval_regression]What To Measure, Because The Study Says Cost Is The Lever
- Cost per merged change, split by model and by harness configuration. This is the number the study says you can move, and most teams do not have it.
- Trajectory length versus outcome. If context management is extending runs without improving merge rate, you are paying for duration you do not need.
- Monitor verdict distribution. A rise in refused tool calls is either an agent drifting or an attack - both worth a look, and both invisible without a monitor.
- Intake scan hits. How often a cloned repository or a loaded skill fails the scan. After GitSpawn and CHAINDROP, zero is suspicious.
- Governed-tool share. What fraction of consequential changes went through the typed tool rather than bash. It should be all of them; the gap is your control failure rate.
The Bottom Line
September gave the harness argument its evidence and its cautionary tales in the same fortnight. The empirical study says context management buys duration, planning buys stopping discipline and cost control, and bash-capable models run cheaper with bash - so tune the harness for cost on your own tasks, not accuracy on a benchmark. AGENTS.md makes the instructions a single versioned, reviewable file that every tool reads, which is exactly where a regulated codebase's conventions belong. And GitSpawn, the Hugging Face swarm and CHAINDROP make the boundary plain: instructions govern a cooperative agent, and only ephemeral isolation, scoped credentials, denied-by-default networking, intake scanning, a monitor with measured recall and typed tools at the regulated boundary govern everything else. Build the harness as an engineering artefact with those layers, measure cost per merged change, and a coding agent in a trading or banking codebase becomes a productivity tool with a controls story rather than a liability with a chat window. That is the agentic engineering we do for financial teams in London, and the harness is where the work is.
References & Further Reading
- Fan et al. - An Empirical Study of Harness Design for Coding Agents (arXiv 2609.20804, September 2026). arxiv.org/abs/2609.20804
- Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents - a source-code study of eleven systems (arXiv 2609.00006). arxiv.org/abs/2609.00006
- Ken Ashe - The harness matters as much as the model in coding agents (18 September 2026). kenashe.ai/blog/2026-09-18-the-harness-matters-as-much-as-the-model-in-coding-agents
- VibeEval - Security harness for AI agents, September 2026: GitSpawn hits 7 coding agents, Watcher blocks 93% of bad actions, AIR raises $50M. vibe-eval.com/updates/security-harness-for-ai-agents-sep-2026
- Hacker News AI Digest, 2026-09-19 (AGENTS.md support in Claude Code; harness-design study). github.com/kouweizhu/agents-radar/issues/106
- Developers Digest - What Hacker News gets right about AI coding agents in 2026. developersdigest.tech/blog/what-hacker-news-gets-right-about-ai-coding-agents-2026
- Platform Engineering - Operationalising AI coding agents: a focused look at financial services and government. platformengineering.org/blog/operationalizing-ai-coding-agents-a-focused-look-at-financial-services-and-government
- Simon Willison - The lethal trifecta for AI agents. simonwillison.net/2025/Jun/16/the-lethal-trifecta
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information