Skip to content
Deployment & Production

'The Problem Is Not The AI Code, It Is That Nobody Knows Anything Anymore': Architecture Intent As Code - ADRs, Fitness Functions And Agent Briefs For Regulated Engineering Teams

Three essays dominated engineering Hacker News this week. Simon Späti argued that the crisis in AI-assisted development is not code quality but organisational knowledge loss: teams that 'ask Claude' instead of understanding their systems, shipping without a plan, with maintenance as 'the final boss'. Glyph asked what a serious AI product would look like and answered: mistake-checking as a first-class feature, prominent citations, and no auto-mode. Cal Newport called for investigating the labs. Underneath all three is a senior-engineering question a CTO can actually act on: when the code is cheap and the agents are fast, how does a team keep the intent - the why - alive and enforced? Our answer is to treat architecture intent as code: decision records that agents read, fitness functions in CI that fail when the intent is violated, and a brief that tells every coding agent what this system is for. For a bank, this is also how you keep the audit trail of why. With code.

AlchmAI Engineering14 min read

349

Hacker News points for Simon Späti's essay arguing the real problem is not AI code but teams no longer understanding their systems

597

Points for Cal Newport's 'It's time to investigate the AI labs' by 30 September - the week's most-discussed AI opinion piece

0

Tools most AI products give users to check for the mistakes their own disclaimers warn about, per Glyph's 'serious AI product' essay

3

Artefacts that keep intent alive and enforced: decision records, fitness functions and an agent brief

Simon Späti's essay landed because it named something senior engineers had been feeling. The problem, he argued, is not that AI writes average code. It is that teams have stopped knowing why their systems are shaped the way they are: they ask the model instead of building understanding, they ship without a plan, and the bill arrives at maintenance time - 'the final boss'. What still matters, in his telling, is design, architecture and intent; the model has no conviction, so the humans must supply the direction. Glyph's essay approached from the product side: a serious AI tool would treat mistake-checking as a first-class feature, show citations as large, dated, attributed objects rather than 'barely legible' footnotes, drop the anthropomorphic chatter, and never run in auto-mode without a reviewed plan. Cal Newport's went further up the stack, asking what is going on inside the labs at all.

For a CTO in financial services the useful question is narrower than any of the three: when agents write most of the code, what artefacts keep the intent of the system alive, and how do you make them enforceable rather than aspirational? Documentation that nobody reads does not survive a coding agent that never reads it either. The answer we have converged on is to make intent a first-class input to both humans and agents, and to make its violation a build failure.

1. Decision Records Agents Can Read

The classic ADR - context, decision, consequences - is nearly right. Two additions make it work in an agentic codebase: a machine-readable header so tooling and agents can find the decisions that apply to a path, and an explicit 'enforced by' link to the fitness function that checks it. A decision with no test is a wish.

markdowndocs/adr/0007-calculations-never-in-llm-layer.md
---
id: ADR-0007
status: accepted
date: 2026-06-14
applies_to: [src/answers/**, src/llm/**]
enforced_by: [fitness/no-arithmetic-in-llm-layer.test.ts, .dependency-cruiser.cjs#llm-cannot-import-calc]
supersedes: []
---

# Financial calculations never happen in the LLM layer

## Context
Customer-facing answers include tax, pension and repayment figures. Language models
produce plausible but unreliable arithmetic (see the Saturn study: 57% of money answers wrong).
We are a regulated firm under Consumer Duty.

## Decision
All figures come from src/calc (deterministic, versioned, tested against HMRC examples).
The LLM layer may only render figures it receives; it may never compute or restate them.

## Consequences
- src/llm must not import src/calc directly; it receives figures via typed results.
- Prompts must not contain instructions to compute.
- Any figure in an answer must carry a calc reference for audit.

## What would change this decision
A model with formally verified arithmetic, or a regulator accepting model-computed figures.

2. Fitness Functions: Decisions As Failing Tests

Fitness functions - a term from evolutionary architecture - are automated checks that the architecture still has the properties you decided it should have. In a TypeScript codebase, dependency rules are the highest-value ones and dependency-cruiser expresses them well; language-level rules can be small custom tests. Every ADR names the checks that enforce it; every check names the ADR it enforces, so a failing build tells the agent - or the human - which decision it just violated and where to read why.

javascript.dependency-cruiser.cjs
module.exports = {
  forbidden: [
    {
      name: "llm-cannot-import-calc",
      comment: "ADR-0007: the LLM layer renders figures; it never computes them.",
      severity: "error",
      from: { path: "^src/llm" },
      to: { path: "^src/calc" },
    },
    {
      name: "domain-has-no-framework-deps",
      comment: "ADR-0003: domain code is framework-free so it can be tested and reasoned about alone.",
      severity: "error",
      from: { path: "^src/domain" },
      to: { path: "node_modules/(react|express|fastify|next)" },
    },
    {
      name: "no-direct-broker-calls-outside-gateway",
      comment: "ADR-0011: only the order gateway talks to broker SDKs (pre-trade checks live there).",
      severity: "error",
      from: { pathNot: "^src/gateway" },
      to: { path: "node_modules/@broker/" },
    },
  ],
  options: { tsConfig: { fileName: "tsconfig.json" }, doNotFollow: { path: "node_modules" } },
};
typescriptfitness/no-arithmetic-in-llm-layer.test.ts
import { readFileSync } from "node:fs";
import { globSync } from "glob";

// ADR-0007: prompts must not ask the model to compute. Cheap, blunt, effective.
const FORBIDDEN = [/calculate the/i, /work out the/i, /compute the/i, /add up/i, /what is \d+ ?[%x*+/-]/i];

test("LLM layer prompts contain no arithmetic instructions (ADR-0007)", () => {
  const files = globSync("src/llm/**/*.{ts,md,txt}");
  const hits: string[] = [];
  for (const f of files) {
    const text = readFileSync(f, "utf8");
    for (const re of FORBIDDEN) if (re.test(text)) hits.push(f + " matches " + re);
  }
  expect(hits).toEqual([]);   // failure output names the file; ADR-0007 explains why
});

3. The Agent Brief

The final artefact is the one Späti's essay implies and Glyph's demands: a plain statement of what this system is for and what it must never do, at the root of the repository, read by every coding agent before it touches a file. Most agent tooling now reads a root instruction file. Ours has the same four sections in every regulated repository.

markdownAGENTS.md
# What this system is
A customer-facing financial answers service for a UK-regulated firm. Answers are grounded:
figures come from src/calc; the model explains, it never computes (ADR-0007).

# What you must never do
- Compute or restate a financial figure in src/llm or in a prompt.
- Add a dependency from src/domain to any framework (ADR-0003).
- Call a broker SDK outside src/gateway (ADR-0011).
- Delete or rewrite tests to make a build pass. Fix the code or stop and ask.
- Run destructive git or database commands. Propose them in the plan instead.

# How to work here
1. Read docs/adr/ for the paths you will touch (the front-matter lists applies_to).
2. Write a short plan naming the ADRs that apply before editing.
3. Run npm run fitness before proposing a change. A failure names the ADR you violated.
4. In your summary, list every file changed and every check you ran. Do not claim checks you did not run.

# Where things live
docs/adr/ decisions - fitness/ architecture tests - src/calc deterministic finance - src/gateway the only path to brokers

The last instruction in 'How to work here' comes straight from the week's other big story: OpenAI shelved a model partly because it misreported the work it had done. A brief that requires agents to list what they changed and what they ran - and a CI that checks it - is the same discipline at team scale.

“Documentation is what you hope people read. Intent as code is what the build refuses to ship without.”


Why This Matters More In A Bank

  • Model risk and change management want the 'why' of a system, not just its diff. ADRs with enforcement links are the audit trail regulators actually ask for.
  • AI-generated code arriving faster than review capacity - the CI bottleneck we wrote about this month - is only safe when the architecture rules are automated, because humans will not catch a layering violation in a 900-line PR.
  • Knowledge loss is an operational-resilience risk. If nobody can explain the payments path, nobody can fix it at 03:00. Intent as code is the cheapest insurance against that.

The Bottom Line

The week's viral essays agree on the diagnosis: AI makes code cheap and understanding scarce, and products that cannot check their own mistakes are not serious. The cure for an engineering team is to make intent a first-class, enforced input: decision records with machine-readable scope and named enforcement, fitness functions in CI that turn each decision into a failing test, and an agent brief every coding agent reads that says what the system is for, what it must never do and how to report honestly. For regulated firms that is also the audit trail of why. It is how we build and hand over systems as an AI engineering agency in London, and it is the most direct answer we know to 'nobody knows anything anymore'.

References & Further Reading

AI Agency Developer Londonarchitecture decision recordsfitness functionsAI coding agentsWorkflow Automation Agency architectureAI Automation London codeengineering leadership
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information