'Agents Don't Need Memory, They Need Documentation': Why The Week's Most-Argued Engineering Essay Is Right For Finance, And How To Build Documentation-As-Memory For Agents That Must Be Audited
Kevin Liao's essay reached the top of Hacker News twice this week - 348 points on 5 October and 371 the next day, with more than 500 comments - by attacking the industry's favourite agent feature. Memory plugins, he argues, misread the problem: they capture conversation snippets, embed them and inject whatever is most similar, which surfaces by similarity rather than correctness, strips the context that made a decision make sense, treats stale history as truth, cannot help an agent search for what it does not know it is missing, and is impossible to audit. His alternative is documentation the agent consults before working and updates afterwards: requirements, decisions and constraints in lean Markdown files with a catalogue for discovery, committed and reviewed like code. For financial firms the argument is stronger than he makes it, because 'which memory influenced this decision' is a question a regulator will ask. This is the education piece: the argument, where it is wrong, and the documentation-as-memory architecture we use for agents in banks - with code.
AlchmAI Engineering14 min read
371
Hacker News points for 'Agents don't need memory, they need documentation' on 6 October, after 348 the day before - more than 500 comments in total
5
Failure modes the essay identifies in RAG-style memory: similarity not correctness, lost context, stale history as truth, unsearchable unknowns, unauditable influence
4 steps
The workflow it proposes instead of prompt, build, forget: prompt, consult, build, update
100%
Of an agent's context that a regulated firm should be able to reproduce when asked why the agent did what it did
Agent memory is the feature every platform is shipping, and Kevin Liao's essay - published on 3 October, updated on the 5th, and on the Hacker News front page both days - says most of it is solving the wrong problem. The standard design captures snippets of past conversations, embeds them and injects the most similar ones into the next prompt. Liao's critique is precise: similarity does not guarantee correctness or relevance; a snippet stripped of its surroundings loses the motivation and environment that made it true; history is treated as truth even though the codebase - or the book, or the policy - has since changed; an agent cannot search for what it does not know it is missing; and nobody can say which memories influenced which behaviour. His summary of the flawed premise is the quotable line: the industry decided that because agents forget, the fix is to remember more, index better and retrieve smarter.
His alternative is unglamorous and, we think, correct. Agents should work from documentation: requirements, decisions, constraints and architectural reasoning that the code - or the data - cannot express, kept in lean Markdown files, discovered through a small catalogue rather than a vector search, consulted before work and updated immediately afterwards while the agent still has the context, and committed so it is diffable, reviewable and shareable across a team. The workflow becomes prompt, consult, build, update. No database, no embeddings, no infrastructure: files and well-tuned instructions.
Where The Essay Is Wrong, Or At Least Incomplete
- Facts about the world still need retrieval. Documentation captures what the team knows and decided; it does not contain today's prices, last night's settlements or the client's current positions. Those come from systems of record through tools, at run time. Documentation-as-memory governs context, not data.
- Some per-user preference is genuinely memory. 'This analyst wants summaries in bullet points' is not a decision worth a document, and a small, explicit, user-visible preference store is fine. The point is that it should be explicit and visible, not inferred from similarity.
- Documentation rots too. The essay's answer - the agent updates docs immediately with full context - is right, but in a regulated firm those updates need review before they become truth, which is the one place a workflow step must be added.
The Architecture We Use
The shape is a repository directory the agent is instructed to read and maintain, a catalogue that makes discovery cheap, a consult step that records which documents were used, and an update step that proposes changes through the normal review path. Every piece is a file, so every piece is auditable with tools the firm already has.
# Agent documentation catalogue
Read this file first. Open only the documents relevant to the task. Update them when you learn something durable.
| id | file | covers | last_verified |
|-------|-----------------------------------|------------------------------------------------------------|---------------|
| D-001 | domain/settlement-cycles.md | T+1 vs T+2 by market, holiday handling, what counts as late | 2026-09-30 |
| D-014 | decisions/break-tolerance.md | When an amount break may be auto-matched; who decided; why | 2026-09-22 |
| D-022 | constraints/data-boundaries.md | Which data classes may reach which model; redaction rules | 2026-10-01 |
| D-031 | domain/counterparty-aliases.md | Known name variants for top 200 counterparties | 2026-09-28 |
| D-040 | decisions/escalation-policy.md | What must go to a human and how to phrase the hand-off | 2026-09-15 |
Rules: one topic per document; state the decision, the reason and what would change it; never paste transcripts.---
id: D-014
status: accepted
owner: settlements-ops
last_verified: 2026-09-22
---
# Amount-break auto-match tolerance
**Decision:** An amount break of 100 minor units or less may be auto-matched when both legs reference the same trade id
and the counterparty is in D-031. Anything larger, or any break on a tier-1 counterparty, goes to a human (see D-040).
**Why:** Analysis of Q2 breaks showed 94% under 100 minor units were rounding or fee differences. The 100 threshold was
chosen by ops leadership on 2026-07-08; the model-risk team reviewed it on 2026-08-02.
**Would change this:** a counterparty dispute traced to an auto-match; a change in fee structure; a regulator question.
**Do not infer from this:** that breaks under 100 are unimportant. They are still logged and sampled weekly.import re, pathlib, hashlib, subprocess
DOCS = pathlib.Path(".agent-docs")
def catalog():
rows = []
for line in (DOCS / "CATALOG.md").read_text().splitlines():
m = re.match(r"\| (D-\d+)\s*\| ([^|]+?)\s*\| ([^|]+?)\s*\| ([^|]+?)\s*\|", line)
if m: rows.append({"id": m.group(1), "file": m.group(2), "covers": m.group(3), "last_verified": m.group(4)})
return rows
def select(task: str, llm) -> list:
"""Bounded decision: the model picks document ids from the catalogue; it cannot invent one."""
ids = [r["id"] for r in catalog()]
chosen = llm.choose(question="Which documents are relevant to this task?", allowed=ids, context={"task": task, "catalog": catalog()})
return [r for r in catalog() if r["id"] in chosen]
def consult(task: str, llm) -> dict:
docs = select(task, llm)
commit = subprocess.check_output(["git", "rev-parse", "--short", "HEAD"], text=True).strip()
context, used = [], []
for d in docs:
text = (DOCS / d["file"]).read_text()
context.append(f"# {d['id']} ({d['file']})
{text}")
used.append({"id": d["id"], "sha256": hashlib.sha256(text.encode()).hexdigest()[:12], "commit": commit})
# 'used' goes into the run record: exactly which documents, at exactly which version, shaped this run.
return {"context": "
".join(context), "used": used}def propose_update(doc_id: str, new_text: str, reason: str, run_id: str, git):
"""The agent does not write truth; it proposes it. Review happens in the normal PR flow."""
branch = f"agent-docs/{doc_id.lower()}-{run_id[:8]}"
git.checkout_new(branch)
path = next(r["file"] for r in catalog() if r["id"] == doc_id)
(DOCS / path).write_text(new_text)
git.commit_all(f"docs({doc_id}): {reason}", trailers={"AI-Assisted": "true", "Run-Id": run_id, "Doc-Id": doc_id})
return git.open_pr(branch, reviewers=["settlements-ops"], title=f"Agent proposes update to {doc_id}", body=reason)
# Freshness without rot: the agent updates immediately, while it still has the context; a human merges.What You Get
- Reproducibility: a run record names the documents and versions the agent consulted. Replay the run with the same documents and you get the same context - the thing a model-risk review needs.
- Reviewability: every change to the agent's knowledge is a diff with a reviewer, not a row in a vector store nobody reads.
- Portability: the same documents work for any model or harness. Switch providers and the agent still knows your settlement cycles and break tolerances.
- Discipline: a catalogue with one topic per document, and a rule against pasting transcripts, keeps context lean - which also keeps it cheap and cacheable.
“The question a regulator asks is not 'did the agent remember?' It is 'what did the agent know, and who approved it?' Documentation answers that. Memory, as usually built, cannot.”
The Bottom Line
The week's most-argued engineering essay says agents need documentation, not memory, because similarity-based recall surfaces the wrong things, strips context, treats stale history as truth and cannot be audited - and proposes lean, catalogued, committed documents the agent consults before work and updates after. For financial firms it is right for a reason the essay understates: 'what did the agent know and who approved it' is a compliance question, and files with owners, versions and reviewers answer it where a vector store cannot. The architecture is a documented catalogue, bounded selection of documents, a run record of exactly what was consulted at which version, and agent-proposed updates that go through review - with live data still fetched from systems of record through tools. That is how we build agents that can be audited as an AI agency in London, and the essay has just made the case to a wider audience than we ever could.
References & Further Reading
- Kevin Liao - Agents don't need memory, they need documentation (3 October 2026, updated 5 October). liao.gg/blog/agents-dont-need-memory
- Hacker News - discussion of 'Agents don't need memory, they need documentation' (6 October 2026). news.ycombinator.com/item?id=49945933
- Michael Nygard - Documenting architecture decisions. cognitect.com/blog/2011/11/15/documenting-architecture-decisions
- AGENTS.md - an open format for guiding coding agents. agents.md
- Bank of England / PRA - SS1/23 Model risk management principles for banks. bankofengland.co.uk/prudential-regulation/publication/2023/may/model-risk-management-principles-for-banks-ss
- Liu et al. - Lost in the Middle: How Language Models Use Long Contexts. arxiv.org/abs/2307.03172
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information