Context Engineering For Financial AI: Token Budgets, Lost-In-The-Middle, And Why Your Market-Data RAG Is Quietly Time-Travelling
Context engineering is the discipline that replaced prompt engineering in 2026, and every senior engineer has now read the general version: write, select, compress, isolate. What almost nobody writes about is what changes when the context is financial. Two things, and both are expensive. First, unit economics - an agent burning 50K tokens a turn when 5K would do is 10x the cost per decision, and in finance you are running it per instrument, per document, per day. Second, and far worse, point-in-time correctness: a retrieval layer that returns today's restated figures for a question about last March has silently introduced look-ahead bias into a system somebody will trade on. This is an educational deep-dive for engineers and CTOs on getting both right, with code.
AlchmAI Engineering14 min read
10x
Cost multiple of a poorly-contexted agent using 50K tokens per turn where 5K would do - the unit economics that decide whether a financial AI product is viable
4 strategies
The canonical context toolkit formalised by LangChain: write, select, compress, isolate
Middle
Where models reliably lose information - content at the start and end of a long context is recalled far more dependably than content buried between them
Bitemporal
The retrieval property financial RAG needs and general-purpose RAG ignores: what was known, as at when
Context engineering earned its place as the defining discipline of 2026 for a sound reason: agents fail differently from chatbots. A chatbot that answers badly had a bad prompt. An agent that fails in production almost always had a state-management failure - it was working from the wrong information, too much information, or information arranged so that the part that mattered was in the place models are worst at reading. Andrej Karpathy's framing of context engineering as the deliberate design of what the model sees on every inference call stuck because it named something every team building agents had independently discovered by suffering.
The general guidance is now well covered: LangChain's write, select, compress and isolate is a genuinely useful taxonomy, the lost-in-the-middle finding is well-replicated, and everyone has internalised that more tokens means more cost, more latency and worse recall. We are not going to re-litigate that here. What we want to write about is the part that is missing from every general treatment, and that we deal with on every financial engagement: financial context has a time dimension that ordinary context does not, and if you ignore it you will build a system that is confidently, invisibly wrong about the past.
Part One: Token Budgets Are A Product Decision
Start with economics, because it determines what is even possible architecturally. The often-quoted example - an agent using 50K tokens per turn where 5K would suffice being 10x more expensive - understates the problem in finance, because financial workloads fan out. You are not running one agent conversation. You are running it across every instrument in a universe, every filing in a quarter, every alert in a surveillance queue. A 10x context inefficiency on a single chat is a rounding error; the same inefficiency across four thousand instruments daily is the difference between a product with a margin and a science project with a bill.
The discipline that fixes this is treating the context window as a budget with named line items, allocated explicitly rather than filled opportunistically. In practice that means a budgeter component that sits between retrieval and the model call, with a fixed allocation per section and a defined eviction order:
from dataclasses import dataclass, field
@dataclass
class Section:
name: str
priority: int # 1 = never evict, higher = evicted sooner
min_tokens: int # floor below which the section is dropped entirely
max_tokens: int
content: list[str] = field(default_factory=list)
class ContextBudget:
"""Explicit allocation beats opportunistic filling. Every section declares
what it needs and what it can survive on; the budgeter resolves the rest."""
def __init__(self, total_tokens: int, count_tokens):
self.total = total_tokens
self.count = count_tokens
self.sections: list[Section] = []
def add(self, section: Section) -> None:
self.sections.append(section)
def resolve(self) -> list[Section]:
# Reserve floors first - a section that cannot meet its floor is
# dropped whole rather than truncated into uselessness. Half a
# balance sheet is worse than no balance sheet: the model will
# happily reason over the half it can see.
ordered = sorted(self.sections, key=lambda s: s.priority)
remaining, kept = self.total, []
for s in ordered:
if remaining < s.min_tokens:
continue
take = min(s.max_tokens, remaining)
kept.append(self._fit(s, take))
remaining -= self.count(kept[-1].content)
return kept
def _fit(self, s: Section, budget: int) -> Section:
out, used = [], 0
for item in s.content: # pre-ranked by the retriever
t = self.count([item])
if used + t > budget:
break
out.append(item)
used += t
return Section(s.name, s.priority, s.min_tokens, s.max_tokens, out)
# A surveillance-triage agent's allocation, as an example. Note that the
# instrument reference data is priority 1: the model may lose the news
# before it loses the definition of the instrument it is reasoning about.
budget = ContextBudget(total_tokens=24_000, count_tokens=tokenizer_count)
budget.add(Section("task_and_policy", priority=1, min_tokens=800, max_tokens=1_500))
budget.add(Section("instrument_ref", priority=1, min_tokens=300, max_tokens=800))
budget.add(Section("alert_payload", priority=2, min_tokens=1_000, max_tokens=6_000))
budget.add(Section("order_and_fill_log", priority=3, min_tokens=1_000, max_tokens=8_000))
budget.add(Section("comms_excerpts", priority=4, min_tokens=500, max_tokens=5_000))
budget.add(Section("market_context", priority=5, min_tokens=400, max_tokens=2_000))Compaction That Preserves Numbers
Compression is the most valuable of the four strategies and the most dangerous in finance. The standard technique - summarise older turns or retrieved documents with a smaller model - works beautifully for conversational context and destroys financial context, because summarisation is lossy in exactly the dimension that matters. A summariser asked to condense a filings extract will faithfully preserve the narrative and round, drop or approximate the figures, and it will do so without flagging that it did.
Our rule is a split representation: narrative gets summarised, quantities get extracted into a structured block that is never summarised, only filtered. The structured block is cheap - numbers are token-dense in the good direction - and it is the part the model will actually be asked to reason over.
interface Fact {
metric: string; // "revenue", "net_interest_margin", ...
value: number;
unit: string; // "GBP_millions", "percent", "bps"
period: string; // "FY2025", "Q2_2026"
asOf: string; // when this value was PUBLISHED - see part two
sourceSpan: string; // document id + character range, for verification
}
interface Compacted {
narrative: string; // lossy, summarised, cheap
facts: Fact[]; // lossless, filtered but never rewritten
}
export async function compact(docs: Document[], focus: string): Promise<Compacted> {
// Facts are extracted once at ingestion and cached - not re-derived on
// every query. Extraction is the expensive step; filtering is free.
const facts = docs
.flatMap((d) => d.extractedFacts)
.filter((f) => isRelevant(f.metric, focus));
// Only the prose goes near a summariser, and it is explicitly told that
// the numbers are handled elsewhere so it does not try to be helpful.
const narrative = await summarise(docs.map((d) => d.prose).join("\n\n"), {
instruction:
"Summarise the qualitative narrative only. Do not restate figures; " +
"numeric values are supplied separately in structured form.",
maxTokens: 900,
});
return { narrative, facts };
}Placement: Working With Lost-In-The-Middle Instead Of Against It
The lost-in-the-middle effect - models recall content at the beginning and end of a long context far more reliably than content buried between them - is usually presented as a warning. It is more useful as a layout constraint. If you know recall is U-shaped, you stop treating context assembly as concatenation and start treating it as typesetting.
- 01Open with the task, the output schema and the hard constraints. These are what the model must not lose, and the start position is the strongest slot available.
- 02Place the bulk retrieved material in the middle, ordered so that the least critical sits dead centre. If your retriever ranks by relevance, do not lay the results out in rank order - interleave so that ranks one and two land at the edges of the block rather than at its start.
- 03Close with the specific question, the current state, and a restatement of the output contract. The end position is the second-strongest slot, and it is where you put the thing the model must do right now.
- 04Keep tool definitions stable and early. Their position and content are part of your cacheable prefix, and moving them around defeats prompt caching - which for a fan-out financial workload is a direct and substantial cost increase.
Part Two: Point-In-Time Correctness, The Thing General RAG Ignores
Now the part that makes financial context different in kind rather than degree. In a general knowledge base, a document has one meaningful timestamp and the current version is the right version. In finance, essentially every fact has two independent times: the period it describes, and the moment it became known. Earnings are restated. Economic series are revised - sometimes substantially, months later. Index constituents change, and a universe reconstructed from today's membership silently contains only companies that survived. Ratings are upgraded. A company's reported figure for Q1 as published in April is not the figure for Q1 as it stands today, and both are correct answers to different questions.
Ordinary RAG collapses this. You embed documents, you retrieve by similarity, you return the chunk. The chunk carries whatever version happened to be indexed - usually the latest, because that is what the pipeline ingested most recently. Ask 'what did the market know about this company in March?' and you get an answer built from information that did not exist in March. The model has no way to detect this. Your citations will look impeccable.
“If your retrieval layer cannot answer 'as at what date?', it is not a financial retrieval layer. It is a search box with good manners.”
The fix is bitemporal retrieval: every indexed fact carries both a valid-time (the period it describes) and a knowledge-time (when it became available), and every query carries an as-of date that filters on knowledge-time. This is old news in market-data engineering - it is why point-in-time databases exist - and it is almost entirely absent from LLM retrieval stacks, which were designed for documentation and inherited documentation's assumptions.
from dataclasses import dataclass
from datetime import date
@dataclass(frozen=True)
class Chunk:
doc_id: str
text: str
period_start: date # valid time: what the content is ABOUT
period_end: date
known_from: date # knowledge time: when this became available
known_until: date | None # None = still current; set when superseded
revision: int
embedding: list[float]
def retrieve(query_vec, as_of: date, index, k: int = 12) -> list[Chunk]:
"""as_of is MANDATORY. The default is not 'today' - there is no default.
A retrieval call without an as-of date is a bug, and making it a required
positional argument is the cheapest possible way to enforce that."""
candidates = index.search(query_vec, k=k * 6) # over-fetch, then filter
visible = [
c for c in candidates
if c.known_from <= as_of and (c.known_until is None or c.known_until > as_of)
]
# Among revisions of the same document, keep the one in force AS AT as_of -
# which is frequently not the highest revision number.
latest_as_of: dict[str, Chunk] = {}
for c in visible:
prev = latest_as_of.get(c.doc_id)
if prev is None or c.revision > prev.revision:
latest_as_of[c.doc_id] = c
return sorted(latest_as_of.values(), key=lambda c: -similarity(query_vec, c.embedding))[:k]- Make as_of a required argument with no default. Every default you provide will eventually be wrong, and the wrongness will be invisible. Required arguments turn a silent data-correctness bug into a compile-or-call-time error.
- Never delete or overwrite a superseded chunk. Close it by setting known_until and insert the revision as a new row. Deleting history makes yesterday's answers unreproducible, and reproducing yesterday's answer is a routine request in regulated work.
- Carry the as-of date through into the model's context explicitly, and instruct the model to state it. An answer that says 'as at 14 March 2026, the reported figure was...' is auditable. An answer without a date is a claim about nothing in particular.
- Apply the same discipline to the universe, not just the documents. Reconstructing a peer group from today's index membership is survivorship bias, and it is the same class of error wearing different clothes.
- Log the as-of date on every retrieval. When someone disputes an answer six months later - and in finance they will - the question is always 'what did it see?', and this is the only cheap way to answer it.
Isolation: Separate Contexts Beat One Big One
The fourth strategy, isolate, deserves a specific financial reading too. The general advice is to give different agents separate contexts to stop them contaminating each other. In financial systems the stronger reason is entitlement. Research, positions, client data and order flow have genuinely different permission models, and a single agent context holding all of them is an information-barrier problem wearing an architecture diagram. The moment one context can see both a client's order and the research desk's unpublished view, you have built something a compliance officer will have strong feelings about.
Architecturally this pushes you towards several narrow agents with disjoint retrieval scopes and a coordinating layer that passes structured results rather than raw context between them. That is more engineering than one large agent. It is also the only version that survives a controls review, and it has a pleasant side effect: narrow contexts are small contexts, so the isolation that compliance requires is the same isolation that fixes your token economics.
A Checklist For Senior Engineers And CTOs
- 01Can you state your token budget per decision, and does anyone own it? If cost per agent turn is not on a dashboard, it is growing.
- 02Is retrieval point-in-time, with a mandatory as-of parameter, and can you prove it with a test that asserts a March query never returns a June revision?
- 03Are numbers extracted structurally at ingestion rather than summarised at query time?
- 04Is your cacheable prefix genuinely stable, and have you checked your cache hit rate rather than assumed it?
- 05Does context assembly account for positional recall, or is it concatenation in retriever rank order?
- 06Do agent contexts respect the firm's information barriers, or does one context span several entitlement domains?
- 07Is every retrieval logged with its as-of date, its result ids and its token count, so that any answer can be reconstructed later?
The Bottom Line
Context engineering deserves its promotion to first-class discipline, but the general playbook stops one step short of where financial systems live. The economics half is a known problem with known answers: budget explicitly, compact without destroying quantities, lay out for positional recall, keep your prefix cacheable. The correctness half is the one that separates teams who have built financial AI from teams who have built AI and pointed it at financial data. Facts in this domain have two times, not one, and any retrieval layer that cannot answer 'as at when?' will eventually produce a beautifully-cited answer containing information that did not exist when the question is about. Make as_of mandatory, never overwrite history, carry the date into the answer, and log what was seen. Do that, and the rest of context engineering behaves the way the general literature promises it will. That is the difference between an AI feature and a financial system, and building the second is what we do as an AI and fintech developer in London.
References & Further Reading
- LangChain - Context engineering for agents (write, select, compress, isolate). blog.langchain.com/context-engineering-for-agents
- Liu et al. - Lost in the Middle: How Language Models Use Long Contexts (TACL). arxiv.org/abs/2307.03172
- Anthropic - Effective context engineering for AI agents. anthropic.com/engineering/effective-context-engineering-for-ai-agents
- LogRocket - The LLM context problem in 2026: strategies for memory, relevance and scale. blog.logrocket.com/llm-context-problem-strategies-2026
- Sourcegraph - Context engineering: a practical guide for AI agents (2026). sourcegraph.com/blog/context-engineering
- Mem0 - Context engineering AI: how to build smarter LLM agents in 2026. mem0.ai/blog/context-engineering-ai-agents-guide
- Anthropic - Prompt caching documentation (stable-prefix design). docs.claude.com/en/docs/build-with-claude/prompt-caching
- Richard Snodgrass - Developing Time-Oriented Database Applications in SQL (foundational reference on bitemporal data). www2.cs.arizona.edu/~rts/tdbbook.pdf
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information