Skip to content
RAG & LLM Engineering

Cache Reads Now Cost 5% Of Input: Re-Architecting Financial RAG Prompts Around The Prompt Cache After The September Price Cuts

The headline prices from 22 September hid the bigger change for retrieval-heavy workloads. Claude Opus 5.5 cut cache reads by 60% to $0.20 per million tokens - 5% of its $4 input price - while five-minute cache writes cost $5 and one-hour writes $8. OpenAI credited GPT-6 Sol and Luna's price cuts of 50% or more partly to caching improvements. For financial RAG, where the same filings, policies and tool definitions are sent again and again, the prompt cache is now the largest cost lever available - but only if prompts are built so the expensive part is identical, byte for byte, on every call. This is the engineering: the stable-prefix layout, where to place cache breakpoints, the break-even maths for five-minute versus one-hour writes, the silent cache-busters (including changing effort mid-session), and how to measure hit rate in production, with code.

AlchmAI Engineering14 min read

$0.20

Opus 5.5 cache read per million tokens - down 60%, and 5% of the $4 input price (22 September)

$5 / $8

Opus 5.5 cache writes per million tokens for five-minute and one-hour lifetimes

1 reuse

Break-even for a five-minute cache write; a one-hour write pays back on the second reuse

50%+

GPT-6 Sol and Luna price cuts, which OpenAI attributed to inference and caching improvements

Most commentary on the 22 September launches compared input and output prices. For retrieval-augmented systems in finance, the more important line was further down Anthropic's price table: Opus 5.5 cache reads fell from $0.50 to $0.20 per million tokens, while input fell to $4. A cached token now costs one-twentieth of a fresh one. OpenAI, launching GPT-6 Sol and Luna the same day at half the previous prices or less, credited improvements to inference and caching. The economic message from both labs is the same: reuse is cheap, novelty is expensive.

Financial RAG is unusually well suited to this. A research assistant answering questions about one company sends the same filing sections dozens of times. A compliance assistant sends the same policy manual on every call. An agent sends the same tool definitions on every step. If those parts are identical byte for byte and placed first, they cost 5% after the first call. If a timestamp, a user name or a reordered retrieval result sits in front of them, they cost full price every time.

The Stable-Prefix Layout

  1. 01Tools and system instructions: identical for every call in a deployment. Version them; never interpolate dates or user details here.
  2. 02Corpus block: the documents this session is about - a company's filings, a policy manual - in a deterministic order (sorted by document ID and section), with stable formatting.
  3. 03Session context: the user's role, the portfolio or case under discussion.
  4. 04Conversation and per-request retrieval: the question, fresh search hits, today's market data.
pythonrag/cached_prompt.py
import anthropic

client = anthropic.Anthropic()

def build_request(tools, system_text, corpus_docs, session_ctx, history, question, hits):
    # 1) Deterministic corpus ordering: same docs => same bytes => cache hit.
    corpus = "".join(
        "<doc id='" + d["id"] + "' section='" + d["section"] + "'>" + d["text"] + "</doc>"
        for d in sorted(corpus_docs, key=lambda d: (d["id"], d["section"]))
    )
    return dict(
        model="claude-opus-5-5",
        max_tokens=2000,
        tools=tools,                                   # stable, versioned
        system=[
            {"type": "text", "text": system_text},     # stable
            {"type": "text", "text": corpus,
             # Long-lived corpus: one-hour cache, breakpoint AFTER the corpus.
             "cache_control": {"type": "ephemeral", "ttl": "1h"}},
            {"type": "text", "text": session_ctx,
             "cache_control": {"type": "ephemeral"}},  # 5-minute, per session
        ],
        messages=history + [{
            "role": "user",
            # Volatile content LAST: fresh hits, live prices, the question.
            "content": "Fresh results:" + format_hits(hits) + " Question: " + question,
        }],
    )

resp = client.messages.create(**build_request(...))
u = resp.usage
metrics.record(read=u.cache_read_input_tokens,
               written=u.cache_creation_input_tokens,
               fresh=u.input_tokens)

Five Minutes Or One Hour? The Break-Even Maths

At Opus 5.5 prices, a fresh input token costs $4 per million, a five-minute write $5 (1.25x), a one-hour write $8 (2x), and a read $0.20 (0.05x). For a prefix reused n times after it is written, the cost relative to sending it fresh every time is:

pythonrag/breakeven.py
def relative_cost(n_reuses: int, write_multiplier: float, read_multiplier=0.05) -> float:
    uncached = 1 + n_reuses                      # send it fresh every time
    cached = write_multiplier + read_multiplier * n_reuses
    return cached / uncached

for n in (0, 1, 2, 5, 20):
    print(n, round(relative_cost(n, 1.25), 3), round(relative_cost(n, 2.0), 3))
# n=0  1.25   2.0    <- a write nobody reuses is pure overhead
# n=1  0.65   1.025  <- 5-min pays back on the first reuse
# n=2  0.45   0.70   <- 1-hour pays back on the second
# n=5  0.25   0.375
# n=20 0.107  0.143  <- heavy reuse: ~90% saving either way
  • Use the five-minute cache for interactive sessions where follow-up questions arrive within minutes.
  • Use the one-hour cache for shared corpora - a policy manual or a watch-list company's filings - hit by many users or by batch jobs spread across an hour.
  • Do not cache a prefix you expect to use once. A write nobody reads costs more than no cache at all.

The Silent Cache-Busters

  • Timestamps and request IDs in the system prompt. The single most common cause of a 0% hit rate we find in audits.
  • Non-deterministic retrieval order. If the same five chunks come back in a different order, the prefix differs. Sort anything that goes into the cached region.
  • Reformatting on the fly: whitespace normalisation, different JSON key orders, locale-dependent number formatting.
  • Tool definitions regenerated per request - for example, tools filtered by user permission placed before the corpus. Put per-user tool subsets after the cached block, or cache one prefix per permission tier.
  • Changing model settings mid-session. With Opus 5.5, changing the effort level mid-session invalidates the prompt cache, so pick an effort per workload and keep it fixed.
pythonrag/cache_guard.py
import hashlib

class PrefixGuard:
    """Alert when the cached region's bytes drift for the same logical corpus."""
    def __init__(self, alert):
        self.seen, self.alert = {}, alert

    def check(self, corpus_key: str, cached_text: str) -> None:
        digest = hashlib.sha256(cached_text.encode()).hexdigest()
        prev = self.seen.setdefault(corpus_key, digest)
        if prev != digest:
            self.alert("cache prefix drift for " + corpus_key +
                       " - same corpus, different bytes; check ordering/formatting")
            self.seen[corpus_key] = digest

def hit_rate(read: int, written: int, fresh: int) -> float:
    total = read + written + fresh
    return 0.0 if total == 0 else read / total

What It Looks Like In Practice

In the equity-research assistants we build, the filings and instructions that repeat across a session are usually the large majority of input tokens. Get the layout right and those tokens are billed at 5% after the first question, while the retrieval hits and question that change each turn pay full price. The exact saving depends on reuse patterns, so measure it: track hit rate per workload, cost per successful answer, and latency, since cached prefixes are also faster to process. A falling hit rate after a deploy almost always means someone put something volatile in front of the breakpoint.

“Caching is not a setting you turn on. It is a prompt architecture, and the architecture is mostly about deciding what is allowed to change.”


The Bottom Line

The September price cuts made the prompt cache the biggest cost lever in retrieval-heavy financial AI: Opus 5.5 cache reads at $0.20 per million tokens are 5% of fresh input, and OpenAI credited caching for GPT-6 Sol and Luna's cuts. Capturing it takes architecture rather than configuration - tools and instructions first, a deterministically ordered corpus behind a breakpoint, session context next and volatile retrieval last; five-minute writes for interactive sessions, one-hour writes for shared corpora; no timestamps, reordering or mid-session setting changes in the cached region; and hit rate monitored per workload. That is the RAG and LLM engineering we do as an AI agency for fintech teams in London, and after this week it is the fastest cost reduction most of them can ship.

References & Further Reading

prompt cachingfinancial RAGAI Automation London codeAgentic AI finance codeAI Agency Developer LondonLLM cost optimisationClaude Opus 5.5
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information