Skip to content
RAG & LLM Engineering

$70bn Or $50bn? The OpenAI Revenue Confusion Was A Data Problem: Extracting Financial Metrics With Their Basis, Period And Source So Your Models Never Compare Gross With Net Again

On 8 October a gauge of chip stocks fell 3.4% after the Financial Times reported that OpenAI had told investors its annualised revenue was approaching $50bn - not the roughly $70bn in circulation. Most of the gap was definitional: Anthropic books cloud-partner sales gross, OpenAI books them net, and somewhere between investor decks, press reports and models, the basis got lost. Every system that ingested '$70bn' was wrong in the same direction at the same time. For anyone building LLM pipelines that extract figures from filings, transcripts, decks and news into ratings, signals or research, this is the failure mode to engineer out. Here is how we do it: an extraction schema where basis, period, annualisation and source are mandatory, an LLM extractor constrained to that schema, a validator that refuses unsourced or ambiguous figures, and a comparability layer that will not compare unlike numbers - in code.

AlchmAI Engineering15 min read

~$50bn

OpenAI's annualised revenue as told to investors at end-September, per the FT on 8 October - versus roughly $70bn in circulation

-3.4%

A gauge of chip stocks the same afternoon; CoreWeave fell about 8% and Oracle more than 5%

Gross vs net

The basis difference behind most of the gap: Anthropic records cloud-partner sales gross, OpenAI records its share

4 fields

Every extracted figure must carry: basis, period, annualisation method and source span

The OpenAI episode is a case study in how a number loses its meaning on the way into a model. The company told investors its annualised revenue was approaching $50bn. The roughly $70bn figure that had circulated - including in earlier reporting - appears to have come from investors putting OpenAI on the same basis as Anthropic, whose annualised revenue was reported at about $65bn and which records the full value of sales made through cloud providers as revenue. OpenAI records only its share. Neither convention is wrong; IFRS 15 and ASC 606 both ask whether a company is the principal or an agent in a sale, and judgement is involved. But a gross figure and a net figure are different quantities, and once '$70bn' and '$65bn' sat side by side in spreadsheets, signals and summaries without their definitions, the comparison was meaningless and the market priced it anyway.

LLM extraction pipelines make this failure easier, not harder. A model asked to 'extract revenue' from a press article will return a number; it will rarely volunteer that the number is annualised from one month's run rate, reported by an unnamed investor, on a gross basis. Unless the schema forces those fields to exist, they will not.

1. The Schema

pythonextract/schema.py
from enum import Enum
from typing import Optional
from pydantic import BaseModel, Field

class Basis(str, Enum):
    gross = "gross"            # includes amounts passed through to partners (principal accounting)
    net = "net"                # only the company's share (agent accounting)
    unknown = "unknown"

class Annualisation(str, Enum):
    reported_period = "reported_period"       # actual figure for a stated period
    run_rate_month = "run_rate_month"         # one month x 12
    run_rate_quarter = "run_rate_quarter"     # one quarter x 4
    company_forecast = "company_forecast"
    unknown = "unknown"

class SourceType(str, Enum):
    filing = "filing"; company_statement = "company_statement"; earnings_call = "earnings_call"
    press_named = "press_named_source"; press_unnamed = "press_unnamed_source"; analyst = "analyst_estimate"

class FinancialMetric(BaseModel):
    entity: str
    metric: str = Field(description="e.g. revenue, annualised_revenue, arr, operating_loss")
    value: float
    currency: str = Field(min_length=3, max_length=3)
    unit: str = Field(description="units, thousands, millions, billions")
    basis: Basis
    period_end: Optional[str] = Field(description="ISO date the figure relates to")
    annualisation: Annualisation
    source_type: SourceType
    source_doc: str = Field(description="document id")
    source_span: str = Field(description="verbatim sentence the figure came from")
    attributed_to: Optional[str] = Field(description="who stated it: the company, an investor, an analyst")

2. Constrained Extraction

pythonextract/extract.py
import json

SYSTEM = (
    "Extract financial figures from the document. For every figure you MUST fill basis, annualisation, "
    "source_type and source_span. Use 'unknown' when the document does not state the basis or annualisation; "
    "never infer it from another company's convention. source_span must be copied verbatim from the document. "
    "If a figure is attributed to people familiar with the matter or investors, use press_unnamed_source."
)

def extract(doc_id: str, text: str, llm) -> list:
    tool = {
        "name": "emit_metrics",
        "description": "Return every financial figure with its basis and source.",
        "input_schema": {"type": "object", "properties": {"metrics": {"type": "array", "items": FinancialMetric.model_json_schema()}}, "required": ["metrics"]},
        "strict": True,
    }
    res = llm.call(system=SYSTEM, messages=[{"role": "user", "content": "DOC " + doc_id + ":\n\n" + text}], tools=[tool])
    raw = res.tool_input["metrics"]
    return [FinancialMetric(**m) for m in raw]

3. Validate Before Anything Downstream Sees It

pythonextract/validate.py
def validate(m: FinancialMetric, doc_text: str) -> list:
    problems = []
    if m.source_span not in doc_text:
        problems.append("source_span is not verbatim in the document - possible fabrication")
    span = m.source_span.replace(",", "")
    if not any(tok in span for tok in {str(int(m.value)), str(m.value), f"{m.value:g}"}):
        problems.append("value does not appear in the quoted source sentence")
    if m.metric in ("annualised_revenue", "arr") and m.annualisation == "reported_period":
        problems.append("annualised metric labelled as a reported-period figure")
    if m.source_type in ("press_unnamed_source", "analyst_estimate"):
        problems.append("secondary source - not eligible for signals until confirmed by filing or company statement")
    return problems

# Figures with problems go to a review queue, not into the feature store.

4. Refuse To Compare Unlike Numbers

The comparability layer is where the OpenAI error would have stopped. Any comparison - a ratio, a ranking, a growth rate, a peer table - first checks that both figures share basis, annualisation method and a recent enough period. If they do not, it either applies an explicit, labelled normalisation the analyst has approved, or it refuses.

pythonextract/compare.py
from datetime import date

class NotComparable(Exception): ...

def days_between(a: str, b: str) -> int:
    return (date.fromisoformat(a) - date.fromisoformat(b)).days

def comparable(a: FinancialMetric, b: FinancialMetric, max_period_gap_days=45) -> None:
    if "unknown" in (a.basis, b.basis):
        raise NotComparable(f"basis unknown for {a.entity if a.basis == 'unknown' else b.entity}")
    if a.basis != b.basis:
        raise NotComparable(f"{a.entity} is {a.basis}, {b.entity} is {b.basis} - normalise explicitly first")
    if a.annualisation != b.annualisation:
        raise NotComparable(f"annualisation differs: {a.annualisation} vs {b.annualisation}")
    if a.period_end and b.period_end and abs(days_between(a.period_end, b.period_end)) > max_period_gap_days:
        raise NotComparable("periods too far apart")

def normalise_gross_to_net(m: FinancialMetric, partner_share: float, approved_by: str) -> FinancialMetric:
    """Explicit, labelled estimate - never silent."""
    est = m.model_copy(update={"value": m.value * (1 - partner_share), "basis": "net",
                               "source_type": "analyst_estimate",
                               "attributed_to": "normalised by " + approved_by + f" at partner share {partner_share:.0%}"})
    return est

# compare(openai_net_50bn, anthropic_gross_65bn) -> NotComparable: 'OpenAI is net, Anthropic is gross'

“The $20bn did not disappear on Thursday. It was never there - it was a definition missing from a field nobody had made mandatory.”


Where This Matters Most

  • AI signals and stock ratings that ingest private-company metrics from news - the fastest-growing and least-audited input in many quant and research pipelines.
  • Research assistants that answer 'how does X compare with Y' from documents; they will happily compare gross with net unless the retrieval layer carries basis metadata.
  • Credit and counterparty analysis, where run-rate revenue is routinely presented as if it were reported revenue.
  • Board and investor reporting on your own AI programme, where the same gross-versus-net choice applies to resold cloud and AI services.

The Bottom Line

OpenAI's roughly $50bn annualised revenue, set against a $70bn figure assembled on a different basis, knocked 3.4% off chip stocks because a definition was lost between source and model. LLM extraction makes that loss the default unless the schema prevents it. The engineering is straightforward: mandatory basis, period, annualisation, source type and verbatim source span on every figure; constrained extraction that allows 'unknown'; validation that rejects unverbatim spans and keeps unnamed-source figures out of signals; and a comparability layer that refuses to compare gross with net or run rate with reported unless an approved, labelled normalisation is applied. That is the RAG and data engineering we build for research and trading teams in London, and it is the cheapest insurance against the next $20bn headline.

References & Further Reading

AI Automation London codefinancial data extractiondata lineageAI Agency Developer LondonAgentic AI finance codestructured outputsAI signal integration
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information