Shipping LLM Trade Surveillance: Cutting False Positives 30-50% Without Losing The Audit Trail
Trade surveillance is the clearest ROI story in financial AI, and the numbers are now solid rather than promotional: AI-augmented programmes at broker-dealers and investment banks are cutting false-positive alert volumes by 30-50%, some reported deployments considerably more, and front-running detection has become the most mature LLM surveillance use case. But surveillance is also the least forgiving place to put a non-deterministic component, because every output is potentially evidence. This is the production architecture that gets both: a deterministic detection layer that never changes, an LLM triage layer that only ever ranks and explains, pinned model versions, replayable decisions, and an evaluation harness built from real production failures rather than synthetic examples.
AlchmAI Engineering14 min read
30-50%
False-positive alert volume reduction reported by AI-augmented trade surveillance programmes at broker-dealers and investment banks
73%
Of asset managers above $5bn AUM deploying LLM surveillance capability by 2026, with front-running the most mature use case
2 Aug 2026
EU AI Act high-risk obligations in force - risk management, documentation, human oversight and post-market monitoring
0 changes
To the deterministic detection layer. The LLM ranks and explains; it never decides what is an alert
Ask a compliance technology team what their biggest problem is and you will get the same answer you would have got a decade ago: the alerts are almost all noise. Static rules-based surveillance generates enormous false-positive volumes, analysts spend their days closing alerts that were never going to be anything, and the genuine signal - the cross-asset, cross-venue pattern that a threshold rule cannot express - is buried in the pile. It is a textbook case for machine intelligence, and in 2026 the results are finally concrete rather than aspirational: AI-augmented surveillance programmes at broker-dealers and investment banks are reducing false-positive volumes by 30% to 50%, with some deployments reporting substantially more, and front-running detection has emerged as the most mature LLM use case, with 73% of asset managers above $5 billion in assets deploying some form of the capability.
We build compliance and regulatory systems for investment banks - trade surveillance platforms comparable in structure to the market-standard systems - and we want to be precise about why these deployments succeed when so many financial LLM projects stall. It is not model quality. It is that the successful ones made a specific architectural choice early: the LLM was never allowed to decide whether something is an alert. It ranks, it explains, it drafts. The detection remains deterministic, unchanged, and fully replayable. Teams that instead let a model replace detection logic spend a year in model validation and never ship, because they have accidentally proposed to make their regulatory obligation non-deterministic.
The Two-Layer Architecture
The shape is straightforward once the boundary is clear. Layer one is your existing deterministic detection: the rules, thresholds and statistical tests that fire alerts. You do not touch it, you do not tune it down, and crucially you do not reduce its sensitivity because the AI is now handling triage. Layer two consumes the alerts layer one produces and does three things a rule cannot: assembles the surrounding context, assesses whether the pattern is consistent with abuse or has an innocent explanation, and produces a ranked queue with a written rationale for each item.
- 01Detection (deterministic, unchanged). Fires on the same conditions it always did. Full recall is preserved and provable. Every alert is recorded regardless of what happens downstream.
- 02Context assembly (deterministic). Gathers the order and fill sequence, the prevailing book, the relevant news window, the trader's recent behaviour baseline, related instruments, and any linked communications. This step is pure data retrieval with point-in-time correctness - it must reconstruct what was true at alert time, not what is true now.
- 03LLM triage (non-deterministic, bounded). Produces a structured verdict against a fixed schema: a likelihood band, the specific factors supporting and undermining an abuse hypothesis, and citations to the evidence it used. It has no ability to close, suppress or delete an alert.
- 04Human review (unchanged authority). The analyst sees the ranked queue and the rationale, and retains complete discretion. Their agreement or disagreement with the triage is captured as the training and evaluation signal for layer three.
from dataclasses import dataclass
from typing import Literal
@dataclass(frozen=True)
class TriageVerdict:
"""A fixed schema, not free text. The model fills in a form; it does not
write a memo. Every field below is queryable, chartable and evaluable."""
likelihood: Literal["low", "medium", "high"]
supporting_factors: tuple[str, ...]
mitigating_factors: tuple[str, ...]
evidence_refs: tuple[str, ...] # ids into the assembled evidence pack
rationale: str # <= 200 words, shown to the analyst
abstained: bool # explicit "I cannot assess this"
def triage(alert, evidence_pack, cfg) -> TriageVerdict:
"""Non-deterministic step, made reproducible by construction."""
result = model.complete(
model=cfg.model_snapshot, # dated snapshot, never a floating alias
temperature=0.0,
seed=cfg.seed,
system=cfg.prompt_template, # versioned in the repo, hashed in cfg
input=render(alert, evidence_pack),
response_schema=TRIAGE_SCHEMA, # structured output, validated on return
)
verdict = TriageVerdict(**result.parsed)
# Every evidence reference must resolve into the pack we supplied. A model
# citing something we did not give it is a hallucinated citation, and in
# surveillance that is a defect, not a quirk - it gets the alert routed to
# a human with the triage suppressed rather than shown.
unknown = [r for r in verdict.evidence_refs if r not in evidence_pack.ids]
if unknown:
audit.write(kind="triage.invalid_citation", alert_id=alert.id, refs=unknown)
return TriageVerdict(
likelihood="medium", supporting_factors=(), mitigating_factors=(),
evidence_refs=(), rationale="Automated triage unavailable for this alert.",
abstained=True,
)
return verdictReproducibility: Every Verdict Must Be Replayable
Surveillance output is potentially evidence. If a regulator asks in 2028 why a particular alert was ranked low in March 2026, 'the model said so' is not an answer. The system must be able to reconstruct the verdict: the exact evidence pack as assembled at the time, the exact model snapshot, the exact prompt version, the exact parameters. This is the same decision-fingerprint discipline we apply to agents on the trading desk, and in surveillance it is not optional.
- Pin the model to a dated snapshot. A provider silently upgrading a floating alias changes the behaviour of a regulated control with no change record. Treat a model version bump exactly as you would treat a change to a detection threshold: proposed, evaluated, approved, dated.
- Store the evidence pack itself, not a pointer to a query that regenerates it. Regenerating a pack in 2028 from a query written in 2026 gives you the data as it stands in 2028, which is a different pack and therefore a different question.
- Set temperature to zero and record the seed. This does not give you bit-exact determinism across provider infrastructure, and you should not claim that it does - but it removes deliberate variance and makes replay comparisons meaningful.
- Version the prompt template in the repository under the same review as the detection rules. A prompt that changes triage rankings is a control change.
- Write the verdict, the fingerprint and the analyst's eventual disposition into one immutable record. That record is simultaneously your audit trail and your evaluation dataset, which is a pleasant efficiency.
The Evaluation Harness Is The Product
The teams shipping AI successfully in 2026 are not the ones with the best models but the ones with the best evaluation infrastructure, and surveillance is the clearest illustration of that claim. Your golden dataset should be built from real production failures rather than synthetic examples - which in surveillance means every alert where the analyst disagreed with the triage, every confirmed abuse case you have ever had, and every alert that escalated after being ranked low.
def evaluate(candidate_cfg, golden_set) -> dict:
"""Run before ANY config change ships: model snapshot, prompt, evidence
assembly, ranking weights. A change that improves average ranking while
demoting a single confirmed case is a change that does not ship."""
results = [ (case, triage(case.alert, case.evidence_pack, candidate_cfg))
for case in golden_set ]
confirmed = [(c, v) for c, v in results if c.outcome == "confirmed_abuse"]
benign = [(c, v) for c, v in results if c.outcome == "closed_no_action"]
return {
# THE metric. Any confirmed case ranked low is a hard failure,
# regardless of what the aggregate numbers look like.
"confirmed_ranked_low": sum(1 for _, v in confirmed if v.likelihood == "low"),
"confirmed_recall_high": pct(confirmed, lambda v: v.likelihood == "high"),
"benign_demoted": pct(benign, lambda v: v.likelihood == "low"),
"abstention_rate": pct(results, lambda v: v.abstained),
"invalid_citations": count_invalid(results),
"analyst_agreement": agreement_rate(results),
"p95_latency_ms": p95([r.latency for r in results]),
"cost_per_alert_gbp": total_cost(results) / len(results),
}
GATES = {
"confirmed_ranked_low": 0, # zero tolerance, non-negotiable
"invalid_citations": 0,
"benign_demoted": ">= baseline", # the whole point of the system
"analyst_agreement": ">= baseline",
}“A surveillance model that improves your average ranking while demoting one confirmed case has not improved. It has traded a measurable efficiency gain for an unmeasurable regulatory risk, and that is a trade no compliance function would authorise if it were stated plainly.”
Where The False-Positive Reduction Actually Comes From
It is worth being specific about the mechanism, because it is not magic and understanding it tells you where to invest. Rules fire on patterns. An LLM with a properly assembled evidence pack can recognise the innocent explanation that the rule had no way to see, and the reductions cluster in a small number of recognisable shapes:
- Publicly explained moves. A price and volume spike that triggered a manipulation rule, occurring two minutes after a scheduled announcement that the rule had no access to. The commonest single category by some distance.
- Legitimate mechanical flow. Index rebalancing, month-end, option expiry, a fund flow, a known hedging programme - all of which look like informed trading to a threshold and are obvious in context.
- Behavioural baselines. A trader whose order-to-trade ratio triggers a layering rule but whose pattern is unchanged from their consistent two-year baseline in that instrument. Rules are near-blind to per-trader normality; a model with the baseline in context is not.
- Related-instrument explanations. A cash equity move explained by the futures basis or an ETF creation, which a single-instrument rule cannot express and a cross-asset context pack makes obvious.
- Duplicate and cascading alerts. One event firing six rules across three systems. Clustering these into a single reviewable case is, on its own, a large share of the headline reduction - and notably it requires no model judgement at all, which is why we always build it first.
Deploying It Under The 2026 Regulatory Frame
Two regulatory realities shape deployment. For firms with EU exposure, the AI Act's high-risk obligations have been in force since 2 August 2026, bringing risk management, technical documentation, data governance, human oversight and post-market monitoring requirements. In the UK, the FCA has been actively interested in this specific problem - its Market Abuse Surveillance TechSprint explored how AI and machine learning could detect evolving forms of market abuse, and solutions that combined efficient screening of trade and order data across asset classes with LLM assessment of potential violations were exactly the pattern described here.
The encouraging part is that a well-built triage architecture satisfies most of this by construction. Human oversight is not a bolt-on because the analyst always had and retains decision authority. Technical documentation is largely generated from the evaluation harness. Post-market monitoring is the metrics you are already computing. Data governance is the evidence-pack lineage. What teams typically have to add is explicit: a written risk assessment, a record of the human oversight design and why it is meaningful rather than nominal, and a monitoring plan with named owners and thresholds. That is a few weeks of documentation on top of a system built the right way, versus a rebuild for a system that was not.
A Realistic Delivery Sequence
- 01Weeks 1-3: instrument the existing alert flow. Measure true volumes, duplicate rates, analyst time per alert and current disposition mix. You cannot claim a 40% reduction without a defensible baseline, and most firms do not have one.
- 02Weeks 3-6: deterministic deduplication and event clustering. Immediate measurable reduction, no model risk, no validation burden.
- 03Weeks 5-10: evidence pack assembly with point-in-time correctness. This is the largest engineering component and the one that determines the ceiling on everything after it. Underinvest here and no model will rescue you.
- 04Weeks 8-14: triage in shadow mode. The model produces verdicts, nothing is shown to analysts, and every verdict is compared to the analyst's independent disposition. This is where your golden set is really built.
- 05Weeks 14-18: ranked queue in production, with the analyst UI showing rationale and citations, and a one-click disagreement capture that is genuinely one click - if it takes four, you will not get the data.
- 06Continuous: golden set grows from every disagreement and every escalation; evaluation gates block config changes; abstention rate, agreement rate and cost per alert on a dashboard with an owner.
The Bottom Line
Trade surveillance is where financial AI is delivering its most defensible returns, with 30-50% false-positive reductions now well documented and front-running triage mature enough that most large asset managers have deployed something. The architecture that gets there is disciplined rather than clever: leave deterministic detection completely alone so recall stays provable, invest heavily in point-in-time evidence assembly because it sets the ceiling, constrain the model to a fixed verdict schema with validated citations and a cheap abstention path, pin and version everything so any verdict can be replayed years later, build your golden set from real analyst disagreements rather than synthetic cases, and gate every change on zero confirmed cases ranked low. Do the deterministic deduplication first, because a large share of the benefit is sitting there requiring no model at all. Get that sequence right and you end up with a system that is simultaneously cheaper to run, better at finding real abuse, and easier to explain to a regulator than the rules-only system it replaced - which is a rare combination, and the reason this remains the strongest use case in the category. That is the compliance and surveillance engineering we do for financial institutions in London.
References & Further Reading
- FCA - Market Abuse Surveillance TechSprint. fca.org.uk/publications/techsprints/market-abuse-surveillance
- SteelEye - Harnessing AI for market abuse detection: takeaways from the FCA's TechSprint. steel-eye.com/news/harnessing-ai-for-market-abuse-detection-takeaways-from-fcas-techsprint
- LSEG - Market surveillance in 2026: a Q&A with LSEG experts on regulation, data and technology. lseg.com/en/insights/fx/market-surveillance-in-2026-a-q-and-a-with-lseg-experts-on-regulation-data-and-technology
- Emerj - Market surveillance and AI: two use cases. emerj.com/market-surveillance-and-ai-two-use-cases
- Trapets - AI and machine learning in trade surveillance. trapets.com/resources/blog/ai-machine-learning-trade-surveillance
- European Commission - The EU Artificial Intelligence Act (high-risk obligations, in force 2 August 2026). artificialintelligenceact.eu
- Confident AI - LLM agent evaluation metrics in 2026: tool calling, task completion, reasoning and trace-based evals. confident-ai.com/blog/llm-agent-evaluation-complete-guide
- Towards Data Science - Building an evaluation harness for production AI agents: a 12-metric framework from 100+ deployments. towardsdatascience.com/building-an-evaluation-harness-for-production-ai-agents-a-12-metric-framework-from-100-deployments
- FCA - Market Watch 79: market abuse surveillance failures and peer review of front-running model testing across nine investment banks. fca.org.uk/publications/newsletters/market-watch-79
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information