Skip to content
RAG & LLM Engineering

Open Weights Reached The Frontier In One Week: Beam (501B) And Mistral Large 4 (1.05T) Make A Self-Hosted Model Inside A Bank's Perimeter Realistic. Here Is The Sizing, Serving And Evaluation Plan

Two releases this week change the arithmetic for firms that cannot send certain data to a model provider. On 5 October Reflection AI introduced Beam, a 501-billion-parameter sparse mixture-of-experts model with 23 billion active parameters, a one-million-token context, Apache 2.0 weights due later in October, and benchmarks it says approach models four times its size at a third to a quarter of the hardware. On 6 October Mistral introduced Large 4: 1.05 trillion parameters with 49 billion active, a one-million-token context, weights due on 27 October, trained end to end in European data centres, priced at $0.68 per million input tokens, and - per Mistral - the strongest open-weight model from the US or Europe on finance, legal and cybersecurity workloads. Sparse experts mean these run on far less hardware than their parameter counts suggest. This is the engineering plan for a regulated firm: which workloads justify self-hosting, how to size GPUs for a sparse MoE, how to serve it, how to evaluate it against the API models you use today, and what to keep off-premises anyway. With code.

AlchmAI Engineering16 min read

501B / 23B

Beam's total and active parameters: a sparse MoE with a 1M-token context, Apache 2.0 weights due later in October (announced 5 October)

1.05T / 49B

Mistral Large 4's total and active parameters, 1M context, weights due 27 October, trained on 3,800 GPUs in Mistral's European data centres

$0.68 / $2.09

Mistral Large 4 API pricing per million input and output tokens, undercutting the $2 / $10 tier set by GPT-6.1 Sol and Gemini 4 Argon

3-4x

Less inference compute Beam claims versus GLM-5.2 on comparable reasoning benchmarks, with 80.1 on Terminal-Bench 2.1 and 77.2 on SWE-Bench Pro v2-Hard

For two years the answer to 'can we run a frontier-class model ourselves' was no, or not without a data-centre budget. This week it became 'probably, for the right workloads'. Reflection AI's Beam, introduced on 5 October, is a 501-billion-parameter sparse mixture-of-experts with 23 billion active parameters per token, 52 layers of interleaved local and global attention, a context extended to one million tokens, and training on 23.8 trillion tokens followed by more than 100 million reinforcement-learning rollouts. It reports 80.1 on Terminal-Bench 2.1, 77.2 on SWE-Bench Pro v2-Hard and 90.5 on GPQA Diamond, and claims three to four times less inference compute than GLM-5.2 at comparable quality. Weights arrive under Apache 2.0 later in October with the stack to run, evaluate and fine-tune it.

Mistral's Large 4, introduced the next day in Abu Dhabi, is bigger and more pointed. A 1.05-trillion-parameter mixture-of-experts with 49 billion active parameters and a vision encoder, a one-million-token context, training end to end on 3,800 GPUs in Mistral's own European data centres, multilingual data across more than 160 languages including every official EU language, and weights scheduled for 27 October. Mistral calls it the best open-weight model from the United States or Europe and says it leads open models on enterprise workloads including cybersecurity, finance and legal. API pricing of $0.68 per million input tokens and $2.09 output undercuts the $2 and $10 floor that GPT-6.1 Sol and Gemini 4 Argon set a week ago.

1. Which Workloads Justify Self-Hosting

  1. 01Data your provider contracts exclude: client-identifiable records, unredacted transaction data, material non-public information, anything your DPIA says may not leave the perimeter.
  2. 02Residency-bound workloads: UK or EU processing requirements for which a provider's region guarantees are insufficient or unverifiable.
  3. 03Continuity: a fallback for critical pipelines when a provider has an incident or changes terms - evaluated in advance so switching is routine.
  4. 04High-volume, predictable load where amortised GPU cost beats per-token pricing. Below a few hundred million tokens a month this is rarely true; above a few billion it often is.
  5. 05Not: customer-facing chat for a small team, experimental workloads, or anything where the frontier API's capability lead measurably matters. Self-hosting is a control decision first and a cost decision second.

2. Sizing A Sparse MoE

pythoninfra/size_moe.py
from dataclasses import dataclass

@dataclass
class MoeSpec:
    total_params_b: float; active_params_b: float; layers: int
    context_tokens: int; kv_bytes_per_token_per_layer: int = 2 * 2 * 1024 * 8  # k+v, bf16, 1024 dims, 8 kv heads (adjust from the model card)

def memory_plan(spec: MoeSpec, bytes_per_param: float, concurrent_seqs: int, avg_ctx: int, gpu_mem_gb: int):
    weights_gb = spec.total_params_b * bytes_per_param           # ALL experts must be resident
    kv_gb = concurrent_seqs * avg_ctx * spec.layers * spec.kv_bytes_per_token_per_layer / 1e9
    overhead_gb = 0.15 * weights_gb                              # activations, CUDA graphs, fragmentation
    total = weights_gb + kv_gb + overhead_gb
    gpus = -(-total // (gpu_mem_gb * 0.9))                       # ceil, 90% usable
    return {"weights_gb": round(weights_gb), "kv_cache_gb": round(kv_gb), "total_gb": round(total), "gpus_needed": int(gpus)}

# Beam-class: 501B total. bf16 = 2 bytes/param -> ~1,000 GB of weights alone; FP8 -> ~500 GB.
print("beam  fp8  :", memory_plan(MoeSpec(501, 23, 52, 1_000_000), 1.0, concurrent_seqs=32, avg_ctx=16_000, gpu_mem_gb=141))
# ML4-class: 1.05T total. FP8 -> ~1,050 GB; INT4 -> ~525 GB.
print("ml4   int4 :", memory_plan(MoeSpec(1050, 49, 64, 1_000_000), 0.5, concurrent_seqs=32, avg_ctx=16_000, gpu_mem_gb=141))
# Compute per token is that of a 23B / 49B dense model: throughput is bounded by memory bandwidth
# and expert-routing balance, not by the headline parameter count.

Two practical consequences. First, quantisation decides the node count: FP8 for Beam fits an eight-GPU node of 141-gigabyte accelerators with room for a modest KV cache; Large 4 at FP8 does not, and needs INT4 weights or two nodes with expert parallelism. Second, long contexts are expensive in KV cache even when weights fit; a million-token context is a capability, not a default, and your serving config should cap it per workload.

3. Serving

bashinfra/serve.sh
# vLLM with tensor parallelism across the node and expert parallelism for the MoE layers.
# Flags follow vLLM's documented options; confirm names against the version you deploy.
vllm serve reflection-ai/beam   --tensor-parallel-size 8   --enable-expert-parallel   --quantization fp8   --max-model-len 131072   --max-num-seqs 64   --gpu-memory-utilization 0.90   --enable-prefix-caching   --served-model-name internal-reasoner   --api-key "$INTERNAL_LLM_KEY"   --host 0.0.0.0 --port 8000

# Behind it: mTLS from the gateway only, no public ingress, egress denied (weights do not phone home,
# but your policy should not depend on that), OpenTelemetry GenAI spans exported to the same collector
# as your API-model calls so cost, latency and quality compare like for like.

4. Evaluate Before You Believe

Vendor benchmarks tell you a model is plausible, not that it suits your workloads. Run the same golden sets you use for API models - extraction, summarisation, classification, agentic tasks - with the same pass criteria and the same safety cases, and compare cost per successful task using your amortised GPU cost rather than a list price. Mistral's finance and legal claims and Beam's coding results are the hypotheses; your evaluation is the test.

pythonevals/compare_hosted.py
PROFILES = {
    "internal-reasoner": {"in": 0.35, "out": 1.10},   # YOUR amortised cost per 1M tokens at target utilisation
    "claude-sonnet-5-5": {"in": 2.00, "out": 10.00},
    "gpt-6.1-sol":       {"in": 2.00, "out": 10.00},
    "mistral-large-4":   {"in": 0.68, "out": 2.09},   # Mistral's API, for the residency-permitted workloads
}

def cost_per_success(runs, price):
    spend = sum((r.input_tokens * price["in"] + r.output_tokens * price["out"]) / 1e6 for r in runs)
    wins = sum(1 for r in runs if r.passed)
    return float("inf") if not wins else spend / wins

def decide(workload, results, floor=0.97, max_safety_regressions=0):
    table = {}
    for model, runs in results.items():
        pass_rate = sum(r.passed for r in runs) / len(runs)
        safety = sum(r.safety_regression for r in runs)
        table[model] = {"pass_rate": pass_rate, "safety": safety,
                        "cps": cost_per_success(runs, PROFILES[model]),
                        "eligible": pass_rate >= floor and safety <= max_safety_regressions}
    eligible = {m: t for m, t in table.items() if t["eligible"]}
    choice = min(eligible, key=lambda m: eligible[m]["cps"]) if eligible else None
    return {"workload": workload, "choice": choice, "table": table}

# Residency rule applied before cost: if the workload's data class is 'confidential', only
# 'internal-reasoner' is a candidate, whatever the table says.

5. What To Keep Off-Premises Anyway

  • Workloads where the frontier API's measured lead is real and the data is permitted - do not pay a capability tax for sovereignty you do not need.
  • Burst and experimentation. GPUs you own are a fixed cost; the API is elastic.
  • Model updates. Open weights do not patch themselves; owning the model means owning its evaluation, its safety behaviour and its upgrade cycle. Budget the engineering time, not just the hardware.

“Open weights at the frontier do not mean every bank should run its own model. They mean every bank can now choose which data never leaves - and have a model worth running on it.”


The UK And European Angle

Mistral's choice to train entirely in European data centres and to lead with sovereignty is a direct pitch to regulated European buyers, and a UK bank can take it up through the API for permitted data or through the 27 October weights for the rest. For the UK's sovereign-AI ambitions, the lesson is sharper: a credible frontier-class open model trained in Europe exists; one trained in Britain does not. The Sovereign AI Unit's compute and the City's demand could change that, and the firms that build the evaluation and serving discipline above will be the ones ready to use it when it does.

The Bottom Line

Beam's 501-billion-parameter sparse MoE with Apache 2.0 weights and Mistral Large 4's 1.05-trillion-parameter model trained in Europe with weights due on 27 October put open weights at the frontier in a single week, and sparse experts mean they serve on a node or two rather than a data centre. For a regulated firm the decision is about control: self-host the workloads whose data cannot leave or whose residency cannot be verified, size by memory with quantisation deciding node count, serve behind the gateway with the same telemetry as your API calls, evaluate on your own golden sets at amortised cost, and keep the elastic, permitted and frontier-dependent workloads on the APIs. That is the LLM engineering we do for financial firms in London, and this week made the self-hosted option real.

References & Further Reading

open-weight modelsself-hosted LLMAI Agency Developer LondonMistral Large 4Reflection BeamAI Automation London codedata residency
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information