Skip to content
AI Integration

Sonnet 5.5, GPT-6.1 Sol And Ember-1 In One Week: The Mid-Tier Is Now The Default For Financial Workloads - Token-Efficiency Evals, The Five Sonnet API Changes And A Self-Hosted Fallback, In Code

The mid-tier stopped being the compromise this week. Anthropic's Claude Sonnet 5.5 (28 September) kept Sonnet 5's $2/$10 pricing while jumping from 10.3% to 70.6% on Terminal-Bench 4.0, running 30%-plus faster and using markedly fewer tokens, with Opus-grade $0.20 cache reads, a 1M context, 128K output - and five breaking API changes. OpenAI's GPT-6.1 Sol (29 September) matched GPT-6 Astra on DeepSWE 1.1 at roughly one-fifth of the cost, at the same $2/$10 with $0.10 cached input. Fireworks' Ember-1 research preview, built on Kimi K3, delivers comparable quality with 35-50% fewer reasoning tokens and opens a credible self-hosted route. For a fintech, the consequence is that most production workloads - extraction, summarisation, agentic loops, code review - should now default to a mid-tier model and reserve the frontier for what measurably needs it. This is the migration: token-efficiency as a first-class eval metric, the exact Sonnet 5.5 changes, and a provider layer with a self-hosted exit, in code.

AlchmAI Engineering15 min read

70.6%

Sonnet 5.5 on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, at the same $2/$10 per million tokens (28 September)

1/5

GPT-6.1 Sol's cost relative to GPT-6 Astra for Astra-level results on DeepSWE 1.1; $2/$10 with $0.10 cached input (29 September)

35-50%

Fewer reasoning tokens for Ember-1 versus its Kimi K3 base at comparable quality, per Fireworks' research preview

5

Breaking API changes when moving from Sonnet 5 to Sonnet 5.5, from thinking configuration to advisor-model constraints

A week after the flagships, the mid-tier answered - and the answer changes procurement. Claude Sonnet 5.5 arrived on 28 September at the same $2 per million input and $10 per million output as Sonnet 5, with cache reads at $0.20 (matching Opus 5.5), a 1M-token context and 128K output, 30%-plus faster generation and, Anthropic says, far fewer tokens for most work. The benchmark jump is unusual for a same-price release: Terminal-Bench 4.0 from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5%. GPT-6.1 Sol followed on 29 September at $2/$10 with $0.10 cached input: Astra-level on DeepSWE 1.1 at roughly a fifth of the cost, 6.4 points above GPT-6 Sol at lower reasoning effort, and within 2.1 points of Astra on OSWorld 2.0's offline set at about a seventh of the cost per task. And Fireworks Research's Ember-1, built on Kimi K3 and available as a research preview, keeps the base model's quality while cutting reasoning tokens by 35-50% - in live A/B tests, about 35% fewer tokens per task.

Read together: for the workloads that make up most of a financial firm's AI spend, the question is no longer whether a mid-tier model is good enough but whether the frontier is measurably better enough to justify five times the price. Often it will not be. The engineering response has three parts: measure tokens as carefully as accuracy, migrate correctly, and keep an exit that does not depend on any single provider.

1. Make Tokens Per Success A Reported Metric

pythonevals/token_efficiency.py
from dataclasses import dataclass
from statistics import median

@dataclass
class Run:
    model: str
    workload: str
    passed: bool
    input_tokens: int
    cached_tokens: int
    output_tokens: int          # includes reasoning/thinking tokens where billed
    latency_ms: int

def summarise(runs: list) -> dict:
    by = {}
    for r in runs:
        by.setdefault((r.model, r.workload), []).append(r)
    out = {}
    for key, rs in by.items():
        wins = [r for r in rs if r.passed]
        total_tokens = sum(r.input_tokens + r.output_tokens for r in rs)
        out[key] = {
            "pass_rate": len(wins) / len(rs),
            "tokens_per_success": float("inf") if not wins else total_tokens / len(wins),
            "median_output_tokens": median(r.output_tokens for r in rs),
            "cache_share": sum(r.cached_tokens for r in rs) / max(1, sum(r.input_tokens for r in rs)),
            "p50_latency_ms": median(r.latency_ms for r in rs),
        }
    return out

# Rank candidates: quality floor first, then tokens_per_success, then latency.
# A model that passes 2% more but uses 40% more tokens is usually the wrong choice
# for extraction and summarisation - and often the right one for multi-step agents.

2. The Five Sonnet 5.5 Changes, As A Diff

Sonnet 5.5 is a same-price upgrade with a real migration. The published changes: disabled thinking is no longer accepted and becomes a 'between tools' mode; forced tool use gives way to automatic tool choice; thinking blocks are bound to the model and account that produced them; the older computer-use tool version is replaced by the 2026-08-01 toolset; and advisor-model configurations are restricted to Opus 5 and 5.5, Sonnet 5.5, or newer models. Isolate them in your profile layer so application code never sees them.

typescriptllm/profiles/sonnet-5-5.ts
import type { ModelProfile } from "../profiles";

export const SONNET_5_5: ModelProfile = {
  id: "claude-sonnet-5-5",
  provider: "anthropic",
  price: { in: 2, out: 10, cacheRead: 0.2 },
  context: 1_000_000,
  maxOutput: 128_000,
  supportsForcedTool: false,                 // was true on Sonnet 5
  thinking: "between_tools",                 // was "configurable"; disabled is rejected
  computerUseTool: "computer_toolset_20260801",
  advisorModelsAllowed: ["claude-opus-5", "claude-opus-5-5", "claude-sonnet-5-5"],
  replayThinkingAcrossModels: false,         // thinking blocks are model- and account-bound
};

// Request builder: one place where the differences become parameters.
export function buildRequest(p: ModelProfile, req: AppRequest) {
  const base: Record<string, unknown> = { model: p.id, max_tokens: Math.min(req.maxTokens, p.maxOutput) };
  base.thinking = p.thinking === "between_tools"
    ? { type: "between_tools" }              // Sonnet 5.5
    : req.wantThinking ? { type: "enabled", budget_tokens: req.thinkingBudget } : { type: "disabled" };
  if (req.schemaTool) {
    base.tools = [p.supportsForcedTool ? req.schemaTool : { ...req.schemaTool, strict: true }];
    base.tool_choice = p.supportsForcedTool ? { type: "tool", name: req.schemaTool.name } : { type: "auto" };
  }
  // Never replay prior thinking blocks into a different model/account.
  base.messages = p.replayThinkingAcrossModels ? req.messages : stripThinkingBlocks(req.messages);
  return base;
}
  • Run the contract tests per workload before flipping any traffic: extraction pipelines that relied on forced tool choice are the most likely to need a validation step now that the model chooses.
  • Fix effort and thinking mode per workload. Mid-session changes invalidate the cache, and Sonnet 5.5's cheaper cache reads only pay if the prefix stays identical.
  • Re-baseline token budgets. A model that uses a third fewer tokens will finish agent loops sooner; your loop caps and budget guards should be re-tuned, not left at last month's numbers.

3. Keep A Self-Hosted Exit

Ember-1 matters less for its benchmark numbers than for what it represents: a token-efficient reasoning model on an open base that a regulated firm could run inside its own perimeter, in the UK, for data classes the provider contracts do not cover - and as a fallback when a provider has an incident or changes terms. The provider layer should treat 'self-hosted open model' as a first-class target with the same profile shape, the same evals and the same routing, so switching is a configuration change.

typescriptllm/profiles/self-hosted.ts
export const SELF_HOSTED_REASONER: ModelProfile = {
  id: "ember-1",                              // or any open-weight model you have evaluated
  provider: "openai-compatible",              // vLLM / TGI endpoint speaking the OpenAI API shape
  endpoint: "https://llm.internal.yourbank.co.uk/v1",
  price: { in: 0.35, out: 1.1, cacheRead: 0 }, // YOUR amortised infra cost per 1M tokens, not a list price
  context: 256_000,
  maxOutput: 32_000,
  supportsForcedTool: true,
  thinking: "configurable",
  dataResidency: "UK",
  allowedDataClasses: ["public", "internal", "confidential"],   // the reason it exists
};

// Routing rule: any workload whose data class exceeds a provider's contract routes here;
// any workload can fail over here when the provider breaker opens.
export function eligible(profile: ModelProfile, dataClass: string): boolean {
  return profile.allowedDataClasses?.includes(dataClass) ?? dataClass === "public";
}

“The frontier is for the tasks that fail without it. Everything else just got a third cheaper in tokens and a fifth cheaper in price - if you measure, migrate and keep an exit.”


Where Each Model Lands For Finance Workloads

  • Document extraction, classification, summarisation: Sonnet 5.5 or GPT-6.1 Sol by default; Luna-class models where the quality floor allows. Token efficiency dominates cost here.
  • Multi-step agents and code changes: Sonnet 5.5's Terminal-Bench jump and cheap cache reads make it a serious default; escalate to Opus 5.5 or Astra only on measured failure.
  • Confidential and client-identifiable data: a self-hosted, UK-resident model for the data classes your provider contracts exclude - Ember-1-style efficiency makes that affordable.
  • Customer-facing figures: still never computed by any model. Grounded calculation is unchanged by this week.

The Bottom Line

Sonnet 5.5 at an unchanged $2/$10 with a seven-fold Terminal-Bench improvement, faster output and fewer tokens; GPT-6.1 Sol at Astra-level results for a fifth of the price; and Ember-1 showing 35-50% fewer reasoning tokens on an open base: the mid-tier is now the default for most financial workloads, with the frontier reserved for tasks that measurably need it. The engineering to capture that is token efficiency as a reported eval metric, a profile layer that isolates Sonnet 5.5's five API changes and every future one, per-workload contract tests before traffic moves, and a self-hosted, UK-resident exit that shares the same profile shape and routing. That is the model-operations work we do as an AI agency for fintech teams in London, and this was the week it started paying for itself at the mid-tier.

References & Further Reading

AI Agency Developer LondonClaude Sonnet 5.5GPT-6.1 Soltoken efficiencyAI Automation London codeAI Agency fintechmodel migration
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information