Skip to content
AI Integration

The Agents API Can Now Click: Putting Computer-Use Agents On Legacy Bank Systems With A Harness That Watches Every Screen, Approves Every Consequential Click And Keeps The Evidence

DevDay on 29 September gave OpenAI's Agents API computer use - agents that operate software through its graphical interface, running in a managed sandbox with multi-agent coordination, tool search and context compaction, plus a limited preview of Bedrock Managed Agents for firms that live in AWS. Google's Gemini 4 Argon, announced the next day for vetted cyber defenders first and API customers later at $2/$10 introductory pricing, leads on AutomationBench and offers a million-token output limit. Apple, meanwhile, said it will tighten macOS Full Disk Access because agents are reading everything on users' machines. For banks the opportunity is obvious and specific: decades of legacy systems with no API, operated by people clicking through screens. The risk is equally specific, and the labs just demonstrated it. Here is the harness we build for computer-use agents on legacy systems - screen-state assertions, a consequential-click approval gate, read-only by default, a recording for every run - with code.

AlchmAI Engineering15 min read

29 Sept

OpenAI's Agents API gained computer use, multi-agent coordination, tool search and context compaction in a managed sandbox

51.3%

Gemini 4 Argon on AutomationBench, with a 1M-token output limit and $2/$10 introductory pricing (announced 30 September, cyber defenders first)

0 APIs

On the legacy bank systems where computer use adds the most value - mainframe terminals, thick clients, vendor portals operated by clicking

Every click

That changes state in a system of record should pass a consequential-action gate the agent cannot approve itself

Computer use went from research demo to platform feature this week. OpenAI's Agents API now lets an agent operate software through its graphical interface, with multi-agent coordination, tool search and context compaction, in a sandbox OpenAI runs; the company said harness work had halved computer-use latency, and announced Bedrock Managed Agents in limited preview for firms whose estate lives in AWS. Google followed with Gemini 4 Argon on 30 September: top of AutomationBench at 51.3%, a million-token output limit, strong results on finance and legal knowledge work, $2/$10 introductory pricing, and a release sequence that puts vetted cyber defenders first and API customers later - because the model can also find and patch vulnerabilities autonomously. And Apple said on 2 October it would tighten macOS Full Disk Access, after reports of an agent reading a journalist's private messages, warning that 'as AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially'.

For a bank the pull is specific. Every institution has systems that predate APIs: mainframe green screens, thick-client applications, vendor portals, spreadsheets that are really databases. They are operated by people who read a screen, type into fields and click buttons, often thousands of times a day. Computer-use agents promise to do that work without a re-platforming programme. The push is equally specific: an agent that can click anything can click the wrong thing, and this month's incidents - agents reaching systems they should not have, a model shelved for acting beyond its authorisation - were all agents with more reach than oversight.

1. Screen-State Assertions

The agent should never act on a screen it has not proved it is on. Each step in a workflow declares the screen it expects - identified by stable anchors such as a title, a field label layout or a known control - and the harness asserts the anchors are present before any input is sent. A failed assertion stops the run; it does not let the model improvise.

pythonharness/screens.py
from dataclasses import dataclass, field

@dataclass(frozen=True)
class ScreenSpec:
    name: str
    anchors: list                     # text that must be visible, e.g. ["PAYMENT ENQUIRY", "Account No", "F3=Exit"]
    forbidden: list = field(default_factory=list)   # text that must NOT be visible, e.g. ["CONFIRM RELEASE"]

class ScreenAssertionError(Exception): ...

def assert_screen(ocr_text: str, spec: ScreenSpec) -> None:
    missing = [a for a in spec.anchors if a not in ocr_text]
    present = [f for f in spec.forbidden if f in ocr_text]
    if missing or present:
        raise ScreenAssertionError(f"expected {spec.name}; missing={missing} forbidden_present={present}")

# A workflow is a list of (ScreenSpec, action). The harness asserts BEFORE every action and
# re-asserts the expected next screen AFTER it, so a mis-click surfaces immediately.
PAYMENT_ENQUIRY = ScreenSpec("payment_enquiry", ["PAYMENT ENQUIRY", "Account No", "F3=Exit"], forbidden=["CONFIRM RELEASE"])
PAYMENT_DETAIL  = ScreenSpec("payment_detail", ["PAYMENT DETAIL", "Beneficiary", "Status"], forbidden=["CONFIRM RELEASE"])

2. The Consequential-Click Gate

Classify every action the agent can take as read, navigate or mutate. Reads and navigation run freely. Mutations - typing into a field that persists, pressing a button that releases, approves, deletes or sends - require a token minted by a separate approval service for that exact action on that exact record, and the agent's identity cannot call the approval service. This is the same propose-and-approve pattern we use for trading agents, applied to clicks.

pythonharness/gate.py
import hashlib, json

MUTATING_ACTIONS = {"press:F10", "press:ENTER_on_confirm", "click:Release", "click:Approve", "click:Delete", "type:persisted_field"}

def action_hash(screen: str, action: str, record_id: str, value: str = "") -> str:
    return hashlib.sha256(json.dumps([screen, action, record_id, value], sort_keys=True).encode()).hexdigest()

class ConsequentialGate:
    def __init__(self, approvals):   # approvals = client for the approval service; agent creds cannot reach it
        self.approvals = approvals

    def check(self, screen: str, action: str, record_id: str, value: str = "") -> None:
        if action not in MUTATING_ACTIONS:
            return
        h = action_hash(screen, action, record_id, value)
        if not self.approvals.has_valid_token(h, max_age_seconds=120):
            raise PermissionError(f"mutating action requires approval: {action} on {record_id} at {screen}")

# Flow: the agent PROPOSES the mutation (screen, action, record, value) -> the harness shows a human
# the screenshot + proposal -> the human approves in a separate app -> a token bound to the hash is
# minted -> only then does the harness send the keystroke. The agent never sees the token.

3. Read-Only By Default, Recorded Always

  • Start every legacy workflow in read-only mode: the agent may navigate and extract, never type into persisted fields. Most value - lookups, reconciliations, report assembly - needs nothing more.
  • Record a screenshot before and after every action, with the OCR text, the assertion result and the action taken, into an append-only store. A run is a film, not a log line.
  • Run inside a sandbox with egress limited to the target system and the model provider, and with the agent's own credentials - never a human's session - on the legacy system.
  • Apple's Full Disk Access change is the consumer version of the same rule: an agent should see the application it is working in, not the whole machine.
pythonharness/run.py
def run_step(agent, harness, step, record_id):
    before = harness.screenshot()
    text = harness.ocr(before)
    assert_screen(text, step.expected_screen)                       # 1. prove where we are
    proposal = agent.next_action(text, step.goal)                   # model proposes {action, value}
    harness.gate.check(step.expected_screen.name, proposal.action, record_id, proposal.value)  # 2. gate mutations
    if harness.read_only and proposal.action in MUTATING_ACTIONS:
        raise PermissionError("read-only run attempted a mutation")
    harness.send(proposal)                                          # the only place input is sent
    after = harness.screenshot()
    assert_screen(harness.ocr(after), step.next_screen)             # 3. prove where we landed
    harness.evidence.append({"step": step.name, "record": record_id, "before": before.id, "after": after.id,
                             "action": proposal.action, "value_hash": action_hash(step.expected_screen.name, proposal.action, record_id, proposal.value)})

Where Computer Use Pays Off First In A Bank

  1. 01Reconciliation lookups across systems with no API - read-only, high volume, immediate value.
  2. 02Vendor and regulator portals with manual data entry - mutating, but low value per action and easy to approve in batches.
  3. 03Legacy report assembly: navigating screens to extract figures that today are re-keyed into spreadsheets.
  4. 04Migration discovery: agents that walk a legacy application and document its screens and fields - the cheapest way to scope a modernisation Barclays-style.

“Computer use lets an agent operate the systems you never managed to give an API. The harness decides whether that is the most useful automation you have ever deployed or the most dangerous.”


Choosing The Stack

OpenAI's Agents API with computer use and Bedrock Managed Agents suit firms that want the sandbox and orchestration managed; Gemini 4 Argon's AutomationBench lead and large output limit make it a strong candidate once it reaches API customers, with the caveat that its first audience is cyber defenders for a reason. Whichever model drives the clicks, the harness above is yours: the assertions, the gate, the read-only default and the evidence store are what a bank's control functions will ask about, and they do not change when the model does.

The Bottom Line

The Agents API can now operate software through its interface, Gemini 4 Argon leads the automation benchmarks with a million-token output limit, Bedrock Managed Agents bring the pattern to AWS, and Apple is tightening disk access because agents already read too much. For banks the prize is automating legacy systems that have no API without re-platforming them; the condition is a harness that proves which screen the agent is on before every action, gates every mutating click behind an approval the agent cannot grant, defaults to read-only and records every step as evidence. That is the AI integration work we do for financial institutions in London, and this week made it buildable on mainstream platforms.

References & Further Reading

AI Agency Developer Londoncomputer use agentslegacy modernisationAgentic AI finance codeAI Automation London codeGemini 4 ArgonAI Agency fintech
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information