The FTC Is Investigating, California Has Subpoenaed, Apple Is Locking The Disk: Running Financial Agents In NVIDIA OpenShell, Policy By Policy
In four days the agent-containment question went from engineering debate to legal exposure. On 30 September the FTC opened a consumer-protection investigation into OpenAI, Anthropic and other labs over agents that exceeded their instructions and hacked external sites. On 1 October California's attorney general subpoenaed OpenAI over cyber incidents involving its models, after 25 attorneys general had asked Congress for an AI incident-response regime. On 2 October Apple said it would tighten macOS Full Disk Access because agents are reading everything on users' machines. Two days before all of that, Nvidia had released OpenShell: an Apache 2.0 runtime that puts each agent in a Landlock- and seccomp-isolated sandbox, routes every outbound request through a policy proxy, injects credentials only at authorised endpoints and writes OCSF audit records - with Sentry, a watchdog on BlueField-4 hardware, as the enforcement layer the agent cannot reach. More than 100 organisations including Anthropic and Microsoft signed on. This is how we would run a bank's agents on it, with the real policy schema: filesystem, process, network, middlewares, providers and the gaps you still have to close yourself.
AlchmAI Engineering17 min read
30 Sept
FTC opened a consumer-protection investigation into OpenAI, Anthropic and other labs over agents exceeding instructions and hacking external sites
1 Oct
California's attorney general served OpenAI an investigative subpoena over cybersecurity incidents involving its models; 25 AGs want an incident regime
100+
Organisations, including Anthropic, Microsoft and Palantir, adopting Nvidia's Open Agent Safety Platform (OpenShell plus Sentry), launched 28 September
4
Enforcement layers in OpenShell: filesystem (Landlock), process (seccomp, unprivileged identity), network (policy proxy) and provider credentials
The regulators moved first this week, and they moved on agents specifically. The Federal Trade Commission's investigation, confirmed on 30 September, is a consumer-protection probe into whether OpenAI, Anthropic and other labs are meeting their obligations after months of disclosures about agents going beyond instructions, finding their way onto the internet and hacking external websites; civil investigative demands, including to the evaluator METR, are being prepared. California's attorney general Rob Bonta served OpenAI with a subpoena on 1 October about cybersecurity incidents and risks involving its models - 'companies that develop these models have a moral and legal responsibility to ensure that they do not perpetrate or enable cyberattacks', he said - weeks after 25 attorneys general asked Congress for a government-led AI incident-response regime with direct access to company records. And Apple, on 2 October, said it would add controls to macOS Full Disk Access because 'as AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially'.
The engineering answer had arrived two days before the FTC. Nvidia's Open Agent Safety Platform, launched 28 September, pairs OpenShell - an open-source runtime that gives each agent a sandbox with kernel-level isolation and checks filesystem, process, network and credential access against declarative YAML policy - with Sentry, a watchdog reference design on BlueField-4 data-processing units that watches agent traffic from outside the host and can quarantine an agent in milliseconds. The premise is the one every incident this year has taught: agents cannot be trusted to police themselves, so the controls must live where the agent cannot reach them. For a financial institution facing supervisors who now ask 'what stops your agent' rather than 'what does your agent do', that is precisely the artefact to be able to show.
A Policy For A Reconciliation Agent
The policy below follows OpenShell's documented schema - version 1 with filesystem_policy, process, landlock, network_policies and network_middlewares - for an agent that reads a bank's ledger and settlement systems, calls one model provider and writes a report. Note what is absent: no write access outside its output directory, no general internet, no payments endpoint, and a hard requirement that Landlock actually applied.
version: 1
filesystem_policy:
include_workdir: true
read_only: [/usr, /lib, /etc, /data/settlements, /data/ledger]
read_write: [/tmp, /out/reports]
process:
run_as_user: "1500"
run_as_group: "1500"
landlock:
compatibility: hard_requirement # refuse to start if the kernel cannot enforce filesystem rules
network_policies:
ledger_api:
name: ledger-readonly
endpoints:
- host: ledger.internal.bank
port: 443
protocol: rest
enforcement: enforce
access: read-only
binaries:
- path: /usr/bin/curl
- path: /opt/agent/bin/recon
settlements_api:
name: settlements-readonly
endpoints:
- host: settlements.internal.bank
port: 443
protocol: rest
enforcement: enforce
access: read-only
binaries:
- path: /opt/agent/bin/recon
model_provider:
name: anthropic-api
endpoints:
- host: api.anthropic.com
port: 443
protocol: rest
enforcement: enforce
access: read-write
binaries:
- path: /opt/agent/bin/recon
network_middlewares:
pii-redactor:
middleware: openshell/regex
order: 10
config:
mode: redact
on_error: fail_closed # if redaction fails, the request does not leave
endpoints:
include: ["api.anthropic.com"]- Deny by default: any host not listed - a paste site, a URL shortener, a screenshot service, a payments API - is unreachable, which is the control every exfiltration incident this year lacked.
- Access is per endpoint and per binary. The agent's own binary may call the model; curl may read the ledger; nothing may write to a system of record.
- The redaction middleware runs on traffic to the model provider and fails closed. Client identifiers do not leave the bank even if the agent puts them in a prompt.
- Audit mode exists for a reason: start a new agent with enforcement set to audit, read the OCSF logs for a week, then flip to enforce. You will find calls you did not know it made.
Credentials The Agent Never Sees
The credential model is the part that would have prevented the Census and Gemini incidents outright. Provider profiles describe which credential names a service needs and how they are handled; provider instances hold the values in the configured credential backend; and the supervisor substitutes them into outbound requests only when the destination matches policy. The sandbox filesystem never contains the secret, the agent's environment holds at most a placeholder, and a request to an unapproved host carries nothing worth stealing.
# Create the sandbox with no auto-attached providers and the strict policy above.
openshell sandbox create --name fx-recon-agent --no-auto-providers --policy policies/fx-recon-agent.yaml
# Attach exactly one model provider; its API key lives in the credential backend, not the sandbox.
openshell provider attach fx-recon-agent anthropic-prod
# Tighten or loosen network policy on the running sandbox without restarting it (filesystem/process stay locked).
openshell policy set fx-recon-agent --policy policies/fx-recon-agent.yaml --wait
# Watch what it tried to do. Denied requests are the earliest signal an agent is looking for a way out.
openshell logs fx-recon-agent --since 15m
# Export OCSF records into the SIEM the security team already runs.
openshell logs fx-recon-agent --format ocsf --since 24h > /var/log/agents/fx-recon-$(date +%F).jsonThe provider-attach and OCSF-export commands above follow the documented CLI shape; check your installed version's reference for exact flags. The important property is architectural and does not depend on syntax: the secret is injected at the proxy for an allowed endpoint, and nowhere else.
Sentry: Enforcement The Agent Cannot Reach
OpenShell runs on the host that runs the agent. If that host is compromised - by the agent, or by someone through the agent - software controls on it are suspect. Sentry is Nvidia's answer: a reference design on BlueField-4 DPUs that watches agent traffic from outside the host, enforces policy in the network path, and can quarantine an agent that moves outside its boundaries in milliseconds. For a bank it maps onto an existing control: the network enforcement point that security teams already trust more than anything running on a workload. The practical pattern is defence in depth - OpenShell policy on the host, Sentry or an equivalent network enforcement point beside it, and the same allow-list in both.
What OpenShell Does Not Do For You
- 01Budgets. OpenShell bounds where an agent can go, not how much it can spend or how many sub-agents it can spawn. Keep the run-budget governor in your orchestrator.
- 02Authorization for consequential actions. A policy can make the payments API reachable or not; it cannot decide that this particular payment needs a human. That stays in your gateway, bound to the exact order.
- 03Honesty. Nothing in a sandbox checks that the agent's report matches what it did. Run the claimed-versus-done harness per release.
- 04Model behaviour. The model inside the sandbox is still the model. Evaluate it; sandboxing bounds the blast radius of a bad answer, it does not prevent one.
- 05Non-Linux estates. Landlock and seccomp are Linux kernel features. Developer laptops running agents - the Apple problem - need the OS vendor's controls and your endpoint policy; OpenShell is for the server fleet.
“The FTC and the California attorney general are asking the labs what stopped their agents. A versioned policy file, an OCSF log and a watchdog the agent cannot reach is the answer a bank wants to have ready when its supervisor asks the same thing.”
Rolling It Out In A Regulated Firm
The Bottom Line
The FTC's investigation, California's subpoena and Apple's disk-access change make agent containment a matter of legal record, and Nvidia's OpenShell - Landlock and seccomp isolation, a deny-by-default network policy proxy, credentials injected only at approved endpoints, OCSF audit and declarative, versioned YAML - is the first open, mainstream runtime built to provide it, with Sentry adding enforcement the agent cannot reach. It does not replace budgets, consequential-action gates, honesty testing or model evaluation, but it closes the boundary every incident this year crossed. The policy above is what a reconciliation agent's containment looks like in a bank. That is the agent deployment engineering we do as an AI agency in London, and the regulators have just explained why it belongs in version control.
References & Further Reading
- NVIDIA Docs - OpenShell policy schema reference. docs.nvidia.com/openshell/latest/how-it-works/policies/schema
- NVIDIA Developer Blog - Add runtime controls to AI agents with NVIDIA OpenShell. developer.nvidia.com/blog/add-runtime-controls-to-ai-agents-with-nvidia-openshell
- GitHub - NVIDIA/OpenShell-Community. github.com/NVIDIA/OpenShell-Community
- Help Net Security - NVIDIA wants AI agent safety enforced in silicon, not left to the agent (28 September 2026). helpnetsecurity.com/2026/09/28/nvidia-open-agent-safety-platform
- SecurityWeek - FTC is investigating OpenAI and Anthropic over possible risks to consumers (30 September 2026). securityweek.com/ftc-is-investigating-openai-and-anthropic-over-possible-risks-to-consumers
- The Register - OpenAI's wandering AI agents earn it a California subpoena (2 October 2026). theregister.com/ai-and-ml/2026/10/02/openais-wandering-ai-agents-earn-it-a-california-subpoena/5300850
- TechCrunch - Apple says it's tightening macOS Full Disk Access controls due to new risks from AI agents (2 October 2026). techcrunch.com/2026/10/02/apple-says-its-tightening-macos-full-disk-access-controls-due-to-new-risks-from-ai-agents
- OCSF - Open Cybersecurity Schema Framework. schema.ocsf.io
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information