Skip to content
AI Integration

Skills, MCP Or CLI? The Three-Layer Agent Stack, And How To Decide Which Layer A Capability Belongs In

Every engineering leader building with agents in 2026 has hit the same question and mostly resolved it by accident: when a new capability arrives, does it become an MCP server, an Agent Skill, or a command-line tool the agent shells out to? The answer is not a preference, and treating it as one is why so many agent platforms end up as a pile of overlapping servers nobody can reason about. The three primitives solve genuinely different problems - connectivity, procedure and execution - and the decision has real consequences for token cost, auditability, blast radius and who is allowed to change what. This is the decision framework we use, written for senior engineers and CTOs who have to own the result.

AlchmAI Engineering14 min read

110M+

Monthly MCP SDK downloads by later 2026, up from 97M in March - growing faster than React did at the same stage

13,000+

Public MCP servers in the ecosystem, after the protocol moved to the Linux Foundation's Agentic AI Foundation in roughly 14 months

68%

Of production agent deployments have adopted MCP or an equivalent standardised tool layer

3 layers

Skills for procedure, MCP for governed connectivity, CLI for token-cheap local execution - complementary, not competing

A pattern we now see on most agent platform reviews: a firm has eleven MCP servers, four of which wrap tools that already have a perfectly good command-line interface, two of which exist only to encode a procedure rather than to provide access to anything, and at least one that nobody can name an owner for. Nothing is broken. It is simply that every new capability arrived as a server because a server was how the last one arrived, and eighteen months of that produces an architecture by accretion.

The industry has now converged on a cleaner way to think about it, and it is worth adopting deliberately before the accretion sets. There are three primitives, they solve different problems, and a well-built agent platform uses all three: Skills carry reusable domain capability and procedural know-how, MCP provides governed connectivity with rich semantics, authorisation and auditability, and CLI or local execution provides token-efficient, composable execution for tool-heavy work. Production-ready agents are built by matching each task to the layer that fits, rather than by choosing a favourite.

Layer One: MCP, For Governed Connectivity

MCP earned its position. It crossed 97 million monthly SDK downloads in March 2026, has since passed 110 million, went from an Anthropic internal experiment to the Linux Foundation's Agentic AI Foundation in roughly fourteen months, and now hosts over 13,000 public servers, with 68% of production agent deployments having adopted it or an equivalent standardised tool layer. For a financial institution the relevant property is not the ecosystem size, though. It is that MCP gives you a place to put authorisation and audit that is neither the model nor the underlying system.

A capability belongs in MCP when it has at least one of these characteristics:

  • It crosses a trust or entitlement boundary. Market data licensed per user and per venue, positions behind an information barrier, client records behind consent. The server is where you evaluate entitlement per call against the acting user's identity - never the server's own.
  • It needs an audit record per invocation. If a regulator or an internal reviewer might one day ask what the system accessed and on whose behalf, that question is answered at the MCP layer or not at all.
  • It is remote, authenticated and shared across teams. One server, many agents, one place to revoke.
  • Its schema is the contract. When you want the model to see typed, validated parameters rather than compose a shell command, the tool schema is doing real work.

Layer Two: Skills, For Procedure

Agent Skills are the newest of the three and the least understood by the engineering leaders we talk to, partly because the implementation is almost disappointingly plain: Markdown with a small amount of YAML metadata and some optional scripts. A Skill is a portable, versionable unit of procedural knowledge that teaches an agent a specific, repeatable task, independent of any live data connection. The ecosystem moved quickly - Skills.sh popularised an npm-style install in January 2026, and a well-built skill published to a community directory can be installed across every compatible platform with a single command.

The reason this matters in financial services is that an enormous proportion of what makes an institution's work correct is procedural rather than technical. How this firm scopes a suitability assessment. Which checks precede releasing a payment. What a complete break investigation includes. Which disclaimers a client-facing note requires and in what order. None of that is access - the agent can already reach every system involved. It is knowledge about how the job is done here, and before Skills it lived in three places, all bad: buried in a system prompt that nobody versioned, hardcoded into an orchestration script, or in a Confluence page the agent never read.

yamlskills/break-investigation/SKILL.md
---
name: break-investigation
description: >
  Investigate a settlement break end to end to this firm's standard.
  Use when a break is assigned for triage. Not for pricing disputes.
version: 3.2.0
owner: ops-engineering
# Skills declare what they need; they do not themselves grant access.
requires_tools:
  - positions.read
  - settlement.read
  - counterparty.read
---

# Break investigation

Work in this order. Do not skip step 2 even when the cause looks obvious
in step 1 - the most common false conclusion is a timing difference that
presents as a quantity mismatch.

1. **Classify.** Quantity, price, settlement date, SSI, or counterparty
   reference. Record the classification before investigating further.
2. **Rebuild both sides as at trade date.** Use point-in-time reads, not
   current state. A break investigated against today's data will find a
   difference that did not exist when the break was raised.
3. **Check the known-cause list** in references/known-causes.md before
   escalating. Roughly 60% of breaks match a documented pattern.
4. **Draft the finding** using templates/finding.md. Cite the specific
   records compared, by identifier, with the as-at timestamp.

## Escalate immediately, without completing the steps above, when
- the notional exceeds the desk's escalation threshold, or
- the counterparty appears on the restricted list, or
- the same break reference has been raised more than twice this month.

Three properties of that file are worth drawing out, because they are the argument for the layer. It is reviewable by the operations lead who owns the procedure, not only by an engineer - which means the person with the domain knowledge can actually maintain it. It is versioned in the repository and moves through the same review as code, so a change to how the firm investigates breaks is a pull request rather than a prompt edit. And it is loaded on demand rather than occupying context permanently, which is the key difference from putting the same content in a system prompt.

“MCP tells the agent what it may touch. A Skill tells it how your firm does the job. Conflating them produces servers that encode policy and prompts that encode access - both of which are owned by the wrong people.”

Layer Three: CLI, For Token-Cheap Execution

The most under-used layer, and the one that most often produces an immediate cost improvement. A great deal of agent work is mechanical: search a codebase, transform a file, run a query, parse a large result, filter thousands of rows down to the twelve that matter. Done through tool calls, every intermediate result passes through the model's context and you pay for all of it. Done through a command, the work happens outside the context window and only the result comes back.

bashcompare.sh
#!/usr/bin/env bash
# Reconciliation triage for a day's fills. Run by the agent as ONE tool call.
#
# Through MCP this would be: fetch internal fills (40k rows into context),
# fetch broker fills (40k rows into context), then ask the model to compare.
# Here the model sees only the breaks - typically a few dozen lines.
set -euo pipefail

DATE="$1"
duckdb -c "
  COPY (
    SELECT COALESCE(i.trade_id, b.trade_id) AS trade_id,
           i.qty AS internal_qty, b.qty AS broker_qty,
           i.price AS internal_price, b.price AS broker_price,
           CASE
             WHEN i.trade_id IS NULL THEN 'MISSING_INTERNAL'
             WHEN b.trade_id IS NULL THEN 'MISSING_BROKER'
             WHEN i.qty <> b.qty     THEN 'QTY_MISMATCH'
             WHEN abs(i.price - b.price) > 0.0001 THEN 'PRICE_MISMATCH'
           END AS break_type
    FROM read_parquet('internal/${DATE}/*.parquet') i
    FULL OUTER JOIN read_parquet('broker/${DATE}/*.parquet') b USING (trade_id)
    WHERE break_type IS NOT NULL
    ORDER BY break_type, trade_id
  ) TO '/dev/stdout' (FORMAT JSON);
"

The Decision Framework

Put together, here is the sequence we run when a new capability is proposed. It takes about two minutes per capability and prevents most of the architecture-by-accretion problem.

  1. 01Does it cross an entitlement or trust boundary, or require a per-call audit record? If yes, it is an MCP server, regardless of anything else. Access control and auditability are not negotiable and belong in one layer.
  2. 02Is it knowledge about how this firm does something, rather than access to something? If yes, it is a Skill. The test: could a competent outsider with the same system access still get it wrong? If so, you are encoding procedure.
  3. 03Does it process substantially more data than the model needs to see? If yes, it is a command. Measure it: if the intermediate result is more than a few thousand tokens and the final answer is small, the ratio decides it.
  4. 04Is it needed by one workflow only? Then it is not a server. A single-consumer MCP server is a function with a network hop and a standing context cost.
  5. 05Who owns changes to it? If the right owner is a domain expert rather than an engineer, that pushes strongly toward a Skill - Markdown they can edit beats a server they must file a ticket against.

The Security Dimension, Briefly

One cross-cutting point that senior engineers should carry into this decision: the three layers have different exposure to prompt injection, which remains the number one entry on the OWASP LLM Top 10 and is now treated as an unsolved architectural problem rather than a bug. MCP servers return content into the model's context, and tool descriptions themselves enter context supplied by the server - a compromised or malicious server can put instructions into a description field. Skills are content you author and version, so their trust level is your own review process. Commands are the narrowest: the agent constructs an invocation, the output comes back as text, and nothing about the command layer grants standing access.

The practical consequence is that the CLI layer is often the safest home for a capability that processes untrusted content, because it can be run in a sandbox with no credentials and no network, returning only a bounded result. That is a security argument for a layer most teams are choosing on cost grounds alone.

A Migration Path For Platforms That Already Sprawled

  1. 01Inventory every server with three columns: who owns it, which sessions load it, and whether it crosses an entitlement boundary. The third column usually shrinks the list immediately.
  2. 02Retire single-consumer servers into the calling code or a command. This is normally the largest and easiest win.
  3. 03Extract procedure from prompts into versioned Skills, and hand ownership to the function that owns the procedure. Do this before adding any new capability, or you will encode the next procedure in a prompt too.
  4. 04Stop loading everything into every session. Assemble the tool set per workflow. Measure tool-selection accuracy before and after - the improvement is usually larger than the token saving, and it is the one your users notice.
  5. 05Move bulk data work behind commands and measure tokens per completed task. Make that metric visible on a dashboard with an owner; it is the number that tells you whether the platform is getting better or merely bigger.

The Bottom Line

MCP won the connectivity argument decisively - 110 million monthly downloads, 13,000-plus public servers, Linux Foundation stewardship, 68% of production deployments - and the natural consequence has been that everything now arrives as a server, including a great deal that should not. The three primitives solve different problems and a well-built platform uses all three deliberately: MCP where a capability crosses an entitlement boundary or needs a per-call audit record, Skills where the thing being encoded is how your firm does a job rather than what it can reach, and commands where the agent would otherwise pay to read data it does not need to reason over. Ask the five questions before each new capability, keep session tool sets small because context is a recurring cost and tool-selection accuracy degrades with breadth, and give procedural knowledge to the people who own the procedure. That is the difference between an agent platform that gets better as it grows and one that just gets larger - and it is the architecture review we run most often for financial engineering teams in London.

References & Further Reading

Agent SkillsMCPAI Agency fintechAgentic AI finance codeAI Automation London codeagent architectureAI Agency Developer London
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information