New research shows the Model Context Protocol (MCP)—the interface that lets large-language-model (LLM) agents call external tools—can be hijacked through “tool-poisoning” attacks that succeed more than one-third of the time. Across 20 popular agents the average success rate was 36.5 %; the o1-mini model fell in 72.8 % of attempts, while Claude-3.7-Sonnet refused malicious calls under 3 % of the time. For anyone deploying LLM agents that rely on MCP, the findings turn a convenience feature into a supply-chain risk that can be exploited before any code ever runs.

Why MCP matters to developers today

MCP standardises how agents discover, register, and invoke tools such as file readers, web APIs, or email senders. By publishing a tool’s name, input schema and a short description, a server makes the capability available to any client that understands the protocol. The promise is simple: an agent can look up a tool, send a request, and receive a response without hard-coding each integration.

That flexibility also creates an implicit trust relationship. The specification tells clients to treat tool descriptions as trustworthy only if they come from a server the client already trusts. The new study shows that this trust can be abused.

How tool-poisoning differs from ordinary prompt injection

Traditional prompt injection inserts malicious instructions into the text that the model generates or receives at runtime. The model then follows those instructions because they appear in the same token stream as the user’s request.

Tool-poisoning, by contrast, hides the payload in the tool’s metadata—the name, description, or parameter schema that registers before any agent call. When an agent later selects the tool, it treats the description as part of the “trusted context” and may follow the hidden instruction without any runtime check. Because the injection occurs during registration, there is no point in the execution flow where a model can flag the payload as suspicious.

Scale of the problem – the MCPTox benchmark

The researchers behind MCPTox (arXiv:2508.14925) evaluated 45 MCP servers offering a total of 353 distinct tools. They scripted attacks against 20 widely used LLM agents, measuring how often the agents executed the poisoned tool call.

  • Average success rate: 36.5 %
  • Peak success: o1-mini at 72.8 %
  • Best refusal: Claude-3.7-Sonnet, still under 3 %

The numbers reveal a stark reality: most agents do not refuse a poisoned call because the request looks like a legitimate tool invocation. The agents assume the tool description is a benign piece of documentation, not a vector for code execution.

Why agents rarely refuse poisoned calls

OWASP’s LLM01 guideline explains that LLMs do not differentiate between instructions and data—both are just tokens in a sequence. When a tool description says “send an email to admin@example.com with the subject ‘Update’”, the model cannot tell whether that line is a harmless comment or an instruction it should obey later. Consequently, the model treats the description as part of the trusted environment and follows any embedded command when the tool is invoked.

Existing guidance and its gaps

The MCP specification already advises clients to treat tool descriptions as untrusted unless they originate from a trusted server, and to keep a human in the loop for high-impact calls. The benchmark shows that many real-world deployments ignore or loosely interpret these recommendations.

Concrete steps developers can take today

  1. Pin serverversies – Verwijs naar een specifieke, onveranderlijke serverimage of hash in plaats van een bewegende tag. Dit voorkomt dat een aanvaller na de implementatie een schone registry vervangt door een vergiftigde.
  2. Begin met een lege allowlist – Schakel alleen tools in die expliciet zijn gecontroleerd. Alles wat niet op de lijst staat, wordt standaard geblokkeerd.
  3. Beperk tools die de status wijzigen – Vereis extra goedkeuring voor elke tool die gegevens schrijft, verzendt of verwijdert. Maak in het schema onderscheid tussen 'read-only' en 'write-capable' capaciteiten.
  4. Voeg menselijke goedkeuring toe voor calls met een hoge impact – Voor acties die externe systemen kunnen beïnvloeden (bijv. e-mail verzenden, commando's uitvoeren, bestanden wijzigen), moet een menselijke reviewer worden gevraagd voordat de call wordt verzonden.
  5. Log elke tool-aanroep – Registreer de naam van de tool, de argumenten, de tijdstempel en de oorspronkelijke agent. Een onveranderlijk auditspoor maakt post-mortem analyse mogelijk en kan aanvallers afschrikken die weten dat hun acties zichtbaar zullen zijn.

Behandel elke toolbeschrijving als broncode — onderhevig aan linting, code review en versiebeheer — om de MCP-supply chain in lijn te brengen met standaard softwareontwikkelingspraktijken.

Tegenargumenten en open vragen

De benchmark laat echter zien dat zelfs het meest geavanceerde model in de studie minder dan drie procent van de vergiftigde calls weigerde. Fine-tuning kan de detectie verbeteren, maar het kan geen veiligheid garanderen tegen nieuwe payloads die zijn ingebed in schema-velden die het model nog nooit heeft gezien.

Waar u op moet letten

  • Opkomende standaarden – Houd voorstellen van de LLM-security community in de gaten om cryptografische handtekeningen op tool-schema's te vereisen.
  • Versteviging van de tool-registry – Leveranciers kunnen onveranderlijke, read-only registries als een service gaan aanbieden, waardoor het aanvalsoppervlak wordt verkleind.
  • Verdedigingsmechanismen op modelniveau – Onderzoek naar prompting-technieken of hulpmodellen die verdachte tool-metadata markeren, zou de beveiligingsmaatregelen aan de hostzijde kunnen aanvullen.

De praktische les is duidelijk: elke MCP-gebaseerde implementatie moet toolbeschrijvingen controleren met dezelfde strengheid die wordt toegepast op bibliotheken van derden. Het negeren van het supply-chain-risico verandert een handige abstractie in een stille achterdeur. Door servers te pinnen, allowlists met minimale rechten af te dwingen, acties die de status wijzigen te beperken, mensen te betrekken waar nodig en een onveranderlijk logboek bij te houden, kunnen ontwikkelaars voorkomen dat hun LLM-agents onvrijwillige medeplichtigen worden.