New research shows the Model Context Protocol (MCP)—the interface that lets large-language-model (LLM) agents call external tools—can be hijacked through “tool-poisoning” attacks that succeed more than one-third of the time. Across 20 popular agents the average success rate was 36.5 %; the o1-mini model fell in 72.8 % of attempts, while Claude-3.7-Sonnet refused malicious calls under 3 % of the time. For anyone deploying LLM agents that rely on MCP, the findings turn a convenience feature into a supply-chain risk that can be exploited before any code ever runs.
Why MCP matters to developers today
MCP standardises how agents discover, register, and invoke tools such as file readers, web APIs, or email senders. By publishing a tool’s name, input schema and a short description, a server makes the capability available to any client that understands the protocol. The promise is simple: an agent can look up a tool, send a request, and receive a response without hard-coding each integration.
That flexibility also creates an implicit trust relationship. The specification tells clients to treat tool descriptions as trustworthy only if they come from a server the client already trusts. The new study shows that this trust can be abused.
How tool-poisoning differs from ordinary prompt injection
Traditional prompt injection inserts malicious instructions into the text that the model generates or receives at runtime. The model then follows those instructions because they appear in the same token stream as the user’s request.
Tool-poisoning, by contrast, hides the payload in the tool’s metadata—the name, description, or parameter schema that registers before any agent call. When an agent later selects the tool, it treats the description as part of the “trusted context” and may follow the hidden instruction without any runtime check. Because the injection occurs during registration, there is no point in the execution flow where a model can flag the payload as suspicious.
Scale of the problem – the MCPTox benchmark
The researchers behind MCPTox (arXiv:2508.14925) evaluated 45 MCP servers offering a total of 353 distinct tools. They scripted attacks against 20 widely used LLM agents, measuring how often the agents executed the poisoned tool call.
- Average success rate: 36.5 %
- Peak success: o1-mini at 72.8 %
- Best refusal: Claude-3.7-Sonnet, still under 3 %
The numbers reveal a stark reality: most agents do not refuse a poisoned call because the request looks like a legitimate tool invocation. The agents assume the tool description is a benign piece of documentation, not a vector for code execution.
Why agents rarely refuse poisoned calls
OWASP’s LLM01 guideline explains that LLMs do not differentiate between instructions and data—both are just tokens in a sequence. When a tool description says “send an email to admin@example.com with the subject ‘Update’”, the model cannot tell whether that line is a harmless comment or an instruction it should obey later. Consequently, the model treats the description as part of the trusted environment and follows any embedded command when the tool is invoked.
Existing guidance and its gaps
The MCP specification already advises clients to treat tool descriptions as untrusted unless they originate from a trusted server, and to keep a human in the loop for high-impact calls. The benchmark shows that many real-world deployments ignore or loosely interpret these recommendations.
Concrete steps developers can take today
- Kunci versi server – Gunakan referensi ke citra (image) atau hash server tertentu yang tidak dapat diubah, alih-alih menggunakan tag yang terus berubah. Hal ini mencegah penyerang menukar registri yang bersih dengan registri yang telah diracuni setelah penerapan (deployment).
- Mulai dengan daftar izin (allowlist) kosong – Aktifkan hanya alat yang telah diperiksa secara eksplisit. Apa pun yang tidak ada dalam daftar akan diblokir secara default.
- Batasi alat pengubah status (state-changing tools) – Perlukan persetujuan tambahan untuk alat apa pun yang menulis, mengirim, atau menghapus data. Pisahkan kemampuan "hanya baca" (read-only) dari kemampuan "bisa menulis" (write-capable) dalam skema.
- Tambahkan persetujuan manusia untuk panggilan berdampak tinggi – Untuk tindakan yang dapat memengaruhi sistem eksternal (misalnya, mengirim email, mengeksekusi perintah, memodifikasi file), minta peninjau manusia sebelum panggilan dikirim.
- Catat setiap pemanggilan alat – Rekam nama alat, argumen, stempel waktu (timestamp), dan agen asal. Jejak audit yang tidak dapat diubah membuat analisis pasca-kejadian (post-mortem) menjadi layak dilakukan dan dapat mencegah penyerang yang tahu bahwa tindakan mereka akan terlihat.
Perlakukan setiap deskripsi alat seperti kode sumber—yang tunduk pada linting, tinjauan kode (code review), dan kontrol versi—untuk menyelaraskan rantai pasokan MCP dengan praktik pengembangan perangkat lunak standar.
Argumen kontra dan pertanyaan terbuka
Namun, tolok ukur (benchmark) menunjukkan bahwa bahkan model paling canggih dalam studi tersebut menolak kurang dari tiga persen panggilan yang diracuni. Penyetelan halus (fine-tuning) mungkin dapat meningkatkan deteksi, tetapi tidak dapat menjamin keamanan terhadap muatan (payload) baru yang tertanam dalam bidang skema yang belum pernah dilihat oleh model tersebut.
Hal yang perlu diperhatikan selanjutnya
- Standar yang muncul – Perhatikan proposal dari komunitas keamanan LLM untuk mewajibkan tanda tangan kriptografis pada skema alat.
- Penguatan registri alat – Vendor mungkin mulai menawarkan registri read-only yang tidak dapat diubah sebagai layanan, sehingga mengurangi permukaan serangan (attack surface).
- Pertahanan tingkat model – Penelitian tentang teknik prompting atau model tambahan yang menandai metadata alat yang mencurigakan dapat melengkapi perlindungan di sisi host.
Kesimpulan praktisnya jelas: setiap penerapan berbasis MCP harus mengaudit deskripsi alat dengan ketelitian yang sama seperti yang diterapkan pada pustaka pihak ketiga. Mengabaikan risiko rantai pasokan mengubah abstraksi yang nyaman menjadi pintu belakang (backdoor) yang senyap. Dengan mengunci server, menerapkan daftar izin dengan hak istimewa minimum (least-privilege), membatasi tindakan pengubah status, melibatkan manusia jika diperlukan, dan menyimpan log yang tidak dapat diubah, pengembang dapat mencegah agen LLM mereka menjadi kaki tangan yang tidak disengaja.
