Three in five deployed MCP tools tell an agent nothing about what they do
A September 10, 2026 probability sample of the MCP registry finds 58.8% of deployed tools carry no safety annotation — and that hand-curated samples understate the gap by 17 points.
What is this?
On September 10, 2026, independent researcher Haseeb Mohammed Afsar published a measurement of the Model Context Protocol server population that does something the existing literature does not: it draws a probability sample and refuses to fix anything.
From a census of the official MCP registry — 24,135 servers at the 2026-08-22 sweep, up from 16,548 on 2026-07-14 — the study builds a frame of 7,258 npm-published, stdio-declared active entries, draws 400 with a published seed, and probes each one exactly once. No credentials, no retries, no repair. Every draw is recorded with an outcome.
Two numbers matter for security posture. Only 48.8% completed an initialize handshake. And of the 2,766 tools advertised by the 195 servers that did run, 1,626 — 58.8% — carry no safety annotation at all.
How it works
MCP’s ToolAnnotations, shipped in the 2025-03-26 spec revision, are four optional booleans that let a server tell a client what a tool does before the client calls it: readOnlyHint, destructiveHint, idempotentHint, openWorldHint. They are the protocol’s entire pre-call risk vocabulary.
The spec is careful about them in two directions at once. First, they are hints, not contracts — a March 16, 2026 post from the MCP maintainers restates that clients must treat annotations from untrusted servers as untrusted, because a server can claim readOnlyHint: true and delete your files anyway. Second, the defaults are deliberately pessimistic: an unannotated tool is to be assumed non-read-only, potentially destructive, non-idempotent, and open-world.
That default is the load-bearing part. It only produces safe behavior if clients honor it, and honoring it at a 58.8% omission rate means prompting the user on roughly three of every five tools an agent can reach.
The measurement also shows where the omission lives, and it is not spread thin. Of the 194 servers advertising at least one tool, 72 annotate every tool and 122 annotate none. Not one partially annotated server appeared in the sample. The authors decline to state that as an absolute — zero observations out of 194 supports a one-sided 95% upper bound of 1.53% on partial annotation, not a denial — and they report that an earlier unreleased run found four partial servers and that the present run does not reproduce it.
drawn from registry frame 400 servers
├─ handshake completed 195 48.8% ← the only tier anyone measures
├─ never started 150 37.5%
├─ needs credentials 53 13.3%
└─ package unavailable 2 0.5%
of the 195 that ran → 2,766 advertised tools
├─ fatal JSON Schema violations 0 0.0%
└─ no safety annotation 1,626 58.8% (curated frame: 41.5%)
The schema line is worth noting because it cuts against a common assumption: zero of 2,766 tools carried a fatal JSON Schema violation. MCP tool schemas are not, in practice, malformed. The variance is entirely in the optional metadata.
Why it matters
The finding with the longest reach is not the 58.8%. It is that curation flatters ecosystem health in both directions at once. Run through the same instrument, a 24-server hand-curated frame of reference and popular servers gave a 66.7% start rate (17.9 points better) and a 41.5% annotation omission rate (17.3 points better). Reference servers annotate; curated frames are full of reference servers. Any posture statistic sampled from a popularity list, a reference set, or a pipeline that repairs servers until they run is measuring a population that selected itself for working.
That has a direct consequence for how you read the rest of the MCP security literature. Nicolás Padilla’s July 31, 2026 dynamic assessment of 414 internet-facing servers — 68 reportable findings, 91.8% with no OAuth, and 41.6% of confirmed servers gone within three days between runs — corroborates the churn from the remote side. Two independent instruments now agree that a large fraction of what the registry advertises is not a durable, running thing.
For a client, an unannotated tool leaves exactly two options, both bad at this rate. Honor the pessimistic default and you generate approval prompts on most of the tool surface, which is how you train users to click through — the same fatigue we have covered in the gap between what an approval dialog blocks and what it shows. Ignore the default and you have silently auto-approved tools whose behavior nobody declared. The maintainers’ own post notes that no MCP client today lets users filter tools by annotation value, and none surface annotations in approval prompts, so in practice most of this vocabulary is not reaching the decision it was designed to inform.
The 37.5% that never start is a second-order problem. A registry where two in five published entries are inert is a registry full of names that resolve to nothing — and a name that resolves to nothing is a name an attacker can make resolve to something. That is the install-path hygiene issue behind description rug-pulls and behind the Deadbugz campaign’s runtime-gated metadata, where the hostile behavior only appears after approval.
One last result, aimed at anyone benchmarking agents: held to one similarity method, real MCP tools show 2.8% near-duplication at cosine 0.70 and 0.0% across independent authors at every threshold tested. BFCL v4 shows 16.7%, of which 16.4 points sit between independently presented tasks. Before deduplication, 68.8% of raw BFCL rows and 85.6% of raw UltraTool rows are exact name-plus-description repeats, against 0.4% for real MCP. UltraTool after deduplication is cleaner than real tools at 0.3%, so this is a property of one corpus and not of synthetic corpora as a class — the paper says so explicitly. The operational point: a statistic computed over these releases without global deduplication measures repetition, not tools.
Defenses
Treat a missing annotation as a finding, not as a default. If you run an internal MCP gateway, enumerate tools/list across every registered server and report the annotation omission rate per server. The all-or-nothing distribution makes this cheap: you are classifying servers, not tools.
Do not let an absent hint become an implicit approval. Whatever your client does with destructiveHint: true, it should do the same for a tool with no hints at all. That is what the spec’s default says, and it is the case most clients exercise least.
Never let a hint from an untrusted server relax a control. Annotations are server-authored and unverified. Use them to tighten posture — an openWorldHint: true tool marks the session as carrying untrusted content — and never to skip a gate. Where you need a guarantee rather than a hint, put it in the authorization layer, the transport, or the sandbox. Enforcement belongs where it can be enforced.
Pin the install path, not the registry name. Given a 37.5% non-start rate, most registry entries are not things you depend on — but the ones you do depend on should be pinned by version and integrity hash, resolved from a mirror you control, and re-checked on every reconnect rather than trusted from the first handshake.
Re-read your own posture numbers. If a vendor report, an internal audit, or a scanner benchmark sampled from popular or reference servers, assume it is roughly 17 points optimistic on both start rate and annotation coverage. Ask what the sampling frame was before you use the number in a risk register.
Annotate what you ship. If you author a server, set readOnlyHint: true on read-only tools, destructiveHint: false on purely additive ones, and openWorldHint: false on closed-domain ones. It costs four booleans and it is the only signal a client gets before the call.
Status
| Item | Detail |
|---|---|
| Publication | arXiv:2609.10962v1 [cs.SE], September 10, 2026 |
| Author | Haseeb Mohammed Afsar (independent researcher) |
| Census sweeps | 2026-07-14 (16,548 servers) and 2026-08-22 (24,135 servers), both complete |
| Sample | 400 npm/stdio servers, seed 20260819, frame pinned by SHA-256 |
| Headline | 48.8% start rate; 58.8% of 2,766 tools unannotated; 0 fatal schema violations |
| Curation effect | +17.9 pts start rate, −17.3 pts annotation omission on a 24-server curated frame |
| Corroboration | arXiv:2608.00150 (July 31, 2026), 414 internet-facing servers, 41.6% gone within three days |
| Scope limits | npm/stdio only (30.7% of the population and falling); single probe, no retry — 48.8% is a lower bound |
| Reproducibility | Instrument, seed, per-draw outcomes and scripts released; Zenodo concept DOI 10.5281/zenodo.21347997 |