The observed result matched the expectation for this test.
About the data
Methodology & limitations
How results enter the Observatory and how to interpret them.
MCP Failure Observatory publishes structured results from MCP Failure Lab. It is an evidence index, not a certification program or a complete measure of implementation quality.
How records are produced
- A defined failure scenario is run against a named implementation version and transport.
- The observed behavior, timing, and source report are retained as evidence.
- A Test Run records the outcome for that specific combination.
- Material observations are reviewed before becoming verified Findings.
- When appropriate, the Finding links to a new upstream issue or to evidence added to an existing issue.
Test status meanings
The observed result did not match the expectation.
The result is recorded, but a person still needs to confirm its meaning.
The implementation or transport does not support the capability required by the scenario.
No completed test result was produced.
The scenario does not apply to this implementation or transport.
Upstream reporting meanings
The project opened a new upstream issue for the finding.
A relevant issue already existed, so the project added its evidence there.
No upstream report is linked yet.
The linked upstream problem is recorded as resolved.
The behavior is documented but does not represent an upstream defect to report.
Sources and review
Test Runs and evidence originate in the public MCP Failure Lab reports. Implementation repository links identify the codebase tested. Findings retain their supporting evidence and, where available, links to upstream issues and comments. A “verified” Finding means a person reviewed the recorded evidence; it does not mean the upstream maintainers accepted the conclusion.
Limitations
- Results apply only to the recorded version, environment, transport, and scenario.
- A missing matrix cell means there is no recorded result, not that an implementation failed.
- New releases may behave differently from the version shown.
- Timing results can vary by machine and should not be treated as general performance benchmarks.
- The dataset is selective and does not cover every MCP feature, implementation, or deployment environment.