MCP Verdict

Neutrality is the product

Everyone else scores the packaging. We run the server.

MCP Verdict was built to test MCP servers independently, with no vendor paying to move a score. The public site now retains eight passing filesystem evaluations as a historical snapshot.

A score built from a project's own metadata is a score its author controls.

Several products now put a number on an MCP server. Read how they get the number and the pattern is the same: GitHub stars, manifest completeness, declared permissions, dependency hygiene, maintenance recency. All of it is packaging, and most of it is filled in by the person being rated. There is already a public guide titled "how to hit 100" on one of these scores. A metric you can study for and grind toward is measuring completeness, not quality.

MCP Verdict scores a different thing. We install the server, run every tool it advertises, hit it repeatedly to see if it holds, time it, and check what it can actually reach on the host. Then a person writes a verdict that takes a side. The number comes from behavior. The verdict comes from judgment. Neither comes from the server's npm page.

A referee with a horse in the race is not a referee.

Every marketplace and directory that ranks these tools has a reason to rank them a certain way. A model vendor's store exists to make its platform look good. A directory with a connector gateway has an incentive toward the servers that route through it. None of them can be the neutral judge, not because they are dishonest, but because they cannot prove they are not. The incentive never goes away, so the trust never arrives. MCP Verdict has no such incentive to explain away. That gap is the reason it exists.

Anyone can write the word neutral. We make it checkable.

The method is published and versioned, so you can read exactly how each retained score was produced. Every public evaluation behind a score is shown in full. The result is tied to the package version and test date shown on the entry; it is not continuous monitoring or current certification.

A finished snapshot, not a roadmap.

The project is no longer expanding into new categories, community ratings, continuous monitoring, or a broader protocol. Submissions are closed and no re-test schedule is offered. Keeping the scope explicit is part of keeping the remaining claims honest.

The proof of all this is the method itself.