jev-mcp Packages Agent Judgments as Tools

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
Use jev-mcp as a cheap second pass, not as proof. In jkudish's jev-mcp, the project packages TypeSafe's Jev model as MCP tools for the small judgments an agent often skips: checking a claim against a passage, screening incoming context, ranking candidates, classifying items, and reviewing whether a patch matches the task. For Claude Code MCP users, it looks less like a chat helper and more like an AI code review instrument you can place before a generated report or patch is treated as done.
As of 2 October 2026, the repository had 490 GitHub stars, an MIT license, mainly JavaScript code, and a last push on that date. The README describes the judgments as typed results with probabilities, plus confidence scores for most tools, returning in roughly 150 to 500 ms for a fraction of a cent. Those are author-reported numbers, not a benchmark I ran, but they explain the design: make the check cheap enough that an agent can call it on many small decisions instead of saving verification for the end.
Read the tool surface before the promise
The interesting part is not that jev-mcp exposes one verifier. It exposes twelve judgment tools, each aimed at a narrow decision.
jev_verify checks whether supplied evidence supports a claim. jev_screen judges content before it enters context. jev_noul returns a calibrated probability for a stated proposition. jev_find and jev_rerank choose or sort candidates by meaning. jev_classify, jev_decide, jev_compare, jev_extract, and jev_audit cover bounded classification, comparison, extraction, and source checking. jev_review and jev_gate bring the same idea to proposed diffs and completion claims.
MCP matters here because the judgments are exposed as callable tools rather than hidden inside a custom app. An agent can retrieve evidence, call a specific judgment tool, and then decide whether the draft is ready to continue. That is close to the direction in typesafe-mcp Adds Typed Agent Judgments, but jev-mcp is more explicit about the individual checks it wants an agent to make.
Put verification between evidence and prose
The cleanest place to try jev-mcp is after retrieval and before writing. That is where the agent has enough source material to be checked, but before a weak sentence becomes polished prose.
Take a release-note workflow. Claude Code drafts, "The package now requires Node 22," after reading a diff. Before that sentence enters the changelog, pass the claim and the relevant package.json engines passage to jev_verify. If the result is weak or low confidence, the next action is not to rewrite the claim more fluently. The next action is to fetch better evidence or remove the sentence.
This fits the Review step in our methodology: do not ask the model whether the whole answer feels right, ask a smaller tool whether one claim is supported by one piece of evidence. That is a better boundary for agentic coding work because it keeps the source passage in charge.
Keep the source passage in charge
A typed probability is not an audit trail. jev-mcp can only judge the claim against what you provide, so a missing source, stale documentation page, or irrelevant snippet still breaks the workflow.
That limit is easy to miss because the tool surface is broad. jev_rerank can sort candidates, but it cannot make a bad candidate set complete. jev_extract can pull values with regex plus judgment, but the value still needs the right source. jev_gate can review a patch and completion claims in one call, but it should not become a way to avoid reading the diff.
The right mental model is a fast mechanical check. It catches the kinds of unsupported claims and mismatched candidates that agents produce when they move too quickly. It does not replace source selection, code review, or a human decision on risk.
Try it on one small repo
Start with a low-risk repo where the evidence is easy to inspect, such as documentation, changelog generation, or dependency update notes. Avoid production write paths for the first experiment. You want to learn whether the judgments are useful, not whether your permissions are too loose.
Connect jev-mcp using the repository's current README instructions and your MCP client of choice. Give the agent a task that already has source material, for example: summarize the last five merged changes, list changed runtime requirements, or draft a short migration note from a diff and docs folder.
Then add one rule to the workflow: before the final answer, each factual claim with a source passage must be checked with jev_verify. Do not tune ten tools at once. If claim verification does not improve the output on a small task, the broader tool surface will be harder to judge.
What changed and what to test
Copy this as the first experiment note. It is deliberately small.
| jev-mcp tool | Use it when | Test with | Watch for |
|---|---|---|---|
jev_verify |
A generated claim cites a passage | One changelog sentence and one source snippet | Low confidence on vague or incomplete evidence |
jev_screen |
New content is about to enter context | Incoming issue comments, docs pages, or pasted logs | Useful rejection criteria, not just generic quality scoring |
jev_find or jev_rerank |
The agent has several plausible files or passages | Candidate docs pages for one migration question | Whether the right passage is in the candidate set at all |
jev_extract plus jev_audit |
The agent pulls structured values from text | Version numbers, flags, dates, or config values | Extracted values that are correct but unsupported by the cited source |
jev_review or jev_gate |
A patch is about to be called done | A small diff with an explicit completion claim | False comfort on tasks that require domain review |
Keep the output from failed checks. The failures are more useful than the successes because they show whether jev-mcp is catching mistakes your normal review flow misses.
Further reading
Next step
Pick one generated changelog or research note and add exactly one jev_verify call between evidence retrieval and final prose. If it catches unsupported claims, expand from there to ranking or patch review.