typesafe-mcp Adds Typed Agent Judgments

By Rogier Muller10.01.26
typesafe-mcp Adds Typed Agent Judgments

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.

Use typesafe-mcp when an agent needs a small branchable judgment, not when it needs another paragraph of advice. itsmostafa’s Go MCP connector exposes System One models through an evaluate tool, so a coding agent can ask a typed question, receive probabilities, and let application code inspect the result before choosing a path.

The author’s README frames the problem plainly: agents often ask an LLM whether something is urgent, who owns a request, or which route fits a task, then try to infer a decision from prose. typesafe-mcp changes that surface into an MCP server call where the caller supplies the state, the question, and the allowed answers.

Read the release as a routing primitive

The interesting part is not that the agent can call another model. The interesting part is that the result is shaped for a branch.

The README says evaluate connects clients such as Claude Code, Claude Desktop, Codex, Hermes, and Pi to TypeSafe’s Jev model, and can also run other System One models such as CLM, Laya, and d1. The examples in the README are deliberately small: “is this urgent?” and “which team owns this?” are questions where a probability distribution is more useful than a fluent explanation.

That matters in Cursor work because many agent workflows already have a decision point hidden inside a prompt. A triage helper might need to choose bug, docs, feature, or support before drafting a reply, and the useful output is not a paragraph. It is the label, the confidence, and the runner’s rule for what happens when confidence is low.

Keep the answer space boring

The connector gives you access to the model. It does not define your labels for you.

For a support repo, I would start with historical GitHub issues that already have final labels. Send each issue title and body to evaluate, ask one question, and restrict the options to the labels you already use. Then compare Jev’s probabilities with the known labels, especially on issues that were relabeled or closed as unclear.

That is also where a lot of AI code review work goes wrong. If the answer choices are vague, the model can look precise while the workflow stays ambiguous. “Urgent” needs a definition, and the branch after “urgent: 0.62” needs a threshold or a human review path.

Try it inside one Cursor workflow

For Cursor users, the safe experiment is a narrow MCP call around a decision you already make by hand. Do not wire it into merges, deploys, or customer replies on day one.

A concrete workflow: take 50 closed support issues from a repo, strip private data, and ask evaluate to classify each one into your existing issue labels. Record the returned probabilities next to the actual final label. If the model is confident on clean cases and uncertain on messy ones, you have something worth exploring.

This belongs in the Review part of our methodology, not the Build step. I want the agent to propose a structured judgment, then I want code or a person to decide whether that judgment is allowed to affect the next action.

What changed and what to test

What typesafe-mcp adds What I would test before trusting it
An MCP evaluate tool for typed questions Whether your client passes enough state without leaking private data
Probability scores over allowed answers Whether the top answer and confidence match historical labels
A fixed output shape instead of prose parsing Whether your application handles low-confidence results explicitly
Access to Jev and other System One models named by the author Whether the model you choose is available in your hosting setup
A small decision surface for agents Whether the downstream branch is reversible or reviewed

I would not start with a broad “judge this pull request” prompt. Start with a smaller decision such as “does this issue belong to docs, bug, feature, or support?” and measure misses before the agent is allowed to act on the result.

If you are tracking Jev more broadly, my earlier note on jev-ultrafast and browser agent actions is the nearby piece: same family of idea, narrower action spaces, less reliance on prose.

Further reading

Run one label experiment

Pick one existing classification task in a repo, run typesafe-mcp against known examples, and write down the confidence threshold where the result must go to a person.