JevForAgents中文English
MCP 与 Skills · 编码 Agent · Agent 与模型路由

Jev Agent Toolkit

这是 Thin-Illustrator-912 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。

这条案例记录了什么

场景

MCP 与 Skills、编码 Agent、Agent 与模型路由

以可调用工具的形式返回结构化判断。

证据

社区公开项目或作者演示

原始来源:Thin-Illustrator-912 via Reddit。作者自述,本站未独立复现。

时间与作者

Thin-Illustrator-912

记录日期:2026-09-21。日期与身份应以原始资料为准。

怎样核对这个项目

  1. 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
  2. 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
  3. 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
  4. 真实调用入口和参数。
  5. 密钥及权限边界。
  6. 工具失败时的行为。
  7. 仓库状态和测试是否完整。

原始文字与技术细节

以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。

展开英文项目摘要与原帖

项目摘要

The author built a portable Agent Skill plus an optional MCP server exposing jev_evaluate. Coding agents can use Jev for bounded routing, ranking and code-review judgments while retaining their main LLM for generation and deterministic code for exact checks. The public repository contains both the Skill and MCP implementation.

来源原文

I’ve been experimenting with a question that kept coming up while working with coding agents: Does every uncertain decision really need another large LLM call? For tasks like generating code, debugging, planning or explaining something, a general-purpose LLM obviously makes sense. But coding agents also make a lot of much narrower decisions: - Which subsystem does this bug belong to? - Which of these files is most likely relevant? - Does this diff appear to widen permissions? - Is this test failure more likely a regression, a flaky test, or an environment issue? - How risky does this change look? - Does this retrieved piece of context actually help answer the question? These aren't really generation problems. They’re bounded judgment problems. So I built Jev Agent Toolkit, an open-source Agent Skill + optional MCP bridge that lets Claude Code, Codex, Cursor and other agents use TypeSafe’s Jev/System One model specifically for those decisions. GitHub: https://github.com/reiswaffel78/jev-agent-toolkit The basic architecture is intentionally simple: Coding Agent | +----------------+----------------+ | | | deterministic generative bounded work reasoning judgment | | | code/tests main LLM Jev Choice Score Noul Jev does not generate code or prose. The main coding agent still owns reasoning, implementation and orchestration. Deterministic code still owns things like parsing, arithmetic, builds, tests, thresholds and control flow. Jev is only used when the answer space is bounded and a probability distribution is useful. For example, an agent could send one batched request asking: subsystem: auth | billing | frontend | other change_risk: low | moderate | high | critical logs_secrets: probability true/false adds_network_call: probability true/false widens_permissions: probability true/false The agent can then apply deterministic thresholds in code: low confidence -> investigate further high risk -> require additional review otherwise -> continue workflow One design principle I tried to keep strict is: Don't use an AI model where deterministic code can answer the question exactly. So Jev isn't used for compilation, parsing, counting, date comparisons, applying diffs, dependency resolution, test assertions, etc. The toolkit currently includes: - Portable Agent Skill - Claude Code / Codex / Cursor support - Optional MCP server exposing `jev_evaluate` - Direct TypeSafe API / SDK usage - Choice / Score / Noul patterns - batched evaluations - confidence-gated routing - candidate ranking - retrieval relevance - code-review property checks - multi-agent orchestration guidance - Blender / Unreal / external-tool orchestration guidance It's MIT licensed. One important caveat: I do not currently have evidence that this makes coding agents objectively better, cheaper or faster overall. The architecture works and the integrations have been tested, but the interesting next step is benchmarking whether separating bounded judgment from general-purpose reasoning actually improves real coding-agent workflows. That is also why I'm posting this here. I'm particularly interested in criticism of the architecture: Where would you actually use a bounded judgment model inside a coding-agent workflow? And equally important: Where do you think adding another probabilistic model just creates unnecessary complexity?

原记录的限制

  • The author explicitly reports no benchmark showing that full coding-agent runs are better, cheaper or faster overall.
  • Parsing, compilation, tests, thresholds and other exact operations remain deterministic code paths.

继续浏览

返回中文案例目录 · 阅读相关应用场景