Jev Benchmark
这是 Mike Moore 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。
这条案例记录了什么
场景
安全与权限门禁、Agent 评估
提供风险或支持程度的判断信号;执行权限仍归应用规则。
证据
社区公开项目或作者演示
原始来源:Community X post and GitHub repository。尚未独立核实。
时间与作者
Mike Moore
记录日期:2026-09-20。日期与身份应以原始资料为准。
怎样核对这个项目
- 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
- 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
- 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
- 高风险动作是否单独授权。
- 误判后的阻断和恢复。
- 规则是否由程序执行。
- 是否有可观察的正确答案。
原始文字与技术细节
以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。
展开英文项目摘要与原帖
项目摘要
The project evaluates Jev as a pre-execution risk classifier for agent tool calls, including destructive commands, credential leaks, and unauthorized database operations.
来源原文
Agree the loop: GrokBot → Jev decides → GrokBot executes. Cheap/fast routing only matters if confidence is calibrated. On our 60-case agent tool-call risk run, Jev never returned 1.000 and was wrong. https://github.com/themsquared/jev-benchmark https://webofmike.com/jev-benchmark/
原记录的限制
- The reviewed post does not expose enough methodology to validate the headline result.
- A benchmark score is not an authorization policy; deterministic permissions must remain in code.
- Risk labels and attack examples need to be refreshed as agent tools change.