JevForAgents中文English
安全与权限门禁 · Agent 评估

Jev Benchmark

这是 Mike Moore 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。

这条案例记录了什么

场景

安全与权限门禁、Agent 评估

提供风险或支持程度的判断信号;执行权限仍归应用规则。

证据

社区公开项目或作者演示

原始来源:Community X post and GitHub repository。尚未独立核实。

时间与作者

Mike Moore

记录日期:2026-09-20。日期与身份应以原始资料为准。

怎样核对这个项目

  1. 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
  2. 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
  3. 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
  4. 高风险动作是否单独授权。
  5. 误判后的阻断和恢复。
  6. 规则是否由程序执行。
  7. 是否有可观察的正确答案。

原始文字与技术细节

以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。

展开英文项目摘要与原帖

项目摘要

The project evaluates Jev as a pre-execution risk classifier for agent tool calls, including destructive commands, credential leaks, and unauthorized database operations.

来源原文

Agree the loop: GrokBot → Jev decides → GrokBot executes. Cheap/fast routing only matters if confidence is calibrated. On our 60-case agent tool-call risk run, Jev never returned 1.000 and was wrong. https://github.com/themsquared/jev-benchmark https://webofmike.com/jev-benchmark/

原记录的限制

  • The reviewed post does not expose enough methodology to validate the headline result.
  • A benchmark score is not an authorization policy; deterministic permissions must remain in code.
  • Risk labels and attack examples need to be refreshed as agent tools change.

继续浏览

返回中文案例目录 · 阅读相关应用场景