JevForAgents中文English
工具选择 · Agent 评估

Jev on WebMCP Benchmark

这是 idan levin 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。

这条案例记录了什么

场景

工具选择、Agent 评估

选择工具标识;参数校验、权限和执行仍由程序负责。

证据

社区公开项目或作者演示

原始来源:Idan Levin X post。作者自述,本站未独立复现。

时间与作者

idan levin

记录日期:2026-09-18。日期与身份应以原始资料为准。

怎样核对这个项目

  1. 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
  2. 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
  3. 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
  4. 工具列表和版本。
  5. 参数与权限检查。
  6. 不执行或人工复核分支。
  7. 是否有可观察的正确答案。

原始文字与技术细节

以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。

展开英文项目摘要与原帖

项目摘要

Benchmark evaluation testing Jev paired with Mercury 2.5 against GPT-6 Astra. Jev picked the required tool from WebMCP, while the fast LLM generated arguments, achieving 100% completion at a fraction of standard cost.

来源原文

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using screenshot-based computer use, the model cost was 245× lower (!). We also compared Jev operating the browser with and without WebMCP. We used Browser Use’s open-source Ultrafast, with some improvements to the harness to make it more reliable across the benchmark. Jev’s browser-control accuracy on its own was not amazing - adding WebMCP nearly doubled the number of solved tasks, from 25/49 to 49/49, while reducing model cost by 18% (more on why below). The benchmark and methodology are fully open and reproducible. Full results:

原记录的限制

  • Metrics are author-reported from the initial release unless independently verified.
  • Requires access to the respective agent framework or runtime environment.
  • Generative model execution remains external to the Jev decision step.

继续浏览

返回中文案例目录 · 阅读相关应用场景