Jev on WebMCP Benchmark
这是 idan levin 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。
这条案例记录了什么
工具选择、Agent 评估
选择工具标识;参数校验、权限和执行仍由程序负责。
社区公开项目或作者演示
原始来源:Idan Levin X post。作者自述,本站未独立复现。
idan levin
记录日期:2026-09-18。日期与身份应以原始资料为准。
怎样核对这个项目
- 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
- 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
- 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
- 工具列表和版本。
- 参数与权限检查。
- 不执行或人工复核分支。
- 是否有可观察的正确答案。
原始文字与技术细节
以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。
展开英文项目摘要与原帖
项目摘要
Benchmark evaluation testing Jev paired with Mercury 2.5 against GPT-6 Astra. Jev picked the required tool from WebMCP, while the fast LLM generated arguments, achieving 100% completion at a fraction of standard cost.
来源原文
We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using screenshot-based computer use, the model cost was 245× lower (!). We also compared Jev operating the browser with and without WebMCP. We used Browser Use’s open-source Ultrafast, with some improvements to the harness to make it more reliable across the benchmark. Jev’s browser-control accuracy on its own was not amazing - adding WebMCP nearly doubled the number of solved tasks, from 25/49 to 49/49, while reducing model cost by 18% (more on why below). The benchmark and methodology are fully open and reproducible. Full results:
原记录的限制
- Metrics are author-reported from the initial release unless independently verified.
- Requires access to the respective agent framework or runtime environment.
- Generative model execution remains external to the Jev decision step.