JevForAgents中文English
浏览器 Agent · Agent 评估

Stencil QA Agent

这是 Dan Leshem 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。

这条案例记录了什么

场景

浏览器 Agent、Agent 评估

从当前页面允许的候选动作中选择下一步;浏览器工具负责执行。

证据

社区公开项目或作者演示

原始来源:Dan Leshem via X。作者自述,本站未独立复现。

时间与作者

Dan Leshem

记录日期:2026-09-24。日期与身份应以原始资料为准。

原始演示视频

视频来自此案例记录的原始媒体;播放内容和作者声明不等于本站复现。

怎样核对这个项目

  1. 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
  2. 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
  3. 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
  4. 页面状态与候选动作是否同步。
  5. 失败、停止与人工接管条件。
  6. 权限和不可逆操作的独立审批。
  7. 是否有可观察的正确答案。

原始文字与技术细节

以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。

展开英文项目摘要与原帖

项目摘要

Dan Leshem reports rebuilding Stencil's QA agent around Jev: a frontier model plans the test, Jev chooses browser actions, and vision checks the result. The shown run signs in, opens an app, triggers a paywall, checks a spacing fix, and passes five checkpoints in 28.3 seconds with a screenshot and trace for each step.

来源原文

we rebuilt a QA agent around @typesafeai’s Jev at Stencil. a frontier model plans the test. Jev chooses browser actions. vision handles the visual checks. here it signs in, opens an app, triggers the paywall, and checks a spacing fix. 5 checkpoints passed in 28.3 seconds. every checkpoint has a screenshot, every decision has a trace. i can see Jev driving decisions across our software factory: which issues to prioritize, what to test, when to retry, and when a change needs human review.

原记录的限制

  • The post does not link a public repository or a runnable reproduction.
  • The 28.3-second result is an author-reported run, not a JevForAgents measurement.

继续浏览

返回中文案例目录 · 阅读相关应用场景