Stencil QA Agent
这是 Dan Leshem 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。
这条案例记录了什么
浏览器 Agent、Agent 评估
从当前页面允许的候选动作中选择下一步;浏览器工具负责执行。
社区公开项目或作者演示
原始来源:Dan Leshem via X。作者自述,本站未独立复现。
Dan Leshem
记录日期:2026-09-24。日期与身份应以原始资料为准。
原始演示视频
视频来自此案例记录的原始媒体;播放内容和作者声明不等于本站复现。
怎样核对这个项目
- 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
- 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
- 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
- 页面状态与候选动作是否同步。
- 失败、停止与人工接管条件。
- 权限和不可逆操作的独立审批。
- 是否有可观察的正确答案。
原始文字与技术细节
以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。
展开英文项目摘要与原帖
项目摘要
Dan Leshem reports rebuilding Stencil's QA agent around Jev: a frontier model plans the test, Jev chooses browser actions, and vision checks the result. The shown run signs in, opens an app, triggers a paywall, checks a spacing fix, and passes five checkpoints in 28.3 seconds with a screenshot and trace for each step.
来源原文
we rebuilt a QA agent around @typesafeai’s Jev at Stencil. a frontier model plans the test. Jev chooses browser actions. vision handles the visual checks. here it signs in, opens an app, triggers the paywall, and checks a spacing fix. 5 checkpoints passed in 28.3 seconds. every checkpoint has a screenshot, every decision has a trace. i can see Jev driving decisions across our software factory: which issues to prioritize, what to test, when to retry, and when a change needs human review.
原记录的限制
- The post does not link a public repository or a runnable reproduction.
- The 28.3-second result is an author-reported run, not a JevForAgents measurement.