Computer Use Without Screenshots
这是 Milind S 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。
这条案例记录了什么
浏览器 Agent
从当前页面允许的候选动作中选择下一步;浏览器工具负责执行。
社区公开项目或作者演示
原始来源:Milindlabs X post。尚未独立核实。
Milind S
记录日期:2026-09-17。日期与身份应以原始资料为准。
原始演示视频
视频来自此案例记录的原始媒体;播放内容和作者声明不等于本站复现。
怎样核对这个项目
- 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
- 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
- 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
- 页面状态与候选动作是否同步。
- 失败、停止与人工接管条件。
- 权限和不可逆操作的独立审批。
原始文字与技术细节
以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。
展开英文项目摘要与原帖
项目摘要
An operating system control loop that replaces heavy screenshot uploads with local edge segmentation. CoreML detects bounding boxes, local OCR reads labels, and Jev chooses the next click from text candidates alone.
来源原文
Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all Jev gets. It returns a probability across those elements and tells me the best one to click. Then it clicks, re-runs detection, and decides again. In a loop until the goal is done. ~90ms per decision. Faster than any LLM computer use I've tried. Blazing fast computer use, without any latency @typesafeai is building something really interesting
原记录的限制
- Metrics are author-reported from the initial release unless independently verified.
- Requires access to the respective agent framework or runtime environment.
- Generative model execution remains external to the Jev decision step.