JevForAgents中文English
编码 Agent · 工具选择 · Agent 与模型路由

CasaJev

这是 Blocboiven 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。

这条案例记录了什么

场景

编码 Agent、工具选择、Agent 与模型路由

对模型选择、片段可见度或工具风险提供有界判断。

证据

本站直接测试

原始来源:Blocboiven via Reddit。作者自述,本站未独立复现。

时间与作者

Blocboiven

记录日期:2026-09-23。日期与身份应以原始资料为准。

怎样核对这个项目

  1. 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
  2. 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
  3. 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
  4. 仓库状态和测试是否完整。
  5. 敏感命令的程序门禁。
  6. 完整任务的费用与结果。
  7. 工具列表和版本。

原始文字与技术细节

以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。

展开英文项目摘要与原帖

项目摘要

CasaJev asks Jev for a bounded next step. When no registered tool fits, a coding agent proposes and implements a narrow tool contract; CasaJev checks it in an isolated worker, versions and registers it, then resumes the original task so later tasks can reuse the tool.

来源原文

I'm the author of CasaJev, an experimental local harness for TypeSafe's Jev. The problem I'm exploring: when an agent solves a task by building a small tool, can the next similar task reuse that tool instead of starting from scratch? The loop asks Jev for a bounded next step. If no registered tool fits, a coding agent proposes and implements a narrow tool contract. CasaJev runs checks in an isolated worker, registers the versioned tool, and resumes the original task. Later tasks can select the registered tool directly. The repo includes the chat/desktop UI, task persistence, tool registry, tests, and benchmark inputs/results. In one live example, building a new tool took 29.29 s; reusing it on a different input took 2.07 s. In a separate six-task synthetic pilot with existing tools, the Jev path had a 1.80 s median versus 13.01 s for Sol and 16.84 s for Terra through the measured integrations. These are small local measurements, not evidence that CasaJev is generally faster than GPT. Graph hints showed no measured speed benefit in that pilot. It's Apache-2.0 and still a prototype. macOS has been exercised end to end; targeted Windows browser/desktop tests pass, but full Windows validation is pending because Docker was unavailable on the test machine. Linux end-to-end validation is also open. Repo: https://github.com/8endit/CasaJev I'd welcome contributors or critical feedback on independent tool verification, larger-scale tool retrieval, and a benchmark with real recurring tasks. What failure case would you test first?

原记录的限制

  • The post explicitly describes the project as a prototype and says the measurements are small local or synthetic tests.
  • Linux end-to-end validation remains open; full Windows validation was pending because Docker was unavailable on the test machine.
  • The author reports that graph hints showed no measured speed benefit in the six-task pilot.
  • Tool-contract or API semantic drift can still escape detection when descriptors and output schemas do not change.

继续浏览

返回中文案例目录 · 阅读相关应用场景