Agent Pattern

Agent Evaluation

Judge completed runs, traces, outputs, and patches with typed scores and confidence.

The decision

Did the agent complete the task correctly and safely?

Why Jev fits

Noul + Score + Choice compose into repeatable evaluation workflows.

Common implementations

  • Trace review
  • Patch review
  • Quality scoring
  • Human-review triage

Builds using this pattern