Agent Performance · Early access

Make agent work worth the effort.

See whether your coding agent delivered a verified, accepted result. Find evidence of wasted effort and choose one change to test next.

Local by default Codex + Claude Code No Refario subscription required
Fig. 01A useful outcome needs evidence
Agent Performance / exampleIllustrative
Outcome

Verified.
Accepted.

Both checks matter.

Objective verification
Passed
Human acceptance
Yes
Observed failed tools
2
Monetary cost
Unavailable
One change to test

Clarify the failed tool’s usage instructions, then compare the next similar task.

Sample evidence, not a customer result or measured improvement.

Start where you already work
CodexClaude CodeManual capture scripts
01 / The outcome

Finishing a task is only the beginning.

A completed command does not tell you whether the work is correct or useful. Refario connects execution, verification, and your acceptance decision.

01

Verified outcomes

Keep command completion, objective verification, and human acceptance separate.

02

Effort and timing

Record task-to-first-stop and wall-clock duration when the host exposes the required boundaries.

03

Evidence of waste

Inspect observed failed tools and rework without collecting prompts or tool contents.

04

One next experiment

Get a deterministic recommendation and test one change on a comparable task.

02 / The workflow

One task. One review. One change to test.

Use real work to understand what helps. Keep task scope and configuration comparable, and include failed attempts in the comparison.

01

Capture a task

Start in the right Git repository and work on a real, bounded coding task.

02

Verify and review

Run an objective check and explicitly accept, reject, or partially accept the result.

03

Test one change

Use the recommendation to guide the next similar task. Compare outcomes before drawing conclusions.

03 / Who it’s for

For people building with coding agents.

Start with your own work. Build a habit of checking the outcome before optimising the process.

Developers
  • Review real agent tasks
  • See verification and acceptance together
  • Keep detailed evidence local
Technical founders
  • Understand repeated rework
  • Test clearer instructions
  • Build evidence before changing the stack
Agent builders
  • Compare similar tasks manually
  • Record the configuration used
  • Keep unsupported cost claims out
04 / Available now

A local plugin, with a broader direction.

Capture session evidence, run verification, record acceptance, and get a deterministic recommendation. Cross-task comparisons are currently manual. Team benchmarks and automatic cost attribution are not part of this early release.

Agent Performance / early access

Make the next task a useful experiment.

Capture the evidence, verify the result, and decide what to try next.