Agent Performance / early accessRefario / 2026

Understand the outcome behind the activity.

A local plugin for reviewing coding-agent work in Codex and Claude Code. Capture what happened, verify the result, and decide what to test next.

Agent Performance / exampleIllustrative
Outcome

Verified.
Accepted.

Both checks matter.

Objective verification
Passed
Human acceptance
Yes
Observed failed tools
2
Monetary cost
Unavailable
One change to test

Clarify the failed tool’s usage instructions, then compare the next similar task.

Sample evidence, not a customer result or measured improvement.

01 / Evidence

Measure the work you can verify.

Refario records configuration when available, duration, Git change attribution, sanitized tool evidence, verification, and your acceptance decision.

01

Verified outcomes

Keep command completion, objective verification, and human acceptance separate.

02

Effort and timing

Record task-to-first-stop and wall-clock duration when the host exposes the required boundaries.

03

Evidence of waste

Inspect observed failed tools and rework without collecting prompts or tool contents.

04

One next experiment

Get a deterministic recommendation and test one change on a comparable task.

02 / Practice

A review you can act on.

Recommendations follow deterministic rules. They identify a next experiment; they do not claim an improvement has already happened.

01

Capture a task

Start in the right Git repository and work on a real, bounded coding task.

02

Verify and review

Run an objective check and explicitly accept, reject, or partially accept the result.

03

Test one change

Use the recommendation to guide the next similar task. Compare outcomes before drawing conclusions.

03 / Boundaries

Know what the evidence can tell you.

Missing host data stays missing. Monetary cost is unavailable without authoritative usage evidence. Active duration is task start to first stop, not CPU time or a sum of later rework.

Early-access capture limits

Use one measured task per repository at a time, or isolate its capture directory. Explicit identity for concurrent tasks is still being developed. Comparisons between sessions are manual.

04 / Agent Economics

Connect effort to economics over time.

Agent Performance is the starting point. Refario’s existing platform also supports attributed AI cost and customer margin; broader outcome economics remain a product direction.

Explore customer economics
Agent Performance / early access

Make the next task a useful experiment.

Capture the evidence, verify the result, and decide what to try next.