Score-only pipelines
Failed grades without approved alternatives.
Automated eval vs human review
Automated graders scale volume. Human review captures the approved correction your agent should follow next time. Production agent improvement needs both—connected.
A red score tells you something failed. It rarely stores the trajectory or response that should replace the failure in production.
Failed grades without approved alternatives.
Policy, tool judgment, and escalation paths need humans.
Lab scenarios miss production language and tool state.
Review findings stay in dashboards—not in the agent.
Automated graders filter volume; human reviewers approve replacements; approved records feed retrieval and re-measurement.
Run rubrics on high-volume traces and regressions.
Experts sign off on corrections, scope, and retirement.
Bring new production failures back into eval and retrieval sets.
Use machines for throughput. Use humans for judgment, policy, and approval.
Regression detection, rubric scoring, duplicate failure clustering.
Ambiguous trajectories, compliance paths, tool escalation decisions.
Only approved corrections enter the experience library.
Reusable resource
CSV template for task outcome, tools, state, recovery, escalation, safety, cost, and latency.
Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.
Rubric design tied to production failure clusters.
Grader workflow + human approval + retrieval prototype.
Baseline vs corrected comparison on one workflow.
We measure whether approved corrections retrieved at runtime improve the same workflow the rubric scores.
This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.
Primary references used for terminology and risk framing:
Automated graders filter volume; human reviewers approve replacements; approved records feed retrieval and re-measurement.
Not first. Approved corrections for exceptions stay outside weights and can be retrieved at runtime. We train a speaker adapter when a frozen 2x2 shows the prompt cannot install identity.
A sample of production traces or conversation exports for one workflow, plus a working definition of success for that workflow.
We compare a fixed baseline against a changed agent on the same workflow using task success, edits, escalations, latency, and cost where available.
Send sample failures. We will map automated vs human review split for one workflow.