Answer-only rubrics
Scoring the last message hides bad tool choices, skipped confirmations, and unsafe shortcuts.
AI agent evaluation
Judge the complete path—intent resolution, tool use, recovery, handoffs, and final outcome—against a production-grounded rubric and stable release baseline.
A polite wrong trajectory can still pass a surface grader. Agent evaluation has to inspect tools, intermediate decisions, and whether the path should have escalated.
Scoring the last message hides bad tool choices, skipped confirmations, and unsafe shortcuts.
Scripted paths rarely match live user language, partial tool results, or messy state.
Shipping another instruction does not prove the agent recovers on the next similar request.
A failed grade without a human-reviewed trajectory cannot teach runtime behavior.
We reconstruct agent traces, score against a workflow rubric, capture approved corrections, and re-measure with retrieval of those experiences.
Judge intent resolution, tool correctness, recovery after errors, and policy steps—not only wording.
Reviewers rewrite the response or action sequence that should have happened, with critique and scope.
Compare task success, edits, escalations, latency, and cost on one focused agent workflow.
Unlike text-only LLM grading, agent evaluation has to treat tools, state, and multi-step plans as first-class evidence—then decide which failures become reusable corrections.
Define pass/fail for the full path: required checks, allowed tools, stop conditions, and when escalation is the correct outcome.
Score whether the agent chose the right tools, handled empty or conflicting results, and recovered without inventing state.
Version scenarios, prompts, tools, and expected trajectories so a release can be compared to a stable agent baseline.
Connect agent scoring to broader LLM evaluation servicesStore approved trajectories as golden datasets for AI agents
Agent evaluation must judge the path, not only the final reply.
Wrong tool, bad arguments, lost state, failed recovery, late escalation, policy miss, and correct-answer-but-unsafe-path.
Score outcome, tools, state, recovery, escalation, safety, cost, and latency with the downloadable CSV.
An agent can produce a fluent final answer after skipping a required verification; that trajectory still fails.
Offline replay of production traces, controlled simulation, and online human review of live runs.
Version scenarios and tools, compare against a stable baseline, and require an approved correction for recurring failures.
Reusable resource
Score task outcome, tool selection and arguments, state, recovery, escalation, safety, cost, and latency.
Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.
Trace sample, failure clusters, and an agent rubric covering trajectory, tools, and escalation.
Human-corrected trajectories plus a retrieval prototype for one agent workflow.
Baseline vs corrected comparison with a clear rollout recommendation.
We do not claim improvement before reviewing the workflow. The first sprint exists to produce measured before/after evidence on one scoped agent path.
This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.
Primary references used for terminology and risk framing:
Bring a sample of production agent runs. We will map trajectory failures and propose one Agent Improvement Sprint.