AI agent evaluation

Evaluate AI agents across tools, traces, and outcomes.

Judge the complete path—intent resolution, tool use, recovery, handoffs, and final outcome—against a production-grounded rubric and stable release baseline.

Final-answer scoring misses how agents actually break.

A polite wrong trajectory can still pass a surface grader. Agent evaluation has to inspect tools, intermediate decisions, and whether the path should have escalated.

Answer-only rubrics

Scoring the last message hides bad tool choices, skipped confirmations, and unsafe shortcuts.

Simulated-only suites

Scripted paths rarely match live user language, partial tool results, or messy state.

Prompt patches as evaluation

Shipping another instruction does not prove the agent recovers on the next similar request.

No approved replacement

A failed grade without a human-reviewed trajectory cannot teach runtime behavior.

Agent evaluation tied to an experience library

We reconstruct agent traces, score against a workflow rubric, capture approved corrections, and re-measure with retrieval of those experiences.

Trajectory-aware scoring

Judge intent resolution, tool correctness, recovery after errors, and policy steps—not only wording.

Human-approved paths

Reviewers rewrite the response or action sequence that should have happened, with critique and scope.

Measured before/after

Compare task success, edits, escalations, latency, and cost on one focused agent workflow.

What agent evaluation must inspect

Unlike text-only LLM grading, agent evaluation has to treat tools, state, and multi-step plans as first-class evidence—then decide which failures become reusable corrections.

Trajectory success criteria

Define pass/fail for the full path: required checks, allowed tools, stop conditions, and when escalation is the correct outcome.

Tool and recovery review

Score whether the agent chose the right tools, handled empty or conflicting results, and recovered without inventing state.

Release comparison units

Version scenarios, prompts, tools, and expected trajectories so a release can be compared to a stable agent baseline.

Trajectory evaluation method

Agent evaluation must judge the path, not only the final reply.

Failure taxonomy

Wrong tool, bad arguments, lost state, failed recovery, late escalation, policy miss, and correct-answer-but-unsafe-path.

Full trajectory rubric

Score outcome, tools, state, recovery, escalation, safety, cost, and latency with the downloadable CSV.

Unsafe-path example

An agent can produce a fluent final answer after skipping a required verification; that trajectory still fails.

Evaluation modes

Offline replay of production traces, controlled simulation, and online human review of live runs.

Release-gate workflow

Version scenarios and tools, compare against a stable baseline, and require an approved correction for recurring failures.

Reusable resource

Download the full-trajectory evaluation rubric

Score task outcome, tool selection and arguments, state, recovery, escalation, safety, cost, and latency.

Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.

Who this is for

  • Teams shipping tool-using agents who need trajectory-level evaluation
  • Support, ops, and voice agents with repeating multi-step failure modes
  • Leaders who need release gates tied to production behavior—not leaderboard scores
  • Builders comparing prompt changes vs retrieving approved experiences

Agent Improvement Sprint for evaluation

A — Audit

Trace sample, failure clusters, and an agent rubric covering trajectory, tools, and escalation.

B — Dataset

Human-corrected trajectories plus a retrieval prototype for one agent workflow.

C — Proof

Baseline vs corrected comparison with a clear rollout recommendation.

Evidence before broader rollout

We do not claim improvement before reviewing the workflow. The first sprint exists to produce measured before/after evidence on one scoped agent path.

Limitations

  • Answer-only graders miss unsafe trajectories that end with a plausible reply.
  • Simulation coverage is incomplete without production language and tool noise.
  • Release gates are weak if scenarios are not versioned with prompts and tools.
  • A failed grade without an approved replacement does not improve the next similar run.

Methodology by LTTAI LTD

This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.

Last updated: 2026-07-29

Primary references used for terminology and risk framing:

Book an AI agent evaluation that starts from your traces.

Bring a sample of production agent runs. We will map trajectory failures and propose one Agent Improvement Sprint.