Training latency
Failure clusters and policies change before a fine-tune ships.
Fine-tuning vs golden datasets
A frozen 2x2 decides the first move. Gold feeds retrieve or train: exceptions stay outside weights; a speaker adapter is for identity the prompt cannot install. Disclose the capability tax before you train.
On an internal 2x2, a 760-character persona prompt made a 27B base shorter and more casual. It still had no speaker. Repeating exceptions need retrieval first; speaker identity needs approved gold, then an adapter judged on the same frozen turns.
Failure clusters and policies change before a fine-tune ships.
Individual approved answers disappear inside weights.
Removing one bad example from trained behavior is expensive.
Compliance teams cannot see which example influenced a decision.
A governed library of approved corrections from production traces—versioned, scoped, and removable—retrieved when similar requests arrive.
Measure retrieval impact on one workflow this sprint—not next train cycle.
Every record ties to source trace, reviewer, scope, and approval status.
Retrieve sparse exceptions. Train a speaker adapter when a 2x2 shows the prompt ceiling. Disclose IFEval/MMLU cost.
Choose the first move based on failure type, audit needs, and the 2x2—not vendor defaults.
Exception-heavy workflows where humans already rewrite agent output.
Register, ownership, and identity still fail after the best system prompt.
Retrieval for exceptions; training for speaker when the frozen 2x2 fails. Capability benches can fall.
Improve agents without fine-tuning firstSpeaker gold for character agents
Reusable resource
See fields for source trace, failure, approved correction, critique, scope, version, and approval status.
Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.
Cluster repeating failures and decide retrieval vs train candidates.
Human-approved experience library + retrieval prototype.
Baseline vs corrected comparison on one workflow.
We do not promise improvement before examining traces. The sprint produces evidence—and a recommendation on whether fine-tuning is still needed.
This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.
Primary references used for terminology and risk framing:
A governed library of approved corrections from production traces—versioned, scoped, and removable—retrieved when similar requests arrive.
Not first. Approved corrections for exceptions stay outside weights and can be retrieved at runtime. We train a speaker adapter when a frozen 2x2 shows the prompt cannot install identity.
A sample of production traces or conversation exports for one workflow, plus a working definition of success for that workflow.
We compare a fixed baseline against a changed agent on the same workflow using task success, edits, escalations, latency, and cost where available.
Send production traces. We will map which failures can improve through approved retrieval first.