Length without extra prompt
Median reply 32 vs 427 characters. The adapter was shorter on 29 of 29 turns.
Character agents
A strong system prompt can make a 27B base shorter and more casual. It still may have no speaker. Approved conversational gold decides retrieve versus train with a frozen 2x2. Speaker-gold sprints are typically $25–80k over 4–6 weeks after a trace review.
Sanitized internal 2x2 on the same human turns (no private quotes). A 760-character persona prompt suppressed assistant-isms. It did not install a person. Late frozen-replay turns can favor the adapter because humans wrote them against that branch.
Median reply 32 vs 427 characters. The adapter was shorter on 29 of 29 turns.
Questions 3 vs 27. Lists 0 vs 6.
Median 50 vs 71 characters. Questions 0 vs 14.
The adapter still had a blank reply, intent misses, and unsupported personal specifics. Style moved. Factual reasoning did not.
Illustrative aggregates from an internal frozen-human replay. Not a public model download. Fairness caveat: later turns in a frozen replay can flatter the adapter.
| Metric | Base 27B | Speaker adapter |
|---|---|---|
| Median reply length | 427 characters | 32 characters |
| Questions asked | 27 | 3 |
| Lists | 6 | 0 |
| Shorter replies | — | 29 / 29 turns |
The same approved turns can be retrieved as few-shots or used to train a speaker adapter. The 2x2 decides which. We do not sell a consumer chat app or a personal companion dump.
Sparse policy, tool, and escalation failures stay outside weights as inspectable gold.
Register, ownership, and identity after the best prompt belong in an adapter judged on frozen turns.
If every row used a long system prompt, blank-system serving is out of distribution. Blank-system identity is not a launch claim until measured.
Installing a speaker can regress instruction-following. That tax is why we sell evaluation with the gold—not a naked weight file.
About −5.4 percentage points on a merged speaker run versus its parent (81.33% → 75.97%).
About −10.5 percentage points versus stock Qwen3.8-27B (73.0% → 62.5%, CI −17.0 to −4.5).
EQ-Bench long-form punished short chat. We will not post it as a human-likeness win. A publishable claim needs a blinded live crossover.
Compare retrieve vs train on the fine-tuning pageKeep the anti-fine-tune loop for exceptions
Buyers receive a retrieve-versus-train memo and governed conversational gold—not a public personal model.
Frozen human turns, prompt on/off × adapter on/off, with a fairness note on late-turn replay.
Register, speaker, ownership, identity, intent, tools, AI-disclosure, length, lists.
Exceptions stay in retrieval; speaker identity trains only when the 2x2 shows the prompt ceiling.
Disclose IFEval/MMLU-class regressions if an adapter is trained.
Typical speaker-gold sprint $25–80k over 4–6 weeks after trace review.
Reusable resource
Score identity, AI-disclosure, length, lists, ownership, and critical failures on frozen turns. Illustrative template—not client data.
Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.
Trace slice, speaker labels (register, identity, ownership, intent, tools), and a retrieve-vs-train memo.
Approved conversational gold plus, when the 2x2 fails, an adapter trained for the serving condition you will actually use.
Frozen 2x2, disclosed capability tax, and a rollout recommendation. No public personal weights.
A GGUF that loads is not the training-time model. Public demo weights, if they ship, will be a separate SFW fictional adapter on stock Qwen after the same behavioral gate. Until then there is no download CTA.
This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.
Primary references used for terminology and risk framing:
Approved turns labeled for register, speaker, ownership, identity, intent, and tools—plus a frozen 2x2 that decides retrieve versus train.
Not first. Exceptions stay outside weights and can be retrieved. We train a speaker adapter when a 2x2 shows the prompt cannot install identity.
Not a personal adapter, and not a file that failed the same-chat holdout. A separate SFW demo on stock Qwen may ship after it passes the behavioral gate. The product is the gold method.
Length, questions, lists, AI-disclosure, identity, and critical failures on frozen turns—plus disclosed IFEval/MMLU cost if you train. Booked proof is a trace review, not Hugging Face likes.
Bring a slice of production conversations. We will return a retrieve-versus-train memo—not a public companion dump.