Character agents

Golden datasets for character agents—when the prompt cannot install a speaker.

A strong system prompt can make a 27B base shorter and more casual. It still may have no speaker. Approved conversational gold decides retrieve versus train with a frozen 2x2. Speaker-gold sprints are typically $25–80k over 4–6 weeks after a trace review.

The prompt made it casual. It still had no speaker.

Sanitized internal 2x2 on the same human turns (no private quotes). A 760-character persona prompt suppressed assistant-isms. It did not install a person. Late frozen-replay turns can favor the adapter because humans wrote them against that branch.

Length without extra prompt

Median reply 32 vs 427 characters. The adapter was shorter on 29 of 29 turns.

Questions and lists

Questions 3 vs 27. Lists 0 vs 6.

Same 760-character prompt

Median 50 vs 71 characters. Questions 0 vs 14.

Not better reasoning

The adapter still had a blank reply, intent misses, and unsupported personal specifics. Style moved. Factual reasoning did not.

Sanitized 2x2 (no extra system prompt)

Illustrative aggregates from an internal frozen-human replay. Not a public model download. Fairness caveat: later turns in a frozen replay can flatter the adapter.

MetricBase 27BSpeaker adapter
Median reply length427 characters32 characters
Questions asked273
Lists60
Shorter replies29 / 29 turns

Retrieve exceptions. Train a speaker when the 2x2 fails.

The same approved turns can be retrieved as few-shots or used to train a speaker adapter. The 2x2 decides which. We do not sell a consumer chat app or a personal companion dump.

Retrieve first

Sparse policy, tool, and escalation failures stay outside weights as inspectable gold.

Train speaker when needed

Register, ownership, and identity after the best prompt belong in an adapter judged on frozen turns.

Serve how you trained

If every row used a long system prompt, blank-system serving is out of distribution. Blank-system identity is not a launch claim until measured.

Speaker adapters cost capability benches

Installing a speaker can regress instruction-following. That tax is why we sell evaluation with the gold—not a naked weight file.

IFEval strict

About −5.4 percentage points on a merged speaker run versus its parent (81.33% → 75.97%).

MMLU-Pro-200

About −10.5 percentage points versus stock Qwen3.8-27B (73.0% → 62.5%, CI −17.0 to −4.5).

Wrong protocol

EQ-Bench long-form punished short chat. We will not post it as a human-likeness win. A publishable claim needs a blinded live crossover.

Speaker-gold engagement artifacts

Buyers receive a retrieve-versus-train memo and governed conversational gold—not a public personal model.

2x2 protocol

Frozen human turns, prompt on/off × adapter on/off, with a fairness note on late-turn replay.

Label schema

Register, speaker, ownership, identity, intent, tools, AI-disclosure, length, lists.

Retrieve vs train memo

Exceptions stay in retrieval; speaker identity trains only when the 2x2 shows the prompt ceiling.

Capability tax note

Disclose IFEval/MMLU-class regressions if an adapter is trained.

Sprint commercial

Typical speaker-gold sprint $25–80k over 4–6 weeks after trace review.

Reusable resource

Get the speaker 2x2 rubric

Score identity, AI-disclosure, length, lists, ownership, and critical failures on frozen turns. Illustrative template—not client data.

Illustrative template—not client data. Adapt it to your policies, privacy controls, and approval process.

Who this is for

  • Founder-led conversation-product companies whose product is the conversation
  • Teams whose agent is often correct and still sounds like a generic assistant
  • Voice-agent SaaS founders who need speaker gold, not another tone dropdown

Who this is not for

  • Teams looking for a Character.AI clone or always-on inference seats from us
  • Anyone expecting a download of a personal adapter or a GGUF that failed the holdout
  • Hospital or payer outbound in the first 90 days—those pages stay for search, not this offer

Speaker-gold sprint ($25–80k, 4–6 weeks)

A — Audit

Trace slice, speaker labels (register, identity, ownership, intent, tools), and a retrieve-vs-train memo.

B — Dataset

Approved conversational gold plus, when the 2x2 fails, an adapter trained for the serving condition you will actually use.

C — Proof

Frozen 2x2, disclosed capability tax, and a rollout recommendation. No public personal weights.

Proof is the 2x2, not a file that loads

A GGUF that loads is not the training-time model. Public demo weights, if they ship, will be a separate SFW fictional adapter on stock Qwen after the same behavioral gate. Until then there is no download CTA.

Limitations

  • Sanitized aggregates are not a license to publish private chats or a personal adapter.
  • Late frozen-replay turns can favor the adapter; they are not a causal human-likeness study.
  • A merged GGUF that loads can still fail the behavioral gate versus training-time serving.
  • Blank-system identity is out of distribution if every training row used a long system prompt.
  • EQ-Bench long-form is the wrong protocol for short conversational chat.

Methodology by LTTAI LTD

This page describes less than three’s production-trace review method: reconstruct failures, capture human-approved corrections, store them outside model weights, and measure whether scoped retrieval improves a defined workflow.

Last updated: 2026-08-31

Primary references used for terminology and risk framing:

Common questions

What does character-agent gold include?

Approved turns labeled for register, speaker, ownership, identity, intent, and tools—plus a frozen 2x2 that decides retrieve versus train.

Do you fine-tune our model?

Not first. Exceptions stay outside weights and can be retrieved. We train a speaker adapter when a 2x2 shows the prompt cannot install identity.

Can we download your GGUF?

Not a personal adapter, and not a file that failed the same-chat holdout. A separate SFW demo on stock Qwen may ship after it passes the behavioral gate. The product is the gold method.

How is success measured?

Length, questions, lists, AI-disclosure, identity, and critical failures on frozen turns—plus disclosed IFEval/MMLU cost if you train. Booked proof is a trace review, not Hugging Face likes.

Book a speaker-gold trace review.

Bring a slice of production conversations. We will return a retrieve-versus-train memo—not a public companion dump.