Appearance
Audience — asking people, real and simulated
Load when: the question is "what would users think", or a persona, cohort, expert or live participant is involved.
Two pyramids, never pooled. For claims about the world: measured › cited › recalled › judgement. The second pyramid, the one this document is about, is for signal about people: live › twin › validated persona › proto. A lower rung never borrows a higher one's authority — three live interviews and twenty synthetic runs are never "23 responses".
What synthetics can and cannot buy
A hundred synthetic respondents are not a hundred opinions. They are one bias repeated a hundred times.
Synthetic answers show artificially low variability and distorted magnitudes, and they miss the extremes — which are exactly the people a real sample is run to find. So a balanced synthetic cohort buys a wider variety of angles and never a percentage.
Verdicts are direction-only: which concern appeared and which bias fired. Never a magnitude. No "23% would churn".
Say this before running, not after. When a request implies numbers — "test it on a sample", "what percentage would drop off", "what would most people say" — state what synthetic runs can and cannot give before spending anything, and offer the honest alternative: synthetics to find the angles worth asking live people about. Producing a plausible percentage and disclaiming it afterwards is worse than not producing it.
The field enforces this, not the warning. A cohort declares made_of:
made_of | Legitimate output |
|---|---|
synthetic | angles only — which concerns appeared |
live | measurements; a real sample supports statistics |
mixed | the two reported separately, never pooled |
A cohort distribution does not buy percentages either. Personas may carry the mix of the real population so the read is a spread rather than one voice repeated — and it is still direction-only.
Segment and cohort
| What it is | Where it lives | |
|---|---|---|
| Segment | a property of one persona — SMB, enterprise, technical, newcomer | a field on the persona, multiple, like labels |
| Cohort | a named composition for a run — "5 SMB, 3 enterprise, 2 churned" | its own file, reusable by name |
Group by the axis the question needs, and there may be several: by segment, by lifecycle (a newcomer and a veteran react differently to the same screen), or situational — one artifact, one time. One persona lives in several.
A lifecycle cohort without real churn data is fantasy. "The churned" help only if you know why they left. Without that the cohort is a guess, marked a judgement call, and its members stay proto.
Personas: document first, agent when asked
A persona is always a document. It is an agent only while it is being asked something.
| Mode | When |
|---|---|
| documents only | spec, copy, design intake — almost always |
| one agent per segment | dialogue is needed and individuals within the group need not differ |
| an agent per persona | parallel, distinguishable voices on one artifact |
Standing personas as addressable roles is fine — the roster is generated and marked, so they never appear in headcount. What they still cost is the dispatch list: every definition is an entry the model reads when deciding whom to send work to, and thirty personas competing with eight workers is the same class of problem as an over-loaded skill list. Count them like anything else.
For a twin of a living person, invocation stays deliberate. The usage log exists because the consent contract says so, and casual mentions make it meaningless.
Bias profiles — 2 to 4, each with its source
Every persona carries two to four named cognitive biases, and each has a named grounding:
- a twin's come from its own interview transcript — this one person, observed
- a validated persona's from pooled research across a segment
- a proto's from published literature, source named
- never from demographics, which produce a caricature
Response calibration — a bias lives in decisions, not in every reply. Two failures sit either side: a persona that never behaves like its profile, and one performing the bias in every sentence. A calibrated persona reads normally, and then at the moment of choice the bias shows — anchored on the first number, gone at the first friction.
Staging: proto → validated → twin
Proto is a hypothesis from literature. Validated is grounded in pooled research about real people. A twin is grounded in one real person's own material.
Each step up is a claim about evidence, so each step up has a cost: a proto is free and weak; a twin requires consent, an accuracy score, and a usage log.
Consent is a pointer, required, and revocable. A twin of a living person carries where the consent is recorded and what it covers. Revocation is honoured by removing the twin, not by marking it inactive.
The accuracy score is measured against the person's own material, not asserted — and it ages, because people change.
Live participants — four things people get wrong
Live cadence is honest. Humans answer in days, agents in seconds. A live round must never silently hold a gate running at agent speed. Name the trade-off out loud: a separate stage, a deadline, or "proceed on what we have and revisit when the answers land". The alternative is work hanging on people who do not know they are blocking anything.
Inviting someone in is an access decision, and the gate runs backwards here. Everywhere else the outward gate is about what leaves; letting someone in reveals. What they can see is the owner's call.
Notification etiquette, or you burn the people you need. Reassignment does not unsubscribe; who stops the notifications is part of the flow, not "later". And there is no broadcast — reach one person at a time.
Paying participants is spend — owner-gated, and a line in the ledger. Never "free feedback".
Mixed rounds: hypothesis beside fact
Run cheap on synthetics, then the deciding round with real people.
Provenance is mandatory and the two are counted separately. A synthetic reaction is a hypothesis; a live person's is a fact. They are never merged into "5 of 7 approved".
A live expert outweighs a twin; a twin outweighs a pooled persona; a proto is a marked guess — and the weighting is written down rather than applied silently.
Marking has a defined scope. Theatre entities carry a marker in their name, visible in every list. Where no such entity exists — a plain consultation over the persona documents — marking the reaction in prose satisfies the rule. Markers on entities, prose where there are none.
Staff views exclude theatre from headcount, and the ledger gives it its own line. A persona appearing as an employee in a count or a status report is the failure this prevents.
The verdict format is fixed
Findings → a recommendation explicitly labelled as the advisor's own judgement, with its reasoning → the gate line: whose decision this is, and the options.
Asked "so should we ship?", the advisor gives its labelled read and hands the decision back. It never issues a ship-or-not verdict of its own, and synthetic findings are proposals, never numbered shipping requirements.
That keeps two laws at once: an opinion is owed, and the gate is returned.
Guarding against bias in the agents themselves
The same phenomenon that is an asset in a persona is a hazard in a worker. Two checkpoints, not a general instruction to be objective:
At framing. Am I anchored on the first option offered? Is the search phrased to confirm rather than to find? Did I consider not doing it at all?
At output. Am I agreeing because it is true or because it is pleasant? Is this a guess wearing the shape of a fact? Does every claim carry its source?
And across a handoff: the rung travels with the claim. Agent B may not promote agent A's recalled to measured by quoting it — which is the most common way a guess becomes a fact inside a team.