What we measure against
We use Twin-2K-500, a public benchmark built from thousands of real people who each answered a large battery of survey questions. Because their real answers are held out, we can build a persona, ask it the same questions, and score it against what the actual person said. Nothing is graded by us or by another model. It is graded against real human responses.
The test is pre-registered: the benchmark, the exact measures, and what we expected to see were all written down before we looked at the results. That is what stops a number from being quietly reshaped after the fact.
Test one: do they answer like real people?
- They behave, they don't just talk. They move through a real product, stall, abandon, and refuse at a price, and you can replay any session and watch it happen. This is the signal we lean on hardest, because it's the one you can check with your own eyes.
- The numbers are public and pre-registered. The benchmark, the measures, and the bar were written down before we looked. They hold up when a skeptic pokes at them, not only when we present them.
Test two: do they catch real breakage?
Answering surveys well is not the job. The job is finding what's broken in your product. So we seeded five realistic bugs into a real open-source app, locked the answer key cryptographically before the first run, and kept the roles separate: one agent seeded, others drove the product, another graded. Nobody grades their own work.
- Across three full rounds, the testers caught two of the five bugs every single round, and four of five in the best round. We report the stable number first and the best round second, not the other way around.
- They also surfaced two real bugs nobody planted. The app is a widely used open-source project, and the testers found genuine defects in it along the way.
- About one flag in five was noise. That false-positive rate is exactly why every finding you receive is labeled before you see it, so noise can't masquerade as fact.
Can they tell two audiences apart?
- Every number here is a percentage of answers. Each answer is compared with what a real person actually said. A multiple choice is right or wrong, a slider gets credit for landing close, and the score is the percentage that matched. It is the same scale as the 83.7% above.
- The two audiences are the two main United States political parties, because that is where this benchmark records the clearest disagreement between large groups of real people. On the questions where those two audiences differ, one persona came back +9.9% and the other +11.1%, against a real gap of +16.4%. Across all 126 questions the same three figures are +3.7%, +2.5% and +4.5%, smaller because most questions are ones both audiences answer the same way, so there is no gap to reproduce.
- Single pass, so there are no confidence intervals on these three numbers yet. The gaps are wide enough to be real, but treat the exact decimals as provisional. The benchmark carries only broad demographic groupings, so it cannot tell us whether a persona for founders differs correctly from a persona for therapists.
- This is a comparison, not a forecast, and it is not a claim about accuracy. It tells you two audiences come back different, which is what you need when you are choosing between them, or checking whether a page lands differently with one. We separately tested whether a persona built for a group answers a question about that group better than simply predicting the overall average, and across twenty groups none of them did.
Where it's currently weaker
It reads the group far better than the individual. Most of the 83.7% comes from getting people-in-general right, and building a simulated customer from one specific person's full history adds little on top of that. "How will customers like this react to my pricing?" is a question it answers well. "What will this one specific person do?" is still the frontier, and individuation stays on the weak list until the data says otherwise.
It's also stronger on actions than on feelings: what people do and where they drop off is solid ground, while the finer emotional read is softer. We treat those softer signals as leads to confirm, not settled facts.
How every finding is labeled
So you always know how much weight a finding can carry, each one is marked:
EXECUTED observed behavior you can reproduce yourself.
INFERRED / TESTIMONY a claim worth checking, not a fact to bank.
We keep testing in the open
This isn't a one-time result. Next on the list: a larger benchmark run, and re-scoring on just the questions where telling one person from another is even possible, to measure individuation where it can actually show up. This page gets the new numbers when they land, whichever direction they move. And when the evidence says an idea shouldn't be built, that is the verdict you get, with the reasons. A clear no in week one is the cheapest thing we sell.