What we measure against
We use Twin-2K-500, a public benchmark built from thousands of real people who each answered a large battery of survey questions. Because their real answers are held out, we can build a persona, ask it the same questions, and score it against what the actual person said. Nothing is graded by us or by another model. It is graded against real human responses.
The test is pre-registered: the benchmark, the exact measures, and what we expected to see were all written down before we looked at the results. That is what stops a number from being quietly reshaped after the fact.
Test one: do they answer like real people?
- They behave, they don't just talk. They move through a real product, stall, abandon, and refuse at a price, and you can replay any session and watch it happen. This is the signal we lean on hardest, because it's the one you can check with your own eyes.
- The numbers are public and pre-registered. The benchmark, the measures, and the bar were written down before we looked. They hold up when a skeptic pokes at them, not only when we present them.
Test two: do they catch real breakage?
Answering surveys well is not the job. The job is finding what's broken in your product. So we seeded five realistic bugs into a real open-source app, locked the answer key cryptographically before the first run, and kept the roles separate: one agent seeded, others drove the product, another graded. Nobody grades their own work.
- Across three full rounds, the testers caught two of the five bugs every single round, and four of five in the best round. We report the stable number first and the best round second, not the other way around.
- They also surfaced two real bugs nobody planted. The app is a widely used open-source project, and the testers found genuine defects in it along the way.
- About one flag in five was noise. That false-positive rate is exactly why every finding you receive is labeled before you see it, so noise can't masquerade as fact.
Where it's currently weaker
It reads the group far better than the individual. Most of the 83.7% comes from getting people-in-general right: a simulated customer built from one specific person's full history only beats a generic one by about two points on this benchmark. "How will customers like this react to my pricing?" is a question it answers well. "What will this one specific person do?" is still the frontier. An earlier, smaller pilot suggested a prompt change had widened that gap to about five points; the larger confirming run did not reproduce that, so two points is the number we stand behind, and individuation stays on the weak list until the data says otherwise.
It's also stronger on actions than on feelings: what people do and where they drop off is solid ground, while the finer emotional read is softer. We treat those softer signals as leads to confirm, not settled facts.
How every finding is labeled
So you always know how much weight a finding can carry, each one is marked:
EXECUTED observed behavior you can reproduce yourself.
INFERRED / TESTIMONY a claim worth checking, not a fact to bank.
We keep testing in the open
This isn't a one-time result. Next on the list: a larger benchmark run, and re-scoring on just the questions where telling one person from another is even possible, to measure individuation where it can actually show up. This page gets the new numbers when they land, whichever direction they move. And when the evidence says an idea shouldn't be built, that is the verdict you get, with the reasons. A clear no in week one is the cheapest thing we sell.