Daydream ← Back to Daydream

Methodology

How we validate the simulated customers.

We run two pre-registered tests against reality. On a public benchmark of real people, our simulated customers answer at 83.7% of the consistency those people show with themselves, confirmed in a second, larger run. And in a seeded-bug study on a real product, they caught planted breakage round after round, plus two real bugs nobody planted. They read a group better than any single individual, and what people do better than how they feel. The rest of this page is the numbers, including the weak ones. Last updated July 11, 2026.

What we measure against

We use Twin-2K-500, a public benchmark built from thousands of real people who each answered a large battery of survey questions. Because their real answers are held out, we can build a persona, ask it the same questions, and score it against what the actual person said. Nothing is graded by us or by another model. It is graded against real human responses.

The test is pre-registered: the benchmark, the exact measures, and what we expected to see were all written down before we looked at the results. That is what stops a number from being quietly reshaped after the fact.

Test one: do they answer like real people?

83.7% It answers like a real person. Our simulated customers reproduce real people's held-out answers at 83.7% of the consistency those people show answering for themselves. First measured in a small pilot, then confirmed in a second pre-registered run of 50 people with every item answered. Underneath the ratio: they match 60% of held-out answers outright, against a 72% human ceiling. The bar we registered in advance was 70% of human consistency, and this clears it.

Test two: do they catch real breakage?

Answering surveys well is not the job. The job is finding what's broken in your product. So we seeded five realistic bugs into a real open-source app, locked the answer key cryptographically before the first run, and kept the roles separate: one agent seeded, others drove the product, another graded. Nobody grades their own work.

Where it's currently weaker

It reads the group far better than the individual. Most of the 83.7% comes from getting people-in-general right: a simulated customer built from one specific person's full history only beats a generic one by about two points on this benchmark. "How will customers like this react to my pricing?" is a question it answers well. "What will this one specific person do?" is still the frontier. An earlier, smaller pilot suggested a prompt change had widened that gap to about five points; the larger confirming run did not reproduce that, so two points is the number we stand behind, and individuation stays on the weak list until the data says otherwise.

It's also stronger on actions than on feelings: what people do and where they drop off is solid ground, while the finer emotional read is softer. We treat those softer signals as leads to confirm, not settled facts.

How every finding is labeled

So you always know how much weight a finding can carry, each one is marked:

EXECUTED observed behavior you can reproduce yourself.

INFERRED / TESTIMONY a claim worth checking, not a fact to bank.

We keep testing in the open

This isn't a one-time result. Next on the list: a larger benchmark run, and re-scoring on just the questions where telling one person from another is even possible, to measure individuation where it can actually show up. This page gets the new numbers when they land, whichever direction they move. And when the evidence says an idea shouldn't be built, that is the verdict you get, with the reasons. A clear no in week one is the cheapest thing we sell.