·Methodology·Minds Team

How to Read Synthetic Audience Accuracy Claims

Nine synthetic audience vendors publish accuracy figures. They use different metrics, different baselines and different samples, so the numbers cannot be ranked against each other. Eight questions separate a checkable claim from an unfalsifiable one.

Synthetic research vendors publish accuracy figures between 85% and 98%. Read quickly, the field looks settled and the differences look small.

They are not comparable. The percentages measure different quantities against different baselines on different tasks, and several do not measure accuracy at all. This article sets out what we found reading the published material of nine vendors in September 2026, including our own, and the eight questions that separate a claim you can check from one you cannot.

We have an interest here. Minds sells synthetic audiences. So the standard below is applied to us on the same terms, and the place where we come off worst is stated as plainly as everyone else's.

The eight questions

A published accuracy claim is checkable when it answers these:

  1. Against what external standard? A named, independent dataset or published study, not an internal sample.
  2. Against what ceiling? Humans do not reproduce their own answers perfectly. Their self-consistency is the realistic maximum.
  3. Against what floor? What the same metric scores with no model behind it.
  4. What does the metric mean? "Accuracy" is not self-defining.
  5. Compared with what baseline? A naive prompt to the same model is the comparison that matters, because that is the free alternative.
  6. Reviewed by whom? Peer review or a preprint invites correction.
  7. Where does it fail? A method with no published failure mode has not been tested hard enough to find one.
  8. Can anyone check it? Method, sample and data available for inspection.

What the field publishes

Compiled from vendors' own public pages and papers, September 2026.

VendorHeadline figureWhat it measures
Simile85% of human self-retest reliabilityReproduction of source individuals' answers, against the human noise floor
Artificial Societies86% distribution accuracyDistribution match, against a stated 91% human ceiling
Articos86% of expert-identified themesTheme recovery vs published expert reports
Evelance89.78% thematic overlapTheme match, 7 personas vs 23 real users
Synthetic Users85–92% "Synthetic Organic Parity"Similarity of synthetic to real interviews
Aaru0.90 median correlationCorrelation on one client study
Evidenza88% accuracy, 0.81 correlationMetric not defined publicly
Fairgen98% sanity pass rateAn internal quality gate, not accuracy
MindsSpearman 0.85, +4% scale biasRank correlation on 54 ads vs a human panel

Two observations follow, and neither is about which product is better.

The highest number is not an accuracy figure. Fairgen's 98% is the share of generated profiles passing their own six-dimension quality gate. It says nothing about whether the output matches human data, and Fairgen does not claim it does. Anyone sorting this table by percentage would rank them first on a measure they never entered.

Only two figures are commensurable. Simile's 85% and Artificial Societies' 86% are both expressed as a share of what humans themselves achieve. They can be compared. Nothing else in the column can be compared to anything, including to each other.

Who publishes failures

This is where the field separates most sharply.

Fairgen states the studies their method is unsuitable for: brand tracking, segmentation, conjoint. Their guidance is to "pressure-test before going to field, not instead". Synthetic Users state that synthetic respondents "without calibration are individually believable, but collectively wrong". Articos and Evelance both say their output is directional and belongs alongside human research rather than in place of it.

Aaru, Evidenza and Artificial Societies publish no failure conditions we could find.

Where we stand, including the gap

Our own studies use four-condition ablations, hold out real answers, seal candidate selections before targets are opened, and report the chance floor next to the result. On one GSS item set a flat distribution with no model behind it already scores 83.71%; our 92.47% closes 54% of the remaining distance. The floor belongs next to the percentage, and it is usually missing.

We publish results that did not work. A registered profile lift of 0.17 points had a 95% interval from -1.00 to +1.31, crossed zero, and failed its promotion criterion. On short everyday prompts in a blinded comparison, our audiences won 13.8% against 30.5% for generic prompts and 34.0% for demographic prompts. That is a loss, and it is published.

What we do not have is peer review. Simile's fine-tuning result appeared at EMNLP 2025 with an open dataset of 2.9 million responses. Artificial Societies has a group-behaviour result in the British Journal of Psychology. We have neither. Our validation work is published on our own site and has not been through external review, and until it has, a reader is right to weigh it differently.

Our partnership with SINUS-Institut is sometimes read as validation. It is not. It is provenance: the milieu model our personas are grounded in comes from decades of their research, which is a statement about inputs rather than about accuracy. Their managing director puts it plainly: the Milieu-Minds "replace neither social research nor our institute's expertise".

How to check any vendor in ten minutes

Ask for the metric's definition, the sample size, the baseline and the failure conditions. Ask what the same measure scores with no model behind it. If a vendor cannot produce those, the percentage is marketing, whoever published it and however favourable it looks.

That test is not comfortable for us either. It is simply the only one that tells a buyer anything.

Sources

Every figure above is taken from the vendor's own public material, retrieved 22 September 2026. Claims change; check the source before relying on it. If we have misread yours, tell us and we will correct it.

Frequently asked questions

How accurate are synthetic audiences?

There is no single number. Published figures range from 85% to 98%, but they measure different things: theme overlap, distribution match, response consistency, quality-gate pass rates and rank correlation are not interchangeable. An accuracy figure is only meaningful alongside its metric, its baseline and the task it was measured on.

Why can't I compare two vendors' accuracy percentages?

Because the denominators differ. One vendor may report the share of themes it recovers from expert reports, another the share of a response distribution it matches, another the proportion of generated profiles passing an internal quality gate. Ranking those numbers against each other produces a meaningless ordering.

What is a human ceiling, and why does it matter?

Humans do not reproduce their own survey answers perfectly. Asked again two weeks later, they differ from themselves. That self-consistency rate is the realistic ceiling for any simulation, so reporting a result as a share of it is more informative than reporting a raw percentage.

What is a chance floor?

The score an approach achieves with no model and no data behind it. On one of our GSS item sets, a flat distribution already scores 83.71%. A reported 92.47% therefore closes 54% of the distance between that floor and a perfect score, which is a very different claim from '92% accurate'.