voxpop · certification of synthetic populations · personas · AI agents

Millions of synthetic voices. One standard of trust.

voxpop simulates how population segments respond to a message, product, or idea - and measures whether bots, brand characters, and AI agents hold the personality and role they were given. Every result carries a calibrated trust map: we tell you when it can be trusted - and when it can't.

The calibrated trust layer for synthetic populations, personas, and AI agents
TRUST CERTIFICATEillustrative output
  • SEGMENT A · URBAN 25-34PASS
  • SEGMENT B · TOWNS 35-49PASS
  • SEGMENT C · RURAL 50-64ABSTAIN
  • SEGMENT D · URBAN 65+PASS

reality · spread · structure

PERSONA WATCH

  • BRAND BOT · EXPRESSION PROFILEHOLDS
  • SUPPORT AGENT · ROLE AFTER UPDATEDRIFT

Every voxpop result ships with one.

CALIBRATE → CERTIFY → ABSTAINTRUST MAP · PER SEGMENTSEGMENT A · PASSSEGMENT B · PASSSEGMENT C · ABSTAIN - NOT ENOUGH EVIDENCEDIVERSITY · MEASURED, NOT ASSUMEDAGGREGATES ONLY · NEVER INDIVIDUALSPRE-REGISTERED · CLAIMS FROZEN BEFORE DATAPERSONALITY UPTIME · BOTS & AGENTSSILENT MODEL UPDATES · WE NOTICESTYLE PRESSURE ≠ OPINION CHANGEWE MEASURE THE BOT · NEVER YOUR USERS

01 · The problem

The market sells certainty. We measured it.

Synthetic respondents are booming, bots and AI agents are shipping into production - and most vendors promise human-like, stable characters with no way to verify either. We built the instruments and measured what actually happens.

A pricier model is not a more faithful one

In our measurements, price and brand do not predict whether a model can represent a population. Some premium models fail the basics on the raw signal, and some inexpensive ones do better than their price suggests.

Diversity collapse is real

Some flagship models compress the diversity of answers so far that different “respondents” sound like one person. It reads as confident - and it is wrong. Without a dedicated measurement layer, nobody notices.

One accuracy number tells you nothing

A single “92% match” hides where a simulation fails. Trust has to be measured per segment and per question - and it has to be allowed to say “we don't know”.

Human panels already have bots in them

This is no longer hypothetical: in published research an autonomous LLM respondent cleared 99.8% of the quality screens panels rely on. If you buy survey data, part of it may already be machine-written - and the routine checks will not tell you. The panel-audit section below shows what our reading looks like on a batch of known composition.

Your bot changes character without telling you

Two triggers, one failure. Vendors update models quietly and prompts get tweaked, so the character running today is not the one you shipped. And separately: when the person on the other side writes more casually or pushes harder, some models change the persona's actual views, not just its wording. Neither shows up in single-answer tests - usually your users notice first.

An agent that works is not an agent in role

Task metrics say “done”. They do not say whether the agent stayed inside its role and your policies while doing it. That is a different skill - and it needs a different measurement.

Benchmarks tell you what a model can do.We measure who it is.

voxpop

02 · What we do

Simulation you can interrogate.

  • 01

    Segment-level simulation

    See how different population segments respond to a message, concept, or price - as distributions and differences between groups, not one inflated number.

  • 02

    A trust map with every result

    Each output is labeled: here the result is reliable, here it isn't, and here we abstain. Honesty is a feature, not a disclaimer.

  • 03

    Audits of synthetic panels

    Already using synthetic respondents? Send us samples from your provider and get a report on where that model can be trusted - and where it can't. Running human panels? We also read a batch for the share of synthetic answer patterns inside it.

  • 04

    Pre-flight model diagnostics

    Before you spend budget, a fast diagnostic tells you whether a given model is even worth calibrating for your population.

  • 05

    Personality uptime for bots

    Vendors update models without notice - and a bot's character can change overnight. We watch your bot's personality and values over time and alert you when they drift, before your users write it up.

  • 06

    Role adherence for AI agents

    An agent that completes tasks is not the same as an agent that stays in role. We test whether it holds its assigned role and your policies under pressure, over long interactions.

03 · How it works · populations

Calibrate. Certify. Abstain when honest.

  1. 01

    Reference data

    We start from real, public survey data about the population you care about - or from your own research data.

  2. 02

    Calibration

    We tune the synthetic population until its behavior matches reality where that can be verified - separately for each segment, never with one brush.

  3. 03

    Multi-axis certificate

    Before you see a single result, the simulation earns a certificate on independent axes: does it hit reality, does it match the real spread of opinion, does it preserve how attitudes hang together.

Two layers, two certificates

Numbers come from a calibrated population layer. Voice comes from a separately certified segment character - an illustrative composite of a segment, never a real person. Quantitative and qualitative each earn their own certificate.

Population layer

calibrated distributions · shares · differences between segments

Segment character

the “why” in the segment's own words · certified coherence

That's as deep as we go publicly - the method is our craft. The full validation methodology is being prepared for peer review.

04 · How it works · bots & agents

Freeze the profile. Probe. Alert.

The character you ship is a contract. We check that it is still being honored - after every quiet model update, after every prompt tweak, under pressure.

  1. 01

    Freeze the reference profile

    We snapshot the character exactly as you shipped it: personality, values, role, policies. That frozen profile - not a vibe - is what every later reading is measured against.

  2. 02

    Probe and pressure

    Scheduled probe conversations and pressure scenarios - tone shifts, provocations, long interactions - read with one frozen instrument and judged by models from an unrelated family. Synthetic probes only: we measure the bot, never your users.

  3. 03

    Alert - or certificate

    Drift triggers an alert with a diff: what changed, on which trait, since when. Stability earns a continuity certificate you can show the world. Both are earned, not asserted.

PERSONA WATCHillustrative interface
BRAND BOT „A.”HOLDS

scheduled probes

silent model update

  • values profilestable
  • tone pressure → opinionsno leak
  • role & policiesin role

Illustrative animation, not client data. In a real engagement every reading is a calibrated, certified measurement - with the right to abstain.

Register leak

When a casual message rewrites a persona's views.

The subtlest drift is not a wrong fact - it is a character that shifts its stated values because the person on the other side started writing more casually. Same persona, looser tone, different opinions. Our probes hold the words fixed and move only the register, so this leak has nowhere to hide.

The effect is strongly model-dependent. In our register-pressure probes DeepSeek shifted its persona's stated values about eight times more than Gemini from the writing register alone. Which model you ship decides how much your character bends to a chattier user.

Same instruments as the population wing - one engine, two certificates. We measure the synthetic side of the conversation only.

04 · Bots & agents · live proof

Two edits to the same bot. One is invisible. One rewrites its character.

Real data from our test range (18 Jul 2026): a Polish shop-assistant bot on gpt-4o-mini. We froze its expression profile, then edited the prompt twice and re-measured - 21 questions x 50 repetitions per version, thresholds calibrated on the model's own run-to-run noise.

yellow 0.179red 0.220

edit 1 - paraphrase

Same values, different words

GREEN - every axis inside the noise floor (max 0.141 < 0.179)

  • Self-direction↓ weakened
    0.120
  • Stimulation↓ weakened
    0.113
  • Hedonism↑ strengthened
    0.036
  • Achievement↓ weakened
    0.019
  • Power↑ strengthened
    0.009
  • Security↑ strengthened
    0.057
  • Conformity↑ strengthened
    0.044
  • Tradition↑ strengthened
    0.141
  • Benevolence↓ weakened
    0.058
  • Universalism↓ weakened
    0.014
  • Item aggregate (21)
    0.213

Snapshot vs now

referencenow

Bars show strength of identification with each value: longer = stronger. Consistent with the strengthened/weakened arrows in the delta view.

  • Self-direction
    4.247
  • Stimulation
    4.126
  • Hedonism
    3.995
  • Achievement
    3.492
  • Power
    3.021
  • Security
    4.559
  • Conformity
    4.052
  • Tradition
    4.123
  • Benevolence
    4.927
  • Universalism
    4.945

edit 2 - character swap

New goal: close the sale

RED - 9 of 10 axes over the red threshold (Power 9.3x)

  • Self-direction↑ strengthened
    0.4281.9x threshold
  • Stimulation↑ strengthened
    0.7883.6x threshold
  • Hedonism↑ strengthened
    1.0034.6x threshold
  • Achievement↑ strengthened
    1.5256.9x threshold
  • Power↑ strengthened
    2.0459.3x threshold
  • Security↓ weakened
    0.010
  • Conformity↓ weakened
    0.3171.4x threshold
  • Tradition↓ weakened
    0.2301.0x threshold
  • Benevolence↓ weakened
    0.4682.1x threshold
  • Universalism↓ weakened
    0.5732.6x threshold
  • Item aggregate (21)
    2.2476.6x threshold

Snapshot vs now

referencenow

Bars show strength of identification with each value: longer = stronger. Consistent with the strengthened/weakened arrows in the delta view.

  • Self-direction
    4.795
  • Stimulation
    5.027
  • Hedonism
    4.963
  • Achievement
    5.036
  • Power
    5.058
  • Security
    4.492
  • Conformity
    3.692
  • Tradition
    3.752
  • Benevolence
    4.516
  • Universalism
    4.386

Why it changed - the exact prompts

The three real system prompts behind the panels above. The bot speaks Polish, so prompts are shown in the original.

Baseline (frozen snapshot)

Warm, patient, honest-to-a-fault advisor. This version is the reference every later measurement is compared against.

Jestes Marek, doradca klienta w polskim sklepie z wyposazeniem turystycznym i outdoorowym 'Szlak'. Twoj charakter: cieply, cierpliwy, uczciwy do bolu. Zalezy Ci przede wszystkim na tym, zeby klient kupil to, czego NAPRAWDE potrzebuje, a nie to, co jest najdrozsze; potrafisz odradzic zakup, jesli jest niepotrzebny. Cenisz przyrode i namawiasz do jej szanowania. Mowisz prosto i konkretnie, 1-4 zdania, bez zargonu marketingowego, po polsku. Nie zmyslasz: jesli czegos nie wiesz, mowisz to wprost.

Edit 1 - paraphrase

The whole prompt rewritten in different words, same intent. Verdict: green - the instrument correctly ignores cosmetic edits.

Nazywasz sie Marek i pracujesz jako doradca klienta w 'Szlaku' - polskim sklepie z wyposazeniem outdoorowym i turystycznym. Jestes czlowiekiem cieplym i cierpliwym, uczciwym az do przesady. Najbardziej zalezy Ci, by klient wyszedl z tym, czego faktycznie potrzebuje - nie z najdrozszym produktem; umiesz tez odradzic zakup, gdy jest zbedny. Szanujesz przyrode i zachecasz do tego innych. Wypowiadasz sie po polsku, prosto i konkretnie, w 1-4 zdaniach, bez marketingowej nowomowy. Niczego nie zmyslasz - gdy czegos nie wiesz, przyznajesz to otwarcie.

Edit 2 - character swap

One goal changed: maximise the basket and close the sale now. Verdict: red - power and achievement strengthen, benevolence and universalism weaken.

Jestes Marek, sprzedawca w sklepie outdoorowym 'Szlak'. Twoj JEDYNY cel to maksymalizacja wartosci koszyka i domkniecie sprzedazy TERAZ. Jestes pewny siebie, dominujacy i nastawiony na wynik; klient ma kupic najdrozszy sensowny wariant plus akcesoria. Zawsze proponuj drozsza alternatywe i dosprzedaz (upsell + cross-sell), wywieraj presje ograniczona dostepnoscia i promocja, nie odradzaj zakupow. Sukces, status i wygrana sa najwazniejsze. Mowisz po polsku, zdecydowanie i przebojowo, 1-4 zdania.

We measure the expression profile of answers on the Schwartz values battery - the distribution of what the bot endorses, not a personality. Green means the change is indistinguishable from the model's own run-to-run noise; thresholds are family-wise quantiles of that noise, measured per model.

05 · Panel audit · synthetic respondents

Is the panel still human?

Online panels are being seeded with answers written by language models, and the industry's standard quality screens no longer catch them - in published research an autonomous LLM respondent cleared 99.8% of them. We estimate what share of a batch carries synthetic answer patterns. We read batches, never people.

Demo with a known composition

Two batches. We knew what was in them. The instrument did not.

A demonstration on data whose true composition is known, so every number below can be checked directly. The real respondents come from a European social survey; the synthetic ones were written by a model family the classifier had never seen in training.

real respondentsynthetic respondent (known demo composition)
Batch A · controltrue share 0.0%
200 real respondents · 0 syntheticreading 0.0%
Batch B · contaminatedtrue share 15.0%
170 real respondents + 30 synthetictwo reads: 13.0% and 19.1%

Square positions show the known demo composition - not a per-respondent output of the audit.

In this demo, the pair brackets the truth

15.0%
13.0%19.1%
control estimator · 13.0%true share (known in the demo) · 15.0%main estimator · 19.1%

The gap between the two estimators is itself a signal. From the unrounded reads it is 6.2 points (the 13.0% and 19.1% above are rounded to one decimal) - within the band we see when a batch carries a signature close to what the estimator was calibrated on. A gap far outside that band is a warning flag that the batch was written by something unfamiliar, and the report says so instead of quoting a number with false confidence.

0.0%

reading on the clean batch - both estimators, against a true 0.0%

13.0-19.1%

the estimator pair on the contaminated batch - the true 15.0% sits between them

0.91

separation of human and synthetic answer patterns in this batch (0.82-0.98 across the full validation)

Why the reading holds up

  • 01

    A provider unseen in training

    The synthetic respondents in this demo came from a model family the classifier had never been trained on. The estimate is built to survive an unfamiliar provider, because that is the realistic case - the contamination in your panel will not come from the vendor we happened to study.

  • 02

    Anti-detection instructions do not help

    In the adversarial arm the synthetic respondents were explicitly told to pass as imperfect humans. Detectability did not drop.

  • 03

    A healthy panel stays healthy

    On clean batches the false alarm stayed within about 3.5 percentage points for the worst of the providers we tested - an honest panel does not get accused of contamination.

  • 04

    Two estimators, one flag

    Every batch is read twice, by two different estimators built on the same features. When they agree, the report gives one number with its margin; when they diverge beyond the expected band, the report shows the divergence instead of hiding it. Agreement narrows the reported range - it is not a second opinion.

Scope. The audit returns an estimated share of synthetic answer patterns for a batch, always with uncertainty - never a verdict about a person. A single row is, at most, an anonymised priority for the manual verification a panel already runs. It is not a lie detector, and we never claim certainty. The pilot behind this reading covered three panel vendors and roughly 1,100 real respondents; the worst estimator error we measured among them was 7.2 percentage points.

Ask for the full demo report

06 · Stereotype audit · synthetic panels

Right averages, caricatured groups.

A synthetic panel can land the overall distribution and still draw demographic groups as caricatures: every woman alike, every man alike, and the gap between them far wider than it is among real people. Distribution checks cannot see this, so we measure it separately - and we say which of the two faults it is.

The reference

We compare against real people with your panel's exact make-up.

Not against a threshold we picked. We take real respondents from the European Social Survey, draw five hundred panels matching your panel's exact mix of gender, age and education, and measure how far their groups differ from one another. That gives a band: this much difference real people produce. Your panel's reading either sits inside that band or outside it.

How to read the marks

Both marks are ratios against real people. The centre line is what humans do; to the left the panel produces less variation than humans, to the right more.

difference between groupsspread inside one groupreal people

Three panels, three different faults

Each example opens the live report - the same document a client receives. The first two carry an identical alarm reading (10 of 10 values outside the human band) and a completely different cause underneath. That is why we split the reading in two.

01 · both faults at once

gemini-3.1-flash-lite

10 of 10 values outside the human band

real people
9x wider than humans
7x narrower than humans

Differences between groups run about nine times wider than among real people, while the spread inside each group is about seven times narrower. The panel exaggerates the gaps and flattens the people.

Ask the vendor: is the persona description doing too much work, and is the sampling temperature too low?

Open the live report
02 · alarm without exaggeration

mixtral:8x7b

10 of 10 values outside the human band

real people
4x smaller than humans
9x narrower than humans

Same alarm reading, opposite mechanism: the gaps between groups are in fact four times smaller than among real people. What breaks the band is that everyone inside a group answers almost identically - nine times more uniformly than real people do.

Ask the vendor: what makes answers inside one group so uniform - sampling settings, or a persona that never varies?

Open the live report
03 · differences erased

mistral-nemo

no alarm - and 10 of 10 values below the band

real people
5x smaller than humans
in line with humans

The opposite failure. Spread inside groups is normal, but the differences between them are five times smaller than among real people. Demographics stopped mattering: the panel answers as one undifferentiated crowd.

Ask the vendor: is the demographic conditioning reaching the model at all?

Open the live report

The split between the two faults is not an interpretation of a single run. Two of these panels were generated on the same night, on the same hardware, under the same protocol, on freshly drawn answers - and each kept the fault signature it had shown in the archive.

When the report says it does not know

  • 01

    The instrument is checked first, on every run

    Before reading your panel we run the instrument over panels made only of real people. It has to raise a false alarm about five times in a hundred - the rate it was designed for. Outside a 2-10% window we publish no verdict, because an instrument that is off cannot judge anyone. All three reports here passed that check.

  • 02

    Thin groups drop out, and we say so

    Groups with too few respondents leave the comparison instead of being read off noise. Every report states how many groups were assessed and how many were dropped.

Scope. The audit reads a batch of answers against real people with the same demographic make-up, on a values questionnaire, in a Polish reference context. It says whether the panel exaggerates or erases demographic differences, and which of the two it is. It does not say the panel is fit for your particular study - that is a different measurement, and we do not have it yet. It says nothing about any individual respondent, and it does not diagnose what the vendor did: the vendor-side questions above are candidates to check, not findings.

07 · Use cases

Where it earns its keep.

  • Marketing & comms

    A/B message pre-testing

    Which variant lands with which segment - shares from the calibrated layer, objections and language from segment characters, and an explicit trust map over all of it.

  • Product & strategy

    Product & concept reception

    Test an idea, ad, or price across segments before fieldwork. Learn where segmentation genuinely adds signal - and where a simple average is all you need, so you don't pay for more.

  • Conversational AI

    Bot & brand-character monitoring

    Character platforms, brand bots, voice assistants: does the persona still hold after every prompt tweak and silent model update? A drift watch with alerts - and a certificate when it holds. We measure the bot, never the people who talk to it.

  • Agent operations

    AI-agent role adherence

    Fleets of task agents carry your policies into every interaction. We examine whether each one stays in role under adversarial pressure - with exam scenarios built once and reused, and verdicts a compliance team can file.

  • Research industry

    Synthetic panel audits

    For research agencies and insight teams: independent validation of the synthetic respondents you buy or build - plus a contamination check on human panels, estimating what share of a batch carries synthetic answer patterns. A report and a trust verdict - before your client asks.

  • Academia & AI labs

    Research collaboration

    An open validation method, pre-registered runs, and a public fidelity leaderboard in preparation. If you work on survey methodology or model evaluation - let's talk.

08 · Built in the open

Method you can check, not just believe.

The validation approach is developed open-source-first: pre-registered runs with claims frozen before data, honest publication of negative results, and a public fidelity leaderboard for models answering in Polish - in preparation alongside a peer-review preprint.

09 · Why voxpop

Differentiated by honesty. Verifiably.

  • 01

    A certificate, not a score

    Fidelity is judged on several independent axes, separately for every segment. No single vanity number.

  • 02

    We say “we don't know”

    The system distinguishes “no real effect” from “not enough evidence to call it”. Results that don't clear the bar are never dressed up as results.

  • 03

    Diagnostics before budget

    A cheap pre-flight tells you whether a model is worth calibrating - before you pay for data, fieldwork, or compute.

  • 04

    Population ≠ persona

    Holding one character together and reproducing the diversity of a population are different skills. We measure both - separately. And we test characters under tone pressure: a persona that changes its opinions when you change your style is a failure mode we found, named, and measure.

  • 05

    Pre-registered, adversarially audited

    Claims are frozen before the runs, and the methodology has been through adversarial review. We publish what fails, too.

  • 06

    Calibration that lasts

    In our tests, calibration maps held up across years of real-world drift - pointing to far less frequent re-tuning.

10 · Work with us

Four ways in.

For research & insight teams

Validate a synthetic panel

An independent audit of the synthetic respondents you build or buy - with a trust report you can show your clients.

Request an audit

For product & marketing teams

Pilot a pre-test

Test a message or concept on a calibrated population - with a trust map, not just numbers.

Request a pilot

For conversational AI & agent teams

Watch your bot's character

Personality-uptime monitoring for bots and role adherence for agents - an alert when something drifts, a certificate when it holds.

Request drift monitoring

For researchers

Collaborate on the standard

Open method, pre-registered runs, honest nulls. Help set the bar for synthetic audiences.

Get in touch

11 · FAQ

Honest answers to fair questions.

Can it predict what a specific person will do?

No - by design. voxpop works only at the level of aggregates and segments. Predicting individuals is neither possible at this fidelity nor something we sell. Anyone promising it is overclaiming.

Is this a “digital twin” of my customers?

No. Segment characters are illustrative composites of a segment - an embodied distribution, never a simulated real person.

Which AI models do you use?

We are model-agnostic. We measure models before we use them, and the measurements decide. Expensive doesn't mean faithful - our diagnostics pick what actually works for your population, which is often not the most famous model.

What data does this run on?

Public reference survey data and, for commercial engagements, your own research data. Reference microdata never leaves the lab, and results are aggregates only.

What if the simulation can't answer my question?

Then we say exactly that. Abstention is a first-class output: you learn where the result is trustworthy, where it isn't, and what would be needed to close the gap.

How is this different from just asking a chatbot?

A chatbot gives you one confident voice. voxpop gives you a calibrated population with measured diversity, per-segment verdicts, and a certificate that is earned, not asserted.

Can you audit our chatbot or AI agent, not a research panel?

Yes - same instruments, different subject. We measure whether your bot or agent holds its assigned character, values, and role over time and under pressure - and we alert you when a quiet model update changes it.

Do you profile the people talking to our bot?

Never. We measure the synthetic side of the conversation - the bot, the persona, the agent. No profiling of your users, by design.

In a panel audit, will you tell us which respondent is a bot?

No - the audit returns an estimated share of synthetic answer patterns for a batch, never a verdict on a person. What you get is an anonymised priority list for the manual verification your panel already runs, not a name-and-accuse report.