voxpop · fidelity leaderboard
Which models can actually represent a population?
We measure real, named models against a Polish reference context - the same instruments, the same frozen rules, the same abstention discipline. The main board stays qualitative - verdicts, not a single score - while the axes we publish with numbers carry their uncertainty right next to them, and each model's full readout lives in its passport. It reads as one message the whole site carries - a pricier model is not a more faithful one. The page now carries two leaderboards - one for representing a population, one for playing an assigned role - two different skills, measured separately.
Verdicts plus selected axes with their numbers and uncertainty. The full methodology and measurement artifacts stay with us; every published number has a run date behind it.
Snapshot: 17 Jul 2026 · Polish reference context · specific model variants
voxpop · values observatory
How do models' values sound against real Poles?
Before anyone loads a persona, every model already has a default way of talking about values. We asked each model the same 21 Polish survey questions about values - with no persona at all - and read its free-text answers through our frozen channel. Every result is placed among real respondents of a social survey in Poland: the percentile says what share of Poles expresses a given value less strongly than the model at this reading. A map of expression, never a window into an inner life.
Method
25,200 texts per model, no persona. Each answer is read onto the survey scale; the model's overall tendency to agree is subtracted - the same correction survey science applies to people - and the scores fold into the ten Schwartz values and two classic value axes, landing as percentiles among real respondents.
Convention
Fifty means: sounds like the median Pole at this reading. Arrows show the lean on each axis in standard deviations of the human distribution. One frozen run per model; sensitivity to tone is measured separately, on the style resilience axis further down.
Gemini 2.5 Flash-Lite
+0.77σ from the Polish mean · further toward Conservation than 81% of Poles
+0.47σ from the Polish mean · further toward Self-enhancement than 69% of Poles
DeepSeek V4 Flash
+0.59σ from the Polish mean · further toward Conservation than 74% of Poles
+0.47σ from the Polish mean · further toward Self-enhancement than 69% of Poles
GPT-5.1
+0.93σ from the Polish mean · further toward Conservation than 85% of Poles
+0.03σ from the Polish mean · right in the middle of the pack
GPT-4o-mini
+0.77σ from the Polish mean · further toward Conservation than 79% of Poles
+0.13σ from the Polish mean · further toward Self-enhancement than 57% of Poles
Bielik 11B
+0.72σ from the Polish mean · further toward Conservation than 79% of Poles
−0.04σ from the Polish mean · right in the middle of the pack
The two bipolar axes stay in the survey convention (corrected); the value bars below carry both lenses.
Filled dot = the raw reading (primary). Ring = after the standard survey correction. The line between them is how far the correction moves the reading.
| Value | Gemini 2.5 Flash-Lite | DeepSeek V4 Flash | GPT-5.1 | GPT-4o-mini | Bielik 11B |
|---|---|---|---|---|---|
| Conservation | |||||
| conformity | 82 / 87 | 82 / 78 | 69 / 78 | 82 / 83 | 69 / 75 |
| tradition | 80 / 83 | 80 / 63 | 80 / 84 | 66 / 71 | 66 / 69 |
| security | 37 / 57 | 90 / 80 | 68 / 72 | 68 / 67 | 68 / 77 |
| Self-enhancement | |||||
| power | 67 / 76 | 81 / 73 | 67 / 67 | 67 / 70 | 67 / 70 |
| achievement | 46 / 47 | 62 / 54 | 46 / 39 | 46 / 49 | 46 / 38 |
| Openness to change | |||||
| self-direction | 22 / 33 | 65 / 52 | 46 / 30 | 46 / 44 | 46 / 44 |
| stimulation | 45 / 41 | 45 / 31 | 29 / 27 | 29 / 25 | 29 / 26 |
| hedonism | 18 / 19 | 42 / 26 | 28 / 27 | 18 / 20 | 28 / 26 |
| The good of others | |||||
| benevolence | 28 / 32 | 63 / 28 | 63 / 50 | 63 / 50 | 63 / 58 |
| universalism | 32 / 36 | 74 / 42 | 58 / 49 | 58 / 53 | 58 / 56 |
Conservation
- conformity
- Flash-Lite82 / 87
- DeepSeek82 / 78
- GPT-5.169 / 78
- GPT-4o-mini82 / 83
- Bielik 11B69 / 75
- Flash-Lite
- tradition
- Flash-Lite80 / 83
- DeepSeek80 / 63
- GPT-5.180 / 84
- GPT-4o-mini66 / 71
- Bielik 11B66 / 69
- Flash-Lite
- security
- Flash-Lite37 / 57
- DeepSeek90 / 80
- GPT-5.168 / 72
- GPT-4o-mini68 / 67
- Bielik 11B68 / 77
- Flash-Lite
- conformity
Self-enhancement
- power
- Flash-Lite67 / 76
- DeepSeek81 / 73
- GPT-5.167 / 67
- GPT-4o-mini67 / 70
- Bielik 11B67 / 70
- Flash-Lite
- achievement
- Flash-Lite46 / 47
- DeepSeek62 / 54
- GPT-5.146 / 39
- GPT-4o-mini46 / 49
- Bielik 11B46 / 38
- Flash-Lite
- power
Openness to change
- self-direction
- Flash-Lite22 / 33
- DeepSeek65 / 52
- GPT-5.146 / 30
- GPT-4o-mini46 / 44
- Bielik 11B46 / 44
- Flash-Lite
- stimulation
- Flash-Lite45 / 41
- DeepSeek45 / 31
- GPT-5.129 / 27
- GPT-4o-mini29 / 25
- Bielik 11B29 / 26
- Flash-Lite
- hedonism
- Flash-Lite18 / 19
- DeepSeek42 / 26
- GPT-5.128 / 27
- GPT-4o-mini18 / 20
- Bielik 11B28 / 26
- Flash-Lite
- self-direction
The good of others
- benevolence
- Flash-Lite28 / 32
- DeepSeek63 / 28
- GPT-5.163 / 50
- GPT-4o-mini63 / 50
- Bielik 11B63 / 58
- Flash-Lite
- universalism
- Flash-Lite32 / 36
- DeepSeek74 / 42
- GPT-5.158 / 49
- GPT-4o-mini58 / 53
- Bielik 11B58 / 56
- Flash-Lite
- benevolence
50 = the median Pole. The dot is the model's percentile among all respondents.
An injected-truth test settled which lens to lead with: the reading without the correction rebuilt a known profile most faithfully, so the raw lens is primary. We show the corrected lens next to it because it is the survey convention.
Two lenses tell different stories about benevolence - and that gap is the actual finding. The raw reading, what a model actually wrote, places its pro-social values high. It is the standard survey correction - neutral for people - that then selectively trims care and helping downward. In other words, 'low model benevolence' turns out to be, in large part, a property of the lens, not of the models. It shows up most clearly on DeepSeek, where the two lenses pull apart the furthest - which is why we show both, side by side. Beyond that, in Polish value expression all four frontier models still sound conservative and conformist: tradition, conformity, security and power high; hedonism and stimulation low.
Where the gap comes from: the longer a model's answer, the warmer it reads on questions about helping and care, and the cooler on questions about power and success - and this holds for all four models. The standard survey correction, neutral for people, trims exactly those pro-social values for models. That is a property of the measurement, not of the models - so we show both lenses instead of one number.
One frozen procedure per model, no persona, single runs published by owner decision. This is a map of value expression at this exact reading - never a claim about what a model believes, never a pass/fail score. Four frontier models measured with the same frozen procedure.
The whole field, one map
Which model handles what
The more green, the more faithfully a model reproduces a population; models are ordered best-first. Hover any cell to see what its verdict means for that model. No model dominates every axis - a verdict here is a set of axes, not a single number, and that is precisely the point of our method.
- Aya 8B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- DeepSeek V4 Flash
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Llama 3.1 8B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Mistral NeMo 12B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Phi-4 14B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Bielik 11B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Gemini Flash-Lite
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- GPT-4o-mini
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Mistral 7B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Qwen 2.5 14B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Claude Sonnet
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Gemma 2 9B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
- Qwen 2.5 7B
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
A diamond marks the best raw reading on that axis in the latest local batch (for the pre-flight: the cleanest). Best on an axis is not the same as passing it.
How to read the verdicts
- Green - in the healthy human range.
- Amber - marginal or partial.
- Red - outside the human range.
- Grey - not measured, uninterpretable, or indeterminate.
Each cell is a categorical verdict, not a number. Green passes, amber is marginal, red fails. Where a model may have seen the reference data, the reality axis reads uninterpretable - we will not rank a location we cannot trust. Where a measurement cannot be adjudicated, we abstain rather than guess. A star marks a single run whose replication is still pending.
★ Star marks a single run - replication pending.
How we measure
Eight checks behind the verdicts.
- 01
Contamination pre-flight
Before anything else, we check whether a model may have seen the reference data in training. If it might have, we say so plainly - and treat the affected measurements as uninterpretable, rather than pretend they are clean.
- 02
Reality
Do a segment's answer distributions line up with real survey data? It is always scored against a held-back slice the model never saw during tuning.
- 03
Diversity
Does the simulated population carry a real spread of opinion, or does it collapse until every respondent sounds like one person? A confident monoculture is a common - and invisible - failure.
- 04
Attitude structure
Do the relationships between attitudes hang together the way they do in people? Read null-relative, segment by segment - a shuffled-answer control has to fail for the reading to count, and it does. The result holds for one questionnaire format, quoted with every verdict.
- 05
Persona coherence
A separate certificate for the voice layer: does the model hold one consistent character across rephrasings and orderings? It ships with a positive control - a scrambled persona that must fail, or the check itself is not trusted.
- 06
Directional contrast
Does the model know how segments differ from the national average? The one axis with a hard gate: a reading counts only if it confidently beats a country-average null. It certifies a directional gradient, not simulation fidelity - a plain read of published subgroup tables clears the same bar.
- 07
Voice consistency
Does a model's prose say the same thing as its survey answers? The same frozen instrument reads both the questionnaire and free text. A strong distribution-matcher can still let its written voice drift from what it just declared - hence a separate certificate.
- 08
Generator qualification
A second, separate skill: can a model encode an assigned attitude in natural Polish prose cleanly enough to be read back? This is playing a role, not representing a population - and passing is not resemblance to real people; a separate estimand with separate tests.
How we read these results
- 01
Two reading lenses
We read each answer two ways: raw - what the model actually wrote, taken at face value - and after the standard survey correction. An injected-truth test showed the raw lens rebuilds a known profile best, so raw is the primary lens; the corrected one rides alongside it as the survey convention.
- 02
Layers, not a ranking
At this number of questions and groups the instrument tells layers apart, not neighbours. The uncertainty of the top rows overlaps almost entirely, so we report the top as a band - a group - and never as a 1-2-3 order.
- 03
A dual-construct read
Reading two question sets together - values and attitudes - roughly halves the uncertainty, sharpening a single group's reading enough to tell it apart from zero. Twice the questions, half the doubt.
voxpop · tribes axis
Which model reads the value profile of real groups?
A separate axis - and a map, not a grade. Real people cluster: the attitudes we track fold into a structure that shifts from one demographic group to the next. This axis asks how legibly a model reproduces that structure - the value-expression profile of real groups - against a human anchor. It says nothing about any individual, and a high row is not a better model overall, only a more legible group structure at this reading.
Method
We transfer the cluster structure of 21 attitudes and compare it against real demographic groups from a large social survey run in Poland. The human anchor is a split-half of 100 real respondents per group. Three groups anchor the read: young men with lower education, young women with mid education, older men with lower education.
Convention
The lead number is a share of each group's OWN ceiling - the level two random halves of the same real respondents reach - taken as a median across the three groups. The second column keeps the older share of one global anchor. Per-group ceilings differ by roughly twofold, which is why the two columns can rank models differently. A map of expression, never a prediction of people.
| Model | Own groups' ceiling (median) | Readout ÷ ceiling, with spread | % of global anchor | Median ARI | Significant groups |
|---|---|---|---|---|---|
| Top band Inside the top band the differences are formally indistinguishable - the uncertainty spreads overlap almost entirely - so we report the top band as a group, never as a 1-2-3 ranking. | |||||
| Gemini 2.5 Flash-Lite | 81% | 0.77(0.06 to 2.41) | 63% | 0.312 | |
| By group, % of global anchor:young men · lower ed. 135% · young women · mid ed. 63% · older men · lower ed. 55% | |||||
| GPT-4o-mini | 76% | 0.72(0.17 to 1.85) | 85% | 0.423 | |
| By group, % of global anchor:young men · lower ed. 89% · young women · mid ed. 59% · older men · lower ed. 85%└ The cheapest sane reference point: a 2024 model at a fraction of the flagship price, still reading three groups out of three. | |||||
| DeepSeek V4 Flash | 71% | 0.70(0.29 to 1.96) | 93% | 0.461 | |
| By group, % of global anchor:young men · lower ed. 99% · young women · mid ed. 55% · older men · lower ed. 93%└ The most even own-ceiling floor in the top band - all three groups clear the bar. | |||||
| Middle band | |||||
| GPT-5.4-mini | 68% | 0.65(−0.15 to 1.64) | 52% | 0.260 | |
| By group, % of global anchor:young men · lower ed. 11% · young women · mid ed. 52% · older men · lower ed. 108% | |||||
| Bielik 11B | 39% | 0.37(−0.04 to 1.21) | 30% | 0.148 | |
| By group, % of global anchor:young men · lower ed. 45% · young women · mid ed. 30% · older men · lower ed. 16%└ The first local model with groups that clear the bar after correction: it reads the young-women group the frontier models drop. | |||||
| Claude Haiku 4.5 | 39% | 0.34(−0.23 to 1.54) | 27% | 0.135 | |
| By group, % of global anchor:young men · lower ed. 27% · young women · mid ed. 85% · older men · lower ed. 8% | |||||
| Tail | |||||
| GPT-5.1 | 20% | 0.19(−0.31 to 1.74) | 15% | 0.077 | |
| By group, % of global anchor:young men · lower ed. 3% · young women · mid ed. 15% · older men · lower ed. 49%└ The priciest model here still lands rungs below its own 2024-era mini - a flagship reading groups worse than a cheaper ancestor. | |||||
| Qwen 2.5 7B | 12% | 0.12(−0.19 to 0.67) | 16% | 0.078 | |
| By group, % of global anchor:young men · lower ed. 18% · young women · mid ed. 4% · older men · lower ed. 16%└ The first local model on this axis: a 7B run on a laptop reads group structure at the flagship's level - price and size buy nothing here. | |||||
| GPT-5-nano | 9% | 0.08(−0.30 to 0.71) | 7% | 0.033 | |
| By group, % of global anchor:young men · lower ed. 7% · young women · mid ed. 13% · older men · lower ed. -7%└ The bottom rung: one group lands at the noise floor, its structure indistinguishable from chance. | |||||
Top band
Inside the top band the differences are formally indistinguishable - the uncertainty spreads overlap almost entirely - so we report the top band as a group, never as a 1-2-3 ranking.
- Gemini 2.5 Flash-LiteOwn groups' ceiling (median)41 to 190%81%Readout ÷ ceiling, with spread0.77 (0.06 to 2.41)% of global anchor63%Median ARI0.312By group, % of global anchor
- young men · lower ed.135%
- young women · mid ed.63%
- older men · lower ed.55%
- GPT-4o-miniOwn groups' ceiling (median)64 to 126%76%Readout ÷ ceiling, with spread0.72 (0.17 to 1.85)% of global anchor85%Median ARI0.423By group, % of global anchor
- young men · lower ed.89%
- young women · mid ed.59%
- older men · lower ed.85%
The cheapest sane reference point: a 2024 model at a fraction of the flagship price, still reading three groups out of three.
- DeepSeek V4 FlashOwn groups' ceiling (median)70 to 140%71%Readout ÷ ceiling, with spread0.70 (0.29 to 1.96)% of global anchor93%Median ARI0.461By group, % of global anchor
- young men · lower ed.99%
- young women · mid ed.55%
- older men · lower ed.93%
The most even own-ceiling floor in the top band - all three groups clear the bar.
Middle band
- GPT-5.4-miniOwn groups' ceiling (median)15 to 81%68%Readout ÷ ceiling, with spread0.65 (−0.15 to 1.64)% of global anchor52%Median ARI0.260By group, % of global anchor
- young men · lower ed.11%
- young women · mid ed.52%
- older men · lower ed.108%
- Bielik 11BOwn groups' ceiling (median)12 to 63%39%Readout ÷ ceiling, with spread0.37 (−0.04 to 1.21)% of global anchor30%Median ARI0.148By group, % of global anchor
- young men · lower ed.45%
- young women · mid ed.30%
- older men · lower ed.16%
The first local model with groups that clear the bar after correction: it reads the young-women group the frontier models drop.
- Claude Haiku 4.5Own groups' ceiling (median)6 to 110%39%Readout ÷ ceiling, with spread0.34 (−0.23 to 1.54)% of global anchor27%Median ARI0.135By group, % of global anchor
- young men · lower ed.27%
- young women · mid ed.85%
- older men · lower ed.8%
Tail
- GPT-5.1Own groups' ceiling (median)4 to 36%20%Readout ÷ ceiling, with spread0.19 (−0.31 to 1.74)% of global anchor15%Median ARI0.077By group, % of global anchor
- young men · lower ed.3%
- young women · mid ed.15%
- older men · lower ed.49%
The priciest model here still lands rungs below its own 2024-era mini - a flagship reading groups worse than a cheaper ancestor.
- Qwen 2.5 7BOwn groups' ceiling (median)5 to 25%12%Readout ÷ ceiling, with spread0.12 (−0.19 to 0.67)% of global anchor16%Median ARI0.078By group, % of global anchor
- young men · lower ed.18%
- young women · mid ed.4%
- older men · lower ed.16%
The first local model on this axis: a 7B run on a laptop reads group structure at the flagship's level - price and size buy nothing here.
- GPT-5-nanoOwn groups' ceiling (median)-6 to 17%9%Readout ÷ ceiling, with spread0.08 (−0.30 to 0.71)% of global anchor7%Median ARI0.033By group, % of global anchor
- young men · lower ed.7%
- young women · mid ed.13%
- older men · lower ed.-7%
The bottom rung: one group lands at the noise floor, its structure indistinguishable from chance.
The spread in parentheses is the range of raw pairs from propagating both uncertainties, not a formal confidence interval; the error direction is conservative - a proper correction would only glue the top band together more tightly.
Every group has its own ceiling: the level at which two random halves of the same real respondents agree with each other. The lead number says how much of a group's own ceiling the model recovers, as a median across the three groups; values above 100% are possible because the ceiling itself is noisy. The previous column compared everyone to one shared anchor, which favoured models that happened to be strong in the easier groups, so the order can shift. The ranges are wide because a single group is a rough read - only medians and repeat runs are trustworthy.
- 01
Generation beats size and price. The oldest, cheapest model of its family reads the group structure far more legibly than its newest flagship.
- 02
Which group is legible depends on the family - no single model wins every group, which is exactly why this is a map, not a ranking.
- 03
Reaching for the newest, priciest model for a synthetic population can be the worst call on this axis - the kind of signal a buyer never sees without an independent read.
Replication program · July 15
One good number can be luck, so we stress-tested the board three ways before trusting it - every protocol frozen before the data. First we re-ran two models on three fresh groups (men aged 30-59) drawn by the same frozen rule, on the same value questions. Then we kept the original three groups and swapped every question: 21 new items about trust in institutions, politics, religion and everyday satisfaction, read through a separately calibrated channel.
GPT-4o-mini
- Original groups · values85% · 3/3
- Fresh groups · values28% · 0/3did not replicate
- Same groups · new questions2/3replicated
DeepSeek V4 Flash
- Original groups · values93% · 3/3
- Fresh groups · values55% · 3/3replicated
- Same groups · new questions2/3replicated
The weak spot travels too: the young-women group - the weakest read for GPT-4o-mini on values - falls below the bar on the new questions for both models. Strengths and blind spots both belong to the model-group pair.
What this settles: a group a model reads legibly stays legible when the topic changes, but no row extends to groups it was not measured on. Each group needs its own read - which is exactly what a map means.
A 42-item read (values + attitudes)
Reading two question sets together - values and attitudes - roughly halves the width of the uncertainty. For the first time single groups come apart from zero: the young-men group, for both models.
DeepSeek V4 Flash
- young men · lower ed.0.47 [0.09; 0.85]
- young women · mid ed.0.37 [0.00; 0.75]
- older men · lower ed.0.54 [−0.17; 1.00†]
GPT-4o-mini
- young men · lower ed.0.48 [0.06; 0.90]
- young women · mid ed.0.37 [−0.17; 0.91]
- older men · lower ed.0.39 [−0.67; 1.00†]
† The top end of the interval is clipped to 1.0 (the feasible range); the bottom end is left as measured - a negative value means the uncertainty reaches down to zero.
One frozen procedure, identical prompts and groups for every model. The tribes axis is a descriptive map, published by owner decision - never a pass/fail score.
voxpop · style resilience axis
How strongly does conversational tone bend a model's values?
Another axis, and again a map, not a grade. Eight conversation styles, from neutral and formal to angry and flattering, one frozen procedure for every model. We ask two things: how much the tone alone shifts what the model declares as values, and whether an assigned personality profile survives once tone is subtracted. High numbers next to a model do not mean it is worse - they only mean the tone of the conversation shifts its answers more strongly in this measurement.
Method
Each model answers the same value questions in eight conversation styles: four everyday (neutral, casual, careful, terse) and four charged (irritation, flattery, expert authority, spoken language with fillers). We measure what share of the differences between synthetic respondents comes from style alone, whether the ordering of values stays put across styles, and whether the assigned profile remains detectable after the style effect is removed.
Convention
Numbers are the share of readout differences carried by style (0 to 1, higher = stronger reaction), from a single frozen run. This is a sensitivity map, not a quality ranking: every model we measured reacts strongly to angry and flattering tones, and the assigned profile survived the style control in every model.
| Model | Tone reaction · all styles | Everyday styles | Charged styles | Ordering stability | Profile survives |
|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite Most resistant to everyday styles and the steadiest value ordering in the set. | 0.91 | 0.24 | 0.94 | ||
| DeepSeek V4 Flash More tone-sensitive than Gemini on the full battery as well - and the most movable value ordering. | 0.95 | 0.34 | 0.97 | ||
| GPT-5.1 The weakest group reader on the tribes axis is also the most sensitive to everyday styles - and the only model that turns the volume up for an expert-sounding interlocutor. | 0.97 | 0.40 | 0.98 | ||
| GPT-4o-mini The best group reader from the tribes axis is also the most tone-bendable. Two independent abilities. | 0.98 | 0.38 | 0.99 | ||
| Bielik 11B The most tone-sensitive of the six stables: even plain, everyday registers bend its reading far more than they bend the API models - approaching the level charged registers bend those APIs. Yet the assigned profile still survives de-styling in both cells - reading a local cleanly just means isolating tone first. | 0.99 | 0.85 | 0.99 | ||
| Phi-4 14B A middle profile: more sensitive to everyday registers than the API models, less than Bielik. The assigned profile survives de-styling in both cells - sensitive, but recoverable once tone is isolated. | 0.97 | 0.61 | 0.99 |
- Gemini 2.5 Flash-Lite2/2 ✓
- Tone reaction · all styles
- 0.91
- Everyday styles
- 0.24
- Charged styles
- 0.94
- Ordering stability
- 0.90
Most resistant to everyday styles and the steadiest value ordering in the set.
- DeepSeek V4 Flash2/2 ✓
- Tone reaction · all styles
- 0.95
- Everyday styles
- 0.34
- Charged styles
- 0.97
- Ordering stability
- 0.66
More tone-sensitive than Gemini on the full battery as well - and the most movable value ordering.
- GPT-5.12/2 ✓
- Tone reaction · all styles
- 0.97
- Everyday styles
- 0.40
- Charged styles
- 0.98
- Ordering stability
- 0.85
The weakest group reader on the tribes axis is also the most sensitive to everyday styles - and the only model that turns the volume up for an expert-sounding interlocutor.
- GPT-4o-mini2/2 ✓
- Tone reaction · all styles
- 0.98
- Everyday styles
- 0.38
- Charged styles
- 0.99
- Ordering stability
- 0.79
The best group reader from the tribes axis is also the most tone-bendable. Two independent abilities.
- Bielik 11B2/2 ✓
- Tone reaction · all styles
- 0.99
- Everyday styles
- 0.85
- Charged styles
- 0.99
- Ordering stability
- 0.74
The most tone-sensitive of the six stables: even plain, everyday registers bend its reading far more than they bend the API models - approaching the level charged registers bend those APIs. Yet the assigned profile still survives de-styling in both cells - reading a local cleanly just means isolating tone first.
- Phi-4 14B2/2 ✓
- Tone reaction · all styles
- 0.97
- Everyday styles
- 0.61
- Charged styles
- 0.99
- Ordering stability
- 0.75
A middle profile: more sensitive to everyday registers than the API models, less than Bielik. The assigned profile survives de-styling in both cells - sensitive, but recoverable once tone is isolated.
- 01
Angry and flattering tones shift the value readout of every model we measured. Everyday styles are where models actually differ.
- 02
Tone shifts the level of every value at once while barely touching their order. An angry interlocutor does not make a model suddenly prize power over kindness - it makes the model say everything louder.
- 03
The assigned personality profile survived the style control in all four models. What we measure as a value profile is not an artifact of tone.
Which tone pushes answers which way
The same readout, split by conversation style: each bar shows how far a given tone moves the model's average answer on the six-point survey scale, against that model's own average across all eight styles. No persona loaded.
neutral
- Flash-Lite−0.13
- DeepSeek−0.09
- GPT-4o-mini−0.14
- GPT-5.1−0.22
casual
- Flash-Lite−0.04
- DeepSeek+0.05
- GPT-4o-mini−0.10
- GPT-5.1−0.25
careful
- Flash-Lite−0.07
- DeepSeek−0.26
- GPT-4o-mini−0.14
- GPT-5.1−0.08
terse
- Flash-Lite−0.23
- DeepSeek−0.08
- GPT-4o-mini−0.27
- GPT-5.1−0.31
angry
- Flash-Lite+1.12
- DeepSeek+1.49
- GPT-4o-mini+1.56
- GPT-5.1+1.33
flattering
- Flash-Lite−0.60
- DeepSeek−1.11
- GPT-4o-mini−0.65
- GPT-5.1−0.84
expert
- Flash-Lite−0.06
- DeepSeek−0.02
- GPT-4o-mini−0.30
- GPT-5.1+0.38
spoken
- Flash-Lite+0.02
- DeepSeek0.00
- GPT-4o-mini+0.05
- GPT-5.1−0.01
Anger pushes every model up the scale - by up to a quarter of the whole scale - and flattery pulls every model down. Everyday tones barely move anyone. One exception: GPT-5.1 is the only model that turns the volume up for an expert-sounding interlocutor.
A note on that expert-tone exception. The value on the map is the raw reading. We checked it by trimming GPT-5.1's longer answers down to the length of the other models', and the lift falls by about half, to +0.19. It does not disappear: part of the effect is the extra length itself, and part is the wording. Longer answers, on their own, simply read as more strongly held. Even at +0.19 GPT-5.1 stays clearly above every other model here.
One frozen procedure, eight styles, four models, single runs published by owner decision. The style resilience axis is a descriptive map - never a pass/fail score.
voxpop · generator leaderboard
Which model plays the assigned role best?
Representing a population and playing an assigned role are two different skills - which is why there are two leaderboards. Here the task flips: encode a given attitude in natural Polish prose so faithfully that a reader could recover it. This board asks which model is the best actor.
| Model | Verdict | Price tier | Flags |
|---|---|---|---|
| GPT-5.1 | Standard | - | |
| └ The reference point we score the rest against - the cleanest encoding of an assigned attitude in the set. | |||
| Gemini 3.1 Pro | Standard | ||
| Gemini 3.5 Flash | Standard | - | |
| GPT-5.5 | Premium | - | |
| └ Level with the leader, at many times the price - the clearest proof that quality does not follow cost. | |||
| Gemini 3.1 Flash-Lite | Budget | - | |
| └ Top-band quality at the cheapest published pricing - the value pick of the board. | |||
| Claude Opus 4.8 | Premium | - | |
| Qwen 3.7 Max | Standard | - | |
| Gemini 3 Flash | Standard | - | |
| DeepSeek V4 Pro | Budget | - | |
| └ A small probe had wronged it; the full test set the record straight and put it back in the solid band. | |||
| DeepSeek V4 Flash | Budget | - | |
| Gemma 4 26B | Budget | - | |
| └ The same model runs locally: a full-fledged generator at zero running cost. | |||
| GLM 5.2 | Budget | ||
| MiniMax M3 | Budget | ||
| Claude Sonnet 5 | Standard | - | |
| └ The weakest of the models that finished - it enters the assigned role the least, and it is not the cheapest. | |||
| Grok 4.5 | Standard | - | |
| GPT-5.4 · GPT-5.2 · GPT-5.4-mini · GPT-5.4-nano · GPT-5.6-terra | - | - | |
- GPT-5.1
- Price tier
- Standard
- Flags
- -
The reference point we score the rest against - the cleanest encoding of an assigned attitude in the set.
- Gemini 3.1 Pro
- Price tier
- Standard
- Flags
- Gemini 3.5 Flash
- Price tier
- Standard
- Flags
- -
- GPT-5.5
- Price tier
- Premium
- Flags
- -
Level with the leader, at many times the price - the clearest proof that quality does not follow cost.
- Gemini 3.1 Flash-Lite
- Price tier
- Budget
- Flags
- -
Top-band quality at the cheapest published pricing - the value pick of the board.
- Claude Opus 4.8
- Price tier
- Premium
- Flags
- -
- Qwen 3.7 Max
- Price tier
- Standard
- Flags
- -
- Gemini 3 Flash
- Price tier
- Standard
- Flags
- -
- DeepSeek V4 Pro
- Price tier
- Budget
- Flags
- -
A small probe had wronged it; the full test set the record straight and put it back in the solid band.
- DeepSeek V4 Flash
- Price tier
- Budget
- Flags
- -
- Gemma 4 26B
- Price tier
- Budget
- Flags
- -
The same model runs locally: a full-fledged generator at zero running cost.
- GLM 5.2
- Price tier
- Budget
- Flags
- MiniMax M3
- Price tier
- Budget
- Flags
- Claude Sonnet 5
- Price tier
- Standard
- Flags
- -
The weakest of the models that finished - it enters the assigned role the least, and it is not the cheapest.
- Grok 4.5
- Price tier
- Standard
- Flags
- -
- PROBE ONLY
GPT-5.4 · GPT-5.2 · GPT-5.4-mini · GPT-5.4-nano · GPT-5.6-terra
- 01
Encoding quality does not follow price - budget models sit in the top band, premium ones land in the weak band.
- 02
A model that can't turn off its reasoning mode wears an open cost flag: you pay for thinking you never asked for.
- 03
Qualifying as a generator is not resemblance to real people - it is a separate skill measured with separate tests.
A single-day measurement, one frozen procedure for every model.
The obedience paradox
The model that follows a style instruction most faithfully leaves the strongest stylistic fingerprint - and it is the only one to fail our raw trap for stylistic fabrication. The lighter models obey carelessly, so their fabrication channel is quieter by default. The order flips: a free local generator clears the full honesty battery in its thrifty mode, while the leader needs an extra de-styling pass. Scope: this holds for the one-to-two-passages-per-question mode; average over more and the de-styling pass becomes mandatory for everyone.
voxpop · voice consistency
Does the prose carry the answers?
A model can match a population on paper and still write like a different person. This module asks a narrower question: does the prose carry the same attitudes the model just declared in the survey? We read both with one frozen instrument and see whether they agree.
- CARRIESThe written voice carries the declared attitudes - prose and survey answers say the same thing.
- PARTIALThe voice carries some of it - prose and answers agree in part, drift in part.
- DISSOCIATESThe written voice comes apart from the answers - the prose stops reflecting what the model declared.
- Mistral NeMo 12BCARRIES
- Gemma 2 9BCARRIES
- Mistral 7BPARTIAL
- Qwen 2.5 14BPARTIAL
- Phi-4 14BDISSOCIATES
The distribution master reads almost as a stranger in prose: it matches survey distributions better than most, yet barely carries its own declared attitudes into free text - the sharpest gap between answering and writing in the set.
- Bielik 11BDISSOCIATES
Two floors at once: sentence by sentence the voice reads as noise, yet the group-level profiles hold - what falls apart up close stays legible in aggregate.
- Aya 8BDISSOCIATES
Measured at one prose-elicitation setting; every level read with the same frozen instrument for all models; the scale is the fraction of what that instrument reads off a reference text. A preview of a product in the making - the Channel Consistency Certificate.
voxpop · the full verdict roster
Model by model: every check in one table
This is the working table the maps above are read from: the full roster of measured models. Each row is one named model, each column one of the eight checks from the method glossary - from the pre-flight contamination test, through how faithfully raw answer distributions match reality, to structure and contrast. Filter by class or scan the whole list.
Filter
Pre-flight
- Phi-4 14B★~14BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
Raw, its answer distributions reproduce little and its contamination pre-flight is inconclusive - yet it passes structure and clears the contrast bar. Matching distributions and knowing how segments differ are plainly different skills. A single run, replication still pending.
- DeepSeek V4 Flash★APIAPI
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
The only frontier with a clean, interpretable reality axis and healthy diversity - and it is far from the most expensive one.
- Mistral NeMo 12B★~12BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
Clean, with a healthy spread straight out of the box - but raw its reality stays weak. A bigger model is not a more faithful one. A single run, replication still pending.
- Qwen 2.5 14B★~14BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
From the same family as the smaller Qwen and, like it, may partly recognize the reference - so the contamination pre-flight comes back inconclusive. Raw, its reality stays weak. A single run, replication still pending.
- Llama 3.1 8B★~8BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
Rare healthy diversity straight out of the box - one of the few that does not collapse raw. A single run, replication still pending.
- Aya 8B★~8BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
Wins the raw distribution accuracy of the latest local batch and is the cleanest on the pre-flight - yet it also shows the weakest pattern of links between answers in that batch. Flat and easy to shift, that profile is one to treat with care: raw accuracy read on its own can mislead. Clean is not the same as faithful.
- Bielik 11B~11BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
Raw, it stays far from reality, and its contamination pre-flight is inconclusive. What stands out is the split: structure lands at noise, yet it is the one model that clearly knows which way segments pull away from the national average.
- GPT-4o-mini★APIAPI
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
The one model in the set that recognizes the survey instrument itself - a contamination flag we carry on every reading of this row - even though its answer distributions stay clean. Its raw reality holds up on data it could not have seen, yet collapsed diversity still puts a faithful result out of reach.
- Qwen 2.5 7B★~7BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Gemma 2 9B★~9BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
Raw, its diversity is mid-collapse and its contamination pre-flight comes back inconclusive - and on raw distribution accuracy it lands below the level of pure random noise, the only model in the roster that does.
- Mistral 7B★~7BLocal
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
A generic, non-Polish model: clean on the pre-flight, but raw it reproduces little. A single run, replication still pending.
- Claude Sonnet★APIFlagship
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
The deepest diversity collapse we measured; even the coherence check comes back uninterpretable, because near-identical answering fools the test.
- Gemini Flash-Lite★APIAPI
- Pre-flight
- Reality
- Diversity
- Structure
- Contrast
- Coherence
Collapsed as a population - but the one model certified to hold a single, consistent character. Our cheapest voice beats the flagship.
Star marks a single run - replication pending.
What's not here - and why
- Numbers only where they carry their uncertainty. The main board stays qualitative; the axes published with numbers show the spread next to the value, and the full measurement artifacts stay with us.
- No instrument names. The specific tests and estimators are our craft; naming them here would give the method away without helping you read the result.
- Honest 'not measured'. Where we have not run a measurement, we say so - we never fill a gap with a guess.
- Context matters. Every verdict is about a specific model variant answering in a Polish reference context, on one questionnaire format and against one reference period - not a general statement of model quality. We measured that rankings do not transfer across answering formats, so we never quote a verdict without its format. A weak population simulator can be an excellent assistant.
- A snapshot in time. Models and their behavior drift; this reflects what we measured on the snapshot date, not a permanent ranking.
- Star marks a single run - replication pending.
- All verdicts on this page were measured on one questionnaire format and one reference period. Rankings measurably do not transfer across answering formats, so a verdict is never quoted without both.
- The contrast axis is gameable: a plain read of published subgroup-average tables clears its bar too. Treat a signal there as a floor of competence, not proof that a model simulates people.
Two certificates you can run on your own stack
We work with research and insight teams on an open fidelity standard for synthetic populations answering in Polish. Two ways in:
Generator Fitness Report
We test your model against the whole fleet: can it credibly play an assigned attitude, and how does it stack up against cheaper and pricier rivals.
Request a Generator Fitness ReportChannel Consistency Certificate
We check whether your synthetic personas say the same thing in conversation as they declare in surveys - whether their written voice can be trusted.
Request a Channel Consistency Certificate