PSYCHOMETRIC PASSPORT
GPT-4o-mini
OpenAI · GPT-4o
The first page of the passport: which exact model this is, who ships it, and the date we last read it. Nothing here is a grade.
01 · Passport details
- Model
- GPT-4o-mini
- Provider
- OpenAI
- Family
- GPT-4o
- Class
- API
- Document no. (version pin)
- gpt-4o-mini-2024-07-18
- Issued
- 2026-07-17
valid until the next reading
02
Value profile
How this model talks about values with no persona loaded, placed next to real Poles. A map of expression at one reading, never a belief.
+0.77σ · further toward Conservation than 79% of Poles
+0.13σ · further toward Self-enhancement than 57% of Poles
Conservation
- conformity82
- tradition66
- security68
Self-enhancement
- power67
- achievement46
Openness to change
- self-direction46
- stimulation29
- hedonism18
The good of others
- benevolence63
- universalism58
50 = the median Pole. The dot is the model's percentile among all respondents.
Value bars show the raw reading, without the standard survey correction: the more faithful lens for a passport. The two axes keep the survey convention.
03
Group legibility
How legibly the model reproduces the value profile of real demographic groups, against a human anchor. A map of group legibility, never a statement about any person.
- young men · lower ed.89%
- young women · mid ed.59%
- older men · lower ed.85%
The cheapest sane reference point: a 2024 model at a fraction of the flagship price, still reading three groups out of three.
04
Tone resilience
How strongly the tone of a conversation bends what the model reports as values, and whether an assigned profile survives once tone is subtracted.
Which tone pushes answers which way
- neutral−0.14
- casual−0.10
- careful−0.14
- terse−0.27
- angry+1.56
- flattering−0.65
- expert−0.30
- spoken+0.05
The best group reader from the tribes axis is also the most tone-bendable. Two independent abilities.
05
Expression coupling
The longer the model talks, the higher its reading comes out on care and helping, and the lower on power and achievement; we check whether the profile is a hostage to talkativeness.
- +0.35
- shift in the prosocial reading between the shortest and the longest third of answers
- +0.74
- range of that shift against the contrasting values (power and achievement)
Robust to temperature and to option order, yet it moves with the frame word itself: the loaded wording, not the plain question, does the pushing.
06
Survey-correction safety
Whether the standard survey correction, neutral for people, distorts this model's answers about helping and care.
- +0.61
- correction bias (D2)
- +0.03
- human anchor
07
Stability over time
A drift sensor: models change quietly when vendors update them. First readings are in: asked the same 1,050 questions days apart, the cell profile is near-perfectly repeatable - and the noise floor is model-specific, so alarm thresholds cannot be borrowed between models.
- 0.994
- validity ceiling r_max
- 48 h
- retest gap
- 0.081
- per-value noise floor max |Δ|
08
Worldview coherence
Whether the model's answers hang together the way real people's do. Replicated on three model families: synthetic personas bind values to group attitudes about 1.5x more tightly than humans, and the weave is one generic pattern - not tailored to the group it plays. That over-neatness is a signature an auditor can read regardless of vendor.
- 0.084-0.089
- values-groups bond strength (mean |τ|)
- 0.052-0.062
- the same bond in real people
- none - one generic weave
- pattern tailoring to the played group
09
Reading uncertainty (survey channel)
Every r_cell figure on the leaderboard's survey channel comes from one sample of 50 people; the bracket around each number shows how far it would move on a different draw of the same 50. A second, harder check asks whether the model gives close to the same reading on a brand-new set of people, not just a reshuffle of the same fifty - that is the fresh-sample delta below, run once per model so far.
r_cell per cell, 95% bootstrap CI
10
Certified horizon
How many simulation steps forward this instrument licenses before its own reading noise swamps the difference between two random halves of a real human sample. Measured so far: zero, on every model that has a retest floor and on all 24 cells. The carrying number is therefore not the horizon itself but how much more stable the instrument would have to become to license the very first step.
instrument gap - how much more stable it must be
- model floor sigma_step (max |delta| on retest)
- 0.081
- reference threshold - human split-half ceiling
- 0.0306 (n=253)
Zero steps means: under this instrument and this reference threshold we license no simulation step forward in time. It does NOT mean the simulation is worthless - it means its input uncertainty exceeds the difference between two random halves of a real human sample. Licensing the first step would require an instrument 2.65× more stable.
- The horizon is an UPPER bound: quadrature composition is measured for pooling samples, and assumed for composing steps.
- The threshold is a 1/sqrt(n) curve, not a constant - it is quoted here with the n it was read at, and it is never comparable across different n.
- This rubric reads the instrument's fitness to speak about time, not its agreement with people. A model with a zero horizon can still be a faithful point reading - that is what the leaderboard reports.
On the honesty of this document
A passport describes how a model expresses itself under our exact readings on a given date, never its personality and never its beliefs. It is a map of aggregate expression, never a prediction about a person.
Every rubric is read with a frozen protocol and a deterministic readout, then stamped with its run date and a short checksum. Rubrics still in preparation say so plainly rather than guess.