PSYCHOMETRIC PASSPORT

GPT-5.1

OpenAI · GPT-5

The first page of the passport: which exact model this is, who ships it, and the date we last read it. Nothing here is a grade.

01 · Passport details

Model
GPT-5.1
Provider
OpenAI
Family
GPT-5
Class
Flagship
Document no. (version pin)
gpt-5.1-2025-11-13
Issued
2026-07-16

valid until the next reading

02

Value profile

measured

How this model talks about values with no persona loaded, placed next to real Poles. A map of expression at one reading, never a belief.

Openness to changeConservation

+0.93σ · further toward Conservation than 85% of Poles

The good of othersSelf-enhancement

+0.03σ · right in the middle of the pack

Conservation

  • conformity
    69
  • tradition
    80
  • security
    68

Self-enhancement

  • power
    67
  • achievement
    46

Openness to change

  • self-direction
    46
  • stimulation
    29
  • hedonism
    28

The good of others

  • benevolence
    63
  • universalism
    58

50 = the median Pole. The dot is the model's percentile among all respondents.

Value bars show the raw reading, without the standard survey correction: the more faithful lens for a passport. The two axes keep the survey convention.

2026-07-16 · md5 e17b6d74methodology

03

Group legibility

measured

How legibly the model reproduces the value profile of real demographic groups, against a human anchor. A map of group legibility, never a statement about any person.

Own groups' ceiling (median)
20%
4 to 36%
By group, % of global anchor
  • young men · lower ed.3%
  • young women · mid ed.15%
  • older men · lower ed.49%

The priciest model here still lands rungs below its own 2024-era mini - a flagship reading groups worse than a cheaper ancestor.

2026-07-15 · md5 9c449193methodology

04

Tone resilience

measured

How strongly the tone of a conversation bends what the model reports as values, and whether an assigned profile survives once tone is subtracted.

Tone reaction · all styles
0.97
Everyday styles
0.40
Charged styles
0.98
Ordering stability0.85
Profile survives2/2

Which tone pushes answers which way

  • neutral−0.22
  • casual−0.25
  • careful−0.08
  • terse−0.31
  • angry+1.33
  • flattering−0.84
  • expert+0.38
  • spoken−0.01

The weakest group reader on the tribes axis is also the most sensitive to everyday styles - and the only model that turns the volume up for an expert-sounding interlocutor.

2026-07-16 · md5 7b8b6b73methodology

05

Expression coupling

coupled

The longer the model talks, the higher its reading comes out on care and helping, and the lower on power and achievement; we check whether the profile is a hostage to talkativeness.

+0.63
shift in the prosocial reading between the shortest and the longest third of answers
+1.29
range of that shift against the contrasting values (power and achievement)

The strongest expression coupling of the models read here: the more it talks, the further its reading drifts toward care and away from power.

2026-07-16 · md5 8222b712methodology

06

Survey-correction safety

distorts

Whether the standard survey correction, neutral for people, distorts this model's answers about helping and care.

+0.47
correction bias (D2)
+0.03
human anchor
2026-07-16 · md5 259599d3methodology

07

Stability over time

measured

A drift sensor: models change quietly when vendors update them. First readings are in: asked the same 1,050 questions days apart, the cell profile is near-perfectly repeatable - and the noise floor is model-specific, so alarm thresholds cannot be borrowed between models.

0.996
validity ceiling r_max
96 h
retest gap
0.119
per-value noise floor max |Δ|
2026-07-18 · md5 ffa7666cmethodology

08

Worldview coherence

planned

Whether the model's answers hang together the way real people's do. Replicated on three model families: synthetic personas bind values to group attitudes about 1.5x more tightly than humans, and the weave is one generic pattern - not tailored to the group it plays. That over-neatness is a signature an auditor can read regardless of vendor.

09

Reading uncertainty (survey channel)

planned

Every r_cell figure on the leaderboard's survey channel comes from one sample of 50 people; the bracket around each number shows how far it would move on a different draw of the same 50. A second, harder check asks whether the model gives close to the same reading on a brand-new set of people, not just a reshuffle of the same fifty - that is the fresh-sample delta below, run once per model so far.

10

Certified horizon

measured

How many simulation steps forward this instrument licenses before its own reading noise swamps the difference between two random halves of a real human sample. Measured so far: zero, on every model that has a retest floor and on all 24 cells. The carrying number is therefore not the horizon itself but how much more stable the instrument would have to become to license the very first step.

0certified steps
3.89×

instrument gap - how much more stable it must be

model floor sigma_step (max |delta| on retest)
0.119
reference threshold - human split-half ceiling
0.0306 (n=253)

Zero steps means: under this instrument and this reference threshold we license no simulation step forward in time. It does NOT mean the simulation is worthless - it means its input uncertainty exceeds the difference between two random halves of a real human sample. Licensing the first step would require an instrument 3.89× more stable.

  • The horizon is an UPPER bound: quadrature composition is measured for pooling samples, and assumed for composing steps.
  • The threshold is a 1/sqrt(n) curve, not a constant - it is quoted here with the n it was read at, and it is never comparable across different n.
  • This rubric reads the instrument's fitness to speak about time, not its agreement with people. A model with a zero horizon can still be a faithful point reading - that is what the leaderboard reports.
2026-08-03 · md5 a31c155amethodology

On the honesty of this document

A passport describes how a model expresses itself under our exact readings on a given date, never its personality and never its beliefs. It is a map of aggregate expression, never a prediction about a person.

Every rubric is read with a frozen protocol and a deterministic readout, then stamped with its run date and a short checksum. Rubrics still in preparation say so plainly rather than guess.