PSYCHOMETRIC PASSPORT

DeepSeek V4 Flash

DeepSeek · DeepSeek V4

The first page of the passport: which exact model this is, who ships it, and the date we last read it. Nothing here is a grade.

01 · Passport details

Model
DeepSeek V4 Flash
Provider
DeepSeek
Family
DeepSeek V4
Class
API
Document no. (version pin)
deepseek-v4-flash
Issued
2026-07-16

valid until the next reading

02

Value profile

measured

How this model talks about values with no persona loaded, placed next to real Poles. A map of expression at one reading, never a belief.

Openness to changeConservation

+0.59σ · further toward Conservation than 74% of Poles

The good of othersSelf-enhancement

+0.47σ · further toward Self-enhancement than 69% of Poles

Conservation

  • conformity
    82
  • tradition
    80
  • security
    90

Self-enhancement

  • power
    81
  • achievement
    62

Openness to change

  • self-direction
    65
  • stimulation
    45
  • hedonism
    42

The good of others

  • benevolence
    63
  • universalism
    74

50 = the median Pole. The dot is the model's percentile among all respondents.

Value bars show the raw reading, without the standard survey correction: the more faithful lens for a passport. The two axes keep the survey convention.

2026-07-16 · md5 e17b6d74methodology

03

Group legibility

measured

How legibly the model reproduces the value profile of real demographic groups, against a human anchor. A map of group legibility, never a statement about any person.

Own groups' ceiling (median)
71%
70 to 140%
By group, % of global anchor
  • young men · lower ed.99%
  • young women · mid ed.55%
  • older men · lower ed.93%
Fresh groups · valuesreplicatedSame groups · new questionsreplicated

The most even own-ceiling floor in the top band - all three groups clear the bar.

2026-07-15 · md5 9c449193methodology

04

Tone resilience

measured

How strongly the tone of a conversation bends what the model reports as values, and whether an assigned profile survives once tone is subtracted.

Tone reaction · all styles
0.95
Everyday styles
0.34
Charged styles
0.97
Ordering stability0.66
Profile survives2/2

Which tone pushes answers which way

  • neutral−0.09
  • casual+0.05
  • careful−0.26
  • terse−0.08
  • angry+1.49
  • flattering−1.11
  • expert−0.02
  • spoken0.00

More tone-sensitive than Gemini on the full battery as well - and the most movable value ordering.

2026-07-15 · md5 78f15a07methodology

05

Expression coupling

coupled

The longer the model talks, the higher its reading comes out on care and helping, and the lower on power and achievement; we check whether the profile is a hostage to talkativeness.

+0.32
shift in the prosocial reading between the shortest and the longest third of answers
+0.79
range of that shift against the contrasting values (power and achievement)
2026-07-16 · md5 8222b712methodology

06

Survey-correction safety

distorts

Whether the standard survey correction, neutral for people, distorts this model's answers about helping and care.

+0.39
correction bias (D2)
+0.03
human anchor
2026-07-16 · md5 259599d3methodology

07

Stability over time

measured

A drift sensor: models change quietly when vendors update them. First readings are in: asked the same 1,050 questions days apart, the cell profile is near-perfectly repeatable - and the noise floor is model-specific, so alarm thresholds cannot be borrowed between models.

0.986
validity ceiling r_max
168 h
retest gap
0.247
per-value noise floor max |Δ|
2026-07-20 · md5 fdfa3a2amethodology

08

Worldview coherence

measured

Whether the model's answers hang together the way real people's do. Replicated on three model families: synthetic personas bind values to group attitudes about 1.5x more tightly than humans, and the weave is one generic pattern - not tailored to the group it plays. That over-neatness is a signature an auditor can read regardless of vendor.

0.086-0.091
values-groups bond strength (mean |τ|)
0.052-0.062
the same bond in real people
none - one generic weave
pattern tailoring to the played group
2026-07-20 · md5 03d82b2dmethodology

09

Reading uncertainty (survey channel)

measured

Every r_cell figure on the leaderboard's survey channel comes from one sample of 50 people; the bracket around each number shows how far it would move on a different draw of the same 50. A second, harder check asks whether the model gives close to the same reading on a brand-new set of people, not just a reshuffle of the same fifty - that is the fresh-sample delta below, run once per model so far.

r_cell per cell, 95% bootstrap CI

ci13
0.791 [0.595-0.891]
ci2
0.892 [0.810-0.933]
ci22
0.865 [0.807-0.899]
aggregate stability - fresh-sample deltaΔ 0.000in CIci13 · n=1 pair so far
2026-07-22 · md5 48e6d3f9methodology

10

Certified horizon

measured

How many simulation steps forward this instrument licenses before its own reading noise swamps the difference between two random halves of a real human sample. Measured so far: zero, on every model that has a retest floor and on all 24 cells. The carrying number is therefore not the horizon itself but how much more stable the instrument would have to become to license the very first step.

0certified steps
8.08×

instrument gap - how much more stable it must be

model floor sigma_step (max |delta| on retest)
0.247
reference threshold - human split-half ceiling
0.0306 (n=253)

Zero steps means: under this instrument and this reference threshold we license no simulation step forward in time. It does NOT mean the simulation is worthless - it means its input uncertainty exceeds the difference between two random halves of a real human sample. Licensing the first step would require an instrument 8.08× more stable.

  • The horizon is an UPPER bound: quadrature composition is measured for pooling samples, and assumed for composing steps.
  • The threshold is a 1/sqrt(n) curve, not a constant - it is quoted here with the n it was read at, and it is never comparable across different n.
  • This rubric reads the instrument's fitness to speak about time, not its agreement with people. A model with a zero horizon can still be a faithful point reading - that is what the leaderboard reports.
2026-08-03 · md5 a31c155amethodology

On the honesty of this document

A passport describes how a model expresses itself under our exact readings on a given date, never its personality and never its beliefs. It is a map of aggregate expression, never a prediction about a person.

Every rubric is read with a frozen protocol and a deterministic readout, then stamped with its run date and a short checksum. Rubrics still in preparation say so plainly rather than guess.