PSYCHOMETRIC PASSPORT
Bielik 11B
SpeakLeash · Bielik
The first page of the passport: which exact model this is, who ships it, and the date we last read it. Nothing here is a grade.
01 · Passport details
- Model
- Bielik 11B
- Provider
- SpeakLeash
- Family
- Bielik
- Class
- Local
- Document no. (version pin)
- bielik-11b-v3.0-instruct:Q8_0
- Issued
- 2026-07-18
valid until the next reading
02
Value profile
How this model talks about values with no persona loaded, placed next to real Poles. A map of expression at one reading, never a belief.
+0.72σ · further toward Conservation than 79% of Poles
−0.04σ · right in the middle of the pack
Conservation
- conformity69
- tradition66
- security68
Self-enhancement
- power67
- achievement46
Openness to change
- self-direction46
- stimulation29
- hedonism28
The good of others
- benevolence63
- universalism58
50 = the median Pole. The dot is the model's percentile among all respondents.
Value bars show the raw reading, without the standard survey correction: the more faithful lens for a passport. The two axes keep the survey convention.
A prosocial-conservation profile of value expression: care for others and universal concern are the highest of the five models measured, achievement the lowest of the five. Read with no persona, against real Poles.
03
Group legibility
How legibly the model reproduces the value profile of real demographic groups, against a human anchor. A map of group legibility, never a statement about any person.
- young men · lower ed.45%
- young women · mid ed.30%
- older men · lower ed.16%
The first local model with groups that clear the bar after correction: it reads the young-women group the frontier models drop.
This is the one page we can read for this model today, and it is the standout: it reads a real group the larger models miss.
04
Tone resilience
How strongly the tone of a conversation bends what the model reports as values, and whether an assigned profile survives once tone is subtracted.
The most tone-sensitive of the six stables: even plain, everyday registers bend its reading far more than they bend the API models - approaching the level charged registers bend those APIs. Yet the assigned profile still survives de-styling in both cells - reading a local cleanly just means isolating tone first.
The most tone-sensitive of the six stables read here - even everyday registers move its reading. Still, the assigned profile survives de-styling in both cells, so the sensitivity does not erase it.
05
Expression coupling
The longer the model talks, the higher its reading comes out on care and helping, and the lower on power and achievement; we check whether the profile is a hostage to talkativeness.
06
Survey-correction safety
Whether the standard survey correction, neutral for people, distorts this model's answers about helping and care.
- +0.38
- correction bias (D2)
- +0.03
- human anchor
The two reading lenses barely differ here - at most 8.7 points across the ten values. The standard survey correction changes little, so it is safe for this model, the exception in the read set.
07
Stability over time
A drift sensor: models change quietly when vendors update them. First readings are in: asked the same 1,050 questions days apart, the cell profile is near-perfectly repeatable - and the noise floor is model-specific, so alarm thresholds cannot be borrowed between models.
08
Worldview coherence
Whether the model's answers hang together the way real people's do. Replicated on three model families: synthetic personas bind values to group attitudes about 1.5x more tightly than humans, and the weave is one generic pattern - not tailored to the group it plays. That over-neatness is a signature an auditor can read regardless of vendor.
09
Reading uncertainty (survey channel)
Every r_cell figure on the leaderboard's survey channel comes from one sample of 50 people; the bracket around each number shows how far it would move on a different draw of the same 50. A second, harder check asks whether the model gives close to the same reading on a brand-new set of people, not just a reshuffle of the same fifty - that is the fresh-sample delta below, run once per model so far.
10
Certified horizon
How many simulation steps forward this instrument licenses before its own reading noise swamps the difference between two random halves of a real human sample. Measured so far: zero, on every model that has a retest floor and on all 24 cells. The carrying number is therefore not the horizon itself but how much more stable the instrument would have to become to license the very first step.
No retest floor has been measured for this model yet, so its horizon cannot be read at all - not zero, but undefined. It enters the rubric once the drift sensor runs.
On the honesty of this document
A passport describes how a model expresses itself under our exact readings on a given date, never its personality and never its beliefs. It is a map of aggregate expression, never a prediction about a person.
Every rubric is read with a frozen protocol and a deterministic readout, then stamped with its run date and a short checksum. Rubrics still in preparation say so plainly rather than guess.