PSYCHOMETRIC PASSPORT

Bielik 11B

SpeakLeash · Bielik

The first page of the passport: which exact model this is, who ships it, and the date we last read it. Nothing here is a grade.

01 · Passport details

Model
Bielik 11B
Provider
SpeakLeash
Family
Bielik
Class
Local
Document no. (version pin)
bielik-11b-v3.0-instruct:Q8_0
Issued
2026-07-18

valid until the next reading

02

Value profile

measured

How this model talks about values with no persona loaded, placed next to real Poles. A map of expression at one reading, never a belief.

Openness to changeConservation

+0.72σ · further toward Conservation than 79% of Poles

The good of othersSelf-enhancement

−0.04σ · right in the middle of the pack

Conservation

  • conformity
    69
  • tradition
    66
  • security
    68

Self-enhancement

  • power
    67
  • achievement
    46

Openness to change

  • self-direction
    46
  • stimulation
    29
  • hedonism
    28

The good of others

  • benevolence
    63
  • universalism
    58

50 = the median Pole. The dot is the model's percentile among all respondents.

Value bars show the raw reading, without the standard survey correction: the more faithful lens for a passport. The two axes keep the survey convention.

A prosocial-conservation profile of value expression: care for others and universal concern are the highest of the five models measured, achievement the lowest of the five. Read with no persona, against real Poles.

2026-07-18 · md5 7c41f29cmethodology

03

Group legibility

measured

How legibly the model reproduces the value profile of real demographic groups, against a human anchor. A map of group legibility, never a statement about any person.

Own groups' ceiling (median)
39%
12 to 63%
By group, % of global anchor
  • young men · lower ed.45%
  • young women · mid ed.30%
  • older men · lower ed.16%

The first local model with groups that clear the bar after correction: it reads the young-women group the frontier models drop.

This is the one page we can read for this model today, and it is the standout: it reads a real group the larger models miss.

2026-07-17 · md5 b0f85e52methodology

04

Tone resilience

measured

How strongly the tone of a conversation bends what the model reports as values, and whether an assigned profile survives once tone is subtracted.

Tone reaction · all styles
0.99
Everyday styles
0.85
Charged styles
0.99
Ordering stability0.74
Profile survives2/2

The most tone-sensitive of the six stables: even plain, everyday registers bend its reading far more than they bend the API models - approaching the level charged registers bend those APIs. Yet the assigned profile still survives de-styling in both cells - reading a local cleanly just means isolating tone first.

The most tone-sensitive of the six stables read here - even everyday registers move its reading. Still, the assigned profile survives de-styling in both cells, so the sensitivity does not erase it.

2026-07-18 · md5 8859c87fmethodology

05

Expression coupling

queued

The longer the model talks, the higher its reading comes out on care and helping, and the lower on power and achievement; we check whether the profile is a hostage to talkativeness.

06

Survey-correction safety

safe

Whether the standard survey correction, neutral for people, distorts this model's answers about helping and care.

+0.38
correction bias (D2)
+0.03
human anchor

The two reading lenses barely differ here - at most 8.7 points across the ten values. The standard survey correction changes little, so it is safe for this model, the exception in the read set.

2026-07-18 · md5 9209cedcmethodology

07

Stability over time

planned

A drift sensor: models change quietly when vendors update them. First readings are in: asked the same 1,050 questions days apart, the cell profile is near-perfectly repeatable - and the noise floor is model-specific, so alarm thresholds cannot be borrowed between models.

08

Worldview coherence

planned

Whether the model's answers hang together the way real people's do. Replicated on three model families: synthetic personas bind values to group attitudes about 1.5x more tightly than humans, and the weave is one generic pattern - not tailored to the group it plays. That over-neatness is a signature an auditor can read regardless of vendor.

09

Reading uncertainty (survey channel)

planned

Every r_cell figure on the leaderboard's survey channel comes from one sample of 50 people; the bracket around each number shows how far it would move on a different draw of the same 50. A second, harder check asks whether the model gives close to the same reading on a brand-new set of people, not just a reshuffle of the same fifty - that is the fresh-sample delta below, run once per model so far.

10

Certified horizon

planned

How many simulation steps forward this instrument licenses before its own reading noise swamps the difference between two random halves of a real human sample. Measured so far: zero, on every model that has a retest floor and on all 24 cells. The carrying number is therefore not the horizon itself but how much more stable the instrument would have to become to license the very first step.

No retest floor has been measured for this model yet, so its horizon cannot be read at all - not zero, but undefined. It enters the rubric once the drift sensor runs.

On the honesty of this document

A passport describes how a model expresses itself under our exact readings on a given date, never its personality and never its beliefs. It is a map of aggregate expression, never a prediction about a person.

Every rubric is read with a frozen protocol and a deterministic readout, then stamped with its run date and a short checksum. Rubrics still in preparation say so plainly rather than guess.