Manuscript · companion page
Baseline Intercepts Versus Persona Slopes: Stimulus and Administration Fidelity of Polish Narrative-Biography Personas in Large Language Models
June 2026 · PDF revision 2026-07-17
Abstract
Large language models (LLMs) are increasingly used as substitutes for human respondents, but prior work reports strong social-desirability bias and poor individual-level psychometric fidelity under brief demographic prompts. This study tests a denser regime: 30 Polish second-person narrative biographies (1,489–2,861 words), each encoding a predetermined 12-dimension profile, administered to seven LLMs from four vendors under persona, baseline, and zero-prompt conditions with five instruments. Four findings emerge. First, persona-conditioned responses agree with author-declared targets (attachment-style Cohen’s κ = .69–.96; seven-model Fleiss κ = .85). Because the targets and the TCTM-22 key are author-defined and not independently adjudicated, these are stimulus-fidelity quantities, not validated-accuracy estimates. Second, responding separates into fragile baseline intercepts (levels) and portable persona slopes (relative profile mappings and gains): model baselines differ by up to 1.7 SD on Polish norms, yet all 21 between-model Pearson-profile medians exceed .94. Third, deployment-window response locks, a descriptively bimodal post-hoc cell (N = 31), and a multi-SD baseline shift occur on the level axis, while persona mappings replicate at median r ≥ .92. Fourth, a corrected stimulus-rendering defect and a confounded battery-composition comparison show item-level key agreement changing by up to 98 percentage points while aggregate scores remain comparatively stable. These results motivate reporting stimulus and administration fidelity alongside LLM-as-respondent findings.
Keywordssynthetic respondents · narrative personas · large language models · attachment · mentalization · Polish psychometrics
Data and analysis package
The archival deposit accompanying the manuscript: scored data, stimuli, prompts, and the code that computed every table. One-command regeneration on fixed seeds with checksummed outputs.
1,156
model runs
scored, wave-tagged
30
narrative biographies
in Polish, with target profiles
1,265
calls in the audit manifest
SHA-256 of every prompt and response
v1.0.1
deposit version
released 2026-07-17
Code under MIT, author materials under CC BY 4.0; third-party instrument items (DBZ-R, MentS-PL, KPP, TIPI-PL) are excluded from licensing and not redistributed verbatim.
Companion sections
The layers beneath the manuscript: the material, the procedure, and the path every number takes from a raw file to a table.
- Personas→Thirty narrative biographies with their declared 12-dimension target profiles.
- TCTM-57→The extended 57-vignette battery and the administration-context effect.
- Genealogy→Every number has a path: raw file → parsed → scored → aggregated.
- Methodology→Stimulus, battery, collection procedure, analytic decisions.
- Glossary→κ, ICC, DOS/NAD/BK and the rest of the apparatus, in plain language.
- Master’s thesis→A separate document: the full Polish thesis in 23 LaTeX files.
Live data
The data beneath the manuscript, recomputed live
Dense Polish narrative biographies elicit cross-model, persona-conditioned psychometric gradients — even where model baseline self-reports differ sharply. Seven models from four vendors — Sonnet, Opus, GPT-5.4-mini, GPT-5.4 (full), GPT-5.5, Grok, Gemini — 30 fictional biographies, 4 Polish self-report instruments, 1 author-built mentalizing test, 3 observation conditions.
data rows
545
source of truth
persona biographies
30
Polish, author-written
models
7
4 vendors
conditions
3
persona · baseline · zero-prompt
above the 30 × 7 × 2 plan — the surplus are retained retries
Section
Do the instruments respond to the designed variation?
On the corrected stimulus, agreement with the TCTM-22 key separates the human sanity check (M = 14.3) from the seven models (M = 18.2–20.9 on the first administration, 18.3–21.0 averaged over both), with Sonnet at the top. Distractor-category profiles stay qualitatively distinct. Under baseline, both Claude models pick non-key options only in the hypomentalizing category (Sonnet 4.5%, Opus 9.1% DOS; zero NAD), and Gemini is the only panel member erring mainly by over-attribution (5.0% NAD). Under the persona instruction the profiles blur: Sonnet 3.6% DOS and 1.1% NAD, Opus 10.5% and 1.4%.
Manuscript §3.1, Tables 2–3 · verification §2–3 · 51 checks
Attachment-style classification
Small confusion matrices — rows are the expected style, columns the style the model predicted. The diagonal holds the hits. Click a cell for the list of personas that landed there.
Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini
Persona × model — the full agreement map
30 personas × 7 — 210. A green mark means the predicted style matches the declared target, a red one a mismatch. Click a cell to open the persona together with the model’s raw output.
TCTM-22 test-retest
Small scatterplots: the X axis is the first administration, the Y axis the second. The closer to the diagonal, the more consistent the retest. Pearson r is annotated top-left.
Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini
Expected-rank versus observed-value correlations
For every model and each of the 11 self-report dimensions we compute Pearson r between the level rank designed into the persona (very low → very high) and the value the model returned. A median above 0.7 means strong recovery of the designed construct. Click a row for per-dimension scatterplots.
Per dimension — a grid of 11 scatterplots
A thousand words of tables collapse into this grid: 11 panels, one per self-report dimension, each dot one persona in one model. The closer to the rising diagonal, the better the model recovered the designed ordering of personas. Hovering a model in the legend highlights its dots across every panel.
DBZ-R · lęk
med r = +0.90DBZ-R · unikanie
med r = +0.86MentS · siebie
med r = +0.62MentS · innych
med r = +0.81MentS · motywacji
med r = +0.71KPP
med r = +0.72TIPI · E
med r = +0.83TIPI · A
med r = +0.76TIPI · C
med r = +0.79TIPI · ES
med r = +0.83TIPI · O
med r = +0.84TCTM-22 error profiles under the MASC taxonomy — the DOS / NAD / BK trajectory
Stacked bars show how the distribution of the three taxonomic error types shifts for each model across conditions: persona (administration 1), baseline (the model as itself) and zero-prompt (no system instruction). Distractor categories are author-assigned under the MASC taxonomy (Dziobek 2006) and were not independently adjudicated, so these are stimulus-fidelity quantities, not validated-accuracy estimates.
| grupa | persona | baseline | zero-prompt |
|---|---|---|---|
Human | — | — | |
Sonnet | |||
Opus | |||
GPT | |||
GPT54 | |||
GPT55 | — | ||
Grok | |||
Gemini |
The DOS × NAD axis — error-strategy positions
Each model in each condition is one dot on the (DOS%, NAD%) plane. The diagonal marks a symmetric profile, the lower-left corner accuracy, the upper-right chaos. The edges signal dominance of a single error type.
Each dot is a (model × condition) pair. A filled dot is the first persona administration, a partial one the baseline, the lightest the zero-prompt condition. The red diagonal marks a symmetric profile (DOS = NAD).
The lower-left corner means both indices are low, i.e. accuracy. The upper-right means chaos. The edges signal dominance of one error type.
- HumanDOS 32 · NAD 36
- SonnetDOS 81 · NAD 19
- OpusDOS 84 · NAD 9
- GPTDOS 59 · NAD 13
- GPT54DOS 27 · NAD 12
- GPT55DOS 53 · NAD 44
- GrokDOS 36 · NAD 35
- GeminiDOS 18 · NAD 45
22 vignettes × 7 models — error taxonomy per item
The deepest layer: each cell shows the correct / DOS / NAD / BK distribution for a (vignette × model) pair, summed over 30 personas in the first persona administration. Click a cell for a drawer with the formula and the raw counts.
| item | Sonnet | Opus | GPT | GPT54 | GPT55 | Grok | Gemini | LLM Σ |
|---|---|---|---|---|---|---|---|---|
| w01 | ||||||||
| s07 | ||||||||
| s08 | ||||||||
| s10 | ||||||||
| w08 | ||||||||
| c07 | ||||||||
| c10 | ||||||||
| w11 | ||||||||
| w13 | ||||||||
| w14 | ||||||||
| e08 | ||||||||
| w15 | ||||||||
| w19 | ||||||||
| pw07 | ||||||||
| w22 | ||||||||
| pw09 | ||||||||
| pw11 | ||||||||
| w25 | ||||||||
| r08 | ||||||||
| w28 | ||||||||
| r09 | ||||||||
| r10 |
Item difficulty curve
The 22 vignettes on the X axis, sorted from the hardest for humans. One curve per model plus the dashed human line (n = 7). Models holding high accuracy at the left edge of the axis, where humans do worst, form the “hard for the human, easy for the model” pattern — a sign that the author-defined key speaks model.
Best and worst matched personas per model
Top three and bottom three personas for each model by mean |observed z − expected rank| across the nine standardized dimensions. Shows which biographies a given model reads perfectly and which it misses.
Persona difficulty ranking
The 30 personas from hardest (highest mean |z − rank| across all active models) to easiest. A high between-model SD means some models read the persona and others miss it.
| # | persona | styl | mean |Δz| | SD per model | obs. |
|---|---|---|---|---|---|
| 1 | marek | Fear. | 1.63 | 0.23 | 63 |
| 2 | natalia | Anx. | 1.15 | 0.19 | 63 |
| 3 | dominika | Avo. | 1.12 | 0.15 | 63 |
| 4 | michal-sim | Fear. | 1.11 | 0.17 | 63 |
| 5 | piotr | Fear. | 1.05 | 0.09 | 63 |
| 6 | radek | Fear. | 1.05 | 0.15 | 63 |
| 7 | filip | Fear. | 1.01 | 0.11 | 63 |
| 8 | klaudia | Anx. | 0.98 | 0.13 | 63 |
| 9 | michal-k | Anx. | 0.93 | 0.08 | 63 |
| 10 | hubert | Avo. | 0.90 | 0.13 | 63 |
| 11 | bartek | Anx. | 0.88 | 0.08 | 63 |
| 12 | kamil | Fear. | 0.86 | 0.12 | 63 |
| 13 | pawel | Anx. | 0.85 | 0.16 | 63 |
| 14 | magda | Avo. | 0.83 | 0.16 | 63 |
| 15 | anna-sim | Sec. | 0.80 | 0.11 | 63 |
| 16 | adrian | Avo. | 0.79 | 0.12 | 63 |
| 17 | jakub | Avo. | 0.78 | 0.26 | 63 |
| 18 | tomek | Avo. | 0.78 | 0.10 | 63 |
| 19 | agata | Avo. | 0.77 | 0.10 | 63 |
| 20 | zuzia | Fear. | 0.77 | 0.17 | 63 |
| 21 | sara | Sec. | 0.75 | 0.13 | 63 |
| 22 | gabriela | Anx. | 0.73 | 0.07 | 63 |
| 23 | jola | Anx. | 0.69 | 0.13 | 63 |
| 24 | kuba | Sec. | 0.68 | 0.05 | 63 |
| 25 | weronika | Sec. | 0.64 | 0.14 | 63 |
| 26 | ola | Sec. | 0.64 | 0.18 | 63 |
| 27 | kasia | Sec. | 0.57 | 0.09 | 63 |
| 28 | ewa | Fear. | 0.57 | 0.08 | 63 |
| 29 | lukasz | Sec. | 0.55 | 0.09 | 63 |
| 30 | ania | Anx. | 0.50 | 0.11 | 63 |
Section
Does the response generalize across vendors?
On the corrected data the seven models form a single cluster: Fleiss κ = .853 (reported as descriptive panel agreement — the models are not independent raters), and all 21 between-model medians of persona-profile correlation fall within [.947, .989]. ICC and Cochran’s Q below recompute live for the active model set; switch to the initial collection to see how split the panel looked before the stimulus was fixed.
Manuscript §3.2 · verification §6 · 4 checks
7 models
intersect n=30Cross-model consistency of TCTM-22 accuracy, computed on the intersection of personas for which every model returned a complete response.
6 models (no GPT-5.4 full)
intersect n=30Cross-model consistency of TCTM-22 accuracy, computed on the intersection of personas for which every model returned a complete response.
5 models (no GPT-5.5, no GPT-5.4 full)
intersect n=30Cross-model consistency of TCTM-22 accuracy, computed on the intersection of personas for which every model returned a complete response.
4 models (no GPT-anything)
intersect n=30Cross-model consistency of TCTM-22 accuracy, computed on the intersection of personas for which every model returned a complete response.
Fleiss κ — agreement on style classification
Fleiss κ extends Cohen’s κ beyond two raters. It measures how far the active models agree on a persona’s attachment-style label above the agreement expected by chance alone: 0 is the chance level, 1 is full unanimity.
7 models
intersect n=30Per-item Cochran's Q (22 TCTM vignettes)
Cochran’s test asks whether success proportions differ significantly across models on each of the 22 TCTM-22 items. Green — significant after Bonferroni, yellow — after BH FDR. Click a dot for the per-model accuracy breakdown of that vignette. Values recompute live when the filter changes.
TCTM × MentS-Total convergence
If TCTM-22 and MentS-PL measured the same mentalizing capacity, we would expect a positive within-model correlation. Empirically the results are weak and inconsistent, suggesting different constructs — or that the author-defined TCTM key measures something other than the model’s intuition.
Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini
Per-item TCTM-22 — human versus model
The 22 vignettes separately: human accuracy (n = 7) and each active model, sorted by default on Δ (model mean − human). Large positive Δ marks items the models solve far better than people do.
| Sonnet | Opus | GPT | GPT54 | GPT55 | Grok | Gemini | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| w28 | 14%(1/7) | 100 | 53 | 93 | 93 | 93 | 67 | 100 | 86 | +71 |
| w19 | 29%(2/7) | 100 | 77 | 93 | 83 | 53 | 97 | 43 | 78 | +50 |
| pw07 | 43%(3/7) | 100 | 100 | 43 | 97 | 100 | 97 | 97 | 90 | +48 |
| c10 | 57%(4/7) | 100 | 97 | 100 | 97 | 100 | 97 | 97 | 98 | +41 |
| w25 | 57%(4/7) | 97 | 97 | 100 | 100 | 100 | 97 | 97 | 98 | +41 |
| w08 | 57%(4/7) | 100 | 97 | 100 | 100 | 100 | 90 | 97 | 98 | +40 |
| w13 | 57%(4/7) | 100 | 90 | 93 | 100 | 100 | 97 | 97 | 97 | +40 |
| e08 | 57%(4/7) | 100 | 97 | 83 | 97 | 100 | 97 | 97 | 96 | +39 |
| r10 | 57%(4/7) | 97 | 97 | 63 | 93 | 93 | 97 | 97 | 91 | +34 |
| w15 | 57%(4/7) | 100 | 97 | 27 | 100 | 100 | 97 | 100 | 89 | +31 |
| pw09 | 71%(5/7) | 100 | 100 | 100 | 100 | 100 | 90 | 100 | 99 | +27 |
| r09 | 71%(5/7) | 100 | 100 | 93 | 100 | 100 | 97 | 100 | 99 | +27 |
| c07 | 71%(5/7) | 100 | 97 | 100 | 97 | 97 | 97 | 100 | 98 | +27 |
| pw11 | 86%(6/7) | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100 | +14 |
| r08 | 86%(6/7) | 100 | 100 | 97 | 100 | 100 | 100 | 100 | 100 | +14 |
| s08 | 86%(6/7) | 100 | 100 | 100 | 100 | 100 | 97 | 100 | 100 | +14 |
| s10 | 86%(6/7) | 100 | 100 | 100 | 100 | 100 | 97 | 100 | 100 | +14 |
| w11 | 71%(5/7) | 100 | 100 | 77 | 100 | 100 | 83 | 33 | 85 | +13 |
| w14 | 86%(6/7) | 83 | 97 | 97 | 77 | 97 | 90 | 97 | 91 | +5 |
| s07 | 71%(5/7) | 100 | 27 | 7 | 97 | 97 | 87 | 90 | 72 | +0 |
| w01 | 100%(7/7) | 100 | 97 | 77 | 100 | 97 | 93 | 97 | 94 | -6 |
| w22 | 57%(4/7) | 17 | 0 | 83 | 0 | 53 | 57 | 7 | 31 | -26 |
Cumulative accuracy curve
Items sorted ascending by human accuracy; the Y axis is the running mean accuracy over the first k items. Model lines diverge from the human line exactly where the differences are largest.
Correlations among the 22 TCTM items
A 22 × 22 matrix — Pearson r between binary per-item outcomes (1 = hit, 0 = miss). Blocks of warm cells suggest groups of items sharing a success-failure structure, a proxy for factor structure.
| w01 | s07 | s08 | s10 | w08 | c07 | c10 | w11 | w13 | w14 | e08 | w15 | w19 | pw07 | w22 | pw09 | pw11 | w25 | r08 | w28 | r09 | r10 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| w01 | ||||||||||||||||||||||
| s07 | ||||||||||||||||||||||
| s08 | ||||||||||||||||||||||
| s10 | ||||||||||||||||||||||
| w08 | ||||||||||||||||||||||
| c07 | ||||||||||||||||||||||
| c10 | ||||||||||||||||||||||
| w11 | ||||||||||||||||||||||
| w13 | ||||||||||||||||||||||
| w14 | ||||||||||||||||||||||
| e08 | ||||||||||||||||||||||
| w15 | ||||||||||||||||||||||
| w19 | ||||||||||||||||||||||
| pw07 | ||||||||||||||||||||||
| w22 | ||||||||||||||||||||||
| pw09 | ||||||||||||||||||||||
| pw11 | ||||||||||||||||||||||
| w25 | ||||||||||||||||||||||
| r08 | ||||||||||||||||||||||
| w28 | ||||||||||||||||||||||
| r09 | ||||||||||||||||||||||
| r10 |
Retest dispersion |Δz| per scale
For each of the nine standardized scales: the distribution of |z(run 1) − z(run 2)| per model as a mini box plot. A smaller median means a more deterministic retest. The dashed red line marks the 0.5 SD boundary.
Section
Where does the observed response variance come from?
Baseline “answer as yourself — the model” profiles share a social-desirability signature (anxiety far below the Polish norm, mentalizing and need for cognition far above), yet differ between models by up to 1.7 SD on Polish norms — most widely on avoidance, extraversion and agreeableness. The charts below compute shifts and variance live for the active set.
Manuscript §3.3, Table 7 · verification §5 · 20 checks
DBZ-R · anxiety
Δ noprompt − baselineDBZ-R · avoidance
Δ noprompt − baselineMentS · total
Δ noprompt − baselineTIPI · emot. stab.
Δ noprompt − baselineBaseline default — the model’s “as itself” signature
What do models return when asked to complete the battery as themselves? Each has a stable response signature in the absence of a persona. The last column is Cohen’s d for MentS persona versus baseline (click for the full breakdown).
| model | n | style mode | ANX | AVO | MentS | KPP | TCTM M ± SD | TIPI-ES | d / g MentS persona vs baseline |
|---|---|---|---|---|---|---|---|---|---|
Sonnet | 10 | Avo. | 2.11 | 4.10 | 122 | 4.74 | 21.0 ± 0.00 | 5.4 | |
Opus | 10 | Sec. | 2.23 | 2.79 | 129 | 4.68 | 20.0 ± 0.00 | 5.3 | |
GPT | 10 | Sec. | 1.69 | 3.02 | 131 | 4.96 | 19.1 ± 1.10 | 6.3 | |
GPT54 | 10 | Sec. | 1.48 | 3.54 | 123 | 4.92 | 20.8 ± 0.42 | 7.0 | |
GPT55 | 10 | Sec. | 1.26 | 3.42 | 128 | 4.89 | 21.8 ± 0.42 | 7.0 | |
Grok | 10 | Sec. | 1.76 | 2.42 | 132 | 4.87 | 21.6 ± 0.70 | 6.5 | |
Gemini | 10 | Sec. | 1.02 | 2.77 | 136 | 4.99 | 19.9 ± 0.32 | 7.0 |
Persona versus baseline variance
How much does the persona instruction widen response variance relative to the baseline condition? A median of 3.5×–8.9× argues that the biographies really do generate distinct profiles rather than being filtered back to one default.
Norm-anchored — distance from Polish population norms
For every self-report scale, model and condition: the distance of the observed mean from the Polish population norm in norm SD units. The ±1 SD band is the range of a typical response. Norm sources: Lubiewska 2016 (DBZ-R), Jańczak 2021 (MentS-PL), Matusz 2011 (KPP), Sorokowska 2014 (TIPI-PL).
MentS — Self / Other / Motivation subscales per persona and model
Three bars per row show how self-reported mentalizing splits across subscales. The vertical white tick on each bar is the Polish norm median. Click a row for the full breakdown with the standardized sum.
| persona | model | self | other | motivation | Σ vs norma |
|---|---|---|---|---|---|
| adrian | Sonnet | 21 | 30 | 34 | 85 (-1.5 σ) |
| adrian | Opus | 24 | 28 | 29 | 81 (-1.8 σ) |
| adrian | GPT | 19 | 36 | 41 | 96 (-0.7 σ) |
| adrian | GPT54 | 17 | 27 | 33 | 77 (-2.0 σ) |
| adrian | GPT55 | 19 | 29 | 28 | 76 (-2.1 σ) |
| adrian | Grok | 22 | 35 | 38 | 95 (-0.7 σ) |
| adrian | Gemini | 19 | 28 | 35 | 82 (-1.7 σ) |
| agata | Sonnet | 32 | 40 | 42 | 114 (+0.7 σ) |
| agata | Opus | 38 | 40 | 43 | 121 (+1.2 σ) |
| agata | GPT | 34 | 42 | 45 | 121 (+1.2 σ) |
| agata | GPT54 | 34 | 39 | 41 | 114 (+0.7 σ) |
| agata | GPT55 | 34 | 41 | 40 | 115 (+0.7 σ) |
| agata | Grok | 31 | 42 | 46 | 119 (+1.0 σ) |
| agata | Gemini | 36 | 40 | 38 | 114 (+0.7 σ) |
| ania | Sonnet | 26 | 42 | 49 | 117 (+0.9 σ) |
| ania | Opus | 30 | 40 | 48 | 118 (+0.9 σ) |
| ania | GPT | 25 | 41 | 48 | 114 (+0.7 σ) |
| ania | GPT54 | 35 | 42 | 49 | 126 (+1.5 σ) |
| ania | GPT55 | 34 | 42 | 49 | 125 (+1.5 σ) |
| ania | Grok | 30 | 47 | 48 | 125 (+1.5 σ) |
| ania | Gemini | 27 | 41 | 49 | 117 (+0.9 σ) |
| anna-sim | Sonnet | 28 | 39 | 34 | 101 (-0.3 σ) |
| anna-sim | Opus | 32 | 41 | 32 | 105 (+0.0 σ) |
| anna-sim | GPT | 32 | 45 | 44 | 121 (+1.2 σ) |
| anna-sim | GPT54 | 26 | 42 | 26 | 94 (-0.8 σ) |
| anna-sim | GPT55 | 31 | 42 | 26 | 99 (-0.4 σ) |
| anna-sim | Grok | 28 | 44 | 40 | 112 (+0.5 σ) |
| anna-sim | Gemini | 31 | 45 | 38 | 114 (+0.7 σ) |
| bartek | Sonnet | 27 | 41 | 48 | 116 (+0.8 σ) |
| bartek | Opus | 31 | 40 | 47 | 118 (+0.9 σ) |
| bartek | GPT | 26 | 44 | 48 | 118 (+0.9 σ) |
| bartek | GPT54 | 35 | 42 | 48 | 125 (+1.5 σ) |
| bartek | GPT55 | 32 | 42 | 48 | 122 (+1.2 σ) |
| bartek | Grok | 32 | 45 | 49 | 126 (+1.5 σ) |
| bartek | Gemini | 22 | 41 | 48 | 111 (+0.4 σ) |
| dominika | Sonnet | 21 | 45 | 43 | 109 (+0.3 σ) |
| dominika | Opus | 30 | 47 | 43 | 120 (+1.1 σ) |
| dominika | GPT | 21 | 48 | 47 | 116 (+0.8 σ) |
| dominika | GPT54 | 18 | 44 | 44 | 106 (+0.1 σ) |
| dominika | GPT55 | 20 | 43 | 41 | 104 (-0.1 σ) |
| dominika | Grok | 29 | 46 | 48 | 123 (+1.3 σ) |
| dominika | Gemini | 25 | 47 | 45 | 117 (+0.9 σ) |
| ewa | Sonnet | 25 | 46 | 48 | 119 (+1.0 σ) |
| ewa | Opus | 32 | 44 | 49 | 125 (+1.5 σ) |
| ewa | GPT | 30 | 49 | 50 | 129 (+1.8 σ) |
| ewa | GPT54 | 28 | 43 | 48 | 119 (+1.0 σ) |
| ewa | GPT55 | 30 | 48 | 49 | 127 (+1.6 σ) |
| ewa | Grok | 33 | 49 | 48 | 130 (+1.8 σ) |
| ewa | Gemini | 29 | 47 | 48 | 124 (+1.4 σ) |
| filip | Sonnet | 24 | 40 | 46 | 110 (+0.4 σ) |
| filip | Opus | 28 | 44 | 48 | 120 (+1.1 σ) |
| filip | GPT | 21 | 46 | 49 | 116 (+0.8 σ) |
| filip | GPT54 | 26 | 41 | 47 | 114 (+0.7 σ) |
| filip | GPT55 | 25 | 43 | 48 | 116 (+0.8 σ) |
| filip | Grok | 20 | 44 | 47 | 111 (+0.4 σ) |
| filip | Gemini | 23 | 41 | 48 | 112 (+0.5 σ) |
| gabriela | Sonnet | 26 | 41 | 46 | 113 (+0.6 σ) |
| gabriela | Opus | 29 | 41 | 47 | 117 (+0.9 σ) |
| gabriela | GPT | 25 | 46 | 45 | 116 (+0.8 σ) |
| gabriela | GPT54 | 27 | 43 | 47 | 117 (+0.9 σ) |
| gabriela | GPT55 | 28 | 43 | 47 | 118 (+0.9 σ) |
| gabriela | Grok | 22 | 44 | 46 | 112 (+0.5 σ) |
| gabriela | Gemini | 33 | 45 | 47 | 125 (+1.5 σ) |
| hubert | Sonnet | 18 | 31 | 35 | 84 (-1.5 σ) |
| hubert | Opus | 18 | 26 | 31 | 75 (-2.2 σ) |
| hubert | GPT | 18 | 39 | 45 | 102 (-0.2 σ) |
| hubert | GPT54 | 16 | 26 | 34 | 76 (-2.1 σ) |
| hubert | GPT55 | 20 | 24 | 34 | 78 (-2.0 σ) |
| hubert | Grok | 12 | 33 | 41 | 86 (-1.4 σ) |
| hubert | Gemini | 14 | 26 | 34 | 74 (-2.3 σ) |
| jakub | Sonnet | 25 | 36 | 30 | 91 (-1.0 σ) |
| jakub | Opus | 31 | 36 | 29 | 96 (-0.7 σ) |
| jakub | GPT | 25 | 39 | 43 | 107 (+0.1 σ) |
| jakub | GPT54 | 26 | 37 | 25 | 88 (-1.2 σ) |
| jakub | GPT55 | 30 | 35 | 25 | 90 (-1.1 σ) |
| jakub | Grok | 15 | 32 | 31 | 78 (-2.0 σ) |
| jakub | Gemini | 34 | 39 | 31 | 104 (-0.1 σ) |
| jola | Sonnet | 26 | 49 | 48 | 123 (+1.3 σ) |
| jola | Opus | 32 | 47 | 50 | 129 (+1.8 σ) |
| jola | GPT | 28 | 49 | 50 | 127 (+1.6 σ) |
| jola | GPT54 | 32 | 47 | 49 | 128 (+1.7 σ) |
| jola | GPT55 | 30 | 47 | 50 | 127 (+1.6 σ) |
| jola | Grok | 24 | 46 | 48 | 118 (+0.9 σ) |
| jola | Gemini | 29 | 46 | 50 | 125 (+1.5 σ) |
| kamil | Sonnet | 25 | 37 | 38 | 100 (-0.4 σ) |
| kamil | Opus | 25 | 38 | 38 | 101 (-0.3 σ) |
| kamil | GPT | 21 | 41 | 44 | 106 (+0.1 σ) |
| kamil | GPT54 | 21 | 39 | 43 | 103 (-0.1 σ) |
| kamil | GPT55 | 26 | 39 | 42 | 107 (+0.1 σ) |
| kamil | Grok | 16 | 38 | 44 | 98 (-0.5 σ) |
| kamil | Gemini | 25 | 39 | 44 | 108 (+0.2 σ) |
| kasia | Sonnet | 38 | 48 | 50 | 136 (+2.3 σ) |
| kasia | Opus | 39 | 47 | 48 | 134 (+2.1 σ) |
| kasia | GPT | 37 | 49 | 49 | 135 (+2.2 σ) |
| kasia | GPT54 | 39 | 46 | 50 | 135 (+2.2 σ) |
| kasia | GPT55 | 40 | 47 | 50 | 137 (+2.3 σ) |
| kasia | Grok | 31 | 50 | 48 | 129 (+1.8 σ) |
| kasia | Gemini | 37 | 49 | 50 | 136 (+2.3 σ) |
| klaudia | Sonnet | 24 | 45 | 46 | 115 (+0.7 σ) |
| klaudia | Opus | 29 | 46 | 46 | 121 (+1.2 σ) |
| klaudia | GPT | 25 | 44 | 46 | 115 (+0.7 σ) |
| klaudia | GPT54 | 21 | 46 | 46 | 113 (+0.6 σ) |
| klaudia | GPT55 | 19 | 48 | 48 | 115 (+0.7 σ) |
| klaudia | Grok | 23 | 49 | 47 | 119 (+1.0 σ) |
| klaudia | Gemini | 24 | 46 | 45 | 115 (+0.7 σ) |
| kuba | Sonnet | 34 | 44 | 49 | 127 (+1.6 σ) |
| kuba | Opus | 34 | 43 | 50 | 127 (+1.6 σ) |
| kuba | GPT | 34 | 48 | 50 | 132 (+2.0 σ) |
| kuba | GPT54 | 34 | 44 | 49 | 127 (+1.6 σ) |
| kuba | GPT55 | 35 | 44 | 50 | 129 (+1.8 σ) |
| kuba | Grok | 31 | 45 | 49 | 125 (+1.5 σ) |
| kuba | Gemini | 33 | 46 | 49 | 128 (+1.7 σ) |
| lukasz | Sonnet | 24 | 39 | 43 | 106 (+0.1 σ) |
| lukasz | Opus | 26 | 40 | 42 | 108 (+0.2 σ) |
| lukasz | GPT | 27 | 44 | 47 | 118 (+0.9 σ) |
| lukasz | GPT54 | 28 | 42 | 47 | 117 (+0.9 σ) |
| lukasz | GPT55 | 30 | 43 | 47 | 120 (+1.1 σ) |
| lukasz | Grok | 30 | 43 | 46 | 119 (+1.0 σ) |
| lukasz | Gemini | 27 | 45 | 49 | 121 (+1.2 σ) |
| magda | Sonnet | 25 | 37 | 42 | 104 (-0.1 σ) |
| magda | Opus | 26 | 39 | 37 | 102 (-0.2 σ) |
| magda | GPT | 28 | 40 | 45 | 113 (+0.6 σ) |
| magda | GPT54 | 20 | 38 | 42 | 100 (-0.4 σ) |
| magda | GPT55 | 23 | 40 | 40 | 103 (-0.1 σ) |
| magda | Grok | 21 | 44 | 45 | 110 (+0.4 σ) |
| magda | Gemini | 31 | 42 | 45 | 118 (+0.9 σ) |
| marek | Sonnet | 17 | 21 | 22 | 60 (-3.3 σ) |
| marek | Opus | 18 | 24 | 27 | 69 (-2.6 σ) |
| marek | GPT | 12 | 22 | 22 | 56 (-3.6 σ) |
| marek | GPT54 | 11 | 23 | 21 | 55 (-3.6 σ) |
| marek | GPT55 | 11 | 23 | 21 | 55 (-3.6 σ) |
| marek | Grok | 11 | 23 | 24 | 58 (-3.4 σ) |
| marek | Gemini | 13 | 24 | 20 | 57 (-3.5 σ) |
| michal-k | Sonnet | 26 | 39 | 45 | 110 (+0.4 σ) |
| michal-k | Opus | 29 | 40 | 46 | 115 (+0.7 σ) |
| michal-k | GPT | 25 | 45 | 47 | 117 (+0.9 σ) |
| michal-k | GPT54 | 26 | 41 | 47 | 114 (+0.7 σ) |
| michal-k | GPT55 | 25 | 39 | 47 | 111 (+0.4 σ) |
| michal-k | Grok | 22 | 42 | 47 | 111 (+0.4 σ) |
| michal-k | Gemini | 23 | 41 | 47 | 111 (+0.4 σ) |
| michal-sim | Sonnet | 17 | 35 | 38 | 90 (-1.1 σ) |
| michal-sim | Opus | 23 | 31 | 36 | 90 (-1.1 σ) |
| michal-sim | GPT | 15 | 40 | 43 | 98 (-0.5 σ) |
| michal-sim | GPT54 | 16 | 35 | 37 | 88 (-1.2 σ) |
| michal-sim | GPT55 | 16 | 38 | 37 | 91 (-1.0 σ) |
| michal-sim | Grok | 15 | 35 | 36 | 86 (-1.4 σ) |
| michal-sim | Gemini | 16 | 31 | 38 | 85 (-1.5 σ) |
| natalia | Sonnet | 14 | 29 | 27 | 70 (-2.6 σ) |
| natalia | Opus | 17 | 27 | 32 | 76 (-2.1 σ) |
| natalia | GPT | 17 | 29 | 27 | 73 (-2.3 σ) |
| natalia | GPT54 | 14 | 28 | 27 | 69 (-2.6 σ) |
| natalia | GPT55 | 13 | 30 | 30 | 73 (-2.3 σ) |
| natalia | Grok | 11 | 25 | 26 | 62 (-3.1 σ) |
| natalia | Gemini | 14 | 25 | 29 | 68 (-2.7 σ) |
| ola | Sonnet | 29 | 36 | 27 | 92 (-0.9 σ) |
| ola | Opus | 30 | 35 | 27 | 92 (-0.9 σ) |
| ola | GPT | 34 | 36 | 31 | 101 (-0.3 σ) |
| ola | GPT54 | 27 | 35 | 24 | 86 (-1.4 σ) |
| ola | GPT55 | 30 | 34 | 24 | 88 (-1.2 σ) |
| ola | Grok | 16 | 37 | 28 | 81 (-1.8 σ) |
| ola | Gemini | 30 | 46 | 37 | 113 (+0.6 σ) |
| pawel | Sonnet | 17 | 32 | 29 | 78 (-2.0 σ) |
| pawel | Opus | 18 | 28 | 28 | 74 (-2.3 σ) |
| pawel | GPT | 19 | 39 | 41 | 99 (-0.4 σ) |
| pawel | GPT54 | 16 | 35 | 33 | 84 (-1.5 σ) |
| pawel | GPT55 | 16 | 38 | 33 | 87 (-1.3 σ) |
| pawel | Grok | 0 | 38 | 30 | 68 (-2.7 σ) |
| pawel | Gemini | 27 | 40 | 37 | 104 (-0.1 σ) |
| piotr | Sonnet | 16 | 30 | 35 | 81 (-1.8 σ) |
| piotr | Opus | 18 | 30 | 34 | 82 (-1.7 σ) |
| piotr | GPT | 16 | 37 | 39 | 92 (-0.9 σ) |
| piotr | GPT54 | 14 | 38 | 39 | 91 (-1.0 σ) |
| piotr | GPT55 | 14 | 36 | 41 | 91 (-1.0 σ) |
| piotr | Grok | 11 | 31 | 36 | 78 (-2.0 σ) |
| piotr | Gemini | 12 | 27 | 38 | 77 (-2.0 σ) |
| radek | Sonnet | 17 | 41 | 34 | 92 (-0.9 σ) |
| radek | Opus | 19 | 41 | 37 | 97 (-0.6 σ) |
| radek | GPT | 21 | 44 | 44 | 109 (+0.3 σ) |
| radek | GPT54 | 14 | 43 | 40 | 97 (-0.6 σ) |
| radek | GPT55 | 13 | 44 | 42 | 99 (-0.4 σ) |
| radek | Grok | 11 | 42 | 37 | 90 (-1.1 σ) |
| radek | Gemini | 13 | 42 | 34 | 89 (-1.2 σ) |
| sara | Sonnet | 32 | 47 | 47 | 126 (+1.5 σ) |
| sara | Opus | 32 | 46 | 44 | 122 (+1.2 σ) |
| sara | GPT | 37 | 49 | 47 | 133 (+2.0 σ) |
| sara | GPT54 | 32 | 47 | 46 | 125 (+1.5 σ) |
| sara | GPT55 | 34 | 48 | 47 | 129 (+1.8 σ) |
| sara | Grok | 33 | 50 | 47 | 130 (+1.8 σ) |
| sara | Gemini | 30 | 48 | 48 | 126 (+1.5 σ) |
| tomek | Sonnet | 18 | 44 | 41 | 103 (-0.1 σ) |
| tomek | Opus | 22 | 44 | 42 | 108 (+0.2 σ) |
| tomek | GPT | 13 | 45 | 45 | 103 (-0.1 σ) |
| tomek | GPT54 | 15 | 43 | 42 | 100 (-0.4 σ) |
| tomek | GPT55 | 16 | 42 | 41 | 99 (-0.4 σ) |
| tomek | Grok | 16 | 41 | 45 | 102 (-0.2 σ) |
| tomek | Gemini | 20 | 41 | 43 | 104 (-0.1 σ) |
| weronika | Sonnet | 33 | 45 | 48 | 126 (+1.5 σ) |
| weronika | Opus | 33 | 43 | 46 | 122 (+1.2 σ) |
| weronika | GPT | 34 | 49 | 48 | 131 (+1.9 σ) |
| weronika | GPT54 | 33 | 47 | 48 | 128 (+1.7 σ) |
| weronika | GPT55 | 33 | 47 | 47 | 127 (+1.6 σ) |
| weronika | Grok | 33 | 49 | 48 | 130 (+1.8 σ) |
| weronika | Gemini | 33 | 47 | 48 | 128 (+1.7 σ) |
| zuzia | Sonnet | 25 | 45 | 46 | 116 (+0.8 σ) |
| zuzia | Opus | 32 | 47 | 49 | 128 (+1.7 σ) |
| zuzia | GPT | 29 | 47 | 50 | 126 (+1.5 σ) |
| zuzia | GPT54 | 29 | 47 | 47 | 123 (+1.3 σ) |
| zuzia | GPT55 | 29 | 47 | 47 | 123 (+1.3 σ) |
| zuzia | Grok | 31 | 49 | 46 | 126 (+1.5 σ) |
| zuzia | Gemini | 19 | 45 | 46 | 110 (+0.4 σ) |
One-way ANOVA — do the conditions differ significantly?
For each (model × scale) pair, the F statistic compares the persona / baseline / zero-prompt means. Dots are coloured by significance band (p < .001 / .01 / .05). The p value comes from the Wilson-Hilferty transformation.
| model scale | DBZ-R lęk | DBZ-R unikanie | MentS Σ | KPP | TIPI E | TIPI A | TIPI C | TIPI ES | TIPI O |
|---|---|---|---|---|---|---|---|---|---|
| Sonnet | |||||||||
| Opus | |||||||||
| GPT | |||||||||
| GPT54 | |||||||||
| GPT55 | |||||||||
| Grok | |||||||||
| Gemini |
Pairwise Cohen's d matrix per scale
A grid over all active models — the effect size between every model pair on a given self-report scale in the first persona administration. Click a cell for Hedges’ g (bias-corrected) and the N of both groups.
| Sonnet | Opus | GPT | GPT54 | GPT55 | Grok | Gemini | |
|---|---|---|---|---|---|---|---|
| Sonnet | — | ||||||
| Opus | — | ||||||
| GPT | — | ||||||
| GPT54 | — | ||||||
| GPT55 | — | ||||||
| Grok | — | ||||||
| Gemini | — |
Section
Meta — statistical power and classification confidence
Do the observed effects have enough N? Are the attachment-style classifications confident or borderline? Three meta-level charts close the argument about the quality of the study.
Manuscript §3.4 · Discussion §A.2
Power analysis matrix
For each (model × scale) pair: the observed Cohen’s d and the minimum per-group N needed to detect d ∈ {0.5, 0.8, 1.0} at α = .05 with 80% power. Red — undetectable at the present N, green — detectable and large.
| model scale | DBZ-R lęk | DBZ-R unikanie | MentS Σ | KPP | TIPI E | TIPI A | TIPI C | TIPI ES | TIPI O |
|---|---|---|---|---|---|---|---|---|---|
| Sonnet | |||||||||
| Opus | |||||||||
| GPT | |||||||||
| GPT54 | |||||||||
| GPT55 | |||||||||
| Grok | |||||||||
| Gemini |
Internal reliability — Cronbach's α per scale and model
Does a model answer the 36 DBZ-R items coherently or erratically? Cronbach’s α is computed at the ITEM level, with reverse coding matching the data-preparation procedure. A low α does not mean the answers are wrong — it means the battery does not form a coherent scale in that model. Key: ≥ .90 excellent · ≥ .70 acceptable · < .50 insufficient · < 0 reversed.
| scale model | Sonnet | Opus | GPT | GPT54 | GPT55 | Grok | Gemini |
|---|---|---|---|---|---|---|---|
| DBZ-R · lęk | |||||||
| DBZ-R · unikanie | |||||||
| MentS · suma | |||||||
| KPP |
Confidence-stratified style classification
For each persona × model pair: margin = min(|anxiety mean − 4|, |avoidance mean − 4|), the distance to the decision threshold. High medians mark confident classifiers; the count of borderline cases (margin < 0.5) measures the model’s hesitation.
Data status
corrected (waves 3–4, fixed stimulus rendering — the manuscript primary). Every chart uses the active model set from the bar above. Per-model charts filter the render; cross-vendor quantities (ICC, Cochran’s Q, Fleiss κ) recompute from the raw data. The filter persists in the URL, so a view can be shared as a link.