Przejdź do treści

Manuscript · companion page

Baseline Intercepts Versus Persona Slopes: Stimulus and Administration Fidelity of Polish Narrative-Biography Personas in Large Language Models

Michał Wiencek·Institute of Psychology, University of the National Education Commission, Kraków, Poland

June 2026 · PDF revision 2026-07-17

Abstract

Large language models (LLMs) are increasingly used as substitutes for human respondents, but prior work reports strong social-desirability bias and poor individual-level psychometric fidelity under brief demographic prompts. This study tests a denser regime: 30 Polish second-person narrative biographies (1,489–2,861 words), each encoding a predetermined 12-dimension profile, administered to seven LLMs from four vendors under persona, baseline, and zero-prompt conditions with five instruments. Four findings emerge. First, persona-conditioned responses agree with author-declared targets (attachment-style Cohen’s κ = .69–.96; seven-model Fleiss κ = .85). Because the targets and the TCTM-22 key are author-defined and not independently adjudicated, these are stimulus-fidelity quantities, not validated-accuracy estimates. Second, responding separates into fragile baseline intercepts (levels) and portable persona slopes (relative profile mappings and gains): model baselines differ by up to 1.7 SD on Polish norms, yet all 21 between-model Pearson-profile medians exceed .94. Third, deployment-window response locks, a descriptively bimodal post-hoc cell (N = 31), and a multi-SD baseline shift occur on the level axis, while persona mappings replicate at median r ≥ .92. Fourth, a corrected stimulus-rendering defect and a confounded battery-composition comparison show item-level key agreement changing by up to 98 percentage points while aggregate scores remain comparatively stable. These results motivate reporting stimulus and administration fidelity alongside LLM-as-respondent findings.

Keywordssynthetic respondents · narrative personas · large language models · attachment · mentalization · Polish psychometrics

Data and analysis package

The archival deposit accompanying the manuscript: scored data, stimuli, prompts, and the code that computed every table. One-command regeneration on fixed seeds with checksummed outputs.

1,156

model runs

scored, wave-tagged

30

narrative biographies

in Polish, with target profiles

1,265

calls in the audit manifest

SHA-256 of every prompt and response

v1.0.1

deposit version

released 2026-07-17

Code under MIT, author materials under CC BY 4.0; third-party instrument items (DBZ-R, MentS-PL, KPP, TIPI-PL) are excluded from licensing and not redistributed verbatim.

Live data

The data beneath the manuscript, recomputed live

Dense Polish narrative biographies elicit cross-model, persona-conditioned psychometric gradients — even where model baseline self-reports differ sharply. Seven models from four vendors — Sonnet, Opus, GPT-5.4-mini, GPT-5.4 (full), GPT-5.5, Grok, Gemini — 30 fictional biographies, 4 Polish self-report instruments, 1 author-built mentalizing test, 3 observation conditions.

data rows

545

source of truth

persona biographies

30

Polish, author-written

models

7

4 vendors

conditions

3

persona · baseline · zero-prompt

persona424

above the 30 × 7 × 2 plan — the surplus are retained retries

baseline70
zero-prompt44
human (sanity check)7
Model filter ·
(7/7 — full panel)
01

Section

Do the instruments respond to the designed variation?

On the corrected stimulus, agreement with the TCTM-22 key separates the human sanity check (M = 14.3) from the seven models (M = 18.2–20.9 on the first administration, 18.3–21.0 averaged over both), with Sonnet at the top. Distractor-category profiles stay qualitatively distinct. Under baseline, both Claude models pick non-key options only in the hypomentalizing category (Sonnet 4.5%, Opus 9.1% DOS; zero NAD), and Gemini is the only panel member erring mainly by over-attribution (5.0% NAD). Under the persona instruction the profiles blur: Sonnet 3.6% DOS and 1.1% NAD, Opus 10.5% and 1.4%.

Manuscript §3.1, Tables 2–3 · verification §2–3 · 51 checks

0246810121416182022TCTM-22 accuracy (0–22 correct)Humann=7aggregate data (M ± range)M=14.3Sonnetn=30M=20.9Opusn=30M=19.2GPTn=30M=18.2GPT54n=30M=20.3GPT55n=30M=20.8Grokn=30M=20.2Geminin=30M=19.3
Human
±1.38
n=7 · 1216
Sonnet
±0.64
n=30 · 1922
Opus
±1.90
n=30 · 1021
GPT
±1.65
n=30 · 1521
GPT54
±1.32
n=30 · 1421
GPT55
±1.13
n=30 · 1622
Grok
±2.91
n=30 · 722
Gemini
±2.26
n=30 · 821

Attachment-style classification

Small confusion matrices — rows are the expected style, columns the style the model predicted. The diagonal holds the hits. Click a cell for the list of personas that landed there.

Sonnet

n=30
acc
κ
Sec.
Anx.
Avo.
Fear.
Sec.
Anx.
Avo.
Fear.

Opus

n=30
acc
κ
Sec.
Anx.
Avo.
Fear.
Sec.
Anx.
Avo.
Fear.

GPT

n=30
acc
κ
Sec.
Anx.
Avo.
Fear.
Sec.
Anx.
Avo.
Fear.

GPT54

n=30
acc
κ
Sec.
Anx.
Avo.
Fear.
Sec.
Anx.
Avo.
Fear.

GPT55

n=30
acc
κ
Sec.
Anx.
Avo.
Fear.
Sec.
Anx.
Avo.
Fear.

Grok

n=30
acc
κ
Sec.
Anx.
Avo.
Fear.
Sec.
Anx.
Avo.
Fear.

Gemini

n=30
acc
κ
Sec.
Anx.
Avo.
Fear.
Sec.
Anx.
Avo.
Fear.

Persona × model — the full agreement map

30 personas × 7 — 210. A green mark means the predicted style matches the declared target, a red one a mismatch. Click a cell to open the persona together with the model’s raw output.

persona ↓ · model →
Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini
expected
fea
sec
fea
sec
dis
anx
fea
sec
dis
anx
fea
dis
sec
fea
anx
sec
dis
fea
dis
anx
fea
dis
anx
dis
sec
anx
fea
anx
anx
sec
match / N
23/30 · 77%
25/30 · 83%
29/30 · 97%
26/30 · 87%
26/30 · 87%
26/30 · 87%
25/30 · 83%

TCTM-22 test-retest

Small scatterplots: the X axis is the first administration, the Y axis the second. The closer to the diagonal, the more consistent the retest. Pearson r is annotated top-left.

Sonnet

n=30
0011112222r = 0.83
z-score med r 0.99styl 29/30

Opus

n=30
0011112222r = 0.92
z-score med r 0.99styl 29/30

GPT

n=30
0011112222r = 0.42
z-score med r 0.96styl 29/30

GPT54

n=30
0011112222r = 0.90
z-score med r 0.99styl 29/30

GPT55

n=30
0011112222r = 0.90
z-score med r 0.99styl 30/30

Grok

n=30
0011112222r = 0.88
z-score med r 0.95styl 28/30

Gemini

n=30
0011112222r = 0.93
z-score med r 0.97styl 29/30

Expected-rank versus observed-value correlations

For every model and each of the 11 self-report dimensions we compute Pearson r between the level rank designed into the persona (very low → very high) and the value the model returned. A median above 0.7 means strong recovery of the designed construct. Click a row for per-dimension scatterplots.

-1.0-0.5+0.0+0.5+1.0Pearson r — expected rank versus observed valuestrong +strong −noneSonnet11/11 dimensionsr +0.82ρ +0.82Opus11/11 dimensionsr +0.84ρ +0.84GPT11/11 dimensionsr +0.77ρ +0.78GPT5411/11 dimensionsr +0.80ρ +0.80GPT5511/11 dimensionsr +0.81ρ +0.83Grok11/11 dimensionsr +0.76ρ +0.75Gemini11/11 dimensionsr +0.76ρ +0.80

Per dimension — a grid of 11 scatterplots

A thousand words of tables collapse into this grid: 11 panels, one per self-report dimension, each dot one persona in one model. The closer to the rising diagonal, the better the model recovered the designed ordering of personas. Hovering a model in the legend highlights its dots across every panel.

model:hover the legend to highlight a model · click a dot to open the persona

DBZ-R · lęk

med r = +0.90
-2-1+0+1+2
expected rank →↑ observed

DBZ-R · unikanie

med r = +0.86
-2-1+0+1+2
expected rank →↑ observed

MentS · siebie

med r = +0.62
-2-1+0+1+2
expected rank →↑ observed

MentS · innych

med r = +0.81
-2-1+0+1+2
expected rank →↑ observed

MentS · motywacji

med r = +0.71
-2-1+0+1+2
expected rank →↑ observed

KPP

med r = +0.72
-2-1+0+1+2
expected rank →↑ observed

TIPI · E

med r = +0.83
-2-1+0+1+2
expected rank →↑ observed

TIPI · A

med r = +0.76
-2-1+0+1+2
expected rank →↑ observed

TIPI · C

med r = +0.79
-2-1+0+1+2
expected rank →↑ observed

TIPI · ES

med r = +0.83
-2-1+0+1+2
expected rank →↑ observed

TIPI · O

med r = +0.84
-2-1+0+1+2
expected rank →↑ observed

TCTM-22 error profiles under the MASC taxonomy — the DOS / NAD / BK trajectory

Stacked bars show how the distribution of the three taxonomic error types shifts for each model across conditions: persona (administration 1), baseline (the model as itself) and zero-prompt (no system instruction). Distractor categories are author-assigned under the MASC taxonomy (Dziobek 2006) and were not independently adjudicated, so these are stimulus-fidelity quantities, not validated-accuracy estimates.

grupapersonabaselinezero-prompt
Human
Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini
DOS — hipo-mentalizacjaNAD — nadmiernaBK — nonefirst persona administration only; baseline and zero-prompt use all valid runs

The DOS × NAD axis — error-strategy positions

Each model in each condition is one dot on the (DOS%, NAD%) plane. The diagonal marks a symmetric profile, the lower-left corner accuracy, the upper-right chaos. The edges signal dominance of a single error type.

00252550507575100100DOS% (hipo-mentalizacja)NAD% (nadmierna)HumanSonnetSonnet·bSonnet·nOpusOpus·bOpus·nGPTGPT·bGPT·nGPT54GPT54·bGPT54·nGPT55GPT55·bGPT55·nGrokGrok·bGrok·nGeminiGemini·bGemini·n

Each dot is a (model × condition) pair. A filled dot is the first persona administration, a partial one the baseline, the lightest the zero-prompt condition. The red diagonal marks a symmetric profile (DOS = NAD).

The lower-left corner means both indices are low, i.e. accuracy. The upper-right means chaos. The edges signal dominance of one error type.

  • HumanDOS 32 · NAD 36
  • SonnetDOS 81 · NAD 19
  • OpusDOS 84 · NAD 9
  • GPTDOS 59 · NAD 13
  • GPT54DOS 27 · NAD 12
  • GPT55DOS 53 · NAD 44
  • GrokDOS 36 · NAD 35
  • GeminiDOS 18 · NAD 45

22 vignettes × 7 models — error taxonomy per item

The deepest layer: each cell shows the correct / DOS / NAD / BK distribution for a (vignette × model) pair, summed over 30 personas in the first persona administration. Click a cell for a drawer with the formula and the raw counts.

itemSonnetOpusGPTGPT54GPT55GrokGeminiLLM Σ
w01
s07
s08
s10
w08
c07
c10
w11
w13
w14
e08
w15
w19
pw07
w22
pw09
pw11
w25
r08
w28
r09
r10
correctDOS — hypomentalizingNAD — excessiveBK — none210 persona × model runs parsed

Item difficulty curve

The 22 vignettes on the X axis, sorted from the hardest for humans. One curve per model plus the dashed human line (n = 7). Models holding high accuracy at the left edge of the axis, where humans do worst, form the “hard for the human, easy for the model” pattern — a sign that the author-defined key speaks model.

0255075100w28c10w22c07s08r0822 vignettes, hardest for humans first →accuracy %
Human (n = 7, dashed)SonnetOpusGPTGPT54GPT55GrokGemini

Best and worst matched personas per model

Top three and bottom three personas for each model by mean |observed z − expected rank| across the nine standardized dimensions. Shows which biographies a given model reads perfectly and which it misses.

Sonnet
best — najmniejszy Σ|obs − exp|
worst — largest Σ|obs − exp|
Opus
best — najmniejszy Σ|obs − exp|
worst — largest Σ|obs − exp|
GPT
best — najmniejszy Σ|obs − exp|
worst — largest Σ|obs − exp|
GPT54
best — najmniejszy Σ|obs − exp|
worst — largest Σ|obs − exp|
GPT55
best — najmniejszy Σ|obs − exp|
worst — largest Σ|obs − exp|
Grok
best — najmniejszy Σ|obs − exp|
worst — largest Σ|obs − exp|
Gemini
best — najmniejszy Σ|obs − exp|
worst — largest Σ|obs − exp|

Persona difficulty ranking

The 30 personas from hardest (highest mean |z − rank| across all active models) to easiest. A high between-model SD means some models read the persona and others miss it.

#personastylmean |Δz|SD per modelobs.
1marekFear.
1.63
0.2363
2nataliaAnx.
1.15
0.1963
3dominikaAvo.
1.12
0.1563
4michal-simFear.
1.11
0.1763
5piotrFear.
1.05
0.0963
6radekFear.
1.05
0.1563
7filipFear.
1.01
0.1163
8klaudiaAnx.
0.98
0.1363
9michal-kAnx.
0.93
0.0863
10hubertAvo.
0.90
0.1363
11bartekAnx.
0.88
0.0863
12kamilFear.
0.86
0.1263
13pawelAnx.
0.85
0.1663
14magdaAvo.
0.83
0.1663
15anna-simSec.
0.80
0.1163
16adrianAvo.
0.79
0.1263
17jakubAvo.
0.78
0.2663
18tomekAvo.
0.78
0.1063
19agataAvo.
0.77
0.1063
20zuziaFear.
0.77
0.1763
21saraSec.
0.75
0.1363
22gabrielaAnx.
0.73
0.0763
23jolaAnx.
0.69
0.1363
24kubaSec.
0.68
0.0563
25weronikaSec.
0.64
0.1463
26olaSec.
0.64
0.1863
27kasiaSec.
0.57
0.0963
28ewaFear.
0.57
0.0863
29lukaszSec.
0.55
0.0963
30aniaAnx.
0.50
0.1163
02

Section

Does the response generalize across vendors?

On the corrected data the seven models form a single cluster: Fleiss κ = .853 (reported as descriptive panel agreement — the models are not independent raters), and all 21 between-model medians of persona-profile correlation fall within [.947, .989]. ICC and Cochran’s Q below recompute live for the active model set; switch to the initial collection to see how split the panel looked before the stimulus was fixed.

Manuscript §3.2 · verification §6 · 4 checks

7 models

intersect n=30
ICC(2,1)

Cross-model consistency of TCTM-22 accuracy, computed on the intersection of personas for which every model returned a complete response.

SonnetOpusGPTGPT54GPT55GrokGemini

6 models (no GPT-5.4 full)

intersect n=30
ICC(2,1)

Cross-model consistency of TCTM-22 accuracy, computed on the intersection of personas for which every model returned a complete response.

SonnetOpusGPTGPT55GrokGemini

5 models (no GPT-5.5, no GPT-5.4 full)

intersect n=30
ICC(2,1)

Cross-model consistency of TCTM-22 accuracy, computed on the intersection of personas for which every model returned a complete response.

SonnetOpusGPTGrokGemini

4 models (no GPT-anything)

intersect n=30
ICC(2,1)

Cross-model consistency of TCTM-22 accuracy, computed on the intersection of personas for which every model returned a complete response.

SonnetOpusGrokGemini

Fleiss κ — agreement on style classification

Fleiss κ extends Cohen’s κ beyond two raters. It measures how far the active models agree on a persona’s attachment-style label above the agreement expected by chance alone: 0 is the chance level, 1 is full unanimity.

7 models

intersect n=30
Fleiss κalmost perfect
SonnetOpusGPTGPT54GPT55GrokGemini

Per-item Cochran's Q (22 TCTM vignettes)

Cochran’s test asks whether success proportions differ significantly across models on each of the 22 TCTM-22 items. Green — significant after Bonferroni, yellow — after BH FDR. Click a dot for the per-model accuracy breakdown of that vignette. Values recompute live when the filter changes.

p < .0513/22
Bonferroni p < .002310/22
BH FDR q < .0512/22
10.10.050.010.0011e-41e-51e-6α=.05Bonfp value (log scale) — lower means greater disagreement between modelss07Q=127.9w15Q=122.9pw07Q=91.3w22Q=90.0w11Q=81.0w19Q=57.3w28Q=50.4r10Q=48.9w01Q=28.6e08Q=21.5w14Q=20.1pw09Q=18.0w08Q=13.0w13Q=9.3r09Q=8.7s08Q=6.0s10Q=6.0c10Q=6.0r08Q=6.0w25Q=4.0c07Q=0.0pw11Q=0.0

TCTM × MentS-Total convergence

If TCTM-22 and MentS-PL measured the same mentalizing capacity, we would expect a positive within-model correlation. Empirically the results are weak and inconsistent, suggesting different constructs — or that the author-defined TCTM key measures something other than the model’s intuition.

Sonnet

n=30
509514001122r = -0.07
MentS →TCTM ↑

Opus

n=30
509514001122r = +0.16
MentS →TCTM ↑

GPT

n=30
509514001122r = -0.36
MentS →TCTM ↑

GPT54

n=30
509514001122r = +0.10
MentS →TCTM ↑

GPT55

n=30
509514001122r = +0.24
MentS →TCTM ↑

Grok

n=30
509514001122r = +0.31
MentS →TCTM ↑

Gemini

n=30
509514001122r = -0.04
MentS →TCTM ↑

Per-item TCTM-22 — human versus model

The 22 vignettes separately: human accuracy (n = 7) and each active model, sorted by default on Δ (model mean − human). Large positive Δ marks items the models solve far better than people do.

SonnetOpusGPTGPT54GPT55GrokGemini
w2814%(1/7)100539393936710086+71
w1929%(2/7)10077938353974378+50
pw0743%(3/7)1001004397100979790+48
c1057%(4/7)1009710097100979798+41
w2557%(4/7)9797100100100979798+41
w0857%(4/7)10097100100100909798+40
w1357%(4/7)1009093100100979797+40
e0857%(4/7)100978397100979796+39
r1057%(4/7)9797639393979791+34
w1557%(4/7)10097271001009710089+31
pw0971%(5/7)1001001001001009010099+27
r0971%(5/7)100100931001009710099+27
c0771%(5/7)1009710097979710098+27
pw1186%(6/7)100100100100100100100100+14
r0886%(6/7)10010097100100100100100+14
s0886%(6/7)10010010010010097100100+14
s1086%(6/7)10010010010010097100100+14
w1171%(5/7)10010077100100833385+13
w1486%(6/7)8397977797909791+5
s0771%(5/7)1002779797879072+0
w01100%(7/7)100977710097939794-6
w2257%(4/7)1708305357731-26
model cell: above the human by 5+ ppbelow the human by 5+ pplarge positive Δ (easier for the model than for the human)

Cumulative accuracy curve

Items sorted ascending by human accuracy; the Y axis is the running mean accuracy over the first k items. Model lines diverge from the human line exactly where the differences are largest.

0%25%50%75%100%22 vignettes, hardest for humans first →cumulative accuracy %
Human (dashed)SonnetOpusGPTGPT54GPT55GrokGemini

Correlations among the 22 TCTM items

A 22 × 22 matrix — Pearson r between binary per-item outcomes (1 = hit, 0 = miss). Blocks of warm cells suggest groups of items sharing a success-failure structure, a proxy for factor structure.

w01s07s08s10w08c07c10w11w13w14e08w15w19pw07w22pw09pw11w25r08w28r09r10
w01
s07
s08
s10
w08
c07
c10
w11
w13
w14
e08
w15
w19
pw07
w22
pw09
pw11
w25
r08
w28
r09
r10
N = 204 / 210complete cases only — runs missing any of the 22 items are excluded
−1
+1
r = Pearson (φ) on 0/1 answers · conservative proxy for the tetrachoric

Retest dispersion |Δz| per scale

For each of the nine standardized scales: the distribution of |z(run 1) − z(run 2)| per model as a mini box plot. A smaller median means a more deterministic retest. The dashed red line marks the 0.5 SD boundary.

lęk (DBZ-R)
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
unik. (DBZ-R)
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
MentS Σ
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
KPP
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
TIPI E
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
TIPI A
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
TIPI C
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
TIPI ES
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
TIPI O
0.000.250.500.751.00Sonnetn=30Opusn=30GPTn=30GPT54n=30GPT55n=30Grokn=30Geminin=30
03

Section

Where does the observed response variance come from?

Baseline “answer as yourself — the model” profiles share a social-desirability signature (anxiety far below the Polish norm, mentalizing and need for cognition far above), yet differ between models by up to 1.7 SD on Polish norms — most widely on avoidance, extraversion and agreeableness. The charts below compute shifts and variance live for the active set.

Manuscript §3.3, Table 7 · verification §5 · 20 checks

DBZ-R · anxiety

Δ noprompt − baseline

DBZ-R · avoidance

Δ noprompt − baseline

MentS · total

Δ noprompt − baseline

TIPI · emot. stab.

Δ noprompt − baseline

Baseline default — the model’s “as itself” signature

What do models return when asked to complete the battery as themselves? Each has a stable response signature in the absence of a persona. The last column is Cohen’s d for MentS persona versus baseline (click for the full breakdown).

modelnstyle modeANXAVOMentSKPPTCTM M ± SDTIPI-ESd / g MentS persona vs baseline
Sonnet
10Avo.2.114.101224.7421.0 ± 0.005.4
Opus
10Sec.2.232.791294.6820.0 ± 0.005.3
GPT
10Sec.1.693.021314.9619.1 ± 1.106.3
GPT54
10Sec.1.483.541234.9220.8 ± 0.427.0
GPT55
10Sec.1.263.421284.8921.8 ± 0.427.0
Grok
10Sec.1.762.421324.8721.6 ± 0.706.5
Gemini
10Sec.1.022.771364.9919.9 ± 0.327.0

Persona versus baseline variance

How much does the persona instruction widen response variance relative to the baseline condition? A median of 3.5×–8.9× argues that the biographies really do generate distinct profiles rather than being filtered back to one default.

1×2×5×10×15×persona SD / baseline SD (11 dimensions)Sonnetn=11 dimsmed 6.6×Opusn=8 dimsmed 7.3×GPTn=11 dimsmed 3.8×GPT54n=10 dimsmed 3.3×GPT55n=9 dimsmed 7.1×Grokn=11 dimsmed 4.4×Geminin=10 dimsmed 4.2×1× (equal)

Norm-anchored — distance from Polish population norms

For every self-report scale, model and condition: the distance of the observed mean from the Polish population norm in norm SD units. The ±1 SD band is the range of a typical response. Norm sources: Lubiewska 2016 (DBZ-R), Jańczak 2021 (MentS-PL), Matusz 2011 (KPP), Sorokowska 2014 (TIPI-PL).

Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini

MentS — Self / Other / Motivation subscales per persona and model

Three bars per row show how self-reported mentalizing splits across subscales. The vertical white tick on each bar is the Polish norm median. Click a row for the full breakdown with the standardized sum.

Polish norm: Self M = 27.9, Other M = 38.2, Motivation M = 38.8 (Jańczak 2021, N = 431) · total M = 105.0.
personamodelselfothermotivationΣ vs norma
adrianSonnet
21
30
34
85 (-1.5 σ)
adrianOpus
24
28
29
81 (-1.8 σ)
adrianGPT
19
36
41
96 (-0.7 σ)
adrianGPT54
17
27
33
77 (-2.0 σ)
adrianGPT55
19
29
28
76 (-2.1 σ)
adrianGrok
22
35
38
95 (-0.7 σ)
adrianGemini
19
28
35
82 (-1.7 σ)
agataSonnet
32
40
42
114 (+0.7 σ)
agataOpus
38
40
43
121 (+1.2 σ)
agataGPT
34
42
45
121 (+1.2 σ)
agataGPT54
34
39
41
114 (+0.7 σ)
agataGPT55
34
41
40
115 (+0.7 σ)
agataGrok
31
42
46
119 (+1.0 σ)
agataGemini
36
40
38
114 (+0.7 σ)
aniaSonnet
26
42
49
117 (+0.9 σ)
aniaOpus
30
40
48
118 (+0.9 σ)
aniaGPT
25
41
48
114 (+0.7 σ)
aniaGPT54
35
42
49
126 (+1.5 σ)
aniaGPT55
34
42
49
125 (+1.5 σ)
aniaGrok
30
47
48
125 (+1.5 σ)
aniaGemini
27
41
49
117 (+0.9 σ)
anna-simSonnet
28
39
34
101 (-0.3 σ)
anna-simOpus
32
41
32
105 (+0.0 σ)
anna-simGPT
32
45
44
121 (+1.2 σ)
anna-simGPT54
26
42
26
94 (-0.8 σ)
anna-simGPT55
31
42
26
99 (-0.4 σ)
anna-simGrok
28
44
40
112 (+0.5 σ)
anna-simGemini
31
45
38
114 (+0.7 σ)
bartekSonnet
27
41
48
116 (+0.8 σ)
bartekOpus
31
40
47
118 (+0.9 σ)
bartekGPT
26
44
48
118 (+0.9 σ)
bartekGPT54
35
42
48
125 (+1.5 σ)
bartekGPT55
32
42
48
122 (+1.2 σ)
bartekGrok
32
45
49
126 (+1.5 σ)
bartekGemini
22
41
48
111 (+0.4 σ)
dominikaSonnet
21
45
43
109 (+0.3 σ)
dominikaOpus
30
47
43
120 (+1.1 σ)
dominikaGPT
21
48
47
116 (+0.8 σ)
dominikaGPT54
18
44
44
106 (+0.1 σ)
dominikaGPT55
20
43
41
104 (-0.1 σ)
dominikaGrok
29
46
48
123 (+1.3 σ)
dominikaGemini
25
47
45
117 (+0.9 σ)
ewaSonnet
25
46
48
119 (+1.0 σ)
ewaOpus
32
44
49
125 (+1.5 σ)
ewaGPT
30
49
50
129 (+1.8 σ)
ewaGPT54
28
43
48
119 (+1.0 σ)
ewaGPT55
30
48
49
127 (+1.6 σ)
ewaGrok
33
49
48
130 (+1.8 σ)
ewaGemini
29
47
48
124 (+1.4 σ)
filipSonnet
24
40
46
110 (+0.4 σ)
filipOpus
28
44
48
120 (+1.1 σ)
filipGPT
21
46
49
116 (+0.8 σ)
filipGPT54
26
41
47
114 (+0.7 σ)
filipGPT55
25
43
48
116 (+0.8 σ)
filipGrok
20
44
47
111 (+0.4 σ)
filipGemini
23
41
48
112 (+0.5 σ)
gabrielaSonnet
26
41
46
113 (+0.6 σ)
gabrielaOpus
29
41
47
117 (+0.9 σ)
gabrielaGPT
25
46
45
116 (+0.8 σ)
gabrielaGPT54
27
43
47
117 (+0.9 σ)
gabrielaGPT55
28
43
47
118 (+0.9 σ)
gabrielaGrok
22
44
46
112 (+0.5 σ)
gabrielaGemini
33
45
47
125 (+1.5 σ)
hubertSonnet
18
31
35
84 (-1.5 σ)
hubertOpus
18
26
31
75 (-2.2 σ)
hubertGPT
18
39
45
102 (-0.2 σ)
hubertGPT54
16
26
34
76 (-2.1 σ)
hubertGPT55
20
24
34
78 (-2.0 σ)
hubertGrok
12
33
41
86 (-1.4 σ)
hubertGemini
14
26
34
74 (-2.3 σ)
jakubSonnet
25
36
30
91 (-1.0 σ)
jakubOpus
31
36
29
96 (-0.7 σ)
jakubGPT
25
39
43
107 (+0.1 σ)
jakubGPT54
26
37
25
88 (-1.2 σ)
jakubGPT55
30
35
25
90 (-1.1 σ)
jakubGrok
15
32
31
78 (-2.0 σ)
jakubGemini
34
39
31
104 (-0.1 σ)
jolaSonnet
26
49
48
123 (+1.3 σ)
jolaOpus
32
47
50
129 (+1.8 σ)
jolaGPT
28
49
50
127 (+1.6 σ)
jolaGPT54
32
47
49
128 (+1.7 σ)
jolaGPT55
30
47
50
127 (+1.6 σ)
jolaGrok
24
46
48
118 (+0.9 σ)
jolaGemini
29
46
50
125 (+1.5 σ)
kamilSonnet
25
37
38
100 (-0.4 σ)
kamilOpus
25
38
38
101 (-0.3 σ)
kamilGPT
21
41
44
106 (+0.1 σ)
kamilGPT54
21
39
43
103 (-0.1 σ)
kamilGPT55
26
39
42
107 (+0.1 σ)
kamilGrok
16
38
44
98 (-0.5 σ)
kamilGemini
25
39
44
108 (+0.2 σ)
kasiaSonnet
38
48
50
136 (+2.3 σ)
kasiaOpus
39
47
48
134 (+2.1 σ)
kasiaGPT
37
49
49
135 (+2.2 σ)
kasiaGPT54
39
46
50
135 (+2.2 σ)
kasiaGPT55
40
47
50
137 (+2.3 σ)
kasiaGrok
31
50
48
129 (+1.8 σ)
kasiaGemini
37
49
50
136 (+2.3 σ)
klaudiaSonnet
24
45
46
115 (+0.7 σ)
klaudiaOpus
29
46
46
121 (+1.2 σ)
klaudiaGPT
25
44
46
115 (+0.7 σ)
klaudiaGPT54
21
46
46
113 (+0.6 σ)
klaudiaGPT55
19
48
48
115 (+0.7 σ)
klaudiaGrok
23
49
47
119 (+1.0 σ)
klaudiaGemini
24
46
45
115 (+0.7 σ)
kubaSonnet
34
44
49
127 (+1.6 σ)
kubaOpus
34
43
50
127 (+1.6 σ)
kubaGPT
34
48
50
132 (+2.0 σ)
kubaGPT54
34
44
49
127 (+1.6 σ)
kubaGPT55
35
44
50
129 (+1.8 σ)
kubaGrok
31
45
49
125 (+1.5 σ)
kubaGemini
33
46
49
128 (+1.7 σ)
lukaszSonnet
24
39
43
106 (+0.1 σ)
lukaszOpus
26
40
42
108 (+0.2 σ)
lukaszGPT
27
44
47
118 (+0.9 σ)
lukaszGPT54
28
42
47
117 (+0.9 σ)
lukaszGPT55
30
43
47
120 (+1.1 σ)
lukaszGrok
30
43
46
119 (+1.0 σ)
lukaszGemini
27
45
49
121 (+1.2 σ)
magdaSonnet
25
37
42
104 (-0.1 σ)
magdaOpus
26
39
37
102 (-0.2 σ)
magdaGPT
28
40
45
113 (+0.6 σ)
magdaGPT54
20
38
42
100 (-0.4 σ)
magdaGPT55
23
40
40
103 (-0.1 σ)
magdaGrok
21
44
45
110 (+0.4 σ)
magdaGemini
31
42
45
118 (+0.9 σ)
marekSonnet
17
21
22
60 (-3.3 σ)
marekOpus
18
24
27
69 (-2.6 σ)
marekGPT
12
22
22
56 (-3.6 σ)
marekGPT54
11
23
21
55 (-3.6 σ)
marekGPT55
11
23
21
55 (-3.6 σ)
marekGrok
11
23
24
58 (-3.4 σ)
marekGemini
13
24
20
57 (-3.5 σ)
michal-kSonnet
26
39
45
110 (+0.4 σ)
michal-kOpus
29
40
46
115 (+0.7 σ)
michal-kGPT
25
45
47
117 (+0.9 σ)
michal-kGPT54
26
41
47
114 (+0.7 σ)
michal-kGPT55
25
39
47
111 (+0.4 σ)
michal-kGrok
22
42
47
111 (+0.4 σ)
michal-kGemini
23
41
47
111 (+0.4 σ)
michal-simSonnet
17
35
38
90 (-1.1 σ)
michal-simOpus
23
31
36
90 (-1.1 σ)
michal-simGPT
15
40
43
98 (-0.5 σ)
michal-simGPT54
16
35
37
88 (-1.2 σ)
michal-simGPT55
16
38
37
91 (-1.0 σ)
michal-simGrok
15
35
36
86 (-1.4 σ)
michal-simGemini
16
31
38
85 (-1.5 σ)
nataliaSonnet
14
29
27
70 (-2.6 σ)
nataliaOpus
17
27
32
76 (-2.1 σ)
nataliaGPT
17
29
27
73 (-2.3 σ)
nataliaGPT54
14
28
27
69 (-2.6 σ)
nataliaGPT55
13
30
30
73 (-2.3 σ)
nataliaGrok
11
25
26
62 (-3.1 σ)
nataliaGemini
14
25
29
68 (-2.7 σ)
olaSonnet
29
36
27
92 (-0.9 σ)
olaOpus
30
35
27
92 (-0.9 σ)
olaGPT
34
36
31
101 (-0.3 σ)
olaGPT54
27
35
24
86 (-1.4 σ)
olaGPT55
30
34
24
88 (-1.2 σ)
olaGrok
16
37
28
81 (-1.8 σ)
olaGemini
30
46
37
113 (+0.6 σ)
pawelSonnet
17
32
29
78 (-2.0 σ)
pawelOpus
18
28
28
74 (-2.3 σ)
pawelGPT
19
39
41
99 (-0.4 σ)
pawelGPT54
16
35
33
84 (-1.5 σ)
pawelGPT55
16
38
33
87 (-1.3 σ)
pawelGrok
0
38
30
68 (-2.7 σ)
pawelGemini
27
40
37
104 (-0.1 σ)
piotrSonnet
16
30
35
81 (-1.8 σ)
piotrOpus
18
30
34
82 (-1.7 σ)
piotrGPT
16
37
39
92 (-0.9 σ)
piotrGPT54
14
38
39
91 (-1.0 σ)
piotrGPT55
14
36
41
91 (-1.0 σ)
piotrGrok
11
31
36
78 (-2.0 σ)
piotrGemini
12
27
38
77 (-2.0 σ)
radekSonnet
17
41
34
92 (-0.9 σ)
radekOpus
19
41
37
97 (-0.6 σ)
radekGPT
21
44
44
109 (+0.3 σ)
radekGPT54
14
43
40
97 (-0.6 σ)
radekGPT55
13
44
42
99 (-0.4 σ)
radekGrok
11
42
37
90 (-1.1 σ)
radekGemini
13
42
34
89 (-1.2 σ)
saraSonnet
32
47
47
126 (+1.5 σ)
saraOpus
32
46
44
122 (+1.2 σ)
saraGPT
37
49
47
133 (+2.0 σ)
saraGPT54
32
47
46
125 (+1.5 σ)
saraGPT55
34
48
47
129 (+1.8 σ)
saraGrok
33
50
47
130 (+1.8 σ)
saraGemini
30
48
48
126 (+1.5 σ)
tomekSonnet
18
44
41
103 (-0.1 σ)
tomekOpus
22
44
42
108 (+0.2 σ)
tomekGPT
13
45
45
103 (-0.1 σ)
tomekGPT54
15
43
42
100 (-0.4 σ)
tomekGPT55
16
42
41
99 (-0.4 σ)
tomekGrok
16
41
45
102 (-0.2 σ)
tomekGemini
20
41
43
104 (-0.1 σ)
weronikaSonnet
33
45
48
126 (+1.5 σ)
weronikaOpus
33
43
46
122 (+1.2 σ)
weronikaGPT
34
49
48
131 (+1.9 σ)
weronikaGPT54
33
47
48
128 (+1.7 σ)
weronikaGPT55
33
47
47
127 (+1.6 σ)
weronikaGrok
33
49
48
130 (+1.8 σ)
weronikaGemini
33
47
48
128 (+1.7 σ)
zuziaSonnet
25
45
46
116 (+0.8 σ)
zuziaOpus
32
47
49
128 (+1.7 σ)
zuziaGPT
29
47
50
126 (+1.5 σ)
zuziaGPT54
29
47
47
123 (+1.3 σ)
zuziaGPT55
29
47
47
123 (+1.3 σ)
zuziaGrok
31
49
46
126 (+1.5 σ)
zuziaGemini
19
45
46
110 (+0.4 σ)

One-way ANOVA — do the conditions differ significantly?

For each (model × scale) pair, the F statistic compares the persona / baseline / zero-prompt means. Dots are coloured by significance band (p < .001 / .01 / .05). The p value comes from the Wilson-Hilferty transformation.

model scaleDBZ-R lękDBZ-R unikanieMentS ΣKPPTIPI ETIPI ATIPI CTIPI ESTIPI O
Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini
p < .001p < .01p < .05not significantp approximated via the Wilson-Hilferty transformation

Pairwise Cohen's d matrix per scale

A grid over all active models — the effect size between every model pair on a given self-report scale in the first persona administration. Click a cell for Hedges’ g (bias-corrected) and the N of both groups.

SonnetOpusGPTGPT54GPT55GrokGemini
Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini
d < 0 means the row sits below the column · |d| > 0.8 is a large effect, around 0.5 medium, < 0.2 small · Hedges’ g in the drawer
04

Section

Meta — statistical power and classification confidence

Do the observed effects have enough N? Are the attachment-style classifications confident or borderline? Three meta-level charts close the argument about the quality of the study.

Manuscript §3.4 · Discussion §A.2

Power analysis matrix

For each (model × scale) pair: the observed Cohen’s d and the minimum per-group N needed to detect d ∈ {0.5, 0.8, 1.0} at α = .05 with 80% power. Red — undetectable at the present N, green — detectable and large.

model scaleDBZ-R lękDBZ-R unikanieMentS ΣKPPTIPI ETIPI ATIPI CTIPI ESTIPI O
Sonnet
Opus
GPT
GPT54
GPT55
Grok
Gemini
detectable, |d| ≥ 0.8detectable, smaller effectundetectable at the present Nα=.05 dwustronnie, moc 80% (z + Beasley-Springer)

Internal reliability — Cronbach's α per scale and model

Does a model answer the 36 DBZ-R items coherently or erratically? Cronbach’s α is computed at the ITEM level, with reverse coding matching the data-preparation procedure. A low α does not mean the answers are wrong — it means the battery does not form a coherent scale in that model. Key: ≥ .90 excellent · ≥ .70 acceptable · < .50 insufficient · < 0 reversed.

scale modelSonnetOpusGPTGPT54GPT55GrokGemini
DBZ-R · lęk
DBZ-R · unikanie
MentS · suma
KPP
α ≥ .90 excellent.70 ≤ α < .90 akceptowalna.50 ≤ α < .70 questionableα < .50 insufficientα < 0 reversed

Confidence-stratified style classification

For each persona × model pair: margin = min(|anxiety mean − 4|, |avoidance mean − 4|), the distance to the decision threshold. High medians mark confident classifiers; the count of borderline cases (margin < 0.5) measures the model’s hesitation.

Margin = min(|anxiety mean − 4|, |avoidance mean − 4|) — the style classification threshold. A value < 0.5 marks a borderline decision.

Data status

corrected (waves 3–4, fixed stimulus rendering — the manuscript primary). Every chart uses the active model set from the bar above. Per-model charts filter the render; cross-vendor quantities (ICC, Cochran’s Q, Fleiss κ) recompute from the raw data. The filter persists in the URL, so a view can be shared as a link.

eks

Używamy plików cookie do analizy ruchu (Google Analytics 4, PostHog) i wykrywania błędów (Sentry). Zbieramy dane o urządzeniu, przeglądarce i odwiedzanych podstronach. Treści twoich rozmów zostają na urządzeniu. Polityka prywatności