Validation
Everything on this page was measured on synthetic personas driven by a language model. No human being has been measured anywhere in this programme. That is the design boundary of the work so far, and it belongs in the first sentence, not a footnote. Whether any result below transfers to people is unestablished, and no amount of further synthetic work can establish it.
This page rests on one run: the pre-registered confirm run of 1–3
September 2026 (calibration/phaseA/wp26/). Every number here comes from
its banked analysis, confirm_result_v1.json, computed once on the finished
corpus with a fixed seed and stamped with the rows' hash. Earlier
generations of the programme — an older corpus engine, an older persona
library, and every figure derived from them — were deprecated on
2026-09-02 and are not cited anywhere on this page. By decision, old and
new data are not compared: the engine, the persona generator and the
persona instruction all changed together, so a "before and after" number
would span several confounded changes at once. A one-line note on those
earlier runs is at the end.
One rule governs the writing: every number names the artifact it comes from. A number without a named artifact does not exist here.
The run
Drawn from the v5 library (3,280 synthetic personas); each plays 50 planned sessions in chains of 20, so later sessions see the persona's own recent past — confirm_result_v1.json
8.45% attrition against the registered 15% gate; 3,822 banked choices — wave_summary.json
$0.067 per session on the batch tier; 39.6 hours wall clock across eight process restarts — checkpoint_spend.json
gemini-3.7-flash plays every persona; the sheet order is randomised per persona and stamped on every row — wave_config.json
The instrument is the production session engine itself, played by synthetic
personas under an isolation harness (replayed model transport, in-memory
storage, stubbed images). Each turn deals four options from a versioned
evidence shelf whose trait loadings are written by the server, never by the
model; the persona's pick is banked as one choice row. The shelf identity
(evidence_table_sha cfbd7786eb4e) and the family-axis version
(v4-facet-axes-2026-08-10, 27 axes over 29 scales) are stamped on all
3,822 rows (verify_smoke.py, 6 of 6 completion checks passed).
Two boundaries on the corpus itself:
- Coverage was complete at the family level and thin at the entry level.
All 27 families were contrasted on at least 294 rows; 312 of the shelf's
366 entries were offered at least once, and the 54 never offered are almost
all reflection-turn entries (
graphq coverage). Family-level claims below are on solid footing; per-entry claims are not made. - Each persona is bound to one driver, and there is only one driver. No claim on this page separates "the trait" from "how this model plays the trait". A second driver family is a named reopening condition in the programme's refutation record, not a footnote.
The pre-registered question, and its answer
The registration (calibration/phaseA/prereg/WP26_instruction_remediation_confirm.md)
asked one question about the twelve families the programme had never been
able to read — the Dark Triad, the emotional-intelligence facets, the moral
foundations, attachment anxiety, resilience, cognitive flexibility:
Do the twelve express above chance, inside this one corpus?
"Express" is model-free: for each persona and family, how far the chosen option's loading sits above its own menu's mean, on menus where the family was genuinely contrasted; then the correlation of that leaning with the persona's planted trait, across the 40 personas; then the mean over the twelve. The bar — pooled mean above 0.0822 — was fixed before the data existed, as the 95th percentile of a shared-relabelling null.
Result: the twelve express above chance. Pooled mean +0.126. Against
this corpus's own null — 5,000 relabellings of which persona each set of
choices belongs to, one relabelling shared across all twelve families so
their real correlation survives — the null averages +0.0005 with a 95th
percentile of +0.078, and the observed value is exceeded by chance in
0.42% of draws (confirm_result_v1.json, H1). The pre-registration
allows exactly two statements from this, and this page makes only them:
- Under the current stack, the twelve do express above chance.
- They sit 0.226 below the five control families — agreeableness,
conscientiousness, extraversion, honesty–humility, openness — which pool
at +0.352 in the same corpus under the same driver and instruction
(
confirm_result_v1.json, H2).
No claim of "improvement" is made: the persona instruction and the driver changed together between generations and cannot be separated.
The third registered item, machiavellianism, was named in advance because it
had read negative before. It reads negative again: −0.117 in expression
and −0.102 in served recovery (confirm_result_v1.json, H3). Its sign chain
has been traced by hand and is correct; the negative number is reported, not
corrected, and the construct is under the semantic re-verification described
on the roadmap.
Family by family
Per-family figures are descriptive, by registration. The design — 40
personas — was sized to answer the pooled question above. Its smallest
detectable per-family effect is about 0.49 in full-depth units for the
median benchmark family and 0.75 for the hardest, so an individual family
that fails to clear zero has not been shown flat; it has not been
resolved (confirm_result_v1.json, declared_limits).
Three statistics per family, all against planted truth across the 40 personas. Expression is the model-free leaning above. Recovery is the estimate the product's own estimator produces — in the configuration a user actually gets, every discrimination pinned to 1 and no nuisance terms — with a 95% interval from resampling the personas. Split-half is consistency: score each persona from half their sessions and from the other half (about 23 each), correlate, Spearman–Brown corrected.
| family | expression | recovery (served) | 95% interval | split-half | confirmed on its own |
|---|---|---|---|---|---|
| honesty–humility | +0.727 | +0.714 | +0.54 … +0.83 | +0.86 | yes |
| psychopathy | +0.600 | +0.640 | +0.45 … +0.79 | +0.13 | yes |
| narcissism | +0.571 | +0.520 | +0.25 … +0.73 | +0.18 | yes |
| self-transcendence | +0.524 | +0.485 | +0.20 … +0.71 | +0.45 | yes |
| openness to change | +0.455 | +0.424 | +0.14 … +0.66 | +0.08 | yes |
| authority / subversion | +0.387 | +0.368 | +0.09 … +0.60 | +0.46 | yes |
| sanctity / degradation | +0.378 | +0.367 | +0.10 … +0.56 | +0.04 | yes |
| agreeableness | +0.344 | +0.322 | +0.04 … +0.55 | +0.19 | yes |
| attachment avoidance | +0.247 | +0.317 | +0.03 … +0.56 | +0.36 | yes |
| stress tolerance | +0.317 | +0.303 | −0.01 … +0.58 | +0.40 | — |
| extraversion | +0.341 | +0.299 | +0.04 … +0.53 | +0.07 | yes |
| openness | +0.273 | +0.256 | −0.05 … +0.54 | +0.07 | — |
| cognitive flexibility | +0.238 | +0.253 | −0.19 … +0.56 | +0.57 | — |
| loyalty / betrayal | +0.235 | +0.233 | −0.01 … +0.46 | +0.01 | — |
| empathy | +0.180 | +0.219 | −0.09 … +0.50 | −0.42 | — |
| care / harm | +0.234 | +0.213 | −0.13 … +0.50 | +0.22 | — |
| fairness / cheating | +0.155 | +0.178 | −0.12 … +0.45 | +0.35 | — |
| emotion granularity | +0.155 | +0.121 | −0.17 … +0.37 | +0.18 | — |
| neuroticism | +0.071 | +0.095 | −0.29 … +0.46 | +0.17 | — |
| emotional awareness | +0.103 | +0.092 | −0.21 … +0.38 | −0.03 | — |
| conscientiousness | +0.077 | +0.058 | −0.31 … +0.36 | +0.22 | — |
| attachment anxiety | +0.124 | −0.009 | −0.30 … +0.23 | −0.01 | — |
| resilience | +0.038 | −0.017 | −0.26 … +0.23 | −0.16 | — |
| impulse control | −0.098 | −0.086 | −0.30 … +0.14 | +0.15 | — |
| machiavellianism | −0.117 | −0.102 | −0.36 … +0.18 | −0.15 | — |
| flexibility (EI) | −0.153 | −0.185 | −0.50 … +0.18 | −0.02 | — |
| liberty / oppression | −0.284 | −0.265 | −0.48 … −0.00 | −0.37 | backwards |
Source: confirm_result_v1.json, per_family_expression_DESCRIPTIVE and
recovery_SERVED_CONFIG. Sorted by served recovery. The same estimator with
fitted discriminations from an earlier research fit reaches a median of
0.234 against 0.233 here — the fitted parameters buy nothing yet, which is
why none are served (recovery_DESCRIPTIVE).
Read as four groups:
- Ten families are confirmed individually — the interval sits clear of zero on the positive side. Two of the three Dark Triad scales are among them. Median served recovery across all 27 is 0.233; 21 of 27 read in the planted direction.
- Thirteen are undetermined: positive or near-zero point estimates, intervals crossing zero. The intervals are about ±0.27 wide, and that width did not narrow between an interim look at 27 sessions and the finish — it is set by the number of personas, not sessions. Resolving these traits one by one is a matter of a larger persona roster, not of engine work.
- Three read clearly backwards without being resolved: EI flexibility,
machiavellianism, impulse control. Their intervals cross zero, but a
repeated negative is a blocking finding under the programme's contracts
(
graph/contracts.yaml,pole_inversion_chain). - One is confirmed backwards: liberty/oppression, whose interval sits entirely at or below zero and which did not move between the interim read and the finish.
Those four families are withheld from the product until their language has been re-verified against their scoring keys. The most striking non-result among the undetermined is conscientiousness: a Big Five family, named in every persona's instruction, contrasted on 38% of all rows — the most-offered family after attachment — and still flat at +0.077. That is neither under-mention nor under-exposure; it points at what the scenes' options actually encode for the trait, and it is under the same re-verification.
Reliability at fifty sessions
Population-level recovery and per-person consistency are different questions, and this instrument is much stronger at the first. The split-half column above answers: does the engine describe the same person the same way twice?
| value | |
|---|---|
| median split-half reliability, 27 families, served configuration | 0.149 |
| families at or above 0.30 | 7 |
| families at or above 0.50 | 2 — honesty–humility 0.86, cognitive flexibility 0.57 |
| families below zero | 7 — empathy −0.42, liberty/oppression −0.37, resilience −0.16, machiavellianism −0.15, and three within rounding of zero |
Source: confirm_result_v1.json, recovery_SERVED_CONFIG.
Stated plainly: at fifty sessions the engine reads a population well on the strong families and reads an individual steadily on only a handful. Two families are not merely uncertain but inconsistent. Every served reading in the product carries its own uncertainty band and a plain confidence word derived from it; the 0.70 bar the programme sets for a decision-grade individual reading is met by one family, honesty–humility.
What went wrong, and what it cost
169 of 2,000 sessions failed. The attrition KPI (≤ 15%) passes at 8.45%; the
fifty-sessions-per-subject milestone does not — completed sessions per
persona were median 46 (min 42, max 49, none at 50) (graphq kpi).
The cause is instrumented rather than assumed. 100 of the 169 died with the same error, and every one sits at chain links 2–7 — exactly the links that were live across the run's eight process restarts. A resumed process rebuilt each link's starting state by re-walking its predecessor; when that re-walk drifted by a byte, the successor's prompt changed, its cached response missed, the scene regenerated, and the banked choice matched no button. The remaining 69 are ordinary validation failures spread across all links (~3.5%). The fix — banking each link's hand-off state so a resume restores it instead of re-deriving it — shipped on 2026-09-03 with a test that reproduces the drift; the milestone is expected to be reachable on the next run without changing anything else.
What this page does not claim
- Nothing about human beings. Every artifact above describes model behaviour against synthetic ground truth.
- Nothing per family beyond description. Only the pooled question was powered; the per-family table is what 40 personas can show, with its detection floor stated.
- No before/after. Earlier-generation figures are deprecated and not compared against.
- No decision-grade individual reading except where the table above says so, and the product's serving policy enforces the same rule.
Earlier generations, for the record
Between May and August 2026 the programme ran an earlier corpus engine
against an earlier persona library (v4), and published figures from those
corpora on this page. Those figures were withdrawn in two rounds — a
primary-artifact audit on 2026-08-19, and the deprecation of every pre-v5
corpus on 2026-09-02 once the engine, generator and instruction had all
changed. The withdrawn numbers are guarded by a build-time watchlist
(graph/claims.yaml) so they cannot quietly reappear on a public surface.
The artifacts remain in the repository's history; nothing on this page rests
on them.