The calibration programme

Every result on this page was measured on synthetic personas driven by a language model. No human being has been measured — by this programme or by any other part of SoulMap. "Detected" means detected in a simulated population whose true trait values we set ourselves.

This page is the run record. What the run showed is on the validation page; how the measurement works is on the methodology page. Here: what was registered, what was executed, what it cost, what broke, and the fences around each artifact.


What "current engine" means, and what was deprecated

On 2026-09-02 every corpus generated before the current engine was deprecated for measurement purposes. Three things had changed at once: the session engine that deals the options, the persona generator (library v5 replaces v4, 3,280 personas each, disjoint identities), and the persona instruction that tells the driver who to be. A figure from an older corpus and a figure from the new one are not comparable, and this programme does not compare them. The deprecated artifacts stay in the repository's history under their original paths; nothing on these pages cites them.

The current instrument is stamped, not assumed. Every row of the confirm corpus carries the evidence-shelf hash (cfbd7786eb4e), the family-axis version (v4-facet-axes-2026-08-10 — 27 measurement axes over 29 library scales; the axis version label is unrelated to the deprecated library version), the persona-sheet version (v2_dossier) and the realised order in which that persona read its trait ladder (verify_smoke.py, 6 of 6 checks).


The confirm run

Registration. calibration/phaseA/prereg/WP26_instruction_remediation_confirm.md, filed before the run. One primary question — do the twelve historically unreadable families express above chance inside the corpus — with the statistic, the null construction and the threshold (pooled mean > 0.0822) fixed in advance. The five classic families as an internal yardstick. Machiavellianism named for explicit report whatever its sign. Per-family figures declared descriptive, with the design's detection floor stated (0.49 median, 0.75 worst, in full-depth units).

Design. 40 personas from the v5 library, stratified by the assignment arm; 50 planned sessions each in chains of 20, so a session starts from the faded state its predecessor left — the engine's continuity machinery sees the persona's own recent past, capped at twenty sessions of history, which is where continuity saturates. Cohorts of eight share dealt scene nodes (one paid generation serves eight personas' decisions, which are their own). One driver model throughout. Sheet order randomised per persona and stamped.

Execution. Batch tier, 1–3 September 2026, 39.6 hours of wall clock across eight process attempts. Two defects were found and fixed during the run, each with the fix proven on the live run before it continued:

  • A decorative field in the option-authoring schema (a two-pole "axis" label) was chaining poles indefinitely until the output cap cut the response mid-string; the whole menu then failed to parse. Request drop rate before the fix: 26% (7 of 27 in the measured round); after deleting the field: 0 of 809, then 21 of 5,416 for the remainder (0.4%).
  • A sixty-minute batch-job deadline, added to stop a stall from consuming the whole night, raised on expiry and killed a healthy five-hour process at 37%. It now leaves the job's requests to resubmit, as its own comment had always promised, and the deadline was widened to three hours after a 47-minute round cleared the old one with thirteen minutes to spare.

Delivered. 1,831 of 2,000 sessions completed (8.45% attrition against the registered 15% gate), 3,822 banked choice rows, $123.16 (wave_summary.json, checkpoint_spend.json). The corpus, config, summary and per-round batch statistics are banked under calibration/phaseA/wp26/runs/; the registered analysis is confirm_analysis.py, and its output confirm_result_v1.json records the rows file's SHA-256 and the git commit it ran from.

Attrition
8.45%

169 of 2,000 sessions; registered gate 15% — passes. wave_summary.json

Completed per persona
median 46

min 42, max 49, none at 50. The fifty-sessions-per-subject milestone misses at 46 — graphq kpi

Shelf entries offered
312 / 366

Never offered: 54, almost all reflection-turn entries. Every family contrasted on ≥ 294 rows — graphq coverage

Cost per session
$0.067

Against a $0.039 estimate carried into the budget; the difference is the deep-history prompts of chained play


Where the 169 sessions went

Two causes, unlike each other, and both instrumented rather than inferred.

100 sessions — replay drift, all at chain links 2 to 7. A resumed process replays banked decisions and rebuilds each chain link's starting state by re-walking its predecessor. On a chained run that is only sound if the re-walk reproduces the original exactly; across eight restarts it did not always, and the failure signature is unambiguous: every one of these 100 sat at the links that were live across a restart, and died when the banked choice no longer matched a regenerated menu. The mechanism had been seen once before on an earlier-generation run; what was new is that resuming a chained run re-triggers it. The fix — bank the faded hand-off state the moment it is computed and restore it on resume, never re-derive it — shipped on 2026-09-03 with a test that reproduces the drift. It is the reason no persona reached 50 sessions, and the milestone is expected to be reachable on the next run without changing anything else.

69 sessions — engine validation, spread across all links. The model produced a menu the engine refused after its repair attempts (~3.5%, the ordinary floor).

Of the 169, 140 still banked their earlier turns as rows stamped incomplete; 29 produced nothing. Attrition never discards evidence that was already paid for.


What each artifact may and may not support

The dataset census (calibration/DATASETS.md) fences every artifact. For the current generation:

ArtifactWhat it isMay supportMay not support
confirm corpus (wp26/runs/, 2026-09-03)40 v5 personas × ~46 completed sessions, chained, one driver, 3,822 rowsThe registered pooled test; descriptive per-family expression and recovery; split-half at full depth; run-integrity and cost figuresPer-family significance claims (detection floor 0.49); any claim about humans; any comparison with a deprecated corpus; separating the trait from how this one driver plays it
confirm_result_v1.jsonThe registered analysis, run once, seeded, hash-stampedEvery number on the validation pageAnything not in it — there is no second analysis
evidence packs (wp27/packs/)Per-family dossiers for the twelve weakest families: bands, shelf actions on both sides with offer counts, run figuresThe semantic re-verification described on the roadmapMeasurement claims of any kind
pre-v5 corpora (r2–r6, wp14, union)Older engine, v4 library, six driversNothing on these pages — deprecated 2026-09-02Any comparison with the current engine

On the record, and open

  • The final holdout is unspent. A 300-persona holdout was frozen by hash on 2026-08-07 and has never been consulted. It belongs to the deprecated library generation and will be re-sealed against v5 before it is spent.
  • One driver. The confirm run was driven by a single model family by decision (earlier multi-model corpora had shown no material driver difference on the statistics of interest, but those corpora are now deprecated, so the finding travels as a design choice, not a result). A second driver family is the named reopening condition for the programme's refuted sign-correction lever.
  • Per-trait certification needs personas, not sessions. The per-family intervals did not narrow between 27 and 50 sessions; roughly 170 personas would pin each family to ±0.15, and ~380 to ±0.10.

What the run showed is on the validation page; what happens next is on the roadmap.