Methodology

Everything on this page was measured on synthetic personas driven by a language model. No human being has been measured. That is the programme's design boundary, stated first because every other claim depends on it. We built an instrument that can be scored against known ground truth — synthetic personas whose generating trait values we hold exactly — so that its failures would be measurable before any human study. Where a limit exists, this page reports it with the same precision as the successes.

One rule governs the writing: every number names the artifact it comes from. A number without a named artifact does not appear here. All numbers on this page come from one run — the pre-registered confirm run of September 2026 (calibration/phaseA/wp26/confirm_result_v1.json). Earlier corpora, generated by an earlier engine over an earlier persona library, were deprecated on 2026-09-02 and are not cited.


One estimator, served and measured

SoulMap has one measurement path, and an earlier version of this page had to explain why the served path and the measured path differed. Since 2026-08-27 they do not: the estimator that produces a user's trait scores is the estimator the calibration programme measures.

It runs in a deliberate uncalibrated configuration: every family's discrimination is pinned to 1 and no nuisance terms are fitted, because no calibration manifest is mounted in production. That is a measured choice, not a placeholder. On the confirm corpus the served configuration recovers planted traits at a median correlation of 0.233 across 27 families; the same estimator with fitted discriminations and nuisance terms from an earlier research fit reaches 0.234 — the same number to within noise (confirm_result_v1.json, recovery_SERVED_CONFIG and recovery_DESCRIPTIVE). Fitted parameters buy nothing yet, so none are served. Everything below describes the instrument in the configuration a user actually gets.


The shelf: the server writes the psychometrics, the model writes the prose

The psychometric core of a session is not the language model. It is a versioned evidence shelf of 366 archetype entries — content hash 7bda5d9bf355 since the wording revision of 3 September 2026, stamped on every session record and every banked corpus row (the confirm corpus below was banked under its predecessor, cfbd7786eb4e; the revision re-worded four option texts and changed no trait loading. A staging cue added to two further entries in the same revision was withdrawn the same day, after a registered re-measurement found it cut how often that pair was offered by about three quarters). Each entry carries authored measurement metadata: signed trait loadings per family, an intent tag, a difficulty rating and a polarity; registered opposite-pole pairs carry a partner.

At each turn the server selects the four answer options by a deterministic, seeded search over the shelf under hard constraints — distinct primary constructs per menu, a cap per registry family on action turns, an opposite-pole gate for the registered pairs — with soft rules that relax down a fixed ladder rather than fail the turn. The language model then writes the scene and the wording of the four buttons; it never assigns psychometrics. After rendering, the server re-attaches each button's trait loadings, intent tag and difficulty from the table. Trait identity on every button is correct by construction, never model-authored.

The model writes the prose. The server writes the psychometrics.

Two honest consequences of this design:

  • The design matrix is the shelf's own declaration of what each button means. A recovery figure is therefore a joint test of the driver and the label registry — not of the driver alone. Where a family reads backwards, the first suspect is the language of its options against its scoring key, and the programme's contract treats a repeated negative as blocking until that has been traced by reading, not by statistics.
  • Delivery counts, not shelf counts. On the confirm corpus, 312 of the 366 entries were offered at least once; the 54 never offered are almost all reflection-turn entries. Every one of the 27 families was genuinely contrasted — offered with differing loadings on the same menu — on at least 294 of 3,822 rows, and on as many as 1,908 (attachment avoidance).

The choice model

The unit of analysis is one dealt menu: four shelf-backed options plus a free-text alternative, and the observed pick. The measurement model is a conditional logit. The utility of option j for person i is:

U(i, j) = Σ over families f of λ_f · s(j, f) · θ(i, f)

where s(j, f) is the shelf's signed loading of option j on family f and θ(i, f) is the person's trait score on that family. In the served configuration every λ_f = 1 and there are no nuisance terms; choice probabilities follow by softmax over the alternatives. A person's θ is the maximum-a-posteriori estimate under a standard-normal prior given their banked choices, with a posterior standard error per family. That standard error is what the product turns into the interval and the confidence word beside every reading.

Two things the served model deliberately does not do, and why:

  • No fitted discrimination. Measured to add nothing on the current engine (0.233 versus 0.234 above); on earlier corpora it measured slightly worse. Pinning λ to 1 also removes a dependency on any fit artifact.
  • No position or difficulty terms. Real position bias exists in synthetic choice; the served model absorbs none of it. This is a known simplification, and the in-run comparison above says it costs nothing measurable at present.

The synthetic-persona programme

Ground truth. Personas come from a library sampled by a Gaussian-copula forge with no language model involved: 3,280 personas in the v5 library (persona_library_v5.jsonl plus a 400-persona mid-grid supplement), each with 29 generating trait scales that project onto 27 measurement axes — the only two-to-one collapses are the two Schwartz bipolar contrasts (family_map_v4.json, the axis map's version label; unrelated to the deprecated library v4). A persona's θ on a family is known exactly, which is the entire point of synthetic ground truth: recovery is scored against truth, not against another questionnaire. The trait dimensions are named after public instruments; no instrument was administered to anyone.

The persona instruction. A deterministic sheet of banded, categorical trait renderings — never the raw trait vector. The v2 sheet names every construct (an earlier sheet named only the top two moral foundations and one attachment word, and its under-mention was measured), reads the trait ladder in a randomised order per persona that is stamped on every row, and had every band's wording placed by blind readers at the level it claims — fourteen of fifteen upper emotional-intelligence bands place exactly (calibration/phaseA/wp24/). Age and gender condition the persona sampling but are never conveyed to the driver.

The driver. One model family plays every persona on the confirm run. Earlier multi-driver corpora are deprecated, so nothing on this page separates "the trait" from "how this model plays it"; a second driver family is the named reopening condition for the programme's refuted sign-correction lever.

Run integrity. 40 personas × 50 planned sessions in chains of 20, so a session begins from the faded state its predecessor left. 1,831 of 2,000 completed (8.45% attrition against a 15% gate; wave_summary.json); completed sessions per persona median 46. The failure that stopped every persona short of 50 was located to the resume path of chained runs, instrumented, fixed and tested (the calibration page has the account).


The testing discipline

Pre-registration. The confirm run's primary question, statistic, null construction and threshold were filed before the run (calibration/phaseA/prereg/WP26_instruction_remediation_confirm.md), and the analysis was run once on the finished corpus with a fixed seed. Two earlier registered tests in this programme missed their thresholds and were published as misses; that is what the discipline is for.

The model-free statistic. To test whether a family expresses — whether the persona's choices lean the planted way at all, before any estimator is involved — the programme uses expression: for each persona and family, how far the chosen option's loading sits above its own menu's mean, on menus where the family was genuinely contrasted, correlated with the planted trait across personas. It touches no fitted parameter, so an estimator cannot flatter it.

The shuffled null. "Could this be chance?" is answered by keeping every scene, option and choice exactly as it was and relabelling only which persona each set of choices belongs to — 5,000 times, one relabelling shared across all twelve families so their real between-family correlation survives. Relabelling each family separately would produce a fictitiously tight bar the corpus clears by construction.

Bootstrap intervals per family. Recovery per family carries a 95% interval from resampling the 40 personas. At this roster the intervals are about ±0.27 wide, and they did not narrow between an interim read at 27 sessions and the finish: their width is set by the number of personas, so per-trait certification is a persona-count question (~170 personas for ±0.15, ~380 for ±0.10), not an engine question.

Split-half consistency. Score each persona from half their sessions and from the other half, correlate the two, and apply the Spearman–Brown correction: does the engine describe the same person the same way twice? Only meaningful at depth — at 27 sessions the halves are too thin — which is why the run was not cut short of 50.

Refutations, on the record. Levers the programme has tested and found dead are recorded with the evidence and a reopening condition (graph/refuted.yaml): adaptive item selection, item purity, item retirement, loading sharpening, statistical sign correction, and content re-authoring for families the persona does not express. Three of them looked large in-sample and reversed sign held out; every gain is cross-validated before it is believed.


What the discipline measured

The twelve, pooled
+0.126

Expression of the twelve benchmark families against a bar of 0.082 fixed in advance; p = 0.0042 against the run's own shuffled null — confirm_result_v1.json

Served recovery
0.233

Median across 27 families; 21 of 27 read in the planted direction; 10 confirmed individually; 1 confirmed backwards and withheld

Split-half
0.149

Median per-person consistency at ~46 sessions; honesty–humility 0.86 is the one family above the 0.70 bar

Detection floor
0.49

Smallest per-family effect the 40-persona design can resolve (0.75 for the hardest family); per-family figures are descriptive by registration

Detection and measurement are different bars, and this instrument is much stronger at the first. Reading a population is well established on the strong families; reading an individual steadily is established on one. The validation page carries the full per-family table.


What this methodology does not claim

  • No human data exists anywhere in the programme. Every θ, every interval, every reliability figure was computed on synthetic personas.
  • No per-family significance. The design was powered for one pooled question; the per-family table is description with its detection floor stated.
  • No before/after. The current engine is not compared with the deprecated corpora, because the engine, the persona generator and the instruction changed together.
  • No conformal coverage. No coverage guarantee has been measured for the served intervals; they are posterior intervals under a standard-normal prior and are labelled as such.
  • No measurement invariance has been tested. It cannot be, before a real-user cohort exists.
  • The persona population's Σ-fidelity check missed its registered bar — relative Frobenius distance ≈0.12 against a 0.10 tolerance — and the measurement stands as a miss.

The calibration page carries the run record; validation carries the results and the withdrawn-figure watchlist.