A pilot measurement of Twyn answer-stability across repeated runs
Twyn / twyns.net · Study S1 (pilot) · draft 2026-06-05 · McKay Consulting
Status: PILOT. This document reports a first, deliberately small execution of the pre-registered Study S1 (
docs/prereg-S1-reproducibility.md). “Pre-registered” means the design and the metrics were written down and committed before we ran anything, so we cannot quietly change the rules after seeing the results — a basic safeguard against fooling ourselves. The pilot runs 4 Twyns × 1 research guide × 30 repetitions × 2 temperature arms = 240 live model calls. (A “Twyn” is one synthetic user; a “guide” is one fixed interview prompt; a “repetition” is one independent re-asking of that prompt; a “temperature arm” is one setting of the model’s randomness dial — more on that below.) A run of this size is enough to establish the method and the direction of the findings; it is not the full locked grid (the larger, final version of the study, which additionally tests whether the Twyn answers the same way when the question is re-worded — “paraphrase robustness” — and adds multiple guides and per-model arms). We flag this openly so the results are read for what they are: a credible first measurement, not a finished proof. Every number here is reproducible — anyone with the files can regenerate it — fromreports/validation/s1-temp07-results.jsonands1-temp0-results.jsonusingmeta/validation/s1-pilot.mjs.

Figure 1 — The live methodology panel (twyns.net/#methodology). Twyn already states, in-product, that its scores are “transparent heuristics over observable signals — not validated predictions.” This white paper is the measurement that begins turning that honesty into evidence.
What Twyn is — and why this paper exists
If you are meeting Twyn for the first time, start here. Twyn is a research tool that builds authentic digital twins of your users: realistic, evidence-grounded AI “users” that a product or research team can interview — each in its own voice, at the speed of a conversation. Instead of guessing in the long, expensive gaps between rounds of fieldwork, a team can put an idea, a message, a price or a design to a panel of these twins and hear how different kinds of user react — the range of views, the points of friction, the language real people use, the objections they raise — before committing to build.
Why the tool exists. Real user research is indispensable but scarce: it is costly, slow and hard to repeat, so teams too often build on untested assumptions about who their users are and what they need. Twyn is designed to make the user’s perspective available continuously and affordably — and, importantly, it is a supplement to talking to real people, never a replacement. It helps you decide where to look and what to ask; the final word still belongs to real research.
Why this paper. A tool that claims to speak for your users only earns trust if it is honest about how well it actually works — so we measure it, and publish what we find, limits included. Twyn’s answers come from a large language model, which is stochastic: ask it the same thing twice and the wording — sometimes the substance — can drift. The first, most basic question any buyer is entitled to ask is therefore about reliability: ask the same twin the same question many times — how stable is the answer? Reliability is a precondition for everything else; a measurement that wobbles when you simply repeat it can’t be trusted, however plausible any single answer sounds. This (pilot) study measures exactly that.
1. The question
This section sets out the single, narrow thing this pilot is trying to answer.
A large language model — the kind of AI that powers Twyn — is stochastic, meaning it has a built-in element of randomness: ask it the same thing twice and the wording — and sometimes the substance — can change. That is normal behaviour for this technology, not a fault. But if Twyn is to be used to inform real research decisions, a buyer is entitled to ask a blunt question: ask the same Twyn the same question 30 times — how stable is the answer? If the answer lurches around on every re-ask, you cannot rely on it; if it lands in the same place again and again, you can.
The distinction that runs through this whole paper is reliability versus accuracy. Reliability means “does it give a consistent answer when repeated”; accuracy (more precisely, validity) means “does that answer match what a real human would say.” These are independent: a Twyn can be perfectly reproducible — saying the same thing every time — and still be wrong about real people. Reliability is the necessary first floor: an instrument that is not even consistent with itself cannot be trusted for anything. Validity is a separate question, tackled in later Studies S2–S3. S1 deliberately isolates reliability and measures it honestly, including the parts where Twyn performs poorly.
2. Method (pilot)
This section describes exactly what we did, so the numbers below can be judged on their setup. In plain terms: we picked four synthetic users, asked each one the same question 30 times, did that twice at two different randomness settings, and recorded the answers in a structured way we could count.
- Twyns (4, frozen): A · Astrid Lund (Denmark), K · Marcus Thompson (United States),
L · Aiko Hasegawa (Singapore), S · Annika Vogel (Germany). “Frozen” means their definitions did not change during the study. They were chosen to span different cultures, so the test is not run on four near-identical personas.
- Stimulus (1 guide): a fixed concept-test prompt (the Twyns’ reaction to “AI digital twins of your
customers”). The prompt asks for a natural, free-text answer plus three specifically tagged fields, so that each answer can be parsed (read by software) and counted consistently rather than interpreted by hand: ADOPT (a yes/no/maybe decision), LIKELIHOOD (a 0–10 rating of how likely they are to adopt), and TOP_CONCERN (their main worry, in free text).
- Repetition: k = 30 independent calls per Twyn. Thirty is a common minimum sample size in
statistics — large enough for averages and spread to stabilise, while keeping a pilot cheap and fast.
- Temperature arms: “temperature” is the model’s randomness dial — higher means more varied wording,
lower means more repetitive. We ran the everyday production setting (0.7) and, as a reference point, a 0 arm (the lowest, most deterministic setting the model allows).
- Prompt: we used the real production interview system prompt — the actual instructions the live
product gives the model — extracted verbatim from the product’s source file index.html (buildInterviewSystemPrompt + describePsychometrics + describeHofstede). This matters because it means we measured the actual shipping product, not a simplified stand-in built for the test.
- Metrics: for each answer type we chose a measure suited to it — explained in plain language as it
comes up in Section 3. For the decision (ADOPT): modal-category agreement and entropy. For the rating (LIKELIHOOD): mean, standard deviation, coefficient-of-variation, and ICC(2,k). For the concern text (TOP_CONCERN): pairwise lexical cosine and Jaccard overlap.
3. Results
This section reports what came back, axis by axis, from strongest to weakest. The headline is that decisions and ratings repeat well, while exact wording does not — and we are explicit about both.
3.1 Structured-output reliability
Before measuring whether the content repeats, we first check the plumbing: did every run come back in the structured, machine-readable shape we asked for? 240 / 240 runs (100%) returned all three fields (ADOPT, LIKELIHOOD, TOP_CONCERN) in parseable form — nothing had to be discarded or hand-fixed. The structured-elicitation format (asking for tagged fields rather than a free essay) is itself highly reliable, which is a practical prerequisite for everything that follows.
3.2 Decision reproducibility (ADOPT) — HIGH
This is the most decision-relevant axis: when you re-ask the same Twyn, does it land on the same yes/no/ maybe verdict? “Modal decision” simply means the most common answer across the 30 runs; “agreement” is the percentage of runs that matched it. Every Twyn returned the same modal decision in 80–100% of runs — i.e. the verdict was stable far more often than not.
| Twyn | Home | Modal | Agreement (0.7) | Agreement (0) |
|---|---|---|---|---|
| A Astrid | Denmark | maybe | 100% | 100% |
| K Marcus | United States | maybe | 100% | 100% |
| L Aiko | Singapore | maybe | 100% | 100% |
| S Annika | Germany | maybe | 87% | 80% |
| Pooled | ~97% | ~95% |
How to read the table: three of the four Twyns gave the identical decision on every single run (100%), and the most variable, Annika, still agreed with herself 87% of the time at the production setting and 80% at temperature 0. Pooled across all four, decisions repeated about 97% of the time (production) and 95% (reference) — high reproducibility by any practical standard.
Caveat (the alpha paradox): there is an honest statistical wrinkle here. Because almost every answer across every Twyn was “maybe”, the spread of categories is degenerate — there is barely any variety to measure. In that situation, a standard agreement statistic called Krippendorff’s α becomes uninformative (it reads 0.08–0.14, which looks terrible) despite the near-perfect agreement we can plainly see in the table. This is a well-known quirk called the prevalence problem: when one answer dominates, α has almost no variation to work with and collapses, even though reliability is genuinely high. So we deliberately report modal agreement instead, which is not fooled by it. This is also a useful product finding in its own right: a single yes/no/maybe item carries very little information, so richer ways of eliciting opinions are needed (see §5).
3.3 Rating reproducibility (LIKELIHOOD 0–10) — HIGH
Here we check whether the 0–10 likelihood rating is stable on repetition. Two things matter: how tightly each Twyn’s own rating clusters across its 30 runs (low scatter = reliable), and how well the four Twyns separate from each other (good separation = the instrument can tell personas apart). Both held up.
| Twyn | Mean ± SD (0.7) | CV (0.7) | Mean ± SD (0) | CV (0) |
|---|---|---|---|---|
| A Astrid | 5.33 ± 0.71 | 0.133 | 5.50 ± 0.63 | 0.114 |
| K Marcus | 5.07 ± 0.45 | 0.089 | 5.17 ± 0.38 | 0.073 |
| L Aiko | 5.87 ± 0.35 | 0.059 | 5.43 ± 0.50 | 0.093 |
| S Annika | 2.90 ± 0.40 | 0.139 | 2.80 ± 0.41 | 0.145 |
Reading the table: “Mean” is each Twyn’s average rating across 30 runs; “SD” (standard deviation) is how much individual runs typically stray from that average; “CV” (coefficient of variation) is the SD expressed as a fraction of the mean, which lets you compare scatter fairly across Twyns with different averages. Lower CV means more consistent.
- Coefficient of variation ≈ 6–15% — in practical terms, the same Twyn’s likelihood rating moves by
roughly ±0.5 points on a 10-point scale from run to run. That is small enough that the rating is meaningful: re-asking does not flip a 6 into a 3.
- ICC(2,k) = 0.872 (temp 0.7) and 0.874 (temp 0) — ICC, the intraclass correlation coefficient, is a
standard reliability score running from 0 (pure noise) to 1 (perfect). By the usual rule of thumb, values above 0.75 are considered “good” and above 0.9 “excellent,” so 0.872 sits firmly in the strong range. It captures two good properties at once: the same Twyn is consistent on repetition (test-retest reliability) and the Twyns are clearly distinguishable from one another (discriminability). One honest qualification: ICC here treats the 4 Twyns as the units being compared, and with only 4 units the figure is indicative rather than definitive — it is to be confirmed on the full grid.
3.4 Wording reproducibility (TOP_CONCERN) — LOW and VARIABLE
This is the honest weak axis, and we lead with that. We are asking the hardest version of the question: does the exact phrasing of the top concern repeat word-for-word? It does not. Across runs, the lexical cosine (a 0–1 measure of how much two texts share the same words, where 1 is identical wording and 0 is no shared words) ranged the full span 0.00 to 1.00, and the number of distinct phrasings was 4–27 out of 30 — i.e. for some Twyns almost every run was worded differently. However, the underlying theme was stable: e.g. Astrid repeatedly circled “synthetic data mistaken for validated insight”, and Annika repeatedly raised “mistaking simulation for customer reality”, even when the sentences themselves differed every time. In short: surface strings change; meaning largely holds. This is precisely why the right way to measure free-text reproducibility is at the theme (semantic) level rather than by string matching, and why the full study will use sentence embeddings (which compare meaning, not words) plus human theme coding. The lexical metric reported here is a deliberately conservative lower bound — it makes Twyn look worse than it really is on meaning, which is the safe direction to err in.
3.5 Temperature 0 does not buy determinism
A natural buyer question is: “can’t you just turn the randomness off and get the same answer every time?” The reference arm answers that directly. Setting temperature to 0 — the lowest setting available — changed reliability barely at all (ICC moved only from 0.872 to 0.874; ratings and modal decisions were essentially unchanged; and wording still varied for 3 of the 4 Twyns). Important honesty point: you cannot eliminate run-to-run variation simply by lowering temperature. Some stochasticity is inherent to how these models compute. The correct engineering stance, therefore, is to treat variation as a property to be measured and disclosed — which is what this whole study does — rather than something that can be switched off.

Figure 2 — Twyn A · Astrid Lund (one of the four Twyns in this study), as she appears in the product: an Authenticity score of 78, a Predictive/consistency estimate of 91%, and plain-language change-rate labels. S1 is the work that makes the Predictive number a measured quantity rather than a heuristic.
4. What this means
This section translates the results into what a buyer should actually take away — and how to use Twyn responsibly given what it does and does not do well.
- Decision- and rating-level reproducibility is strong (modal agreement ~96%, rating ICC ~0.87).
For the kinds of outputs Twyn actually surfaces to a user — a direction (adopt / don’t), a relative rating, a ranked concern — repeated runs are stable enough to be useful in practice, provided results are read at the theme/decision level rather than treated as a precise oracle. This is the core commercial reassurance: re-running an interview will not hand you a contradictory answer.
- Surface wording is not reproducible run-to-run, and we say so. The practical guidance follows
directly: treat verbatim quotes from a Twyn as illustrations, not fixed facts, and trust the recurring theme rather than any one exact sentence. Used that way, the weak axis stops being a trap.
- A measured consistency score replaces a heuristic one. Today the platform’s “Predictive /
consistency” number is an honest but unmeasured estimate (a transparent heuristic). S1 provides the raw material to make it measured — derived from each Twyn’s actual CV, modal agreement, and theme recurrence. This is what closes the “honest-math loop” for that axis (internally tracked as Epic O): the displayed number becomes something we have verified, not something we have asserted.
- An early discriminant-validity signal appeared for free. “Discriminant validity” means the
instrument can tell genuinely different people apart. The German Twyn (Annika) was consistently far more sceptical — an adoption likelihood of 2.8/10 — than the others (~5–6). That clean separation is exactly what Study S4 is designed to test deliberately, so we are not claiming it as proof here. But it is encouraging that the separation emerged on its own, and stably, in a study that was only meant to measure reliability.
5. Limitations (stated plainly)
These are the reasons not to over-read the pilot. We list them prominently and without softening; a buyer should weigh them as seriously as the positive findings.
- Pilot scale: 1 guide, 4 Twyns, 1 model. This is not yet the locked grid — there is no
paraphrase-robustness arm (re-wording the question to see if the answer holds), no multiple guides, and no multiple models. Consequently the ICC and α figures, computed across only 4 units, are indicative only and must be confirmed at full scale.
- Ceiling effect: the particular concept question pushed almost every answer to “maybe”, which
flattened the categorical signal (this is the same issue behind the alpha paradox in §3.2). Future guides must include items that genuinely discriminate, so the decision axis is tested against real variety rather than a near-unanimous “maybe.”
- Lexical, not semantic, free-text metric: as noted in §3.4, the wording metric understates the true
stability of meaning. The upgrade to embeddings plus human coding is required before we would publish any headline free-text reproducibility figure.
- Reliability ≠ accuracy: to be unambiguous, nothing in this paper shows that Twyns match real
people. Consistency is not correctness. Demonstrating that is the job of Studies S2–S3, not this one.
6. Next steps
This section states what we will do to turn the pilot into the full, defensible study.
- Run the full locked S1 grid: add the 10-paraphrase robustness arm (asking the same thing 10
different ways), at least 2 guides containing items that actually discriminate between personas, and a per-model arm; and upgrade the free-text metric to sentence embeddings + human theme coding.
- Wire the measured consistency score into the product surface, so the number a user sees in the
profile drawer (Figure 2) reflects the measurement rather than the heuristic.
- Proceed to S2 (concurrent validity) — a comparison against openly-licensed, segment-labelled real
interview data (a dataset scan is in progress). This is the first study that speaks to accuracy, the separate question this reliability pilot deliberately does not address.
Appendix — reproduce this
For full transparency, anyone can regenerate every figure in this paper by running the two commands below (one per temperature arm) against the same script used to produce the original results: “ cd twyns-net TWYNS=A,K,L,S REPS=30 TEMP=0.7 node meta/validation/s1-pilot.mjs # production arm TWYNS=A,K,L,S REPS=30 TEMP=0 node meta/validation/s1-pilot.mjs # reference arm ` Raw results: reports/validation/s1-temp07-results.json, s1-temp0-results.json`. Total pilot cost: ~$1.50 (240 calls, Sonnet 4.6). Zero parse failures. The low cost and clean parse rate are themselves part of the point: a credible reliability check on this product is cheap and repeatable, so it can be re-run as the product evolves.

Leave a Reply