A pilot back-test against real interviews, and a cultural-transfer (“export”) test
Twyn / twyns.net Β· Studies S2 + S6 (pilot) Β· draft 2026-06-05 Β· McKay Consulting
Status: PILOT, and we lead with the limits. This back-tests Twyns against real interview transcripts (RRING Global Interviews, CC BY 4.0, DOI 10.5281/zenodo.5070359). A back-test means we take interviews that real people actually gave, build a synthetic respondent (a “Twyn”) for each one, ask the Twyn the same questions, and compare the two answers. CC BY 4.0 is an open licence, which means anyone can inspect and re-run this work β the data are not ours to flatter. Two honesty points up front: (1) the Twyns here are built from only four segment fields (country Β· domain Β· stakeholder type Β· gender) β i.e. the weakest, “Concept” tier, with no grounding in real data. “Grounding” means feeding the Twyn the real person’s own words; here we deliberately gave it none, so the Twyn is working from a thin demographic sketch alone. (2) one of our metrics (blinded discrimination) is confounded by transcription style (see Β§5) β that is, the test can be won or lost for reasons that have nothing to do with whether the Twyn understood the topic. Read this as establishing the floor that grounding must beat β not the product’s ceiling. In other words, these numbers are deliberately the worst case; they exist to show how much better the product gets once you feed it real evidence.
Where White Paper 1 asked is a Twyn reproducible? (yes) β i.e. if you run the same Twyn twice, do you get a consistent character β this paper asks the harder question: does a Twyn match a real person? Reproducibility is not accuracy: a Twyn can be perfectly consistent with itself and still be wrong about the real human it is meant to represent. This paper is the accuracy work β the test of whether the Twyn’s answers actually correspond to what a real respondent said.
What Twyn is β and why this paper exists
If you are meeting Twyn for the first time, start here. Twyn is a research tool that builds authentic digital twins of your users: realistic, evidence-grounded AI “users” that a product or research team can interview β each in its own voice, at the speed of a conversation. Rather than guessing in the long, expensive gaps between rounds of fieldwork, a team can put an idea, a message, a price or a design to a panel of these twins and hear how different kinds of user react β range, friction, the language people use, the objections they raise β before committing to build.
Why the tool exists. Real user research is indispensable but scarce β costly, slow and hard to repeat β so teams too often build on untested assumptions. Twyn makes the user’s perspective available continuously and affordably, explicitly as a supplement to talking to real people, never a replacement.
Why this paper. A separate study (White Paper 1) showed Twyn is reliable β ask the same twin the same question repeatedly and the answer holds steady. But reliability is only half the story: a twin could be perfectly consistent and still be consistently wrong. The harder, more important question is validity β do the twins actually resemble real people? This paper puts that to the test in two ways: by comparing twins against real, anonymised interviews on questions they were never shown (a “back-test”), and by checking whether tuning a twin to a different national culture genuinely shifts how it answers. Together these are the evidence behind Twyn’s central claim β that grounding a twin in real data makes it measurably more accurate β and we report the misses as plainly as the hits.
1. Method
- Ground truth: RRING β real semi-structured RRI interviews (RRI = Responsible Research and
Innovation, the topic these interviews covered). “Ground truth” is the standard term for the real-world answer key you measure against β here, the actual words of real interviewees. The same protocol was run in every country, meaning every respondent was asked from the same question script; that consistency is what makes the cultural-transfer test (S6) clean, because differences between countries reflect culture rather than different questions. 13/29 transcripts parse into clean questionβanswer pairs β of the 29 raw transcripts, 13 could be machine-read into tidy question-and-answer form usable for scoring.
- The Twyn under test: for each real respondent we build a **segment-level Twyn from their header
metadata only (gender Β· country Β· stakeholder type Β· domain) using the product’s in-character interview style β then put the same questions the real person was asked** to the Twyn. In plain terms: we tell the Twyn only four facts about the person (no transcript), let it answer in character, and then line its answers up against the real ones. This is the fairest like-for-like comparison: same questions, same person on paper, different source of knowledge.
- Scoring (LLM judge, blinded): an independent AI model acts as the judge, and “blinded” means it is
not told which answer is the human and which is the Twyn, so it cannot be biased by labels. – Discrimination β given the real and synthetic answer in random order, can a judge pick the real human? Target β 50% (indistinguishable). Read this metric backwards: 50% means the judge is guessing (the Twyn passes as human); a high score like 100% means the Twyn is always caught. Lower is better here, and 50% is the goal, not 100%. – Theme recall β fraction of the real respondent’s themes the Twyn also raised. If the real person made 10 distinct points and the Twyn covered 4 of them, recall is 0.40. It measures coverage: how much of what the real person actually said the Twyn managed to capture. Higher is better. – Theme precision β fraction of the Twyn’s themes that have support in the real answers. If the Twyn raised 10 themes and 3 of them genuinely appear in the real person’s answers, precision is 0.30. It measures discipline: how much of what the Twyn said was actually on-target rather than invented. Higher is better. – Hallucinated themes β Twyn themes with no support in the real answers (the metric that most worries a buyer). A “hallucination” is a point the Twyn confidently makes that the real person never raised β plausible-sounding fabrication. For a buyer relying on these answers to make decisions, this is the dangerous number, because it represents content that looks like signal but is noise. Lower is better.
- Harnesses:
meta/validation/s2-backtest.mjs,s6-cultural-transfer.mjs. Raw:
reports/validation/s2-backtest-results.json, s6-cultural-transfer-results.json. (These are the scripts that ran the studies and the raw result files, included so the work can be audited and reproduced.)
2. Study S2 β concurrent-validity back-test
“Concurrent validity” is the social-science term for checking a new instrument against an established one measured at the same time β here, checking the Twyn’s answers against the real interview answers for the same person. This first study uses 6 respondents (UK Γ2, Serbia Γ2, Japan, Singapore), 8 questions each, and uses the weakest, segment-only Twyns described in Β§1.
| Respondent | Country / domain | Discrimination | Theme recall | Theme precision | Hallucinated |
|---|---|---|---|---|---|
| GB02 | UK Β· Bioeconomy | 100% | 0.42 | 0.18 | 20 |
| GB03 | UK Β· ICT | 100% | 0.43 | 0.40 | 16 |
| SRB01 | Serbia Β· Energy | 88% | 0.75 | 0.32 | 13 |
| SRB04 | Serbia Β· Bioeconomy | 100% | 0.39 | 0.25 | 21 |
| JP03 | Japan Β· ICT | 100% | 0.44 | 0.50 | 13 |
| SGP01 | Singapore Β· Energy | 100% | 0.35 | 0.53 | 16 |
| Mean | 97.9% | 0.46 | 0.36 | 16.5 |
How to read the table: each row is one real respondent, identified by a code (e.g. GB03 = a UK respondent) and their domain. Reading across, you can see the judge almost always caught the Twyn (discrimination near 100%), the Twyn covered roughly four-tenths of the real themes (recall around 0.4), and a minority of its own themes were supported (precision), with a sizeable count of invented themes per transcript.
Reading it honestly:
- A segment-only Twyn is easily distinguished from a real respondent (β98%), captures ~46% of
the real themes, and over-generates β only ~36% of its themes are supported, ~16 unsupported themes per transcript. A four-field Twyn is not a stand-in for a real person. Put plainly: when you give the system almost nothing, it produces a plausible but loose impression β it gets some of the gist right, talks about a lot it shouldn’t, and never fools the judge.
- This is not a failure β it is the product’s own honesty framework, measured. A twin built from a
thin segment is exactly what Twyn labels a Concept twin (Authenticity tier β 32). The product already warns users that a Concept-tier Twyn is low-confidence; this study shows that warning is truthful. The back-test shows that low score is earned: low input β low fidelity. The number means what it says β the tier label and the measured accuracy line up, which is exactly what you want from an honest confidence rating.
2b. Study S2b β the GROUNDED arm (the core value-proposition result)
The decisive test: does grounding a Twyn in a real person’s own interview words make it match that person on questions it has not seen β and beat the Concept floor? This is the question that matters commercially, because the product’s promise is that feeding a Twyn real evidence makes it more accurate. Sealed-holdout protocol: split each respondent’s interview into an early grounding set and a later held-out test set; build the twin from only the grounding answers; ask the held-out questions; the held-out real answers are never shown to the twin (no leakage). “Held-out” means we hide part of the real interview from the Twyn and test it only on those hidden questions, so it cannot simply parrot back answers it was given β it has to genuinely generalise. “No leakage” is the guarantee that the test questions’ real answers never reached the Twyn, which is what makes this a fair exam rather than an open-book one. 4 B2B/industry-leaning respondents, 5 held-out questions each.
| Respondent | Arm | Recall | Precision | Hallucinated |
|---|---|---|---|---|
| GB03 UK Β· ICT Β· Industry | Grounded | 0.78 | 0.52 | 10 |
| Concept | 0.62 | 0.38 | 15 | |
| SRB04 Serbia Β· Bioeconomy Β· Industry | Grounded | 0.75 | 0.45 | 11 |
| Concept | 0.33 | 0.08 | 15 | |
| GB02 UK Β· Bioeconomy Β· Research org | Grounded | 0.47 | 0.18 | 14 |
| Concept | 0.25 | 0.08 | 17 | |
| JP04 Japan Β· ICT Β· Research org | Grounded | 0.31 | 0.18 | 14 |
| Concept | 0.31 | 0.14 | 10 | |
| Mean | Grounded | 0.58 | 0.33 | 12.3 |
| Concept | 0.38 | 0.17 | 14.3 |
How to read the table: each respondent appears twice β once as a Grounded Twyn (fed their own early interview words) and once as a Concept Twyn (the four-field version from Β§2, for comparison). Compare the two rows for each person to see what grounding adds. The bold Grounded numbers are the headline result; the Concept rows are the floor they have to beat.
Reading it honestly:
- Grounding is the core value proposition, and it is now a measured number. On held-out questions,
grounding lifts theme recall 0.38 β 0.58 (+53% relative), roughly doubles precision (0.17 β 0.33), and cuts hallucination. In everyday terms: giving the Twyn the person’s real words makes it cover much more of what the person actually thinks, keep far more of its own statements on-target, and invent less. The twin generalises from a person’s real words to questions it was never asked β which is the whole point, since in real use you will ask Twyns new questions, not ones you already have answers to.
- It is strongest exactly where the value lives β B2B/industry. GB03 (UK ICT industry) and SRB04
(Serbia bioeconomy industry) reach 0.75β0.78 recall; SRB04’s precision goes from 0.08 β 0.45. These are the respondents most like a real commercial buyer’s target audience, and they are where grounding pays off most β SRB04 went from almost no on-target content to nearly half of its themes being supported.
- We publish the null too. JP04 showed no lift (0.31 vs 0.31) β grounding does not always help;
here the held-out topics diverged from the grounding material. In other words, when the hidden questions asked about things the person had not touched on in the part we used for grounding, the Twyn had nothing to generalise from and did no better than the thin version. Reporting this is what makes the positive results trustworthy β we are not cherry-picking only the wins.
- The honest ceiling: discrimination stayed ~100% in both arms β grounding makes a twin match the
substance of a real person, not pass as one (and that metric is style-confounded; Β§5). The claim is fidelity of content, not indistinguishability. That is: a grounded Twyn gets the content of a person’s views right, but a judge can still tell it apart from a real transcript β and as Β§5 explains, that is largely because real transcripts contain messy speech artefacts the Twyn doesn’t reproduce, not because the Twyn’s thinking is shallow.
Note on data: RRING’s industry/business respondents are the closest open, license-clean B2B proxy. “License-clean” means we are legally free to publish it; a “proxy” means it stands in for the real target audience we cannot show. Twyn’s production Authentic twins (medtech R/S, B2B-sales T/U) are grounded in private interview data that would demonstrate this more strongly still, but cannot be published; RRING is the publishable proxy, private data the internal confirmation. So the published numbers here are a conservative public stand-in for results we see more strongly in confidential work.
3. Study S6 β cultural-transfer (“export”) validity
Use case: a company that knows its home market (UK) wants to export to a market it knows less. Can the Hofstede culture dial move a home-grounded Twyn toward the target market? The Hofstede dimensions are a long-established framework (six numerical scores per country) used to compare national cultures along axes such as individualism and tolerance of uncertainty; Twyn lets you “dial” a Twyn’s cultural settings from one country toward another. Same real target questions, three conditions, each scored against the real target-country answers β so for each market we test three versions of the Twyn and grade all three against what a real person in the target country actually said.
UK β Japan (ICT), vs real respondent JP03
| Condition | Theme recall | Precision | Culture fit |
|---|---|---|---|
| HOME (UK Hofstede) | 0.12 | 0.08 | 0.05 |
| TUNED (dial β Japan) | 0.05 | 0.10 | 0.20 |
| NATIVE (Japan) | 0.15 | 0.12 | 0.25 |
UK β Singapore (Energy), vs real respondent SGP01
| Condition | Theme recall | Precision | Culture fit |
|---|---|---|---|
| HOME (UK Hofstede) | 0.05 | 0.05 | 0.10 |
| TUNED (dial β Singapore) | 0.05 | 0.05 | 0.20 |
| NATIVE (Singapore) | 0.18 | 0.22 | 0.45 |
How to read these two tables: each has three rows β HOME (a UK Twyn left on UK cultural settings), TUNED (the same UK Twyn with its culture dial moved toward the target country), and NATIVE (a Twyn built natively for the target country). The new column is culture fit, a score for how well the answer’s tone, framing and cultural assumptions match the target market (separate from whether the content is right, which is what recall and precision measure). Reading down the culture-fit column shows whether dialling toward the target actually shifts the Twyn culturally.
Reading it honestly:
- The cultural dial works β for culture.
culture_fitrises monotonically **Home β Tuned β
Native** in both markets (Japan 0.05β0.20β0.25; Singapore 0.10β0.20β0.45). “Monotonically” just means it goes up at every step and never backwards β exactly the pattern you’d hope for if the dial does something real. Moving the Hofstede dial toward the target measurably shifts the Twyn toward that culture β the export proposition holds at pilot scale. (The UKβJapan jump is a large, real cultural distance: UAI 35β92, IDV 89β46 β Japan scores far higher on tolerating uncertainty-avoidance and far lower on individualism than the UK, so this is a genuinely big cultural gap, and the dial still moved the Twyn across it.)
- Tuning culture does not fix substance. Recall/precision stay low and barely move with tuning β
because what a researcher talks about (their actual projects, institutions, constraints) comes from grounding, not from a culture score. Tuning changes how they answer, not what they know. A culture dial can make a Twyn sound Japanese; it cannot tell the Twyn what a specific Japanese researcher is actually working on.
- Native (built from real target data) is consistently best, especially Singapore (culture_fit
0.45). The clear product implication: tune for culture, ground for substance β and grounding in the target’s own evidence is what moves the needle most. The practical takeaway for an exporting company: the culture dial is a useful adjustment, but the biggest gains still come from collecting real evidence from the target market itself.
4. What this means (the honest, and sellable, conclusion)
- The Authenticity tiers are real and measured. A Concept (segment-only) Twyn is demonstrably
low-fidelity. This validates the product’s central honesty claim rather than undermining it β the score predicts the back-test. Put simply, the confidence label the product shows you is not marketing; it correctly forecasts how accurate the Twyn will be, so a low score is a true warning, not a defect.
- The value proposition is grounding β now demonstrated (Β§2b). Feeding a Twyn a real person’s own
interview words lifts held-out theme recall from 0.38 β 0.58 (to 0.75β0.78 for B2B respondents) and roughly doubles precision. This is the “build your own Authentic Twynel” upsell, quantified: the more real evidence you give a Twyn, the more accurately it stands in β on questions it was never asked. This is the central commercial message, and it is now backed by a measured number rather than a promise.
- Cultural tuning is a genuine, demonstrable feature for the export use case β it shifts cultural
fit toward a target market β but it is positioned correctly only as “tune for culture, ground for substance,” never as a substitute for target-market evidence. We are explicit that the dial helps with tone and framing, and we do not claim it replaces real knowledge of the target market.
5. Limitations (stated first, as always)
- Weakest tier only. These Twyns used 4 metadata fields and no grounding; the product’s Authentic
tier is built from rich real data and is not what was tested here. The headline S2 numbers are therefore the floor of the product’s capability, not a representative sample of it.
- Discrimination is style-confounded. Real transcripts carry disfluencies, interruptions and
anonymisation markers (“[removed for anonymization purpose]”); an LLM judge spots those instantly, so ~98% discrimination overstates substantive distinguishability. In plain terms, the judge can tell the real transcript apart mostly because real speech is messy and redacted, not because the Twyn’s ideas are obviously artificial β so that 98% is not a fair measure of how well the Twyn captures the person’s thinking. Theme recall/precision are the more trustworthy substance metrics. A fairer discrimination test must normalise transcription style β i.e. clean both texts to the same style before judging.
- Single LLM judge, pilot n. 6 transcripts (S2) and 2 export pairs (S6); soft judge scores are
directional, not absolute. “Pilot n” means the sample sizes are small, and a single AI judge produces scores that indicate direction (better/worse) but should not be quoted as exact truths. Human coders + multiple judges are required before publishing headline numbers β independent human scoring and several judges are the bar this work must clear before any number here is treated as definitive.
- Theme metrics across S2 and S6 are not directly comparable (different prompts/judge questions);
within S6 the Home/Tuned/Native comparison is internally valid. So do not compare an S2 recall figure against an S6 recall figure as if they were measured the same way; but the three conditions within S6 were measured identically, so comparing them to each other is sound.
6. Next experiments
The grounded arm (Β§2b) β once the “decisive next experiment” β is done and positive: grounding roughly doubles precision and lifts held-out recall to 0.75β0.78 on B2B respondents. The most important open question has now been answered; what remains is to strengthen, generalise and stress-test that result. Remaining:
- S8 β iterative grounding (
docs/iterative-grounding.md): ingest one round’s ground truth, test on
a related-but-different round; does the lift transfer (learning, not memorisation)? This checks whether the Twyn is genuinely learning a person’s perspective rather than just memorising the specific answers it was given.
- S7 β marketing validity (Upworthy “Twyns pick the winning headline” + BRAND;
docs/validation-datasets.md).
A test of whether Twyns can predict real-world marketing outcomes, such as which headline performs best.
- S4 construct, S5 calibration, and the design-partner prospective trial (S3). “Construct”
validity asks whether the Twyn measures what it claims to; “calibration” asks whether its confidence scores match its real accuracy; the prospective trial (S3) is a forward-looking test with a real design partner rather than a back-test.
- Reduce the style confound in discrimination (normalise transcription artefacts) so that metric
measures substance, not anonymisation markers β the fix flagged in Β§5, so the discrimination score reflects whether the Twyn thinks like the person, not whether its text is tidier than a raw transcript.
Attribution (CC BY 4.0): RRING Global Interviews Dataset (WP3), DOI 10.5281/zenodo.5070359. This paper publishes the unflattering floor deliberately: a method that states where it is weak is the one a research-literate buyer can trust on where it is strong.


Leave a Reply