Your Prejudice Map Methodology & validation

Methodology & validation

How the score is built - and how far to trust it.

Every number this instrument produces comes from a chain of well-established measurement methods. This page sets out that chain, cites the source literature, and - just as importantly - states plainly where the method is solid and where this particular implementation is a demonstration rather than a validated test.

1 · The measurement model

Forced choice, on purpose. You are never offered a "they're equal" option. Thurstone's law of comparative judgment3 established the principle the whole instrument leans on: a forced choice between two items is a noisy reading of an underlying, unobservable preference, and enough such readings triangulate that latent value. Removing the neutral option is what forces the fast, automatic system to commit.

From wins and losses to one number. Each trial is treated as a match. The chosen country takes rating points from the rejected one via the Elo update rule5 (K = 32) - the sequential, online relative of the Bradley–Terry model4, which is the standard statistical method for converting paired comparisons into latent "strengths." After a session, each country's rating is its position on your revealed social-distance scale, rendered as the red→green choropleth.

Adaptive pairing. Pairs aren't drawn purely at random: the instrument favours the least-seen country and a partner with a close current rating, because comparisons between near-equals carry the most information - the same logic that makes adaptive testing efficient.

One construct, on purpose: warmth, not competence. Every prompt asks the same kind of thing - your willingness to take a person from a country into closeness with you (as a partner, friend, neighbour, colleague, or someone you'd trust). That is the affective acceptance gradient the Bogardus scale1 measures, and the warmth dimension of the Stereotype Content Model14. We deliberately avoid competence-laden framings - a surgeon, pilot, or boss "from X" - and place-desirability framings ("live in X"). The model shows warmth and competence are orthogonal: you can rate a group highly competent yet feel little warmth toward it (the classic "I'd trust the doctor, but not as a friend"). Folding competence judgments into the same score would measure perceived skill, not prejudice, so the prejudice run keeps strictly to the warmth axis. A separate, clearly-labelled Competence lens (opt-in at the start) lets you measure that other axis on its own - who you'd trust with a high-stakes job. It is scored and stored entirely apart from the prejudice run and never mixed into it; the point of offering both is that the interesting signal is often where they diverge.

2 · From scores to a bias type

The eight "bias types" are a presentation layer on top of the ratings - a legible, personality-style readout, much like the way 16personalities packages the Big Five into 16 types. They are not a validated taxonomy from the literature; the three underlying dimensions are what carry the meaning. Each is computed from your own ratings:

  • Reach (even-handed ↔ selective): the spread between your highest- and lowest-rated nations.
  • Scope (cosmopolitan ↔ partial): how much your preference varies across world regions versus within them.
  • Lean (outward ↔ homeward): the average rating of your own region (your in-group, taken from where you are) minus every other region. It is measured relative to you, not to any fixed bloc.

Graded against the noise floor. Reach and scope can't be read off the raw spread, because with only a handful of comparisons per country, sampling noise alone produces a large spread even for a perfectly even-handed person. So both are scored relative to that expected-by-chance dispersion (a function of how many comparisons each country received): a neutral respondent reads even-handed at any session length, and selectivity only registers once it exceeds what noise would produce. This makes the scores comparable across the short and long runs.

Pool size matters, and is a real limit. Resolving bias needs roughly six to eight comparisons per country. The default run therefore uses a 16-country core (balanced across the six regions) so a short test gives each country enough comparisons; the 39-nation and 165-nation runs trade per-country precision for a fuller map. Each dimension is a 0–100 score split at 50 for the three-letter code. Those cut-points start out heuristic, but become empirical once a version+lens has 40+ sessions: the score is then your percentile against everyone who took it, so 50 is the true median (section 4).

3 · The constructs it borrows from

Social distance. The prompts are a forced-choice adaptation of the Bogardus Social Distance Scale1 - the ~100-year-old instrument that measures prejudice as the closeness of social contact a person will accept (neighbour, colleague, friend, family by marriage). It remains one of the most widely used prejudice measures in sociology2.

Implicit measurement. Conceptually this tool is a cousin of the Implicit Association Test6: both infer attitude from performance under time pressure rather than from self-report, which sidesteps some of the impression-management that distorts what people will openly say on sensitive topics7.

Why the snap choice reveals anything. Under time pressure the mind performs attribute substitution9: asked a hard question ("who would I genuinely trust?") it quietly answers an easier one ("who feels familiar?") and reports the substitute as the original. The forced choice is engineered to catch that substitution in the act.

4 · What is, and isn't, validated

This is the part that matters most, so it's stated bluntly.

Established (the methods)

  • Paired-comparison scaling - Thurstone3, Bradley–Terry4, Elo5 - is mathematically well-founded and used everywhere from psychophysics to chess to machine-learning evaluation.
  • The Bogardus scale has decades of cross-cultural, cross-language use as a prejudice measure2.
  • Performance-based (implicit) measures genuinely avoid some self-report distortion on sensitive topics7.

Not validated (this tool)

  • Sparse data, mitigated. Per-country estimates need many comparisons to stabilise. The 16-country default run is sized to give ~6-8 comparisons each, which is enough for the dimensions to separate real bias from noise; the 39- and 165-nation runs are sparser per country, so their maps are richer but their per-country numbers noisier. The dimensions are graded against the noise floor (section 2) so this can't masquerade as confidence.
  • Reliability + emerging norms. Every session reports its own internal consistency, self-agreement, test–retest and convergent validity (section 5). Scores are also norm-referenced: once a version+lens has 40+ sessions, each dimension is shown as your percentile against that population (a true median split for the types), replacing the simulation-fit cut-points. The honest caveat that remains: the reference group is people who took this, which is self-selected, not a representative or census-weighted sample.
  • Mixed prompts. Rotating prompt types (trust, intimacy, authority) adds construct noise to a single rating.
  • Relative, not absolute. Forced choice removes neutrality by design, so the output is a relative leaning, not a measured attitude strength.
  • Heuristic thresholds. The archetype cut-points are for legibility, not derived from data.

On predicting behaviour - don't. Even the mature, heavily-studied IAT is contested as a predictor of real-world discrimination: meta-analyses range from a modest average correlation of about r = .277 down to "small effects of unknown societal significance"8. A purpose-built demonstration like this one carries far less evidential weight again. A result here says nothing about how you, or anyone, will actually behave.

Bottom line. The scoring rests on validated measurement theory, but the product is a self-reflection demonstration, not a diagnostic test. It reliably shows that your fast, forced choices were uneven - which is interesting and worth sitting with - and it does not, and cannot, certify a verdict about your character. A mirror, not a measurement.

5 · Reliability, measured on every session

Because single-session paired-comparison data is noisy, the instrument now revalidates itself on every run and shows you the result. Three standard psychometric checks, computed entirely in your browser:

Internal consistency

Your answers are split into two halves (even and odd comparisons), each half is scored into a separate set of country ratings, and the two are correlated (Spearman), then Spearman–Brown corrected for the split. This is the paired-comparison analogue of a split-half / Cronbach's-α reliability check34. High means the two halves of your session agree - a stable signal rather than noise.

Self-agreement

A handful of country pairs you have already judged are quietly re-asked later in the session, with a different prompt. The proportion of repeats on which you make the same call is a direct, intuitive consistency measure (and flags intransitive, contradictory choosing).

Test–retest

If you take the test again, this run's country ratings are correlated against your previous run (stored only in your browser). That two-sitting correlation is the textbook test–retest reliability coefficient. A short session will show some honest drift; that is expected, not a fault.

Convergent validity

After the timed quiz, you rate the same countries deliberately, with no timer - an explicit feeling-thermometer14. Correlating that stated measure with your snap Elo is a convergent-validity check: does the fast forced-choice capture what you would say out loud? Asking it after avoids priming the snap choices; the country where the two diverge most is the implicit-vs-explicit gap the whole tool is about.

Adaptive length. The question count is a floor, not a fixed number. When you reach it, the split-half reliability of your answers so far is fed into the Spearman-Brown prophecy formula - which predicts how much longer a measure must be to reach a target reliability - and the session is extended toward that length (up to a cap) only if it's needed. Consistent answerers stop at the floor; noisier sessions get exactly as many more comparisons as the maths asks for. The target is r = 0.70, the conventional bar for acceptable research-grade reliability.

None of this rescues the tool from the limits in section 4 - reliability is necessary, not sufficient, for validity, and norming is still absent. But it does let the instrument state, honestly and live, how internally trustworthy each individual read was.

References

  1. Bogardus, E. S. (1925). Measuring social distance. Journal of Applied Sociology, 9, 299–308. [full text]
  2. Wark, C., & Galliher, J. F. (2007). Emory Bogardus and the origins of the social distance scale. The American Sociologist, 38(4), 383–395. doi:10.1007/s12108-007-9023-9
  3. Thurstone, L. L. (1927). A law of comparative judgment. Psychological Review, 34(4), 273–286. [full text]
  4. Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4), 324–345. doi:10.1093/biomet/39.3-4.324
  5. Elo, A. E. (1978). The Rating of Chessplayers, Past and Present. New York: Arco Publishing.
  6. Greenwald, A. G., McGhee, D. E., & Schwartz, J. L. K. (1998). Measuring individual differences in implicit cognition: The Implicit Association Test. Journal of Personality and Social Psychology, 74(6), 1464–1480. doi:10.1037/0022-3514.74.6.1464
  7. Greenwald, A. G., Poehlman, T. A., Uhlmann, E. L., & Banaji, M. R. (2009). Understanding and using the Implicit Association Test: III. Meta-analysis of predictive validity. Journal of Personality and Social Psychology, 97(1), 17–41. doi:10.1037/a0015575
  8. Oswald, F. L., Mitchell, G., Blanton, H., Jaccard, J., & Tetlock, P. E. (2013). Predicting ethnic and racial discrimination: A meta-analysis of IAT criterion studies. Journal of Personality and Social Psychology, 105(2), 171–192. doi:10.1037/a0032734
  9. Kahneman, D., & Frederick, S. (2002). Representativeness revisited: Attribute substitution in intuitive judgment. In T. Gilovich, D. Griffin, & D. Kahneman (Eds.), Heuristics and Biases: The Psychology of Intuitive Judgment (pp. 49–81). Cambridge University Press.
  10. Allport, G. W. (1954). The Nature of Prejudice. Cambridge, MA: Addison-Wesley. - the contact hypothesis.
  11. Pettigrew, T. F., & Tropp, L. R. (2006). A meta-analytic test of intergroup contact theory. Journal of Personality and Social Psychology, 90(5), 751–783. doi:10.1037/0022-3514.90.5.751
  12. Sidanius, J., & Pratto, F. (1999). Social Dominance: An Intergroup Theory of Social Hierarchy and Oppression. Cambridge University Press. - social-dominance orientation.
  13. Parrillo, V. N., & Donoghue, C. (2005). Updating the Bogardus social distance studies: a new national survey. The Social Science Journal, 42(2), 257–271. [full text]
  14. Fiske, S. T., Cuddy, A. J. C., Glick, P., & Xu, J. (2002). A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. Journal of Personality and Social Psychology, 82(6), 878–902. doi:10.1037/0022-3514.82.6.878 - the warmth/competence (Stereotype Content) model.

References 10–13 ground the explanation guide's "bigger picture" on the historical, cultural and demographic patterns of social distance, and on what reduces prejudice. The home-page "three biases" additionally draw on in-group favouritism (Tajfel et al., 1971, Eur. J. Soc. Psychol.) and mere-exposure (Zajonc, 1968, JPSP).