BIG DATA, DEEP ROOTS — SURNAME AGGREGATES Van Pelt & Kirkegaard (2026), Comparative Sociology. https://doi.org/10.1163/15691330-bja10162 These 10,989 surname records are the processed data used by the site's game and explorer, at the precision stored in the game. They are not the full raw record-level research datasets. No photos or individual records are included. Surname-level estimates describe groups in the source samples, not any individual bearing a name. Coverage and selection into each source differ. CSV: one row per surname, UTF-8, header row. Empty cells are missing, not zero. JSON: metadata and compact surname rows. Missing numeric values are null. Tuple positions (zero-based): 0 surname; 1 ancestry proportions; 2 observed estimates; 3 observed percentiles; 4 ancestry-predicted estimates; 5 predicted percentiles; 6 Census [count, shares] or null; 7 standard errors; 8 lower and 9 upper 95% confidence limits; 10 lower and 11 upper percentile confidence limits; 12 grouped ancestry proportions; 13 ancestry PCA [PC1, PC2]. Array order is defined by ancestries, metrics, censusLabels, and clusters in JSON metadata. CSV METRIC COLUMNS *_estimate, *_percentile, *_predicted, *_predicted_percentile, *_se, *_ci_low, *_ci_high, *_ci_low_percentile, *_ci_high_percentile The prefix identifies a metric. Estimates, standard errors and raw confidence limits use the same metric scale; percentiles use 0–100. Percentiles compare surnames with available observations for that metric, not people. Higher criminal-record percentiles mean MORE representation, not better outcomes. Percentiles and confidence limits are rounded as in the game. METRICS s_factor: composite socioeconomic score combining government salary, occupational prestige, physician licensure and reversed criminal-record representation. Wikipedia is excluded. salary: adjusted standardized government-pay measure, NOT dollars. opr: adjusted government occupational-prestige measure, on the supplied model scale. physician and crime: log10 representation residuals adjusted for surname frequency. 10^estimate gives representation relative to expected counts. These are not individual probabilities. Predicted S factor uses XGBoost; other predictions use linear models based on ancestry. Confidence intervals describe model estimates and do not account for all source-selection biases. ANCESTRY AND CENSUS ancestry_*: 45 source ancestry proportions (0–1); may not sum exactly to 1 due to rounding and source coverage. cluster_*: 23 grouped ancestry proportions used in the game. JSON clusterLabels supplies display names. census_count: surname population count from the source Census surname table, not the number of observations used for each outcome. census_share_1: White (percent, 0–100) census_share_2: Black (percent, 0–100) census_share_3: Hispanic (percent, 0–100) census_share_4: Asian/Pacific Islander (percent, 0–100) census_share_5: American Indian/Alaska Native (percent, 0–100) census_share_6: Two or More Races (percent, 0–100) ancestry_pc1/pc2: PCA coordinates combining standardized ancestry proportions and Census shares, used by the game. These are distinct from the socioeconomic S factor. All original aggregate values are preserved. See the paper and site supplement for methods and limitations.