Big Data, Deep Roots

I scraped 31.6 million records. Ancestry predicted surname status at r = 0.9.

Started
Sep 2026
Finished
Sep 2026
Status
Complete
Authors
Uncorrelated
AI Authors
GPT 6 Astra
AI Tokens
<10M tokens
Supplementary

TL;DR

  • Across roughly 11,000 American surnames, ancestry predicted our composite socioeconomic score at r = 0.899 in cross-validation.
  • The top ten surnames were all Indian. Selective migration is central to understanding the rankings.
  • Surnames capture a surprising amount of social history. Who migrates, who takes a DNA test, and how surnames cross populations all matter.

The long road

Government salaries. Occupation titles. Physician licenses. Criminal records. Wikipedia.

I spent about four months scraping and cleaning this project: 31.6 million records across the processed datasets, plus 11,008 surnames with 45 ancestry estimates each from 23andMe.1 The paper, written with Emil Kirkegaard, is now published in Comparative Sociology. (Van Pelt & Kirkegaard, 2026)

The trick was simple. 23andMe's surname pages displayed only the top three ancestries, but the other 42 were embedded in the page. Extract them, join the surnames to the other datasets, and you can ask how much ancestry predicts about surname-level outcomes.

After matching surnames to ancestry coverage, the datasets included 9.2 million salary records, 8.3 million occupation ratings, 3.4 million criminal records, and 2.0 million physician-license records. For salaries, the median surname had 270 observations. So for many names we had substantial information about average outcomes, rather than a handful of anecdotes.

The collection took months; the publication process took about a year. Here are the findings I found most interesting. There's also a game using the data at the end.

Ancestry gets you to r = 0.9

We combined government salary, occupational prestige, physician licensure, and criminal-record representation into an S factor: a composite score of surname socioeconomic standing. Technically, this is the first principal component: the common pattern across the four measures, with criminality reversed so that higher scores consistently indicate more favorable outcomes. Surnames with higher occupational prestige tended to have higher salaries, more physicians, and fewer criminal records relative to their population size.

Wikipedia was analyzed separately; its coverage was too patchy for the composite. The median surname had just two Wikipedia intellectuals, and roughly one in five surnames lacked a value altogether.

Using ancestry alone, XGBoost predicted this score at r = 0.899, or about 81% shared variance between predictions and observed scores. This was cross-validated across surnames, with similar error on a held-out test set.

Predicted versus observed surname S factor, showing a correlation of 0.899 across 10,989 surnames. Each point is a surname. Predictions come from test splits, rather than fitting and scoring the same observations.

The fancy model helped, but most of the relationship was already linear. A plain additive model reached r = 0.846; adding interactions raised it to 0.889. The remaining gains were small. Squaring the correlations, the linear model captured about 89% of XGBoost's predictive strength.

The relationship also appeared across the separate outcomes. Ancestry models correlated 0.825 with physician licensure, 0.822 with criminal-record representation, 0.737 with occupational prestige, and 0.608 with government salary. Salary was the weakest of these, but hardly unrelated.

That is a strong result for surname averages. It does not mean ancestry explains 81% of an individual's circumstances, or identify how much of the association is genetic rather than social.

Indian surnames at the top

Next we asked what each model predicted for a hypothetical surname with 100% of a given ancestry. The leading clusters were South Indian, Ashkenazi Jewish, North Indian, and Arab.

Predicted surname S factor for ancestry clusters, comparing three models and showing South Indian, Ashkenazi Jewish, North Indian, and Arab clusters at the top. Model estimates at 100% ancestry. These are extrapolations where the data lack nearly homogeneous surnames.

The Indian result also appeared directly in the observed surname rankings. The top ten were Agrawal, Gupta, Iyer, Jain, Patil, Goel, Parikh, Srinivasan, Krishnan, and Rao. Horwitz, an Ashkenazi surname, first appeared at number 16.

So what is going on?

Selection on steroids

Indian Americans are a highly selected slice of India. The discussion cites a review reporting that over 90% of earlier Indian migrants to the US came from upper castes, representing roughly 30% of India's population. Education, resources, and US immigration rules each add another filter. (Alamgir et al., 2022)

The exam data make the scale of selection unusually vivid. Of the top 1,000 scorers in India's 2010 Joint Entrance Exam, 36% migrated abroad. Among the top 100 it was 62%; among the top ten, nine left. The US was the main destination. (Choudhury et al., 2023)

You should expect this to leave a mark on American surname statistics. The selection happens before arrival: access to education, success in a fiercely competitive examination system, and the means and qualifications to move abroad.

The Indian clusters also ranked highly on physician licensure and among the lowest on criminal-record representation. So the result extends beyond which surnames happen to earn more in government employment.

Arabs and Iranians

The Arab and Anatolian/Iranian clusters are another interesting case. Arab ancestry ranked fourth on the S factor and third on occupational prestige.

The discussion points to religious as well as educational selection. Historical estimates cited in the paper put Christians at 63–77% of Arab Americans, a very different composition from the predominantly Muslim populations of many origin countries. The educational literature discussed in the paper documents substantial Christian advantages in Lebanon, making this a plausible part of the migration story. These historical estimates should not be read as a current demographic census. (Van Pelt & Kirkegaard, 2026)

For Iranians, the paper emphasizes students already studying in the US when the 1979 revolution occurred, and educated emigrants who subsequently left. Feliciano's comparison placed Iranian immigrants highest in educational selectivity among the origins she studied. (Feliciano, 2005)

PAAIA's 2025 survey provides another illustration of how distinctive the diaspora sample is: 24% of respondents identified as Muslim, and 79% had a four-year degree or postgraduate education. Those figures describe the survey respondents, not the ancestry of every person carrying an Iranian surname. (Americans, 2025)

The relevant comparison is therefore with the particular people who emigrated. A country's average tells you much less when its emigrants come disproportionately from a small, educated stratum.

Why are British and Irish estimates so low?

This result looked odd to us too. The paper points to a measurement problem.

Take Washington. In the Census, 87.5% of people with that surname identified as Black. In the 23andMe sample, its estimated African American ancestry was only 56%, alongside 23.2% British/Irish ancestry. Self-identified race and genetic ancestry are different measures, and customers who buy DNA tests are also a selected sample.

Comparison of 23andMe ancestry proportions and Census racial identification by surname, illustrating the mismatch between the two samples. The Census and 23andMe do not measure the same thing or sample the same people.

Shared surnames and selective testing make it difficult to separate the British/Irish coefficient from the outcomes of African Americans carrying those names. Iberian ancestry has a related problem with Hispanic surnames. Extrapolating either coefficient to 100% ancestry can give misleading results.

We reduced the extrapolation problem by combining related ancestries into broader clusters. All 45 original ancestries remained in the overall prediction analysis; the grouped categories were used for the 100%-ancestry comparisons. Two clusters, Egyptian and Indonesian/Thai/Khmer/Myanma, were excluded from interpretation because their estimates were too unstable.

The Egyptian example is instructive: it could rank first on physician licensure but eighteenth on occupational prestige. Only 41 surnames exceeded 5% Egyptian ancestry, and the maximum was 39.4%. A model projecting all the way to 100% is being asked to do a lot of guessing.

That is why the discussion matters as much as the headline correlation. Surname data reveal a lot, but interpreting a coefficient requires knowing which population actually produced it.

Read the published paper for the full methods, comparisons, and discussion.

Surname Guesser

How much of this can you predict yourself? Pick among three options, then compare your answer with the surname data. You have three lives.

Test your intuition against 10,989 surnames.

Footnotes

  1. This is a count across processed metric datasets, not 31.6 million distinct people: occupational prestige is derived from the salary records, for example. Matching to ancestry reduced the samples further. The final game and prediction plot cover 10,989 surnames.

References

Van Pelt, D., Kirkegaard, E. O. W. (2026). Big Data, Deep Roots: Correlating Surname Genetic Ancestry and Socioeconomic Status across Millions of Americans. Comparative Sociology. doi:10.1163/15691330-bja10162 ↩¹↩² [supp]
Alamgir, F., Bapuji, H., Mir, R. (2022). Challenges and Insights from South Asia for Imagining Ethical Organizations: Introduction to the Special Issue. Journal of Business Ethics. doi:10.1007/s10551-022-05103-3 [supp]
Choudhury, P., Ganguli, I., Gaulé, P. (2023). Top Talent, Elite Colleges, and Migration: Evidence from the Indian Institutes of Technology. Journal of Development Economics. doi:10.1016/j.jdeveco.2023.103120 [supp]
Feliciano, C. (2005). Educational Selectivity in U.S. Immigration: How Do Immigrants Compare to Those Left Behind?. Demography. doi:10.1353/dem.2005.0001 [supp]
Public Affairs Alliance of Iranian Americans (2025). 2025 National Public Opinion Survey. https://paaia.org/wp-content/uploads/2025/11/2025-National-Survey-Final-Copy.pdf [supp]