Foundation models for social data
Abstract
We explore transformer-based sequence models as foundation model for social science. We pre-train models on Dutch registry data and evaluate them on a set of prediciton tasks.
Full text
Foundation models for social data with Tanzir Pial, Ana Macanovic, Enam Hassan, Dakota Handzlik, Lucas Sage, Steven Skiena, Tom Emery, Arnout van de Rijt Flavio Hafner Odissei conference 04-11-2025
•Foundation models in Natural Language Processing (NLP): BERT Can we build general-purpose machine learning models for social data? Bommasani et al (2022)
•Pre-training: NLP models learn a vector of numbers=embedding oOn sub-words, but we can aggregate to sentences & documents oThe embedding compresses the training data into an easy-to-use set of numbers •Adaptation: Embeddings are used to solve specific tasks oFeed them into a prediction task: sentiment analysis, entity recognition, … oLearn a prediction model (transfer learning) oPatterns specific to the task are extracted from the embeddings Embeddings & representations
Flexibly account for histories in wage regressions -> Vafa et al (2024): gender wage gap Investigate the predictability of life courses and the role of luck Support policy: "find needles in a haystack" -> Example: matching workers and jobs ... Why is this useful for social sciences?
•Using Dutch registry data, we pretrain BERT-like models oBuild on Savcicens et al (2024) oSymbolic language of life sequences ▪Events from job, household, residential mobility ▪Background data: year, place of birth, gender, parental information oTraining objective: "fill the gap" •Experiment with oDifferent model sizes oVariations on the training objective Pre-training: what we do
•Train model up to 2020, predict for 2022-2023 oApproach 1: feed static embeddings into a fully connected NN with 2 layers oApproach 2: Fine-tune on the task = Update the model weights •Prediction tasks oIncome level oDemographic events: marriage, fertility, divorce oLISS survey questions: ethnic self-identification; motor vehicle ownership •Baseline oLinear regression with all variables used for pre-training the sequence model Evaluating the models
Model Income R2 Demographic F1 LISS F1 Baseline 0.30 0.33 0.41 Static embeddings 0.50 0.32 0.64 Finetuned model 0.58 0.39 0.57 Results: prediction tasks
•See Lisa's talk later •We participated after the challenge and had more time •Results when using CBS registry data Strongest baseline: Fertility challenge Model F1 1. GradBoost random forest 0.591 2. Us: Finetuned 0.581 3. GradBoost random forest II 0.561 4. Us: Static embeddings 0.529
•Embeddings do relatively better for events further in the future •Bigger models do not necessarily make better predictions oAround 80M parameters does best •Embedding similarity captures demographic similarity oMen vs women oBirth cohorts Other things