Using Interpretable Network Embeddings to Understand Populist Voting Behavior in Population-Scale Registry Networks
Abstract
Applying machine learning to administrative registry data has challenged the limits of predicting social outcomes. In deep learning, prediction models often rely on embeddings that project complex input data into a latent numerical space. In the social sciences, embeddings can be used to compress social networks, which reduces their dimensionality while preserving information about the network structure. Registry data can be used to construct social networks at population-scale by creating ties between persons who belong to the same administrative entity (e.g., households, neighborhoods). Such ties are understood to represent social opportunities, reflecting potential interactions and shared environments. Using registry data from Statistics Netherlands, we created person-level embeddings for the entire Dutch population solely based on social opportunity ties in five domains . We demonstrate the usefulness of these embeddings by predicting right-wing populist voting in the Dutch parliament election 2023, linking embeddings with data from a representative survey. While embeddings alone showed a predictive signal, they performed worse than established individual covariates of populist voting. Moreover, they did not improve predictions when combined with said covariates. To make the embeddings more interpretable, we applied regularizing auto-encoders that disentangle the embedding dimensions, enforcing sparsity and orthogonality. We found that one of the regularized embedding dimensions was highly predictive of right-wing populist voting. From this embedding dimension, we created a weighted version of the population network that revealed differences in network structure between higher- and lower-educated persons. These structural differences can help explain right-wing populist voting decisions. We see this study as a starting point to create interpretable population-scale social representations that researchers can use to relate social structure and social outcomes. The embeddings are available to the Dutch research community through the Data Storage Facility by Statistics Netherlands and ODISSEI.