scieee AI-readable full text Open interactive document viewer

A Very Brief Introduction to Foundation Models

Sandfeld, Stefan

Abstract

A short presentation given during the CASA Forum (11 Nov 2025) at the Forschungszentrum Jülich GmbH, Germany.

Full text

A Very Brief Introduction to Foundation Models Stefan Sandfeld1,2 1IAS-9: Materials Data Science and Informatics, FZ Jülich GmbH 2Chair for Materials Data Science and Materials Informatics, RWTH Aachen University https://www.doi.org/10.5281/zenodo.17599504 licensed under CC BY-NC 4.0 (https://creativecommons.org/licenses/by-nc/4.0/) Motivation: Specialist Models Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 2/13 Specialist Models in Supervised Learning ... ▶... require labeled data: scarce/expensive in science ▶... suffer under distribution shift ▶... requires 1 supervised model per task/domain ▶... do not transfer to related modalities →task silos ▶sunk cost of compute: training Nmodels vs. one reusable representation. Motivation: Specialist Models Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 2/13 Specialist Models in Supervised Learning ... ▶... require labeled data: scarce/expensive in science ▶... suffer under distribution shift ▶... requires 1 supervised model per task/domain ▶... do not transfer to related modalities →task silos ▶sunk cost of compute: training Nmodels vs. one reusable representation. X Y train model #1 inference, testing X Y train model #2 inference, testing X Y train model #3 inference, testing (Scientific) Data & Generalization Problems Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 3/13 ▶(Scientific) data: heterogeneous, scarce, label-noisy ▶Small, biased datasets; non-i.i.d. sampling ▶Bad Generalization: Ptrain ,Ptest specialist models overfit to Ptrain; covariate shift ▶GIGO = very common in supervised learning (unless you counter it with data curation, strong priors etc) Can we do better than supervised representation learning from labelled data? Northcutt et al., NeurIPS 2021, Track on Datasets and Benchmarks. Unlabeled Data: Why and How (Self-Supervised) Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 4/13 ▶Unlabeled data are abundant ▶Goal: learn generalizable representations from data without task labels. ▶Two big families: generative/masked and contrastive/alignment ▶Payoff: reusable representations ⇒cheap adaptation (linear probe/head, LoRAstyle adapters), better OOD robustness Conrad and Narayan, https://doi.org/10.7554/elife.65894 Unlabeled Data: Why and How (Self-Supervised) Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 4/13 ▶Unlabeled data are abundant ▶Goal: learn generalizable representations from data without task labels. ▶Two big families: generative/masked and contrastive/alignment ▶Payoff: reusable representations ⇒cheap adaptation (linear probe/head, LoRAstyle adapters), better OOD robustness 1. Large-scale, self-supervised pre-training · · · generic model train 2. Transfer to specific tasks by “fine-tuning” · · · specialized model #1 re-train inference, testing · · · specialized model #2 re-train inference, testing 2. Self-Supervised Learning of Nanoparticles: A Generative Approach ▶self-supervised learning on large unlabelled dataset (CEM500K) ▶Knowledge transfer to supervised tasks with limited labelled datasets ▶→HRTEM nanoparticle images: a few labeled images for fine tuning January 18, 2024 Slide 36 Noisy image Denoised image Original image Generator Discriminator The proposed pretraining pipeline 1. Self-supervised pretraining with GANs on CEM500k Pretrained Generator 2. Fine-tuning with domain and task specific datasets Pretrained Generator Real or Fake TEMImageNet for Denoising and Super-resolution HRTEM AU dataset for Semantic Segmentation Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 5/13 Kazimi et al., CVPR 2024 3. Towards Generalist Models: Transformers ≈1982 Hopfield Network (RNN) Hopfield, Rummelhard 1986 Multilayer Perceptron Rummelhard, Hinton ... 1990 LeNet Yann Lecun 1997 LSTMs Hochreiter, Schmidhuber 2012 AlexNet Krizhevsky, Sutskever, Hinton 2014 GANs Goodfellow 2015 ResNet, RNN, LSTM 2016 AlphaGo vs Lee Sedol DeepMind 2017 Transformer Vaswani et al. 2018 GPT-1, BERT, GNNs OpenAI, Google AI, Scarselli ... 2019 GPT-2 OpenAI 2019 GPT-3 OpenAI 2021 AlphaFold 2 DeepMind 2022 Chat GPT Stable Diff. Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 6/13 3. Towards Generalist Models: Transformers Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 7/13 ▶Vaswani et al. (2017): self-attention (Transformer) →fully parallel sequence modeling + long-range dependencies ▶replaced RNNs/LSTMs on many sequence tasks (language modeling, translation) ▶enabled families: BERT (bidirectional), GPT (causal) ▶scales with model/data; flexible tokenization across modalities ▶Vision Transformers (ViT): image patches as tokens; multimodal via aligned token spaces ▶Benefits: in-context learning, efficient adaptation (adapters), modular multi-task Supervised Learning Self-Supervised Learning Transformers Foundation Models State-of-the-Art (DINOv2) Stefan Sandfeld |FZ Jülich : A Very Brief Introduction to Foundation Models 13/13