David Samuel

I am a postdoctoral researcher at Integreat and the Language Technology Group at the University of Oslo. My main focus is data-constrained language modeling – how to train strong language models when text, compute, or native speakers are in short supply. Putting theory into practice, I pre-trained and post-trained a suite of Norwegian language models, including NorBERT and NorMistral. I've also been busy developing the Norwegian LLM dashboard, and contributing to two large-scale European projects, HPLT and OpenEuroLLM.

Selected publications

[ICLR 2026] Fluent alignment with disfluent judges: Post-training for lower-resource languages

David Samuel, Lilja Øvrelid, Erik Velldal and Andrey Kutuzov

A major issue for languages such as Norwegian is the absence of any high-quality post-training data. Researchers usually solve this by training models on machine-translated datasets, which leads to sloppy disfluent models. Our on-policy post-training method preserves fluency even without any native post-training data, enabling preference optimization for low-resource languages.

[ICLR 2026] Dual-objective language models: Training efficiency without overfitting

David Samuel and Lucas Charpentier

The main idea combines autoregressive and masked-diffusion training objectives to get the best of both worlds – training efficiency as well as resilience to overfitting. We show that this leads to much stronger models trained on limited data.

[NoDaLiDa 2025] Small languages, big models: A study of continual training on languages of Norway

David Samuel, Vladislav Mikhailov, Erik Velldal, Lilja Øvrelid, Lucas Charpentier, Andrey Kutuzov and Stephan Oepen

A three-stage continual training approach behind NorMistral 11B. We present a practical approach to training an LLM on a language with a limited amount of data (in this case Norwegian Bokmål, Nynorsk and Northern Sámi).

[NeurIPS 2024] BERTs are generative in-context learners

David Samuel

This paper shows that masked language models can perform in-context learning just as well as autoregressive language models. Looking closer at the results, they reveal complementary strengths between the two training approaches.

Teaching

I've been teaching Developing Large Language Models (IN5640) and Neural Methods in Natural Language Processing (IN5550) at the University of Oslo. I am also supervising Master students; please reach out if you'd be interested in doing a Master thesis at our group!

Projects

Norwegian LLM dashboard

An interactive dashboard for exploring the landscape of Norwegian language modeling. Still in active development.

Generative NorMistral models

Open language models are a critical foundation of sovereign software infrastructure. I developed NorMistral 7B, NorMistral 11B, and co-developed the fully-open NorOLMo 13B language model.

NorBERT suite of models

Large autoregressive language models are very trendy and easy to use, but the real workhorses across many applications are finetuned masked language models. I developed and co-developed two families of such models: NorBERT3 and NorBERT4.

Education

PhD in Informatics, University of Oslo (2021—2026)

My doctoral thesis, Data-constrained language modeling, studies how to train strong language models when text is scarce, going through more efficient training objectives, neural architectures, and post-training methods. The findings are then applied to building strong Norwegian language models.

MSc in Artificial Intelligence, Charles University (2018—2021)

This is when I moved away from general computer science to deep learning and NLP. My Master thesis, Permutation-invariant semantic parsing, introduced a novel graph-based semantic parser.