I am a postdoctoral researcher at Integreat and the Language Technology Group at the University of Oslo. My main focus is data-constrained language modeling – how to train strong language models when text, compute, or native speakers are in short supply. Putting theory into practice, I pre-trained and post-trained a suite of Norwegian language models, including NorBERT and NorMistral. I've also been busy developing the Norwegian LLM dashboard, and contributing to two large-scale European projects, HPLT and OpenEuroLLM.
David Samuel, Lilja Øvrelid, Erik Velldal and Andrey Kutuzov
A major issue for languages such as Norwegian is the absence of any high-quality post-training data. Researchers usually solve this by training models on machine-translated datasets, which leads to sloppy disfluent models. Our on-policy post-training method preserves fluency even without any native post-training data, enabling preference optimization for low-resource languages.
The main idea combines autoregressive and masked-diffusion training objectives to get the best of both worlds – training efficiency as well as resilience to overfitting. We show that this leads to much stronger models trained on limited data.
David Samuel, Vladislav Mikhailov, Erik Velldal, Lilja Øvrelid, Lucas Charpentier, Andrey Kutuzov and Stephan Oepen
A three-stage continual training approach behind NorMistral 11B. We present a practical approach to training an LLM on a language with a limited amount of data (in this case Norwegian Bokmål, Nynorsk and Northern Sámi).
This paper shows that masked language models can perform in-context learning just as well as autoregressive language models. Looking closer at the results, they reveal complementary strengths between the two training approaches.
My first proper NLP paper, a novel solution to graph-based semantic parsing inspired by DETR object detection. Looking back, the method is way too over-engineered, but it was a fun and valuable learning experience.
Open language models are a critical foundation of sovereign software infrastructure. I developed NorMistral 7B, NorMistral 11B, and co-developed the fully-open NorOLMo 13B language model.
Large autoregressive language models are very trendy and easy to use, but the real workhorses across many applications are finetuned masked language models. I developed and co-developed two families of such models: NorBERT3 and NorBERT4.
Education
PhD in Informatics, University of Oslo (2021—2026)
My doctoral thesis, Data-constrained language modeling, studies how to train strong language models when text is scarce, going through more efficient training objectives, neural architectures, and post-training methods. The findings are then applied to building strong Norwegian language models.
MSc in Artificial Intelligence, Charles University (2018—2021)
This is when I moved away from general computer science to deep learning and NLP. My Master thesis, Permutation-invariant semantic parsing, introduced a novel graph-based semantic parser.