Publications
My Google Scholar page should contain everything as well.
2026
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors.
Alexander Hägele, Alejandro Hernández-Cano, Atli Kosson, Martin Jaggi.
arXiv preprint arXiv:2606.25971.
[Blog Post | PDF | bibtex]
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?.
Alexander Hägele, Aryo Pradipta Gema, Henry Sleight, Ethan Perez, Jascha Sohl-Dickstein.
The Fourteenth International Conference on Learning Representations (ICLR 2026).
[Project Page | PDF | Code | Dataset | bibtex]
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments.
Alejandro Hernández-Cano, Alexander Hägele, […], Antoine Bosselut, Martin Jaggi, Imanol Schlag.
Annual Meeting of the Association for Computational Linguistics (ACL 2026). The biggest fully open and compliant training run and LLM to date.
[PDF | Huggingface | bibtex | Pretrain Code | Pretrain Data | Posttrain Code | Posttrain Data | Evals]
2025
Inverse Scaling in Test-Time Compute.
Aryo Pradipta Gema, Alexander Hägele, Runjin Chen, Andy Arditi, Jacob Goldman-Wetzler, Kit Fraser-Taliente, Henry Sleight, Linda Petrini, Julian Michael, Beatrice Alex, Pasquale Minervini, Yanda Chen, Joe Benton, Ethan Perez.
Transactions on Machine Learning Research (TMLR), 2025. J2C Certification (presented at ICLR ‘26) and Featured Certification.
[Project Page | PDF | bibtex | Code]
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler.
Aleksandr Dremov, Alexander Hägele, Atli Kosson, Martin Jaggi.
Transactions on Machine Learning Research (TMLR), 2025. J2C Certification (presented at ICLR ‘26).
[PDF | bibtex | Code]
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training.
Fabian Schaipp, Alexander Hägele, Adrien Taylor, Umut Simsekli, Francis Bach.
Forty-second International Conference on Machine Learning (ICML 2025).
[PDF | bibtex | Code]
2024
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations.
Alexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro Von Werra and Martin Jaggi.
Advances in Neural Information Processing Systems (NeurIPS 2024). Spotlight Award at NeurIPS’24.
Also Spotlight Presentation at the Workshop on Next Generation of Sequence Models (NGSM) and Best Poster Award at the Workshop on Efficient Systems for Foundation Models (ES-FOMO) at ICML’24.
[PDF | bibtex | Code | Slides]
2023
BaCaDI: Bayesian Causal Discovery with Unknown Interventions
Alexander Hägele, Jonas Rothfuss, Lars Lorch, Vignesh Ram Somnath, Bernhard Schölkopf and Andreas Krause.
Proceedings of the 26th International Conference on Artificial Intelligence and Statistics (AISTATS 2023).
Also presented at 1st Workshop on Causal Representation Learning at UAI 2022 (Link).
Oral presentation at AISTATS 2023 (notable paper award), ranked top 32 among 1689 submissions (top 1.9%).
PDF | bibtex | Details | Code | Slides
2021
Robustness Certification with Generative Models
Matthew Mirman, Alexander Hägele, Pavol Bielik, Timon Gehr and Martin Vechev.
Proceedings of the 42nd ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI 2021).
PDF | bibtex | DOI