Sign in

Martin Jaggi

@mjaggi.bsky.social
842 followers 177 following 50 posts

Prof at EPFL AI • Climbing

PostsRepliesMedia
Martin Jaggi @mjaggi.bsky.social · 19/09/2026
In view of recent discussions on risks of AI frontier models, we started a petition for more open science in AI safety. Please consider signing and sharing: make-safety-open.github.io
make-safety-open.github.io
A Call for Open Science in AI Safety
030
Reposted by Martin Jaggi
Apertus @apertusllm.bsky.social · 24/07/2026
Available now! Apertus 1.5 — 8B & 70B models with multimodal text/image/audio input, a 4x longer context window, optional thinking mode, and improved tool use, all built transparently and responsibly. 3.2M+ downloads 🚀 #Apertus #OpenSourceLLM Details+Links+Open roles: apertus-ai.org/articles/202...
Geometric shapes with Apertus 1.5 announcement and EPFL / ETH / CSCS logos.
195
Reposted by Martin Jaggi
EPFL AI Center @epfl-ai-center.bsky.social · 24/07/2026
Apertus 1.5 is out! 🚀 ✅ Multimodal capabilities ✅ Stronger reasoning ✅ A roadmap of regular releases Read more: apertus-ai.org/news Models: huggingface.co/swiss-ai @cscsch.bsky.social @eth-ai-center.bsky.social @icepfl.bsky.social @mjaggi.bsky.social @abosselut.bsky.social
064
Reposted by Martin Jaggi
CSCS - Swiss National Supercomputing Centre @cscsch.bsky.social · 22/07/2026
#Apertus is mentioned —the open large language model developed by ETH Zürich, EPFL and CSCS—which demonstrates how excellent research, cutting-edge infrastructures and international collaboration can contribute to strategically important technologies ⬇️
033
Reposted by Martin Jaggi
EPFL AI Center @epfl-ai-center.bsky.social · 20/07/2026
UNICC and EPFL's Machine Learning and Optimization Laboratory led by @mjaggi.bsky.social have published a white paper presenting a practical framework for evaluating the safety and reliability of LLMs in institutional settings. 👉 Learn more: ai.epfl.ch/epfl-lab-and...
042
Martin Jaggi @mjaggi.bsky.social · 03/07/2026
Searching for something new to read? All ICML 2026 papers are now public: openreview.net/group?id=ICM... (including all 6551 accepted papers and also rejects which opt-in)
openreview.net
ICML 2026 Conference
Welcome to the OpenReview homepage for ICML 2026 Conference
053
Reposted by Martin Jaggi
🤷 Nico Martin @nico.dev · 26/06/2026
Apertus Mini is now running entirely in your browser 🇨🇭 80+ tps for the 1.5B, 60+ tps for the 4B (on my M3). Fully client-side via Transformers.js + ONNX + WebGPU.
152
Reposted by Martin Jaggi
Apertus @apertusllm.bsky.social · 26/06/2026
Three new model weights: 0.5B, 1.5B, 4B are available using new quantization and distillation techniques. Download #Apertus 1.1, read the ICML workshop report, try a new demo on @hf.co - all just a tap away in our latest blog post: apertus-ai.org/articles/202...
apertus-ai.org
APERTVS.ai
Fully Open Foundation Model for Sovereign AI
022
Reposted by Martin Jaggi
ICML Conference @icmlconf.bsky.social · 18/05/2026
Announcing the #ICML2026 invited speakers! Pascale Fung Susan Athey (@susanathey.bsky.social) Sham Kakade (@shamkakade.bsky.social ) Aviv Regev Verena Rieser (@verenarieser.bsky.social) Arvind Narayanan (@randomwalker.bsky.social) Check out the blog post for more info! blog.icml.cc/2026/05/18/a...
0124
Reposted by Martin Jaggi
ICML Conference @icmlconf.bsky.social · 30/04/2026
Ahem, back to business... Decision notifications are being released on OpenReview. There were 23,918 submissions that entered review, roughly double last year. 6,352 papers were accepted, for an acceptance rate of 26.6%. 536 papers (2.2% of submissions) are "spotlights." 1/3
1177
Martin Jaggi @mjaggi.bsky.social · 24/04/2026
Muon: I made a new 3-slides explanation of this amazing optimizer for today's lecture. Let me know what you think
Muon: MomentUm Orthogonalized by Newton-schulzWhy orthogonalize? Balancing per-layer update magnitudesMuon as Frank-Wolfe
240
Reposted by Martin Jaggi
Ethan Mollick @emollick.bsky.social · 22/04/2026
Every system that was regulated, either explicitly or implicitly, by the fact that they were effortful for humans (letters of recommendation, government filings, essays, or, as this paper finds, lawsuits) will break under a wave of AI.
19887219
Reposted by Martin Jaggi
ICML Conference @icmlconf.bsky.social · 18/03/2026
To ensure compliance w peer-review policies, ICML has removed 795 reviews (1% of total) by reviewers who used LLMs when they explicitly agreed to not. Consequently, 497 papers (2% of all submissions) of these (reciprocal) reviewers have been desk rejected Details in blog post 👇
37722
Martin Jaggi @mjaggi.bsky.social · 18/03/2026
www.swissinfo.ch/eng/digital-...
swissinfo.ch
Is Swiss AI a powerhouse for democracy?
US cyber expert Bruce Schneier has high hopes for a Swiss AI model. Optimism is also growing in Switzerland, but not across the board.
000
Reposted by Martin Jaggi
ICML Conference @icmlconf.bsky.social · 14/02/2026
There has been some online discussion of prompt watermarks in ICML submissions. tl;dr: - Yes, this is one of the *conference*'s (several) scientific integrity measures - Yes, it's not infalliable (but it still helps) - No, your paper won't be desk rejected as a result 1/4
1114
Reposted by Martin Jaggi
Rafael Pinto @rcpinto.bsky.social · 13/02/2026
More like every week.
1312
Reposted by Martin Jaggi
Sam Harsimony @harsimony.bsky.social · 04/02/2026
Open models continue to pace closed models on a 9 month lag.
3648
Reposted by Martin Jaggi
Serge Belongie @handle.invalid · 30/01/2026
A factor of 10 billion since 2010 😮 A couple of eye-opening slides form @sloeschcke.bsky.social's presentation at today’s @belongielab.org meeting (1/2)
2118439
Reposted by Martin Jaggi
ICML Conference @icmlconf.bsky.social · 24/01/2026
The #ICML2026 abstract deadline has passed! We're at 33540 active abstracts (and dropping). How many will make it over the finish line? 🏁
1172
Reposted by Martin Jaggi
ICML Conference @icmlconf.bsky.social · 08/01/2026
New blog post (on a shiny new ICML blog!): What's New in #ICML2026 Peer Review Some highlights: - Policies to combat thinly sliced contributions - Cascading desk rejections for peer-review abuse - Reviewer reciprocity - New ways to support authors and reviewers Post: blog.icml.cc/2026/01/08/w...
1238
Reposted by Martin Jaggi
ETH Zurich @ethz.ch · 30/11/2023
A multidisciplinary team of ETH Zurich researchers developed a method of using an autonomous excavator to construct a dry-​stone wall that is six metres high and sixty-​five metres long.
ethz.ch
Autonomous excavator constructs a six-metre-high dry-stone wall
011
Reposted by Martin Jaggi
Nathan Lambert @natolambert.bsky.social · 07/01/2026
We updated the plots we use to measure the open model ecosystem at interconnects, to guide The ATOM Project, and to understand what's happening. We have ~8 plots to summarize what's happening. First, the high level picture showing China's growing adoption lead.
1257
Reposted by Martin Jaggi
ICML Conference @icmlconf.bsky.social · 11/12/2025
Announcing the ICML 2026 policy for LLMs in reviewing! Reviewers and authors both pick either conservative or permissive LLM use, and will be matched accordingly. Importantly: authors on papers who choose conservative must obey the conservative policy as reviewers.
22310
Reposted by Martin Jaggi
🤷 Nico Martin @nico.dev · 21/11/2025
👀 I am working on something pretty cool.. Hopefully, it will soon be possible to try #Apertus 🇨🇭 directly in your browser, powered by Transformers.js 🎉
Experimental Git branch to support Apertus in the browser with Transformers.js
1101
Reposted by Martin Jaggi
Alexander Doria @dorialexander.bsky.social · 26/11/2025
The threshold for consistent English/query understanding is now 3M parameters.
3572
Reposted by Martin Jaggi
Alexander Doria @dorialexander.bsky.social · 10/11/2025
Breaking: we release a fully synthetic generalist dataset for pretraining, SYNTH and two new SOTA reasoning models exclusively trained on it. Despite having seen only 200 billion tokens, Baguettotron is currently best-in-class in its size range. pleias.fr/blog/blogsyn...
318833
Reposted by Martin Jaggi
ICML Conference @icmlconf.bsky.social · 07/11/2025
🎉 ICML 2026 Call for Papers (& Position Papers) is here! 🎉 📅 Key Dates Abstract deadline: Jan 23, 2026 AOE Paper deadline: Jan 28, 2026 AOE A few key changes this year: - Attendance for authors of accepted papers is optional - Originally submitted version of accepted papers will be made public ...
1148
Martin Jaggi @mjaggi.bsky.social · 05/11/2025
so open-weights models are much happier than closed ones i guess, cause they live on in the long run, did i get that right?
020
Martin Jaggi @mjaggi.bsky.social · 14/10/2025
91% of reasoning does not need RL 🤯 arxiv.org/abs/2510.07364
arxiv.org
Base Models Know How to Reason, Thinking Models Learn When
Why do thinking language models like DeepSeek R1 outperform their base counterparts? Despite consistent performance gains, it remains unclear to what extent thinking models learn entirely new reasonin...
180
Reposted by Martin Jaggi
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 07/10/2025
I just tried the official demo for the new Gemini 2.5 Computer Use model and it started by navigating to Google, solving Google's own CAPTCHA and then running a search! simonwillison.net/2025/Oct/7/gemini…
simonwillison.net
Gemini 2.5 Computer Use can solve Google’s own CAPTCHAs
Google just introduced a new Gemini 2.5 Computer Use model, specially designed to help operate a GUI interface by interacting with visible elements using a virtual mouse and keyboard. I …
294
Martin Jaggi @mjaggi.bsky.social · 25/09/2025
We're hiring again for AI research engineering roles: Join the team behind the Apertus LLM, if you share our passion to work on impactful AI that's truly open. careers.epfl.ch/job/Lausanne...
careers.epfl.ch
AI Research Engineers - Swiss AI Initiative
AI Research Engineers - Swiss AI Initiative
254
Reposted by Martin Jaggi
Deniz Bayazit @bayazitdeniz.bsky.social · 25/09/2025
1/🚨 New preprint How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks. #interpretability
2156
Reposted by Martin Jaggi
heise online @heiseonline.flipboard.com.ap.brid.gy · 24/09/2025
Schweizer Sprachmodell Apertus: So sieht EU-konforme, transparente KI aus www.heise.de/hintergrund/Schweizer-… Gepostet in Nachrichten @nachrichten-heiseonline
heise.de
Schweizer Sprachmodell Apertus: So sieht EU-konforme, transparente KI aus
Vielsprachigkeit, Transparenz, Respekt vor geistigem Eigentum: Das offene große Sprachmodell aus Schweizer KI-Schmieden verinnerlicht europäische Werte.
011
Reposted by Martin Jaggi
Sung Kim @sungkim.bsky.social · 07/09/2025
Hugging Face's FinePDFs The largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. - Long context - 3T tokens from high-demand domains like legal and science. - Heavily improves over SoTA
1303
Martin Jaggi @mjaggi.bsky.social · 05/09/2025
you can run the new apertus LLMs fully locally on your (mac) laptop with just 2 lines of code: pip install mlx-lm mlx_lm.generate --model swiss-ai/Apertus-8B-Instruct-2509 --prompt "wer bisch du?" (make sure you have done huggingface-cli login before)
huggingface.co
Apertus LLM - a swiss-ai Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
394
Reposted by Martin Jaggi
Adrienne Fichter @adfichter.eurosky.social · 05/09/2025
Am Schluss müssen sich die Medienverlage eine gesonderte Lösung überlegen, da sie kaum für alle Schweizer Blogger, Firmenwebsites, Künstler:innen,Gesundheitsportalen, eCommerce-Plattformen sprechen können. WBK N will weder Opt Out noch Opt In festschreiben.
Schutz des geistigen Eigentums vor KI-Missbrauch: WBK-N nimmt Motion Gössi in abgeänderter Form an
Die Kommission hat sich an ihrer Sitzung mit der Motion Gössi (24.4596) befasst. In diesem Zusammenhang hat sie Vertreterinnen und Vertreter der Wirtschaft, Forschung, Medien und Kultur sowie Fachleute für Immaterialgüterrecht angehört.

Die Kommission anerkennt, dass beim Schutz des geistigen Eigentums vor Missbrauch durch künstliche Intelligenz (KI) Handlungsbedarf besteht, weshalb sie das Motionsanliegen unterstützt. Sie hält es für wichtig, dass die Schweiz die für den Erhalt der Wettbewerbsfähigkeit ihres Wirtschaftsstandorts und ihrer Innovationskraft notwendigen Bedingungen aufrechterhält, ist aber der Ansicht, dass die Motion in ihrer ursprünglichen Fassung den Handlungsspielraum zu sehr einschränkt. Sie möchte, dass auch andere Lösungsansätze geprüft werden, um sich an künftige Entwicklungen anpassen zu können und sicherzustellen, dass der Schweizer Ansatz mit den Regulierungsbemühungen anderer Staaten und der EU in Einklang steht. Sie hat daher mit 18 zu 6 Stimmen bei 1 Enthaltung beschlossen, ihrem Rat die Annahme der Motion in einer abgeänderten Fassung zu empfehlen. Diese enthält keine konkreten Vorgaben zur Umsetzung der Massnahmen und schafft so mehr Spielraum für die Erarbeitung nachhaltiger Lösungen. Die Minderheit beantragt die Ablehnung der Motion.
071
Reposted by Martin Jaggi
Antoine Bosselut @abosselut.bsky.social · 03/09/2025
The next generation of open LLMs should be inclusive, compliant, and multilingual by design. That’s why we @icepfl.bsky.social @ethz.ch @cscsch.bsky.social ) built Apertus.
2248
Martin Jaggi @mjaggi.bsky.social · 03/09/2025
new extensive evaluation of different optimizers for LLM training arxiv.org/abs/2509.01440
arxiv.org
Benchmarking Optimizers for Large Language Model Pretraining
The recent development of Large Language Models (LLMs) has been accompanied by an effervescence of novel ideas and methods to better optimize the loss of deep learning models. Claims from those method...
042
Reposted by Martin Jaggi
Reto Vogt @rvgt.ch · 02/09/2025
Die Schweiz steigt ins Rennen der grossen Sprachmodelle ein. Unter dem Namen #Apertus veröffentlichen @ethz.ch, @icepfl.bsky.social und das @cscsch.bsky.social das erste vollständig offene, mehrsprachige #LLM des Landes. Fürs MAZ habe ich Apertus kurz analysiert: www.maz.ch/news/apertus...
maz.ch
Apertus: ein neues Sprachmodell für die Schweiz
3277
Reposted by Martin Jaggi
EPFL School of Computer and Communication Sciences @icepfl.bsky.social · 02/09/2025
EPFL, ETH Zurich & CSCS just released Apertus, Switzerland’s first fully open-source large language model. Trained on 15T tokens in 1,000+ languages, it’s built for transparency, responsibility & the public good. Read more: actu.epfl.ch/news/apertus...
15429
Reposted by Martin Jaggi
EPFL AI Center @epfl-ai-center.bsky.social · 09/07/2025
EPFL and ETH Zürich are building together a Swiss made LLM from scratch. Fully open and multilingual, the model is trained on CSCS's supercomputer "Alps" and supports sovereign, transparent, and responsible AI in Switzerland and beyond. Read more here: ai.epfl.ch/a-language-m... #ResponsibleAI
ai.epfl.ch
A language model built for the public good     - EPFL AI Center
ETH Zurich and EPFL will release a large language model (LLM) developed on public infrastructure. Trained on the “Alps” supercomputer at the Swiss National Supercomputing Centre (CSCS), the new LLM ma...
0103
Martin Jaggi @mjaggi.bsky.social · 09/07/2025
huggingface.co/blog/smollm3
huggingface.co
SmolLM3: smol, multilingual, long-context reasoner
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
020
Martin Jaggi @mjaggi.bsky.social · 20/06/2025
060
Reposted by Martin Jaggi
zeynep tufekci @zey.bsky.social · 17/05/2025
Why did Grok suddenly start talking about “white genocide in South Africa” even if asked about baseball or cute dogs? Because someone at Musk’s xAi deliberately did this, and we only found out because they were clumsy. My piece on the real dangers of AI. Gift link: www.nytimes.com/2025/05/17/o...
8244108
Reposted by Martin Jaggi
Angelika Romanou @agromanou.bsky.social · 23/04/2025
If you’re at @iclr-conf.bsky.social this week, come check out our spotlight poster INCLUDE during the Thursday 3:00–5:30pm session! I will be there to chat about all things multilingual & multicultural evaluation. Feel free to reach out anytime during the conference. I’d love to connect!
042
Martin Jaggi @mjaggi.bsky.social · 23/04/2025
Using the 'right' data can hugely speed up LLM training, but how to find the best training data in the vast sea of a whole web crawl? We propose a simple classifier-based selection, enabling multilingual LLMs 🧵
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
182
Martin Jaggi @mjaggi.bsky.social · 19/04/2025
Dion: A Communication-Efficient Optimizer for Large Models (inspired by PowerSGD) arxiv.org/abs/2504.05295
arxiv.org
Dion: A Communication-Efficient Optimizer for Large Models
Training large AI models efficiently requires distributing computation across multiple accelerators, but this often incurs significant communication overhead -- especially during gradient synchronizat...
010
Reposted by Martin Jaggi
Sung Kim @sungkim.bsky.social · 16/04/2025
Prime Intellect's INTELLECT-2 The first decentralized 32B-parameter RL training run open to join for anyone with compute — fully permissionless. www.primeintellect.ai/blog/intelle...
073
Martin Jaggi @mjaggi.bsky.social · 08/03/2025
Anastasia @koloskova.bsky.social recently won the European @ellis.eu PhD award, for her amazing work on AI and optimization. She will be joining University of Zurich as a professor this summer, and hiring PhD students and postdocs. You should apply to her group! Her website: koloskova.github.io
koloskova.github.io
Anastasia Koloskova
Anastasia Koloskova, PhD student in Machine Learning at EPFL.
091
Martin Jaggi @mjaggi.bsky.social · 04/03/2025
The Swiss AI Initiative has launched open calls for disruptive ideas - Democratizing large-scale AI for the benefit of society. Send your idea by end of March 🏃‍♂️‍➡️ , and run on one of the largest public AI clusters globally. Everyone is eligible to apply! swiss-ai.org
Swiss AI Initiative Logo
01711