Martin Jaggi @mjaggi.bsky.social · 19/09/2026In view of recent discussions on risks of AI frontier models, we started a petition for more open science in AI safety. Please consider signing and sharing: make-safety-open.github.iomake-safety-open.github.ioA Call for Open Science in AI Safety 030
Reposted by Martin JaggiApertus @apertusllm.bsky.social · 24/07/2026Available now! Apertus 1.5 — 8B & 70B models with multimodal text/image/audio input, a 4x longer context window, optional thinking mode, and improved tool use, all built transparently and responsibly. 3.2M+ downloads 🚀 #Apertus #OpenSourceLLM Details+Links+Open roles: apertus-ai.org/articles/202... 3106
Reposted by Martin JaggiEPFL AI Center @epfl-ai-center.bsky.social · 24/07/2026Apertus 1.5 is out! 🚀 ✅ Multimodal capabilities ✅ Stronger reasoning ✅ A roadmap of regular releases Read more: apertus-ai.org/news Models: huggingface.co/swiss-ai @cscsch.bsky.social @eth-ai-center.bsky.social @icepfl.bsky.social @mjaggi.bsky.social @abosselut.bsky.social 064
Reposted by Martin JaggiCSCS - Swiss National Supercomputing Centre @cscsch.bsky.social · 22/07/2026#Apertus is mentioned —the open large language model developed by ETH Zürich, EPFL and CSCS—which demonstrates how excellent research, cutting-edge infrastructures and international collaboration can contribute to strategically important technologies ⬇️ 033
Reposted by Martin JaggiEPFL AI Center @epfl-ai-center.bsky.social · 20/07/2026UNICC and EPFL's Machine Learning and Optimization Laboratory led by @mjaggi.bsky.social have published a white paper presenting a practical framework for evaluating the safety and reliability of LLMs in institutional settings. 👉 Learn more: ai.epfl.ch/epfl-lab-and... 042
Martin Jaggi @mjaggi.bsky.social · 03/07/2026Searching for something new to read? All ICML 2026 papers are now public: openreview.net/group?id=ICM... (including all 6551 accepted papers and also rejects which opt-in)openreview.netICML 2026 ConferenceWelcome to the OpenReview homepage for ICML 2026 Conference 053
Reposted by Martin Jaggi🤷 Nico Martin @nico.dev · 26/06/2026Apertus Mini is now running entirely in your browser 🇨🇭 80+ tps for the 1.5B, 60+ tps for the 4B (on my M3). Fully client-side via Transformers.js + ONNX + WebGPU. 152
Reposted by Martin JaggiApertus @apertusllm.bsky.social · 26/06/2026Three new model weights: 0.5B, 1.5B, 4B are available using new quantization and distillation techniques. Download #Apertus 1.1, read the ICML workshop report, try a new demo on @hf.co - all just a tap away in our latest blog post: apertus-ai.org/articles/202...apertus-ai.orgAPERTVS.aiFully Open Foundation Model for Sovereign AI 022
Martin Jaggi @mjaggi.bsky.social · 23/05/2026i’d be supportive for accepting those to arxiv, as by current policy. surely amounts will skyrocket next years, but not necessarily a bad thing for science 030
Reposted by Martin JaggiICML Conference @icmlconf.bsky.social · 18/05/2026Announcing the #ICML2026 invited speakers! Pascale Fung Susan Athey (@susanathey.bsky.social) Sham Kakade (@shamkakade.bsky.social ) Aviv Regev Verena Rieser (@verenarieser.bsky.social) Arvind Narayanan (@randomwalker.bsky.social) Check out the blog post for more info! blog.icml.cc/2026/05/18/a... 0124
Reposted by Martin JaggiICML Conference @icmlconf.bsky.social · 30/04/2026Ahem, back to business... Decision notifications are being released on OpenReview. There were 23,918 submissions that entered review, roughly double last year. 6,352 papers were accepted, for an acceptance rate of 26.6%. 536 papers (2.2% of submissions) are "spotlights." 1/3 1177
Martin Jaggi @mjaggi.bsky.social · 24/04/2026Link to the rest of the course materials: github.com/epfml/OptML_... And to a recent paper on the connection to Frank-Wolfe: arxiv.org/abs/2506.04192github.comGitHub - epfml/OptML_course: EPFL Course - Optimization for Machine Learning - CS-439EPFL Course - Optimization for Machine Learning - CS-439 - epfml/OptML_course 000
Martin Jaggi @mjaggi.bsky.social · 24/04/2026Muon: I made a new 3-slides explanation of this amazing optimizer for today's lecture. Let me know what you think 240
Reposted by Martin JaggiEthan Mollick @emollick.bsky.social · 22/04/2026Every system that was regulated, either explicitly or implicitly, by the fact that they were effortful for humans (letters of recommendation, government filings, essays, or, as this paper finds, lawsuits) will break under a wave of AI. 18887219
Martin Jaggi @mjaggi.bsky.social · 05/04/2026any advice on how to reliably find them? asking for a friend… 110
Martin Jaggi @mjaggi.bsky.social · 18/03/2026using LLMs by authors is allowed -if done responsibly. it is also allowed for reviewers who chose the permissive policy. what we require is that authors who want their paper to be reviewed by humans, must as (reciprocal) reviewer adhere to the same standards. see also here: icml.cc/Conferences/... 130
Reposted by Martin JaggiICML Conference @icmlconf.bsky.social · 18/03/2026To ensure compliance w peer-review policies, ICML has removed 795 reviews (1% of total) by reviewers who used LLMs when they explicitly agreed to not. Consequently, 497 papers (2% of all submissions) of these (reciprocal) reviewers have been desk rejected Details in blog post 👇 37722
Martin Jaggi @mjaggi.bsky.social · 18/03/2026www.swissinfo.ch/eng/digital-...swissinfo.chIs Swiss AI a powerhouse for democracy?US cyber expert Bruce Schneier has high hopes for a Swiss AI model. Optimism is also growing in Switzerland, but not across the board. 000
Reposted by Martin JaggiICML Conference @icmlconf.bsky.social · 14/02/2026There has been some online discussion of prompt watermarks in ICML submissions. tl;dr: - Yes, this is one of the *conference*'s (several) scientific integrity measures - Yes, it's not infalliable (but it still helps) - No, your paper won't be desk rejected as a result 1/4 1114
Reposted by Martin JaggiSam Harsimony @harsimony.bsky.social · 04/02/2026Open models continue to pace closed models on a 9 month lag. 3648
Reposted by Martin JaggiSerge Belongie @serge.belongie.com · 30/01/2026A factor of 10 billion since 2010 😮 A couple of eye-opening slides form @sloeschcke.bsky.social's presentation at today’s @belongielab.org meeting (1/2) 2118339
Reposted by Martin JaggiICML Conference @icmlconf.bsky.social · 24/01/2026The #ICML2026 abstract deadline has passed! We're at 33540 active abstracts (and dropping). How many will make it over the finish line? 🏁 1172
Reposted by Martin JaggiICML Conference @icmlconf.bsky.social · 08/01/2026New blog post (on a shiny new ICML blog!): What's New in #ICML2026 Peer Review Some highlights: - Policies to combat thinly sliced contributions - Cascading desk rejections for peer-review abuse - Reviewer reciprocity - New ways to support authors and reviewers Post: blog.icml.cc/2026/01/08/w... 1238
Reposted by Martin JaggiETH Zurich @ethz.ch · 30/11/2023A multidisciplinary team of ETH Zurich researchers developed a method of using an autonomous excavator to construct a dry-stone wall that is six metres high and sixty-five metres long.ethz.chAutonomous excavator constructs a six-metre-high dry-stone wall 011
Reposted by Martin JaggiNathan Lambert @natolambert.bsky.social · 07/01/2026We updated the plots we use to measure the open model ecosystem at interconnects, to guide The ATOM Project, and to understand what's happening. We have ~8 plots to summarize what's happening. First, the high level picture showing China's growing adoption lead. 1257
Martin Jaggi @mjaggi.bsky.social · 11/12/2025no. is this the same as how en.wikipedia.org/wiki/Royal_S... was created back then?en.wikipedia.orgRoyal Society B - Wikipedia 110
Reposted by Martin JaggiICML Conference @icmlconf.bsky.social · 11/12/2025Announcing the ICML 2026 policy for LLMs in reviewing! Reviewers and authors both pick either conservative or permissive LLM use, and will be matched accordingly. Importantly: authors on papers who choose conservative must obey the conservative policy as reviewers. 22310
Martin Jaggi @mjaggi.bsky.social · 01/12/2025what about Apertus? (seems they missed to add us in that ranking) 010
Reposted by Martin Jaggi🤷 Nico Martin @nico.dev · 21/11/2025👀 I am working on something pretty cool.. Hopefully, it will soon be possible to try #Apertus 🇨🇭 directly in your browser, powered by Transformers.js 🎉 1101
Reposted by Martin JaggiAlexander Doria @dorialexander.bsky.social · 26/11/2025The threshold for consistent English/query understanding is now 3M parameters. 3572
Martin Jaggi @mjaggi.bsky.social · 12/11/2025thanks for the lausanne visit and sharing these super cool results! 030
Reposted by Martin JaggiAlexander Doria @dorialexander.bsky.social · 10/11/2025Breaking: we release a fully synthetic generalist dataset for pretraining, SYNTH and two new SOTA reasoning models exclusively trained on it. Despite having seen only 200 billion tokens, Baguettotron is currently best-in-class in its size range. pleias.fr/blog/blogsyn... 318833
Reposted by Martin JaggiICML Conference @icmlconf.bsky.social · 07/11/2025🎉 ICML 2026 Call for Papers (& Position Papers) is here! 🎉 📅 Key Dates Abstract deadline: Jan 23, 2026 AOE Paper deadline: Jan 28, 2026 AOE A few key changes this year: - Attendance for authors of accepted papers is optional - Originally submitted version of accepted papers will be made public ... 1148
Martin Jaggi @mjaggi.bsky.social · 05/11/2025so open-weights models are much happier than closed ones i guess, cause they live on in the long run, did i get that right? 020
Martin Jaggi @mjaggi.bsky.social · 17/10/2025this seems to become a trend already: arxiv.org/abs/2510.14901arxiv.orgReasoning with Sampling: Your Base Model is Smarter Than You ThinkFrontier reasoning models have exhibited incredible capabilities across a wide array of disciplines, driven by posttraining large language models (LLMs) with reinforcement learning (RL). However, desp... 030
Martin Jaggi @mjaggi.bsky.social · 14/10/202591% of reasoning does not need RL 🤯 arxiv.org/abs/2510.07364arxiv.orgBase Models Know How to Reason, Thinking Models Learn WhenWhy do thinking language models like DeepSeek R1 outperform their base counterparts? Despite consistent performance gains, it remains unclear to what extent thinking models learn entirely new reasonin... 180
Reposted by Martin JaggiSimon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 07/10/2025I just tried the official demo for the new Gemini 2.5 Computer Use model and it started by navigating to Google, solving Google's own CAPTCHA and then running a search! simonwillison.net/2025/Oct/7/gemini…simonwillison.netGemini 2.5 Computer Use can solve Google’s own CAPTCHAsGoogle just introduced a new Gemini 2.5 Computer Use model, specially designed to help operate a GUI interface by interacting with visible elements using a virtual mouse and keyboard. I … 294
Martin Jaggi @mjaggi.bsky.social · 05/10/2025apertus also! (september release, same mission but multilingual) 010
Martin Jaggi @mjaggi.bsky.social · 03/10/2025cool idea. let’s us know how it goes! btw maybe these can be useful github.com/swiss-ai/ape... or, since today, also unsloth and llamacppgithub.comGitHub - swiss-ai/apertus-finetuning-recipesContribute to swiss-ai/apertus-finetuning-recipes development by creating an account on GitHub. 110
Martin Jaggi @mjaggi.bsky.social · 26/09/2025on the engineering track it renews yearly usually, but permanent is possible after some experience & paperwork. on the academic track see e.g. here www.epfl.ch/about/workin...epfl.chFaculty Positions in Computer & Communication Sciences – Learning SciencesThe School of Computer and Communication Sciences (IC) at EPFL invites applications for tenure-track faculty positions in learning sciences and educational technologies, with a focus on computational ... 010
Martin Jaggi @mjaggi.bsky.social · 25/09/2025Link to the first version of the Apertus open-data open-weights LLM - multilingual in >1000 languages, and compliant ethical AI huggingface.co/collections/...huggingface.coApertus LLM - a swiss-ai CollectionDemocratizing Open and Compliant LLMs for Global Language Environments: 8B and 70B open-data open-weights models, multilingual in >1000 languages 010
Martin Jaggi @mjaggi.bsky.social · 25/09/2025Several open positions at EPFL Lausanne and ETH Zurich and, as part of the Swiss AI Initiative. We cover the entire stack of foundation model training. And we're open to international applicants of course (no H-1B required ;)) 120
Martin Jaggi @mjaggi.bsky.social · 25/09/2025We're hiring again for AI research engineering roles: Join the team behind the Apertus LLM, if you share our passion to work on impactful AI that's truly open. careers.epfl.ch/job/Lausanne...careers.epfl.chAI Research Engineers - Swiss AI InitiativeAI Research Engineers - Swiss AI Initiative 254
Reposted by Martin JaggiDeniz Bayazit @bayazitdeniz.bsky.social · 25/09/20251/🚨 New preprint How do #LLMs’ inner features change as they train? Using #crosscoders + a new causal metric, we map when features appear, strengthen, or fade across checkpoints—opening a new lens on training dynamics beyond loss curves & benchmarks. #interpretability 2156
Reposted by Martin Jaggiheise online @heiseonline.flipboard.com.ap.brid.gy · 24/09/2025Schweizer Sprachmodell Apertus: So sieht EU-konforme, transparente KI aus www.heise.de/hintergrund/Schweizer-… Gepostet in Nachrichten @nachrichten-heiseonlineheise.deSchweizer Sprachmodell Apertus: So sieht EU-konforme, transparente KI ausVielsprachigkeit, Transparenz, Respekt vor geistigem Eigentum: Das offene große Sprachmodell aus Schweizer KI-Schmieden verinnerlicht europäische Werte. 011
Martin Jaggi @mjaggi.bsky.social · 18/09/2025funktioniert schon seit letzter woche im neusten LM Studio (mit MLX) huggingface.co/models?searc... GGUF kommt auch bald die tage 100
Martin Jaggi @mjaggi.bsky.social · 13/09/2025no. the commercial models like chatGPT and gemini still can do better swiss german than apertus. 010
Reposted by Martin JaggiSung Kim @sungkim.bsky.social · 07/09/2025Hugging Face's FinePDFs The largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. - Long context - 3T tokens from high-demand domains like legal and science. - Heavily improves over SoTA 1303