Sign in

Raphael Pisoni

@4rtemi5.bsky.social
3.1K followers 467 following 209 posts

Unsupervised multimodal representation of a learning researcher. www.pisoni.ai

PostsRepliesMedia
Raphael Pisoni @4rtemi5.bsky.social · 03/09/2026
Lots of people are smarter than me but my unfair advantage is that I'm dumb enough to try!
020
Raphael Pisoni @4rtemi5.bsky.social · 19/08/2026
What if I told you that Scaled Dot-Product Attention (SDPA) doesn't make a whole lot of sense for ViTs and you can get faster convergence, better performance and near-linear scaling with a few well-steered Gaussians? Let me introduce you to SSOG (Separable Sum Of Gaussians)!
131
Raphael Pisoni @4rtemi5.bsky.social · 17/04/2026
I never really considered how dangerous QK-norm actually is before working on RBF Attention. While solving some obvious issues, it can be the cause of some much less obvious ones.🧵
131
Raphael Pisoni @4rtemi5.bsky.social · 08/04/2026
Neural networks have a fundamental problem. Feed them garbage data and instead of admitting that they are confused, they will confidently hallucinate. I just open-sourced the HALO-Loss to try and fix this. It give the model a mathematically sound *I don't know!* button.🧵
120
Raphael Pisoni @4rtemi5.bsky.social · 01/04/2026
I dove deeper into the rabbit hole of RBF-Attention. I refined the Triton kernel, added register-tokens and developed SuSiE positional embedding as a replacement for RoPE in Euclidean space. Go have a look at the repo or the blogpost in the comments if you're interested! :)
120
Raphael Pisoni @4rtemi5.bsky.social · 28/03/2026
For some reason I decided to swap out standard dot-product attention for a scaled-rbf kernel. Pretty much expected it to fail to converge or be impossibly slow but the scaled-rbf-attention is getting unexpectedly good results?? 👇
Plot of two loss curves were rbf-attention consistently outperforms scaled dot-product attention.
120
Raphael Pisoni @4rtemi5.bsky.social · 26/03/2026
Playing with a new kind of attention. Plots are for the same setup with standard vs. modified attention on one epoch of tiny-stories. Speed is roughly the same as flash-attention. Looking good!🤞
131
Raphael Pisoni @4rtemi5.bsky.social · 31/01/2026
AI isn't coming for your creativity. It's coming for your lack of diligence. People talk a lot about #AGI and "super-intelligence," but the immediate disruption is much simpler: AI is killing "vibe-based" decision-making.
100
Raphael Pisoni @4rtemi5.bsky.social · 23/01/2026
Over the past year Michał Lewandowski and I published a series of papers on Space Folding , and while Michał went to #AAAI to present the latest one, I worked on a blog-post explaining some the central ideas behind the papers. Let me know what you think! www.pisoni.ai/posts/space-...
pisoni.ai
The Shape of Thought: Space Folding in Neural Networks
The mathematical description of deep learning has long been dominated by the language of algebra: matrices, gradients, and optimization landscapes. A parallel and perhaps more intuitive language howev
050
Raphael Pisoni @4rtemi5.bsky.social · 23/01/2026
After a long hiatus I decided to update my blog and write about some of the things I did over the last few years. Come have a look! pisoni.ai
130
Raphael Pisoni @4rtemi5.bsky.social · 01/12/2025
Currently heading to #EurIPS in Copenhagen to present our work on space folding and model interpretability. If you're attending and would like to discuss Representation Learning, SSL, Multimodal LLMs, CV, or other topics that YOU are excited about, feel free to reach out.
040
Reposted by Raphael Pisoni
hardmaru @hardmaru.bsky.social · 07/11/2025
The US government should subsidize Open AI rather than OpenAI
0477
Reposted by Raphael Pisoni
Yuki M Asano @yukimasano.bsky.social · 15/10/2025
On the occasion of the 1000th citation of our Sinkhorn-Knopp self-supervised representation learning paper, I've written a whole post about the history and the key bits of this method that powers the state-of-the-art SSL vision models. Read it here :): docs.google.com/document/d/1...
1225
Raphael Pisoni @4rtemi5.bsky.social · 21/09/2025
We're ready!
000
Raphael Pisoni @4rtemi5.bsky.social · 06/09/2025
The single most undervalued property of neural networks is self-consistency. We should change that!
020
Reposted by Raphael Pisoni
asker the gauche, glycojohn destroyer of carbs @johnbender.bsky.social · 08/08/2025
216022
Raphael Pisoni @4rtemi5.bsky.social · 26/07/2025
You've been researching for a while! Time to have some SOTA! #aislop
030
Raphael Pisoni @4rtemi5.bsky.social · 26/07/2025
You and Adam keep beating Sota? Stop doing that! Poor Sota!
190
Raphael Pisoni @4rtemi5.bsky.social · 26/07/2025
Have some cool idea but only evaluate it on small models? Tough luck buddy. You only get your paper accepted if your experimental results are 0.2% above SOTA and too expensive to falsify! Is academic publishing pay to win yet?
030
Raphael Pisoni @4rtemi5.bsky.social · 23/07/2025
Is there a reason why none of the recent models use RBF-kernel Attention to get rid of the softmax-bottleneck for long context? I tried replacing dot-product attention with the negative squared KQ-distance and was able to remove the softmax without issues and loss in performance!
131
Reposted by Raphael Pisoni
NeurIPS Conference @neuripsconf.bsky.social · 16/07/2025
NeurIPS is endorsing EurIPS, an independently-organized meeting which will offer researchers an opportunity to additionally present NeurIPS work in Europe concurrently with NeurIPS. Read more in our blog post and on the EurIPS website: blog.neurips.cc/2025/07/16/n... eurips.cc
eurips.cc
eurips.cc
A NeurIPS-endorsed conference in Europe held in Copenhagen, Denmark
112439
Raphael Pisoni @4rtemi5.bsky.social · 08/07/2025
Has anyone experimented with "conditional gradients"? Thinking about a setup where, within a specific activation range (e.g., right before a ReLU), you'd only permit positive or negative gradients.
110
Raphael Pisoni @4rtemi5.bsky.social · 29/06/2025
Quick question to the SSL experts out there: Usually you evaluate an ssl-model by freezing it and training a linear probing layer. Would it be fair to somehow learn a final layer with more dimensions than classes and do a nearest-neighbor evaluation?
000
Reposted by Raphael Pisoni
David Picard @davidpicard.eurosky.social · 17/06/2025
There is an oak forest in central France that was planted 400 years ago by Colbert so that France would have quality hard wood by the 2000s to build ships for its navy. This is the type of long term planning that Seldonian predictions can help improving.
172
Reposted by Raphael Pisoni
Nafnlaus 🇮🇸🇪🇺🇺🇦 @nafnlaus.bsky.social · 13/05/2025
New anti-censorship jailbreak just dropped ;)
1317
Raphael Pisoni @4rtemi5.bsky.social · 18/04/2025
Currently on my way to #ICLR in Singapore where we'll present our latest paper on space folding in neural networks. Would be happy to meet some people there so if you're at ICLR as well and want to hang out feel free to pm!🙂
130
Raphael Pisoni @4rtemi5.bsky.social · 16/04/2025
Grok this! What a roller-coaster of emotions...🤪
140
Reposted by Raphael Pisoni
Wissam Antoun @wissamantoun.bsky.social · 14/04/2025
ModernBERT or DeBERTaV3? What's driving performance: architecture or data? To find out we pretrained ModernBERT on the same dataset as CamemBERTaV2 (a DeBERTaV3 model) to isolate architecture effects. Here are our findings:
34315
Reposted by Raphael Pisoni
Dmytro Mishkin @ducha-aiki.bsky.social · 13/04/2025
Just assembled a slide about local feature training time/dataset size. Anything wrong/missing?
5174
Raphael Pisoni @4rtemi5.bsky.social · 12/04/2025
Is the project even still worth doing when wandb runs out of funny names or am I cooked?🫠
110
Reposted by Raphael Pisoni
Jeremy Morrell @jeremymorrell.dev · 05/04/2025
Meta introduced Llama 4 models and added this section near the very bottom of the announcement 😬 “[LLMs] historically have leaned left when it comes to debated political and social topics.” ai.meta.com/blog/llama-4...
Meta
Addressing bias in LLMs

It's well-known that all leading LLMs have had issues with bias-specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet.

Our goal is to remove bias from our Al models and to make sure that Llama can understand and articulate both sides of a contentious issue. As part of this work, we're continuing to make Llama more responsive so that it answers questions, can respond to a variety of different viewpoints without passing judgment, and doesn't favor some views over others.

We have made improvements on these efforts with this release—Llama 4 performs significantly better than Llama 3 and is comparable to Grok:• Llama 4 refuses less on debated political and social topics overall (from 7% in Lama 3.3 to below 2%).
• Llama 4 is dramatically more balanced with which prompts it refuses to respond to (the proportion of unequal response refusals is now less than 1% on a set of debated topical questions).
• Our testing shows that Llama 4 responds with strong political lean at a rate comparable to Grok (and at half of the rate of Llama 3.3) on a contentious set of political or social topics. While we are making progress, we know we have more work to do and will continue to drive this rate further down.
We're proud of this progress to date and remain committed to our goal of eliminating overall bias in our models.
513438
Reposted by Raphael Pisoni
ETH CS Department @csateth.bsky.social · 24/03/2025
🚀Hello, world! We are now live on Bluesky. This is the official account of the Department of Computer Science at ETH Zurich. Follow us for cutting-edge research, the latest innovations, event updates and insights into the future of technology. inf.ethz.ch @csateth.bsky.social @ethzurich.bsky.social
inf.ethz.ch
Department of Computer Science
Computer Science Department at ETH Zurich. The department offers highest quality in computer science research and education and adds to business and industry growth.
1228
Raphael Pisoni @4rtemi5.bsky.social · 24/03/2025
Recently had the pleasure of helping @miclew.bsky.social with a couple of his papers in exchange for him helping me with a couple of mine! This is the first fruit of our common work. We quantify space folding in relu neural networks with a range based measure. Lots of fun to write and read!😉
060
Raphael Pisoni @4rtemi5.bsky.social · 24/03/2025
x''= 0
030
Reposted by Raphael Pisoni
Gabriele Berton @berton-gabri.bsky.social · 18/03/2025
🚀 Paper Release! 🚀 Curious about image retrieval and contrastive learning? We present: 📄 "All You Need to Know About Training Image Retrieval Models" 🔍 The most comprehensive retrieval benchmark—thousands of experiments across 4 datasets, dozens of losses, batch sizes, LRs, data labeling, and more!
24010
Reposted by Raphael Pisoni
Rafael Pinto @rcpinto.bsky.social · 23/02/2025
"no b-but deepseek c-can't tiannamen" Here's Grok for you:
2308
Reposted by Raphael Pisoni
Hank Green @hankgreen.bsky.social · 28/01/2025
The fact that Deepseek R1 was released three days /before/ Stargate means these guys stood in front of Trump and said they needed half a trillion dollars while they knew R1 was open source and trained for $5M. Beautiful.
Trump announces 500B in AI funding. Five days ago. Deepseek r1 release. 8 days ago.
396138071762
Raphael Pisoni @4rtemi5.bsky.social · 21/01/2025
Super interesting new CLIP-loss that takes cross-sample similarities into account to learn consistent representations. Also makes pretraining very data efficient. But i think there is a catch...👇
arxiv.org
$\mathbb{X}$-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs
Learning good representations involves capturing the diverse ways in which data samples relate. Contrastive loss - an objective matching related samples - underlies methods from self-supervised to mul...
140
Raphael Pisoni @4rtemi5.bsky.social · 21/01/2025
Not that I'm super active on social media recently but I still feel like I need a break... 🫣
000
Raphael Pisoni @4rtemi5.bsky.social · 14/01/2025
Another nail in the coffin of cosine similarity! I started disliking cossim some years ago due to multiple reasons such as the non-linearity around 0.0 and the loss of certainty-information due to the normalization of feature vectors but this study seems to give another good reason to abandon it.
shaped.ai
Cosine Similarity: Not the Silver Bullet We Thought It Was | Shaped Blog
In the world of machine learning and data science, cosine similarity has long been a go-to metric for measuring the semantic similarity between high-dimensional objects. However, a new study by resear...
1327
Reposted by Raphael Pisoni
Jeremy Howard @howard.fm · 19/12/2024
I'll get straight to the point. We trained 2 new models. Like BERT, but modern. ModernBERT. Not some hypey GenAI thing, but a proper workhorse model, for retrieval, classification, etc. Real practical stuff. It's much faster, more accurate, longer context, and more useful. 🧵
19620147
Raphael Pisoni @4rtemi5.bsky.social · 29/12/2024
She said YES!🥰
2280
Reposted by Raphael Pisoni
Mark Riedl @markriedl.bsky.social · 28/12/2024
O3 is costly. These numbers are for a single ARC benchmark task
23612
Raphael Pisoni @4rtemi5.bsky.social · 15/12/2024
Fantastic Muse quote on a fantastic NeurIPS poster! Doesn't get much better than that!😅
180
Reposted by Raphael Pisoni
Remi Cadene @remicadene.bsky.social · 15/12/2024
HOT 🔥 fastest, most precise, and most capable hand control setup ever... Less than $450 and fully open-source 🤯 by @huggingface, @therobotstudio, @NepYope This tendon-driven technology will disrupt robotics! Retweet to accelerate its democratization 🚀 A thread 🧵
37327
Raphael Pisoni @4rtemi5.bsky.social · 09/12/2024
Stand by while annual NeurIPS FOMO is loading...
140
Raphael Pisoni @4rtemi5.bsky.social · 05/12/2024
I mean no disrespect but the timing of publishing this before moving to OpenAI, who has changed its ethical standpoint quite frequently and is by now openly deploying its tech on the battlefield is a bit unfortunate. I wish all involved parties the best though so let's hope nobody gets burned.🤞
2160