Sign in

Joan Serrà

@serrjoa.bsky.social
321 followers 151 following 23 posts

Does research on machine learning at Sony AI, Barcelona. Works on audio analysis, synthesis, and retrieval. Likes tennis, music, and wine. serrjoa.github.io

PostsRepliesMedia
Joan Serrà @serrjoa.bsky.social · 08/01/2025
I think I may switch back to Twitter/X. Somehow I feel this site didn't take off and I really don't want to be looking at two feeds all the time...
330
Reposted by Joan Serrà
Dmytro Mishkin @ducha-aiki.bsky.social · 02/01/2025
Image matching and ChatGPT - new post in the wide baseline stereo blog. tl;dr: it is good, even feels like human, but not perfect. ducha-aiki.github.io/wide-baselin...
ducha-aiki.github.io
ChatGPT and Image Matching – Wide baseline stereo meets deep learning
Are we done yet?
2348
Reposted by Joan Serrà
Andrew Gordon Wilson @andrewgwils.bsky.social · 28/12/2024
Many of the greatest papers, now canonical works, have a story of resistance, tension, and, finally, a crucial advocate. It's shockingly common. Why is there a bias against excellence? And what happens to those papers, those people, when no one has the courage to advocate?
1122
Joan Serrà @serrjoa.bsky.social · 23/12/2024
Do you want to work with me for some months? Two internship positions available at the Music Team of Sony AI in Barcelona! 👇
Views from the office window. Photo taken just now.
1124
Joan Serrà @serrjoa.bsky.social · 21/12/2024
I'm happy to have two papers accepted at #ICASSP2025! 1) Contrastive learning for audio-video sequences, exploiting the fact that they are *sequences*: arxiv.org/abs/2407.05782 2) Knowledge distillation at *pre-training* time to help generative speech enhancement: arxiv.org/abs/2409.09357
2150
Joan Serrà @serrjoa.bsky.social · 20/12/2024
Flow matching mapping text to image directly (instead of noise to image): cross-flow.github.io
cross-flow.github.io
Flowing from Words to Pixels: A Framework for Cross-Modality EvolutionFlowing from Words to Pixels: A Framework for Cross-Modality Evolution
Project page for 'Flowing from Words to Pixels: A Framework for Cross-Modality Evolution.'
040
Reposted by Joan Serrà
Alexander Kolesnikov @handle.invalid · 20/12/2024
With some delay, JetFormer's *prequel* paper is finally out on arXiv: a radically simple ViT-based normalizing flow (NF) model that achieves SOTA results in its class. Jet is one of the key components of JetFormer, deserving a standalone report. Let's unpack: 🧵⬇️
2427
Reposted by Joan Serrà
Deep Learning Barcelona @dlbcnai.bsky.social · 19/12/2024
Did you miss any of the talks of the Deep Learning Barcelona Symposyum 2024 ? Play them now from the recorded stream: www.youtube.com/live/yPc-Un3...
youtube.com
YouTube
Share your videos with friends, family, and the world
041
Reposted by Joan Serrà
Jeremy Howard @howard.fm · 19/12/2024
I'll get straight to the point. We trained 2 new models. Like BERT, but modern. ModernBERT. Not some hypey GenAI thing, but a proper workhorse model, for retrieval, classification, etc. Real practical stuff. It's much faster, more accurate, longer context, and more useful. 🧵
19620147
Joan Serrà @serrjoa.bsky.social · 19/12/2024
On pre-acrivation norm, learnable residuals, etc.
010
Reposted by Joan Serrà
Kyle Kastner @kastnerkyle.bsky.social · 18/12/2024
Two great tokenizer blog posts that helped me over the years: sjmielke.com/papers/token... sjmielke.com/comparing-pe... People have mostly standardized on certain tokenizations right now, but there are huge performance gaps between locales with high agglomeration (e.g. common en-us) and ...
sjmielke.com
A simple, reversible, language-agnostic tokenizer — Sabrina J. Mielke
Hi. I'm a PhD student at Johns Hopkins University Center for Language and Speech Processing (JHU CLSP). Machine learning and natural language are fascinating.
1102
Joan Serrà @serrjoa.bsky.social · 15/12/2024
Don't be like Reviewer 2.
030
Reposted by Joan Serrà
Peyman Milanfar @docmilanfar.bsky.social · 14/12/2024
Did Gauss invent the Gaussian? - Laplace wrote down the integral first in 1783 - Gauss then described it in 1809 in the context of least-sq. for astronomical measurements - Pearson & Fisher framed it as ‘normal’ density only in 1910 * Best part is: Gauss gave Laplace credit!
0385
Joan Serrà @serrjoa.bsky.social · 13/12/2024
I already signed up (as a mentor) for this year!
010
Reposted by Joan Serrà
Jörg Franke @jfranke.bsky.social · 09/12/2024
Thrilled to present our work on Constrained Parameter Regularization (CPR) at #NeurIPS2024! Our novel deep learning regularization outperforms weight decay across various tasks. neurips.cc/virtual/2024... This is joint work with Michael Hefenbrock, Gregor Köhler, and Frank Hutter 🧵👇
neurips.cc
NeurIPS Poster Improving Deep Learning Optimization through Constrained Parameter RegularizationNeurIPS 2024
121
Reposted by Joan Serrà
Keenan Crane @keenancrane.bsky.social · 09/12/2024
Entropy is one of those formulas that many of us learn, swallow whole, and even use regularly without really understanding. (E.g., where does that “log” come from? Are there other possible formulas?) Yet there's an intuitive & almost inevitable way to arrive at this expression.
22543128
Reposted by Joan Serrà
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/12/2024
Inventors of flow matching have released a comprehensive guide going over the math & code of flow matching! Also covers variants like non-Euclidean & discrete flow matching. A PyTorch library is also released with this guide! This looks like a very good read! 🔥 arxiv: arxiv.org/abs/2412.06264
110927
Reposted by Joan Serrà
Tanishq Mathew Abraham @iscienceluvr.bsky.social · 10/12/2024
Normalizing Flows are Capable Generative Models Apple introduces TarFlow, a new Transformer-based variant of Masked Autoregressive Flows. SOTA on likelihood estimation for images, quality and diversity comparable to diffusion models. arxiv.org/abs/2412.06329
arxiv.org
Normalizing Flows are Capable Generative Models
Normalizing Flows (NFs) are likelihood-based models for continuous inputs. They have demonstrated promising results on both density estimation and generative modeling tasks, but have received relati...
1549
Reposted by Joan Serrà
Deep Learning Barcelona @dlbcnai.bsky.social · 09/12/2024
That was fast: #DLBCN 2024 was sold out in less than two hours ! New requests will be added to a waiting list. Read the instructions for same day event registration: sites.google.com/view/dlbcn20...
031
Reposted by Joan Serrà
Robert Nowak @rdnowak.bsky.social · 07/12/2024
Past work has characterized the functions learned by neural networks: arxiv.org/pdf/1910.01635, arxiv.org/abs/1902.05040, arxiv.org/abs/2109.12960, arxiv.org/abs/2105.03361. But it turns out multi-task training produces strikingly different solutions! Adding tasks produces “kernel-like” solutions.
17411
Reposted by Joan Serrà
Ferenc Huszár @inference.vc · 06/12/2024
Can language models transcend the limitations of training data? We train LMs on a formal grammar, then prompt them OUTSIDE of this grammar. We find that LMs often extrapolate logical rules and apply them OOD, too. Proof of a useful inductive bias. Check it out at NeurIPS: nips.cc/virtual/2024...
nips.cc
NeurIPS Poster Rule Extrapolation in Language Modeling: A Study of Compositional Generalization on OOD PromptsNeurIPS 2024
71138
Reposted by Joan Serrà
Leo Boytsov @srchvrs.bsky.social · 06/12/2024
"We therefore recommend to consider logistic regression as the first choice for data-scarce applications with tabular data and provide practitioners with best practices for further method selection." arxiv.org/abs/2405.07662
arxiv.org
Squeezing Lemons with Hammers: An Evaluation of AutoML and Tabular Deep Learning for Data-Scarce Classification Applications
Many industry verticals are confronted with small-sized tabular data. In this low-data regime, it is currently unclear whether the best performance can be expected from simple baselines, or more compl...
262
Reposted by Joan Serrà
Nick Stracke @rmsnorm.bsky.social · 04/12/2024
🤔 Why do we extract diffusion features from noisy images? Isn’t that destroying information? Yes, it is - but we found a way to do better. 🚀 Here’s how we unlock better features, no noise, no hassle. 📝 Project Page: compvis.github.io/cleandift 💻 Code: github.com/CompVis/clea... 🧵👇
24210
Reposted by Joan Serrà
David Picard @davidpicard.eurosky.social · 04/12/2024
My eyes, it hurts.
6606
Reposted by Joan Serrà
Ethan Mollick @emollick.bsky.social · 04/12/2024
New paper shows AI art models don't need art training data to make recognizable artistic work. Train on regular photos, let an artist add 10-15 examples their own art (or some other artistic inspiration), and get results similar to models trained on millions of people’s artworks
59913
Reposted by Joan Serrà
Richard McElreath 🐈‍⬛ @rmcelreath.bsky.social · 04/12/2024
Nice to see this article by @mbaldwin.bsky.social surface again. Lots of academics aren't aware how recent and dynamic the peer-review system is. Entire thing worth a read, but this quote always floors me (page 8 of PDF ethos.lps.library.cmu.edu/article/19/g... ):
At the American weekly Science, for instance,
the Editorial Board handled all refereeing in-house for the first half of the twentieth century.
In the 1950s, however, members of the editorial board complained that “the job of refereeing
and suggesting revisions for hundreds of technical papers is neither the best use of their
time nor pleasant, satisfying work,” and agreed to begin sending papers to outside experts.22
Similarly, when the American Journal of Medicine was founded in 1946, its editor Alexander
Gutman wanted to offer his authors fast publication and decided to handle acceptances and
rejections almost entirely on his own. However, as the journal became more popular,
Gutman was unable to keep up with the number of submissions, and by the 1960s he too
had begun sending papers out for external opinions.
26722
Joan Serrà @serrjoa.bsky.social · 04/12/2024
For those of you who don't know (or don't remember), the original paper introducing "attention" was Bahdanau, Cho, & Bengio's "Neural machine translation by jointly learning to align and translate" (2014), not the famous "Attention is all you need" from 2017. arxiv.org/abs/1409.0473
arxiv.org
Neural Machine Translation by Jointly Learning to Align and Translate
Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neur...
180
Reposted by Joan Serrà
Sander Dieleman @sedielem.bsky.social · 02/12/2024
Better VQ-VAEs with this one weird rotation trick! I missed this when it came out, but I love papers like this: a simple change to an already powerful technique, that significantly improves results without introducing complexity or hyperparameters.
18613
Reposted by Joan Serrà
Michael Tschannen @mtschannen.bsky.social · 02/12/2024
Have you ever wondered how to train an autoregressive generative transformer on text and raw pixels, without a pretrained visual tokenizer (e.g. VQ-VAE)? We have been pondering this during summer and developed a new model: JetFormer 🌊🤖 arxiv.org/abs/2411.19722 A thread 👇 1/
415437
Reposted by Joan Serrà
Ana Marasović @anamarasovic.bsky.social · 02/12/2024
I learned about this paper (arxiv.org/abs/2406.09413) when Alexei gave this wonderful talk at the U. They trained 60K diffusion models, each for a different person's visual identity. Sampling weights from this set creates a model for a novel identity.
arxiv.org
Interpreting the Weight Space of Customized Diffusion Models
We investigate the space of weights spanned by a large collection of customized diffusion models. We populate this space by creating a dataset of over 60,000 models, each of which is a base model fine...
1206
Reposted by Joan Serrà
Guillaume Dalle @gdalle.bsky.social · 30/11/2024
A paper a day, episode 15. You liked the matrix cookbook? You’re gonna love this one. 100 statistics inequalities just for your personal enjoyment. As they say in French, moi j’ai Bienaymé cet article ! arxiv.org/abs/2102.07234
arxiv.org
One Hundred Probability and Statistics Inequalities
Herein we present one hundred inequalities culled from various corners of the probability, statistics, and combinatorics literature. We welcome new suggestions.
55910
Reposted by Joan Serrà
Gabriel Peyré @gabrielpeyre.bsky.social · 30/11/2024
I wrote a summary of the main ingredients of the neat proof by Hugo Lavenant that diffusion models do not generally define optimal transport. github.com/mathematical...
523845
Reposted by Joan Serrà
Arthur Douillard @douillard.bsky.social · 29/11/2024
Excellent explanation of RoPE embedding, from scratch with all the math needed: fleetwood.dev/posts/you-could-have-… And with beautiful 3blue1brown's style of animation: github.com/3b1b/manim. Original RoPE paper: arxiv.org/abs/2104.09864
05310
Reposted by Joan Serrà
Deep Learning Barcelona @dlbcnai.bsky.social · 30/11/2024
We are now also on Instagram. We will use this channel to explain what deep learning is to our local community. Follow us at @dlbcn.ai
032
Reposted by Joan Serrà
animal prattle @animal-prattle.bsky.social · 29/11/2024
iNatSounds: new dataset from folks @inaturalist.bsky.social & co-authors; looks to be one of the largest public datasets of animal sounds openreview.net/forum?id=QCY... github.com/visipedia/in... #prattle 💬 #bioacoustics
Examples from dataset, a world map surrounded by spectrograms showing animal sounds from different regions of the worldScatter plot where points are sound data sets, x axis is number of categories in dataset and y axis is duration of dataset in hours

iNatSounds is shown as the largest dataset on both axes
13014
Joan Serrà @serrjoa.bsky.social · 28/11/2024
Very cool/curious finding!
030
Reposted by Joan Serrà
Sander Dieleman @sedielem.bsky.social · 27/11/2024
Amazing blog post on flow matching, stunning visuals! It also makes the connection with normalising flows crystal clear. Incredible effort!
29016
Reposted by Joan Serrà
kyunghyuncho.bsky.social @kyunghyuncho.bsky.social · 26/11/2024
".. these findings are insights that people in this field have intuitively understood without the need for experiments .." reject!
2161
Joan Serrà @serrjoa.bsky.social · 27/11/2024
Scholar-inbox is quite good IMHO.
121
Reposted by Joan Serrà
Sai Prasanna @saiprasanna.in · 26/11/2024
Arxiv sharing reminder pdf ❌ abs ✅
925141
Reposted by Joan Serrà
François Fleuret @francois.fleuret.org · 26/11/2024
My deep learning course at the University of Geneva is available on-line. 1000+ slides, ~20h of screen-casts. Full of examples in PyTorch. fleuret.org/dlc/ And my "Little Book of Deep Learning" is available as a phone-formatted pdf (nearing 700k downloads!) fleuret.org/lbdl/
461253247
Reposted by Joan Serrà
Faro Stöter @faroit.bsky.social · 25/11/2024
Audio bandwidth extension (why do people call this super-resolution?) is making quite some progress! demo: aeromamba-super-resolution.github.io code: github.com/aeromamba-su...
0101
Joan Serrà @serrjoa.bsky.social · 25/11/2024
Hello everyone!
0120