Sign in

Tom Aarsen

@tomaarsen.com
2.7K followers 245 following 680 posts

Sentence Transformers, SetFit & NLTK maintainer Machine Learning Engineer at 🤗 Hugging Face

PostsRepliesMedia
Tom Aarsen @tomaarsen.com · 17/09/2026
The original author of the SPLADE line (THE sparse embedding architecture imo), just released the biggest advancement of Sparse embedding models since SPLADE-v3! - 150M, beats every sparse model up to 1B at search - Reaches sub-ms latency using Seismic (!) Go @linkupplatform.bsky.social !
150
Tom Aarsen @tomaarsen.com · 07/09/2026
🔎 H Company just released NeoMME: 260M & 800M multilingual encoders for text and images. You can load them in Sentence Transformers to search document pages with MultiVectorEncoder, charts & tables included, without an OCR step. Thread 🧵
191
Tom Aarsen @tomaarsen.com · 04/09/2026
🤗 I've just released SetFit v1.2.0! SetFit trains text classifiers from a handful of labeled examples per class by fine-tuning a Sentence Transformer, no prompts or LLMs needed. v1.2 brings support for transformers v5, Sentence Transformers v6 & huggingface_hub v1. Thread 🧵
191
Tom Aarsen @tomaarsen.com · 03/09/2026
🤗💚I'm very excited to continue our open source and open weight journey together with NVIDIA! See the announcement here: blogs.nvidia.com/blog/nvidia-...
blogs.nvidia.com
NVIDIA to Acquire Hugging Face
NVIDIA has agreed to acquire Hugging Face. Together, we will scale Hugging Face’s platform, strengthen its infrastructure and expand access to AI for developers and institutions worldwide.
0263
Tom Aarsen @tomaarsen.com · 02/09/2026
🤗 Tencent recently released WeMM-Embedding: universal multimodal embedding models that embed text, images, videos, and visual documents into one space. Three models on Qwen3.5, all Apache 2.0: - WeMM-Embedding-2B - WeMM-Embedding-4B - WeMM-Embedding-9B 🧵
1173
Tom Aarsen @tomaarsen.com · 31/08/2026
🔧 Sentence Transformers v6.0.1 is out: a patch release fixing multi-vector checkpoints that carry both a [Q]/[D] prefix and a prompt. This only affected you if you used lightonai/ColBERT-Zero via Sentence Transformers, no other models affected AFAIK 🧵
111
Tom Aarsen @tomaarsen.com · 27/08/2026
This is by far the cutest robot I've ever seen. Under $400, with a camera, speaker, LiDAR, NFC, Bluetooth, Wifi, etc. + you can train it yourself with reinforcement learning. Plus it has rollerskates. It's so precious 🦆
2181
Tom Aarsen @tomaarsen.com · 26/08/2026
📈 New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers. As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find. Thread 🧵
2214
Tom Aarsen @tomaarsen.com · 18/08/2026
🚨I've just released Sentence Transformers v6.0! MultiVectorEncoder joins the family: ColBERT-style late interaction models are now a first-class model type, for training, inference & interpretation, alongside dense, sparse & reranker models. Big thread 🧵
1246
Tom Aarsen @tomaarsen.com · 15/08/2026
Big fan of these models, great work to the team at @lightonai.bsky.social
020
Tom Aarsen @tomaarsen.com · 06/08/2026
🤖 I've just released Sentence Transformers v5.7.0! A correctness & performance-focused release: gradient-cached losses got an overhaul (silently wrong gradients fixed, up to 3.9x faster training) & model.compile() now actually speeds up inference. + a long list of fixes. 🧵
1102
Tom Aarsen @tomaarsen.com · 04/08/2026
🤗 Ling-3.0-flash is now also open weighted! - MoE with 124B total & 5.1B active params - Hybrid linear attention for higher efficiency with longer sequences Overall very competitive for its size! Nice work to the Ant Group who worked on this.
1152
Tom Aarsen @tomaarsen.com · 03/08/2026
We now have an MTEB account that posts about new embedding & related models! Give them a follow to stay up-to-date on the latest releases.
0123
Reposted by Tom Aarsen
Daniel van Strien @danielvanstrien.bsky.social · 31/07/2026
A film catalogue tells you what a film is about, not what happens inside it. So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments. Search "typing on a computer keyboard", land on the second it happens.
2144
Tom Aarsen @tomaarsen.com · 23/07/2026
🔧 Sentence Transformers v5.6.1 is out: a patch release fixing silently degraded flash attention embeddings for XLM-R & RoBERTa models. If you encode with flash_attention_2 on transformers v5, you'll want this one 🧵
181
Tom Aarsen @tomaarsen.com · 17/07/2026
🎉 @lightonai.bsky.social just published LightOn-rerank: rerankers that score text passages or document page images against a query. Six models: Qwen3.5 at 0.8B / 2B / 4B, each in a pointwise and a generative listwise variant. Excellent for text <-> image retrieval. 🧵
171
Tom Aarsen @tomaarsen.com · 16/07/2026
I'm very excited to share that NVIDIA just released Nemotron-3-Embed: two multilingual embedding models for retrieval The 8B takes the top spot on RTEB. - Nemotron-3-Embed-1B-BF16 (2048-dim) - Nemotron-3-Embed-8B-BF16 (4096-dim) Both OpenMDW-1.1, ready for commercial use. 🧵
1326
Tom Aarsen @tomaarsen.com · 14/07/2026
Tencent just published R3-Skill, a two-stage retrieval stack purpose-built for a problem RAG-style retrievers weren't designed for: routing LLM agent skills (think Anthropic's SKILLmd format). Two 0.6B models, both Apache 2.0, one embedding model, and one reranker. 🧵
1111
Tom Aarsen @tomaarsen.com · 08/07/2026
NAVER, the original authors of SPLADE, just published V-SPLADE, an inference-free sparse retriever for visual document retrieval. Two models on the same backbone (ModernVBERT, 250M params), both Apache 2.0: - naver/v-splade-quality - naver/v-splade-efficient 🧵
3122
Tom Aarsen @tomaarsen.com · 18/06/2026
💧 Liquid AI released 2 multilingual retrieval models, the first bidirectional members of the LFM family. Both 350M params, 11 languages (ar, de, en, es, fr, it, ja, ko, no, pt, sv): - LFM2.5-Embedding-350M (bi-encoder) - LFM2.5-ColBERT-350M (multi-vector, late interaction) 🧵
161
Tom Aarsen @tomaarsen.com · 16/06/2026
🐛I've just released Sentence Transformers v5.6.0! A correctness- & robustness-focused release, headlined by a fix for a silent scoring bug in causal-LM rerankers (think Qwen3-Reranker), plus a batch of hard-negative mining & loss-correctness fixes. Thread with highlights 🧵
1263
Tom Aarsen @tomaarsen.com · 12/06/2026
The core Massive Text Embedding Benchmark (MTEB) team has heavily updated their leaderboard Space. It's so, so much faster, and there's much more information about the best embedding models (and related) to glean from it. Details in 🧵
260
Tom Aarsen @tomaarsen.com · 11/06/2026
LAION just released VoiceCLAP-Large-v2, a contrastive voice-text embedding model that's essentially CLIP for voice and emotion. 9B params via a rank-16 LoRA on top of LCO-Embedding-Omni-7B (itself based on Qwen2.5-Omni thinker). Apache 2.0. 🧵
182
Tom Aarsen @tomaarsen.com · 19/05/2026
🤗 Announcing the Ettin Reranker family: six new CrossEncoder rerankers from 17M to 1B parameters, state-of-the-art at their respective sizes. Built on the Ettin ModernBERT encoders, with the full training recipe and ~143M-triple training dataset as well. 🧵
1111
Tom Aarsen @tomaarsen.com · 13/05/2026
Fastino Labs just released GLiGuard, an open-source safety moderation model that remembers encoders are king for these kinds of tasks. One model, Apache 2.0: gliguard-LLMGuardrails-300M: 300M params, evaluates multiple safety tasks at a time. 🧵
163
Tom Aarsen @tomaarsen.com · 12/05/2026
🤖 I've just released Sentence Transformers v5.5.0! It's headlined by a new `train-sentence-transformers` Agent Skill: let your AI coding agent train & finetune embedding, reranker & sparse encoder models for you. Plus more, here's a thread with highlights 🧵
2275
Tom Aarsen @tomaarsen.com · 06/05/2026
The excellent zerank-2 reranker model by @zeroentropyai.bsky.social is now fully compatible with Sentence Transformers, no `trust_remote_code=True` needed. It's 4B and cc-by-nc-4.0, and performs very well. I'm quite fond of their training methodology, I'll explain in the 🧵
1111
Tom Aarsen @tomaarsen.com · 01/05/2026
IBM just released the R2 generation of their Granite multilingual embedding models for retrieval, and the jump over R1 is very notable. Two models, both Apache 2.0: - granite-embedding-97m-multilingual-r2 (384-dim) - granite-embedding-311m-multilingual-r2 (768-dim) 🧵
182
Tom Aarsen @tomaarsen.com · 28/04/2026
Dun Zhang just released Prism-Reranker: a new open-source reranker family for improving (agentic) search on Qwen3.5, in four sizes from 0.8B to 9B. MIT licensed. Drop-in CrossEncoder for scores only. Trained on a realistic agent-style data mix. 🧵
170
Tom Aarsen @tomaarsen.com · 24/04/2026
BidirLM-Omni-2.5B-Embedding is live: a single bidirectional encoder that embeds text, images, and audio into the same space! Three modalities, all in one 2048-dim space. 🧵
1141
Tom Aarsen @tomaarsen.com · 18/04/2026
The @flowercomputer.com team brought Static Embedding models to Rust and carried out optimizations to heavily speed up potential inference. They'll be open sourcing the deployment pipeline soon. This looks very nice at first glance! Definitely give the thread a read.
1151
Tom Aarsen @tomaarsen.com · 16/04/2026
📈 New blog post: Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers. As a practical example, I finetuned Qwen3-VL-Embedding-2B for Visual Document Retrieval (matching text queries to document screenshots). Thread with highlights 🧵
191
Tom Aarsen @tomaarsen.com · 15/04/2026
This is what finetuning is all about: hyper-specific use cases tackled using a custom finetuned model to blow the generic baselines out of the water! Nice work to Aresh Tajvar on this San Diego Municipal Code retrieval model.
280
Tom Aarsen @tomaarsen.com · 14/04/2026
🔧 Sentence Transformers v5.4.1 is out, a small patch release patching support for numpy string arrays & improving the safety of activation function loading. Two fixes below 🧵
150
Tom Aarsen @tomaarsen.com · 10/04/2026
🧩 To celebrate yesterday's Sentence Transformers v5.4 release, I went back to update SpanMarker: my Named Entity Recognition project. It's still a solid, extremely efficient option for NER. Here's how it works and what's new 🧵
2130
Tom Aarsen @tomaarsen.com · 09/04/2026
🌐 I've just released Sentence Transformers v5.4: we're going fully multimodal for embeddings & reranking! Also featuring a modular CrossEncoder, and automatic Flash Attention 2 input flattening. Highlights in 🧵
1194
Tom Aarsen @tomaarsen.com · 12/03/2026
⬆️ I've just released Sentence Transformers v5.3.0! This release upgrades training with MultipleNegativesRankingLoss with alternative InfoNCE formulations and hardness weighting, adds two new losses, and more. Details in 🧵
140
Tom Aarsen @tomaarsen.com · 27/02/2026
🤗 Perplexity has released 4 open-weights state-of-the-art multilingual embedding models designed for retrieval tasks! pplx-embed-v1 and pplx-embed-context-v1 Specifically trained for int8 and binary embeddings, they'll be viable for massive search problems. Details in 🧵
1191
Tom Aarsen @tomaarsen.com · 23/02/2026
🚀 LightOn is back with a SOTA late-interaction model for search: ColBERT-Zero! By performing contrastive pre-training directly in the multi-vector setting, it outperforms GTE-ModernColBERT etc. on BEIR, using only public data and reaching 55.43 nDCG@10. Details in 🧵
1110
Tom Aarsen @tomaarsen.com · 20/02/2026
ggml / llama.cpp are joining @hf.co, ensuring it'll stay open, maintained, and up to date for a long long time! 🚀 huggingface.co/blog/ggml-jo...
190
Tom Aarsen @tomaarsen.com · 19/02/2026
👏 Jina AI is back with new state-of-the-art multilingual embedding models for retrieval & more: jina-embedding-v5-text! 2 efficient sizes, 239M & 677M, they outperform Qwen3-embedding, EmbeddingGemma-300m, multilingual-e5-large, etc. Details in 🧵
160
Reposted by Tom Aarsen
Alvaro Bartolome @alvarobartt.com · 17/02/2026
More embedding models and an even more reliable inference engine is what you get with @hf.co Text Embeddings Inference v1.9.0 💥 More in the thread 🧵
143
Tom Aarsen @tomaarsen.com · 17/02/2026
I've just pushed a v5.2.3 update for Sentence Transformers that introduces support training with Transformers v5.2 (apologies for the confusion around the project versions, they just happen to be close now 😆). Updating is only useful if you're training. Details in 🧵
121
Tom Aarsen @tomaarsen.com · 12/02/2026
The folks from @lightonai.bsky.social have released extremely strong & efficient Late Interaction models for code search: LateOn-Code(-edge) Alongside a new tool: ColGrep, to use it with your coding agents straight away, locally & cheap Models, dataset, and training code released 🧵
151
Tom Aarsen @tomaarsen.com · 28/01/2026
A few days back, following a long line of excellent proprietary models, VoyageAI by @mongodb.bsky.social released their 🚨 first ever open-weights embedding model for retrieval 🚨! It's called voyage-4-nano, it's multilingual, and very efficient. Details in 🧵
120
Tom Aarsen @tomaarsen.com · 26/01/2026
🤝 I just released Sentence Transformers v5.2.1, introducing official support for the brand-new Transformers v5.x alongside v4.x. Both versions will be supported for the foreseeable future! Details in 🧵
1201
Tom Aarsen @tomaarsen.com · 26/01/2026
There were very few search rerankers specifically for e-commerce queries. Until now, that is. RexRerankers by Walmart Tech MLEs is a suite of 5 SOTA rerankers ranging from 17M to 0.6B parameters. Details in 🧵:
1101
Tom Aarsen @tomaarsen.com · 22/01/2026
Sentence Transformers 🤝 @unsloth.ai We've collaborated with the fine folks at Unsloth to make your embedding model finetuning ~2x faster and require ~20% less VRAM! The Unsloth team prepared 6 notebooks showing how you can take advantage of it! 🧵
180
Tom Aarsen @tomaarsen.com · 06/01/2026
🏎️ You can perform 200ms search over 40 million texts using just a CPU server, 8GB of RAM, and 45GB of disk space. The trick: Binary search with int8 rescoring. I'll show you a demo & how it works in the 🧵:
1191
Tom Aarsen @tomaarsen.com · 11/12/2025
🔥I've just published Sentence Transformers v5.2.0! It introduces multi-processing for CrossEncoder (rerankers), multilingual NanoBEIR evaluators, similarity score outputs in mine_hard_negatives, Transformers v5 support and more. Details in 🧵
251