Sign in

Tom Aarsen

@tomaarsen.com
2.7K followers 245 following 694 posts

Sentence Transformers, SetFit & NLTK maintainer Machine Learning Engineer at 🤗 Hugging Face

PostsRepliesMedia
Tom Aarsen @tomaarsen.com · 13h
🚨 Use BF16 or FP32, not FP16. The model's activation range exceeds FP16's dynamic range. The card warns of NaNs or silently degraded embeddings. BF16 is recommended where natively supported. Use FP32 elsewhere, including most CPUs.
100
Tom Aarsen @tomaarsen.com · 13h
All modalities share an 8,192-token context window. Default costs: - image: 280 tokens - video frame: 140 tokens - audio: 25 tokens/second That's about 327 seconds of audio with no accompanying text. Mixed inputs draw from the same budget.
110
Tom Aarsen @tomaarsen.com · 13h
You can embed a whole product listing: description, two photos & a demo video, all in one vector. Place <|image|>, <|video|> or <|audio|> markers inside the text to position each media item. Then compare that combined embedding with a text-only search query.
100
Tom Aarsen @tomaarsen.com · 13h
Text embeddings use task prefixes: SearchQuery, QuestionAnswering, CodeRetrieval, Classification, Clustering & more. For retrieval, pair the query prompt with Document for corpus text. Document adds "title: none". Format real titles manually, and no prefix for media.
100
Tom Aarsen @tomaarsen.com · 13h
128d gives 6x smaller vectors, but the quality tradeoff is notable: MMEB v2 overall drops from 59.01 at 768d to 45.65 at 128d. The card positions 128d for text-only workloads. For mixed modalities, the quality impact is much smaller at 256d.
100
Tom Aarsen @tomaarsen.com · 13h
Plus, with Matryoshka training it supports 768, 512, 256 & 128 dimensions. At 256d, vectors take a third of the storage. MTEB multilingual v2 goes from 61.36 to 60.41. In ST: model.encode(..., truncate_dim=256, normalize_embeddings=True), and re-normalize after truncating.
100
Tom Aarsen @tomaarsen.com · 13h
Google's multimodal results at 768d, full precision: - MMEB v2 visual documents: 67.84 NDCG@5 - MMEB v2 video: 50.67 Hit@1 - MSEB audio retrieval: 69.54 MRR@10 The card also reports image & broader audio benchmarks.
100
Tom Aarsen @tomaarsen.com · 13h
Google reports support for 100+ languages, with MTEB multilingual v2 at 61.36 vs 61.15 for EmbeddingGemma 1. The larger gain is code: MTEB code v1 goes from 68.76 to 78.68, roughly a 14% relative improvement. All scores are from the model card, at 768d.
100
Tom Aarsen @tomaarsen.com · 13h
The 740M footprint is modular: - text: 270M - vision: 170M - audio: 300M Disable unused encoders through config_kwargs. Text-only loads 270M, text + images 440M, text + audio 570M. Designed to also work on phones & laptops by loading only the encoders that you need.
100
Tom Aarsen @tomaarsen.com · 13h
A text query can retrieve photos, audio recordings or video clips. Each becomes a dense vector in the same space. In Sentence Transformers, it's the familiar model.encode() & model.similarity() API. The model card includes a short text-search quickstart.
110
Tom Aarsen @tomaarsen.com · 13h
🤗 Google Deepmind is back with EmbeddingGemma 2, which maps text (including code), images, video & audio into one shared 768-dimensional embedding space. 740M total parameters, 100+ languages, Apache 2.0 & Sentence Transformers support. Thread with highlights 🧵
1175
Tom Aarsen @tomaarsen.com · 17/09/2026
For those who haven't heard of sparse embedding models: They're (search) models that embed inputs (often text) into embeddings with huge dimensionalities that are often equal to the vocabulary size. The vast majority (99%) of these positions are 0, i.e. only a few 'active dims'.
100
Tom Aarsen @tomaarsen.com · 17/09/2026
The original author of the SPLADE line (THE sparse embedding architecture imo), just released the biggest advancement of Sparse embedding models since SPLADE-v3! - 150M, beats every sparse model up to 1B at search - Reaches sub-ms latency using Seismic (!) Go @linkupplatform.bsky.social !
150
Tom Aarsen @tomaarsen.com · 07/09/2026
Page images produce a lot of patch vectors. In ST, HierarchicalTokenPooling can cluster a document's vectors & keep the cluster means. pool_factor=2 keeps roughly half the document vectors. Queries stay untouched. Measure the quality tradeoff on your own data.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
The dense & late-interaction heads were trained jointly. In Transformers, NeoMMEForRetrieval returns both in one forward pass. That lets you retrieve candidates with a dense index, then rerank with late interaction using the saved document vectors.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
Want one vector per document? Load the -ST-dense checkpoint with SentenceTransformer. The 260M model supports 128, 256, 512 or 1,024 dimensions via Matryoshka training. The 800M adds 1,792. Set truncate_dim=256 for smaller vectors. Text & page images both work.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
The same checkpoint also retrieves text documents, including across languages. Same encode_query, encode_document & similarity calls. NeoMME uses MeanMaxSim: each query token takes its best document match, then those scores are averaged. Although you should use a fast index.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
H Company reports 0.523 nDCG@10 on ViDoRe v3 for the 260M Retriever, against 0.524 for ColQwen2.5-v0.2 at 3.75B parameters. Roughly 14x fewer parameters for almost the same score on this benchmark. Their 800M model reaches 0.556.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
Text tokens & raw image patches go through one shared bidirectional Transformer, trained from scratch. No separate pretrained vision tower or causal decoder. Both sizes support 16,384 tokens of context, i.e. it's an encoder built for both modalities from the start.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
🔎 H Company just released NeoMME: 260M & 800M multilingual encoders for text and images. You can load them in Sentence Transformers to search document pages with MultiVectorEncoder, charts & tables included, without an OCR step. Thread 🧵
191
Tom Aarsen @tomaarsen.com · 04/09/2026
🤗 I've just released SetFit v1.2.0! SetFit trains text classifiers from a handful of labeled examples per class by fine-tuning a Sentence Transformer, no prompts or LLMs needed. v1.2 brings support for transformers v5, Sentence Transformers v6 & huggingface_hub v1. Thread 🧵
191
Tom Aarsen @tomaarsen.com · 02/09/2026
The ST support: model = SentenceTransformer("tencent/WeMM-Embedding-9B", trust_remote_code=True) model.encode_query(["How is mapo tofu prepared?"]) model.encode_document([ "Mapo tofu is a Sichuan dish...", {"video": "mapo_tofu.mp4", "text": "Represent this video."}, ])
100
Tom Aarsen @tomaarsen.com · 02/09/2026
On the broader MMEB-v3 (190 tasks: text, agent, cross-modal retrieval, and more): - WeMM-Embedding-9B: 59.5 (top of the board) - WeMM-Embedding-4B: 58.2 - WeMM-Embedding-2B: 56.0 The 2B already beats every 7B and 8B model listed.
100
Tom Aarsen @tomaarsen.com · 02/09/2026
On MMEB-v2 (78 datasets), avg score: 8B/9B class: - WeMM-Embedding-9B: 80.6 - DME-Medium (closed): 78.4 - Qwen3-VL-Embedding-8B: 77.8 2B class: - WeMM-Embedding-2B: 77.9 - Qwen3-VL-Embedding-2B: 73.2 The 2B edges out Qwen3-VL-Embedding-8B at a quarter of the size.
100
Tom Aarsen @tomaarsen.com · 02/09/2026
🤗 Tencent recently released WeMM-Embedding: universal multimodal embedding models that embed text, images, videos, and visual documents into one space. Three models on Qwen3.5, all Apache 2.0: - WeMM-Embedding-2B - WeMM-Embedding-4B - WeMM-Embedding-9B 🧵
1173
Tom Aarsen @tomaarsen.com · 31/08/2026
Also, because of the new `requirements` configuration option, you'll get a nice and clear error if you're using Sentence Transformers version lower than v6.0.1. Any model author can add these requirements for any Python package/version, to make sure users get correct outputs.
110
Tom Aarsen @tomaarsen.com · 31/08/2026
🔧 Sentence Transformers v6.0.1 is out: a patch release fixing multi-vector checkpoints that carry both a [Q]/[D] prefix and a prompt. This only affected you if you used lightonai/ColBERT-Zero via Sentence Transformers, no other models affected AFAIK 🧵
111
Tom Aarsen @tomaarsen.com · 27/08/2026
This is by far the cutest robot I've ever seen. Under $400, with a camera, speaker, LiDAR, NFC, Bluetooth, Wifi, etc. + you can train it yourself with reinforcement learning. Plus it has rollerskates. It's so precious 🦆
2181
Tom Aarsen @tomaarsen.com · 26/08/2026
Tool two, HierarchicalTokenPooling: cluster each document's token vectors and keep the cluster means. This is roughly what Omar did. Halving the index costs 0.33 NDCG@10 points and leaves rank-1 accuracy untouched. A quarter of the index still scores 89.9. This stacks, too!
110
Tom Aarsen @tomaarsen.com · 26/08/2026
The fair objection to multi-vector is index size. One vector per token, and these passages are long: 878 vectors each, ~45 GB for 200k passages as raw fp16. But: you should never use these models at fp16. There's two important tools:
110
Tom Aarsen @tomaarsen.com · 26/08/2026
This screenshot gives you a feel of the full recipe, with a pre-supervised checkpoint, 1M domain pairs, in-batch negatives, full document length, and a higher-than-usual learning rate. The blogpost shows each component in more detail separately.
110
Tom Aarsen @tomaarsen.com · 26/08/2026
I evaluated it against 47 model configurations: dense, sparse, lexical & multi-vector, on 1,000 held-out medical questions searching 200,000 passages. 91.4 NDCG@10, ahead of the best zero-shot model of any architecture by +6.2 points.
120
Tom Aarsen @tomaarsen.com · 26/08/2026
📈 New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers. As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find. Thread 🧵
2214
Tom Aarsen @tomaarsen.com · 18/08/2026
Not multi-vector, but the fix I'd most want you to know about 🚨 CrossEncoder.predict now upcasts logits to float32 before the activation. A sigmoid in bfloat16 saturates & ties the top candidates together, randomizing their order. 0.1849 -> 0.6795 NanoBEIR nDCG@10.
110
Tom Aarsen @tomaarsen.com · 18/08/2026
Training works like the other model types: 4 new losses (incl. cached & distillation variants), 5 new evaluators, and a Trainer that takes the same arguments you already know. From a bare ModernBERT-base: 0.1338 -> 0.4831 NanoBEIR mean nDCG@10 in ~25 min on one RTX 3090.
120
Tom Aarsen @tomaarsen.com · 18/08/2026
Because MaxSim is a sum of per-query-token maxima, a ranking decomposes exactly: every point of a score belongs to one query token & one document token. The new interpretability module renders that as the standard ColPali heatmap, aggregated or one map per query token.
110
Tom Aarsen @tomaarsen.com · 18/08/2026
Page images aren't the only non-text modality. ColQwen-Omni takes text, images, audio & video. Retrieving a recorded conversation is the same two calls. Zero-shot, and no transcription step anywhere: the query says "nausea" where the audio says "carsickness".
120
Tom Aarsen @tomaarsen.com · 18/08/2026
Late interaction is the state of the art for visual document retrieval: text queries against page images, charts & tables intact, no OCR step. ColPali-family checkpoints run through the exact same two calls. MaxSim scores query text tokens against image patches.
120
Tom Aarsen @tomaarsen.com · 18/08/2026
HierarchicalTokenPooling (Clavié, Chaffin & Adams) clusters each document's token vectors with Ward linkage & keeps ~1/pool_factor of them. pool_factor=2 halves the index at 100.6% of unpooled BEIR performance. Apply it per call, standalone, or bake it into the model.
120
Tom Aarsen @tomaarsen.com · 18/08/2026
Does it actually help? LightOn trained LateOn (multi-vector) & DenseOn (dense) on the same data, same 149M ModernBERT backbone, differing only in whether they pool. Multi-vector wins 9 of 13 NanoBEIR datasets: 0.6868 vs 0.6764 mean nDCG@10. Same gap on full BEIR.
120
Tom Aarsen @tomaarsen.com · 18/08/2026
Every checkpoint format loads through the same class: PyLate, Stanford-NLP ColBERT (via the HF_ColBERT marker + artifact.metadata), ColPali-style VLMs, or a bare backbone with a fresh projection. Prefixes, query expansion & the punctuation skiplist come from the saved config.
110
Tom Aarsen @tomaarsen.com · 18/08/2026
A dense model compresses a whole text into one vector. A multi-vector model keeps one vector per token & scores query against document with MaxSim: for each query token, take its best match in the document, then sum. Nothing has to be averaged away.
130
Tom Aarsen @tomaarsen.com · 18/08/2026
🚨I've just released Sentence Transformers v6.0! MultiVectorEncoder joins the family: ColBERT-style late interaction models are now a first-class model type, for training, inference & interpretation, alongside dense, sparse & reranker models. Big thread 🧵
1246
Tom Aarsen @tomaarsen.com · 06/08/2026
model.compile() was silently a no-op for inference: encode() & predict() bypassed the compiled forward. Now it applies. With CUDA graphs (mode="reduce-overhead") I measured ~3x faster batch-size-1 inference on bge-small-en-v1.5 in bf16. Plus new torch.compile docs.
110
Tom Aarsen @tomaarsen.com · 06/08/2026
🤖 I've just released Sentence Transformers v5.7.0! A correctness & performance-focused release: gradient-cached losses got an overhaul (silently wrong gradients fixed, up to 3.9x faster training) & model.compile() now actually speeds up inference. + a long list of fixes. 🧵
1102
Tom Aarsen @tomaarsen.com · 04/08/2026
🤗 Ling-3.0-flash is now also open weighted! - MoE with 124B total & 5.1B active params - Hybrid linear attention for higher efficiency with longer sequences Overall very competitive for its size! Nice work to the Ant Group who worked on this.
1152
Tom Aarsen @tomaarsen.com · 23/07/2026
🔧 Sentence Transformers v5.6.1 is out: a patch release fixing silently degraded flash attention embeddings for XLM-R & RoBERTa models. If you encode with flash_attention_2 on transformers v5, you'll want this one 🧵
181
Tom Aarsen @tomaarsen.com · 17/07/2026
Two more things worth noting: 213K training groups, single epoch. Competitors trained on millions of samples or distilled from much larger teachers. Real headroom left in this recipe. Training data is English-only, but French holds up on ViDoRe V3 and beats EN on three domains.
100
Tom Aarsen @tomaarsen.com · 17/07/2026
The blog is very honest about what didn't work, which is the best part: - Tournament scheduling: -27 NDCG - First-token readout: capped at (low) 1.6x gain - Reading a listwise model pointwise: throws away most of the lift - For 2B, the ViT encoder was the slowest, not the decode
100
Tom Aarsen @tomaarsen.com · 17/07/2026
Sentence Transformers doesn't support listwise rerankers that process several docs at once (yet!). The pointwise variants drop in today and hold up well. ViDoRe V3, pointwise 2B: 59.87 - jina-reranker-m0: 59.40 - Qwen3-VL-Reranker-2B: 59.18 And ~3.2x faster than 2B listwise.
100