Sign in

Tom Aarsen

@tomaarsen.com
2.7K followers 245 following 680 posts

Sentence Transformers, SetFit & NLTK maintainer Machine Learning Engineer at 🤗 Hugging Face

PostsRepliesMedia
Tom Aarsen @tomaarsen.com · 17/09/2026
Some more general info on sparse embedding models here: sbert.net/docs/sparse_... All of these snippets will work with the new excellent Linkup-sparseup-embed-v1 as it's Sentence Transformers-compatible. Great work to Thibault Formal and @linkupplatform.bsky.social, a pleasure to colab!
sbert.net
Usage — Sentence Transformers documentation
000
Tom Aarsen @tomaarsen.com · 17/09/2026
This sparsity makes them cheap to store (only store active dims), but also interpretable: if dim 7429 is active, then vocab token ID 7429 is active. In short: your embeddings are human-understandable. They're often trained for retrieval, and strong indexes make them super fast.
100
Tom Aarsen @tomaarsen.com · 17/09/2026
For those who haven't heard of sparse embedding models: They're (search) models that embed inputs (often text) into embeddings with huge dimensionalities that are often equal to the vocabulary size. The vast majority (99%) of these positions are 0, i.e. only a few 'active dims'.
100
Tom Aarsen @tomaarsen.com · 17/09/2026
Very glad to see Linkup publishing open weight models! Much respect, this one looks excellent. Go read the blog/model card for more details! Link: huggingface.co/Linkup-Platf...
huggingface.co
Linkup-Platform/linkup-sparseup-embed-v1 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
110
Tom Aarsen @tomaarsen.com · 17/09/2026
The original author of the SPLADE line (THE sparse embedding architecture imo), just released the biggest advancement of Sparse embedding models since SPLADE-v3! - 150M, beats every sparse model up to 1B at search - Reaches sub-ms latency using Seismic (!) Go @linkupplatform.bsky.social !
150
Tom Aarsen @tomaarsen.com · 07/09/2026
Credit to Tony Wu & Aurélien Lac and H Company for the models. Both sizes are Apache 2.0. Their post covers the architecture, training & retrieval results: huggingface.co/blog/Hcompan... Models: huggingface.co/collections/...
huggingface.co
NeoMME: an efficient Multimodal-native and Multilingual Encoder
A Blog post by H company on Hugging Face
010
Tom Aarsen @tomaarsen.com · 07/09/2026
Page images produce a lot of patch vectors. In ST, HierarchicalTokenPooling can cluster a document's vectors & keep the cluster means. pool_factor=2 keeps roughly half the document vectors. Queries stay untouched. Measure the quality tradeoff on your own data.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
The collection has a lot of names. For retrieval, choose: - -Retriever: both heads in Transformers - -Retriever-ST-late: MultiVectorEncoder - -Retriever-ST-dense: SentenceTransformer The ST versions expose one head each, for inference & independent fine-tuning.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
The dense & late-interaction heads were trained jointly. In Transformers, NeoMMEForRetrieval returns both in one forward pass. That lets you retrieve candidates with a dense index, then rerank with late interaction using the saved document vectors.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
Want one vector per document? Load the -ST-dense checkpoint with SentenceTransformer. The 260M model supports 128, 256, 512 or 1,024 dimensions via Matryoshka training. The 800M adds 1,792. Set truncate_dim=256 for smaller vectors. Text & page images both work.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
The same checkpoint also retrieves text documents, including across languages. Same encode_query, encode_document & similarity calls. NeoMME uses MeanMaxSim: each query token takes its best document match, then those scores are averaged. Although you should use a fast index.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
H Company reports 0.523 nDCG@10 on ViDoRe v3 for the 260M Retriever, against 0.524 for ColQwen2.5-v0.2 at 3.75B parameters. Roughly 14x fewer parameters for almost the same score on this benchmark. Their 800M model reaches 0.556.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
Text tokens & raw image patches go through one shared bidirectional Transformer, trained from scratch. No separate pretrained vision tower or causal decoder. Both sizes support 16,384 tokens of context, i.e. it's an encoder built for both modalities from the start.
110
Tom Aarsen @tomaarsen.com · 07/09/2026
Or look at the models themselves: huggingface.co/collections/...
huggingface.co
NeoMME - a Hcompany Collection
Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual Encoders
110
Tom Aarsen @tomaarsen.com · 07/09/2026
H Company's blog walks through the architecture, training & retrieval results. Read it here, or point your Agent at the URL: huggingface.co/blog/Hcompan...
huggingface.co
NeoMME: an efficient Multimodal-native and Multilingual Encoder
A Blog post by H company on Hugging Face
110
Tom Aarsen @tomaarsen.com · 07/09/2026
🔎 H Company just released NeoMME: 260M & 800M multilingual encoders for text and images. You can load them in Sentence Transformers to search document pages with MultiVectorEncoder, charts & tables included, without an OCR step. Thread 🧵
191
Tom Aarsen @tomaarsen.com · 04/09/2026
Full release notes: github.com/huggingface/... Docs: huggingface.co/docs/setfit pip install setfit==1.2.0
huggingface.co
SetFit · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
010
Tom Aarsen @tomaarsen.com · 04/09/2026
SetFit has been around since 2022 & is still one of the cheapest ways to get a solid text classifier: no prompts, small models that train in minutes, any embedding model from the Hub as the backbone (the intro snippet uses a ModernBERT one) & ONNX/OpenVINO export for deployment.
110
Tom Aarsen @tomaarsen.com · 04/09/2026
Also in v1.2.0: - Python 3.13 support - saving before training no longer fails with codecarbon 3.x - evaluate>=0.4.6 required, older ones can't load metrics with huggingface_hub v1 - docs & notebooks use namespaced dataset ids (stanfordnlp/sst2), bare ids no longer resolve
110
Tom Aarsen @tomaarsen.com · 04/09/2026
Exporting is fixed too. ONNX export works again with torch 2.9+ (the legacy exporter is requested explicitly, as the new dynamo default lacks the needed opsets) & with skl2onnx 1.20, with the scikit-learn head converted in float32. OpenVINO export works with openvino 2026.
110
Tom Aarsen @tomaarsen.com · 04/09/2026
For training with transformers v5, warmup_proportion still works (now fractional warmup_steps), logging_dir is ignored with a warning (set TENSORBOARD_LOGGING_DIR instead) & Trackio & W&B callbacks work out of the box. report_to is now honoured too: "none" really means none.
110
Tom Aarsen @tomaarsen.com · 04/09/2026
Now tested across the range: transformers 4.41 to 5.16, Sentence Transformers 3.0 to 6.0, huggingface_hub 0.24 to 1.30, datasets 2.15 to 5.0 & Python 3.9 to 3.13. The CI runs the newest stack, the oldest supported pins & the last transformers v4 generation as separate presets.
110
Tom Aarsen @tomaarsen.com · 04/09/2026
The ecosystem moved & SetFit had fallen behind. It no longer imported with transformers v5, couldn't load models with huggingface_hub v1, warned on every import with Sentence Transformers v5.4+ & crashed on datasets v4 columns. All fixed, while old versions keep working.
110
Tom Aarsen @tomaarsen.com · 04/09/2026
How it works: SetFit turns a few labeled sentences into positive & negative pairs, fine-tunes a Sentence Transformer contrastively on them & fits a classifier on the embeddings. 8 examples per class rivals RoBERTa-Large fine-tuned on 3k examples. Paper: arxiv.org/abs/2209.11055
arxiv.org
Efficient Few-Shot Learning Without Prompts
Recent few-shot methods, such as parameter-efficient fine-tuning (PEFT) and pattern exploiting training (PET), have achieved impressive results in label-scarce settings. However, they are difficult to...
110
Tom Aarsen @tomaarsen.com · 04/09/2026
🤗 I've just released SetFit v1.2.0! SetFit trains text classifiers from a handful of labeled examples per class by fine-tuning a Sentence Transformer, no prompts or LLMs needed. v1.2 brings support for transformers v5, Sentence Transformers v6 & huggingface_hub v1. Thread 🧵
191
Tom Aarsen @tomaarsen.com · 03/09/2026
🤗💚I'm very excited to continue our open source and open weight journey together with NVIDIA! See the announcement here: blogs.nvidia.com/blog/nvidia-...
blogs.nvidia.com
NVIDIA to Acquire Hugging Face
NVIDIA has agreed to acquire Hugging Face. Together, we will scale Hugging Face’s platform, strengthen its infrastructure and expand access to AI for developers and institutions worldwide.
0263
Tom Aarsen @tomaarsen.com · 02/09/2026
Also: Matryoshka truncation for smaller/faster embeddings, and interleaved inputs (multiple images or videos in one input) Strong release for retrieval over mixed media, especially with video in the corpus. Collection: huggingface.co/collections/... Great work to the Tencent team.
huggingface.co
WeMM-Embedding - a tencent Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
000
Tom Aarsen @tomaarsen.com · 02/09/2026
The ST support: model = SentenceTransformer("tencent/WeMM-Embedding-9B", trust_remote_code=True) model.encode_query(["How is mapo tofu prepared?"]) model.encode_document([ "Mapo tofu is a Sichuan dish...", {"video": "mapo_tofu.mp4", "text": "Represent this video."}, ])
100
Tom Aarsen @tomaarsen.com · 02/09/2026
On the broader MMEB-v3 (190 tasks: text, agent, cross-modal retrieval, and more): - WeMM-Embedding-9B: 59.5 (top of the board) - WeMM-Embedding-4B: 58.2 - WeMM-Embedding-2B: 56.0 The 2B already beats every 7B and 8B model listed.
100
Tom Aarsen @tomaarsen.com · 02/09/2026
Video is where the gap is widest. MMEB-v2 video subset: - WeMM-Embedding-9B: 74.3 - WeMM-Embedding-2B: 70.8 - Qwen3-VL-Embedding-8B: 67.1 The 2B beats an 8B on video by 3.7 points.
100
Tom Aarsen @tomaarsen.com · 02/09/2026
On MMEB-v2 (78 datasets), avg score: 8B/9B class: - WeMM-Embedding-9B: 80.6 - DME-Medium (closed): 78.4 - Qwen3-VL-Embedding-8B: 77.8 2B class: - WeMM-Embedding-2B: 77.9 - Qwen3-VL-Embedding-2B: 73.2 The 2B edges out Qwen3-VL-Embedding-8B at a quarter of the size.
100
Tom Aarsen @tomaarsen.com · 02/09/2026
If you just want a direct link: huggingface.co/collections/... Otherwise, see 🧵
huggingface.co
WeMM-Embedding - a tencent Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
100
Tom Aarsen @tomaarsen.com · 02/09/2026
🤗 Tencent recently released WeMM-Embedding: universal multimodal embedding models that embed text, images, videos, and visual documents into one space. Three models on Qwen3.5, all Apache 2.0: - WeMM-Embedding-2B - WeMM-Embedding-4B - WeMM-Embedding-9B 🧵
1173
Tom Aarsen @tomaarsen.com · 31/08/2026
Full release notes: github.com/huggingface/... pip install sentence-transformers==6.0.1
github.com
Release v6.0.1 - Restore the PyLate prefix on prompted checkpoints, 80 documented multi-vector models · huggingface/sentence-transformers
This patch release fixes a multi-vector loading bug: PyLate checkpoints that carry both a [Q]/[D] prefix and a text prompt lost the prefix, so they were encoded without a marker they were trained w...
010
Tom Aarsen @tomaarsen.com · 31/08/2026
Router.preprocess now also forwards `task` to the routed module, so per-task length caps and query expansion apply behind a Router. No released checkpoint combines the two, so this one is preventive. And the pretrained multi-vector tables grew from 51 checkpoints to 80.
110
Tom Aarsen @tomaarsen.com · 31/08/2026
Also, because of the new `requirements` configuration option, you'll get a nice and clear error if you're using Sentence Transformers version lower than v6.0.1. Any model author can add these requirements for any Python package/version, to make sure users get correct outputs.
110
Tom Aarsen @tomaarsen.com · 31/08/2026
Three checkpoints are affected, all ColBERT-Zero: - lightonai/ColBERT-Zero - lightonai/ColBERT-Zero-supervised - lightonai/ColBERT-Zero-unsupervised Restoring the prefix takes ColBERT-Zero from 0.6569 to 0.6824 NanoBEIR nDCG@10. Indexed with one of these? Re-encode.
110
Tom Aarsen @tomaarsen.com · 31/08/2026
PyLate lets a checkpoint carry a prefix and prompt text. We assumed those were mutually exclusive, dropping the prefix whenever a prompt was present. A model trained on `[CLS] [Q] search_query: ...` was encoded as `[CLS] search_query: ...`, missing a marker it was trained with.
110
Tom Aarsen @tomaarsen.com · 31/08/2026
🔧 Sentence Transformers v6.0.1 is out: a patch release fixing multi-vector checkpoints that carry both a [Q]/[D] prefix and a prompt. This only affected you if you used lightonai/ColBERT-Zero via Sentence Transformers, no other models affected AFAIK 🧵
111
Tom Aarsen @tomaarsen.com · 27/08/2026
Link with details: pollen-robotics.com/microduck/
pollen-robotics.com
Microduck - A tiny biped robot you can teach new tricks | Pollen Robotics
Microduck is a 25 cm biped robot with 15 motors, a camera, LiDAR and a grasping beak. Playable out of the box, and its open-source stack lets you train new behaviours in simulation and run them on the...
030
Tom Aarsen @tomaarsen.com · 27/08/2026
This is by far the cutest robot I've ever seen. Under $400, with a camera, speaker, LiDAR, NFC, Bluetooth, Wifi, etc. + you can train it yourself with reinforcement learning. Plus it has rollerskates. It's so precious 🦆
2181
Tom Aarsen @tomaarsen.com · 27/08/2026
Lean into the hatred, it's okay
010
Tom Aarsen @tomaarsen.com · 26/08/2026
Full post: huggingface.co/blog/train-m... pip install -U "sentence-transformers[train]" If you finetune your own multi-vector model, I'd love to see it!
huggingface.co
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
000
Tom Aarsen @tomaarsen.com · 26/08/2026
Thanks to @lateinteraction.bsky.social for the computed metrics and the discussions. Also, as an extra treat, I've integrated more models with the MultiVectorEncoder, see sbert.net/docs/multi_v... for the list, or just filter by `sentence-transformers` and `multi-vector` on @huggingface .
sbert.net
Pretrained Models — Sentence Transformers documentation
100
Tom Aarsen @tomaarsen.com · 26/08/2026
The model is on the Hub, and the blog post walks through every component: model, dataset, loss, training arguments, evaluator & trainer. huggingface.co/multi-vector...
huggingface.co
multi-vector-encoder/mLateOn-medical · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
100
Tom Aarsen @tomaarsen.com · 26/08/2026
Don't have 14 hours? My scaling experiments put 100k pairs, 75 minutes of training, within 1.2 NDCG@10 points of the full million-pair run. Most of the gain are from the first hour. P.s. I used an RTX 3090, not some 8xH100 cluster, this should be pretty accessible.
100
Tom Aarsen @tomaarsen.com · 26/08/2026
Late interaction is also much stronger than dense models here. Qwen3-Embedding-4B is the strongest dense model on my data, with ~33x the active parameters of mine, and still stops 13 points short. Also, the 8B version scores lower than the 4B.
100
Tom Aarsen @tomaarsen.com · 26/08/2026
And does late interaction actually help? At matched training & matched backbones, yes, a lot. DenseOn & LateOn differ only in whether they pool. The late-interaction sibling wins by +12 points. The multilingual pair replicates it at +13.
100
Tom Aarsen @tomaarsen.com · 26/08/2026
🚨 Importantly, most released checkpoints cap documents at 180-512 tokens, because MS MARCO-style training data rarely goes further. On my 941-token passages, that cap costs up to 24 NDCG@10 points. More than any difference between architectures.
100
Tom Aarsen @tomaarsen.com · 26/08/2026
Those checkpoints sit after contrastive pretraining but before supervised finetuning on general retrieval. They carry the late-interaction structure without the general-purpose tuning that domain training then has to undo. I replicated it across both LateOn and mLateOn.
110