Sign in

Alvaro Bartolome

@alvarobartt.com
719 followers 39 following 40 posts

machine learning + tech lead @hf.co (inference + cloud) opinions, code and mistakes are my own github.com/alvarobartt

PostsRepliesMedia
Alvaro Bartolome @alvarobartt.com · 01/07/2026
I wrote a post so you can try GLM 5.2 on Microsoft Foundry from Codex too. alvarobartt.com/goal-glm-5.2...
010
Alvaro Bartolome @alvarobartt.com · 01/07/2026
GLM 5.2, open frontier-scale intelligence on Microsoft Foundry with AMD MI300X. Running a Codex goal with an open model never felt this good!
150
Reposted by Alvaro Bartolome
Tom Aarsen @tomaarsen.com · 09/04/2026
🌐 I've just released Sentence Transformers v5.4: we're going fully multimodal for embeddings & reranking! Also featuring a modular CrossEncoder, and automatic Flash Attention 2 input flattening. Highlights in 🧵
1194
Alvaro Bartolome @alvarobartt.com · 17/02/2026
And much more additions, improvements and fixes, that couldn't have been possible without the community support and contributions 🙏🏻 github.com/huggingface/...
github.com
Release v1.9.0 · huggingface/text-embeddings-inference
What's changed? 🚨 Breaking changes Default HiddenAct::Gelu to GeLU + tanh in favour of GeLU erf by @vrdn-23 in #753 Default GeLU implementation is now GeLU + tanh approximation instead of exact...
010
Alvaro Bartolome @alvarobartt.com · 17/02/2026
- 💚 NVIDIA Blackwell support, ready for next-gen GPUs as B200, GB200, or RTX 50-series.
100
Alvaro Bartolome @alvarobartt.com · 17/02/2026
- 🔄 Add bidirectional attention support for 3, enabling newer embedding models as Voyage AI by MongoDB.
100
Alvaro Bartolome @alvarobartt.com · 17/02/2026
- 🦙 Add support for Meta Llama 2 and 3 architectures with Flash Attention support, enabling embedding models as NVIDIA Llama Embed Nemotron.
100
Alvaro Bartolome @alvarobartt.com · 17/02/2026
- 🎉 Add support for Microsoft Deberta V2 and V3, for both feature-extraction (and sentence-similarity) and text-classification, enabling models as Meta Llama Prompt Guard.
100
Alvaro Bartolome @alvarobartt.com · 17/02/2026
More embedding models and an even more reliable inference engine is what you get with @hf.co Text Embeddings Inference v1.9.0 💥 More in the thread 🧵
143
Alvaro Bartolome @alvarobartt.com · 05/01/2026
github.com/alvarobartt/...
010
Alvaro Bartolome @alvarobartt.com · 05/01/2026
`hf-mem` is all you need to estimate the required VRAM for inference of any model on @huggingface based on Safetensors metadata. - Written in Python - Lightweight, only depends on `httpx` - Runs w/ `uvx` as `uvx hf-mem ...` - Works with any Safetensors repository - Output inspired by usgraphics.com
101
Alvaro Bartolome @alvarobartt.com · 13/03/2025
🧨 I built something with #Zig! `tokeni.zig` is a std-only implementation of the Byte Pair Encoding (BPE) algorithm in Zig for tokenizing sequences of text, used by OpenAI (among many others) to tokenize the text when pretraining their large language models! github.com/alvarobartt/...
020
Alvaro Bartolome @alvarobartt.com · 10/02/2025
alvarobartt.me/how-to-read-and-pars…
alvarobartt.me
How to read and parse JSON with Zig 0.13
How to read and parse JSON with Zig 0.13 by alvarobartt
010
Alvaro Bartolome @alvarobartt.com · 10/02/2025
For anyone interested in Zig I wrote a small post titled "How to read and parse JSON with Zig 0.13" that explains how to read JSON from a file with keys with different value types and how to access those values.
110
Alvaro Bartolome @alvarobartt.com · 03/02/2025
love this quote "working smarter helps, but the real superpower is resting smarter" a highly recommended read!
030
Alvaro Bartolome @alvarobartt.com · 31/01/2025
Right, the point is that on Rust you end up "refactoring" a lot (at least I do), but seems easier to handle, whilst on Zig I don't feel is as easy, not especially complex either, just more cumbersome
120
Alvaro Bartolome @alvarobartt.com · 31/01/2025
🤗 Here's a simple script that calculates the required VRAM for serving DeepSeek R1 from @huggingface Hub safetensor's metadata! P.S. The result of the script above is: "model_id='deepseek-ai/DeepSeek-R1' requires memory=756.716GB"
040
Alvaro Bartolome @alvarobartt.com · 31/01/2025
hmm refactoring in zig is not as easy as it's in rust, even though seems fairly common too, right? or is it just me? 🤔
110
Alvaro Bartolome @alvarobartt.com · 29/01/2025
stuff that matters takes time
010
Reposted by Alvaro Bartolome
Quentin Gallouédec @qgallouedec.hf.co · 25/01/2025
Last moments of closed-source AI 🪦 : Hugging Face is openly reproducing the pipeline of 🐳 DeepSeek-R1. Open data, open training. open models, open collaboration. 🫵 Let's go! github.com/huggingface/...
github.com
GitHub - huggingface/open-r1: Fully open reproduction of DeepSeek-R1
Fully open reproduction of DeepSeek-R1. Contribute to huggingface/open-r1 development by creating an account on GitHub.
0347
Alvaro Bartolome @alvarobartt.com · 23/01/2025
Check DeepSeek-R1 collection on the Hugging Face Hub, with not just DeepSeek-R1 and DeepSeek-R1-Zero, but also distilled their reasoning patterns to fine-tune smaller models! huggingface.co/collections/...
huggingface.co
DeepSeek-R1 - a deepseek-ai Collection
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
010
Alvaro Bartolome @alvarobartt.com · 23/01/2025
🐐 DeepSeek is not on the @hf.co Hub to take part, they are there to take over! Amazing stuff from the DeepSeek team, ICYMI they recently released some reasoning models (DeepSeek-R1 and DeepSeek-R1-Zero), fully open-source, their performance is on par with OpenAI-o1 and it's MIT licensed!
1101
Alvaro Bartolome @alvarobartt.com · 23/01/2025
you can find so much gold in github gists wow, i was not a big fan because the discoverability doesn't seem great, but been exploring gists lately and so much gold stuff in there!
000
Alvaro Bartolome @alvarobartt.com · 22/01/2025
in case anyone missed it, we're running a certified course on ai agents at hugging face starting on feb 2nd; the course is on how to build you own ai agents for different cool use cases built on top of open source! 👇 you can sign up in the link below, don't miss it! bit.ly/hf-learn-age...
bit.ly
Hugging Face
Hugging Face Email Forms
010
Alvaro Bartolome @alvarobartt.com · 22/01/2025
ok, here we go again 😅
000
Alvaro Bartolome @alvarobartt.com · 28/11/2024
because it's my native language, anyway it was just an idea, not sure I'll do it anyway 🤗
000
Alvaro Bartolome @alvarobartt.com · 27/11/2024
Not quite sure yet about how's following me here, but I may consider not just x-posting but also eventually post more random thoughts + content in Spanish, is that something you'd be interested in?
240
Alvaro Bartolome @alvarobartt.com · 20/11/2024
awesome 🤗
010
Alvaro Bartolome @alvarobartt.com · 20/11/2024
how do I get in there? 🤗
100
Alvaro Bartolome @alvarobartt.com · 20/11/2024
here we go again! i work at hugging face and here you can expect posts about machine learning (llms mainly), some rust, some nvim nerdy stuff and anything related to hugging face 🤗 posting is not easy for me, but i’ll try to do better from now on, support is highly appreciated!
2140
Alvaro Bartolome @alvarobartt.com · 19/11/2024
Read more about the Serverless Inference API in the documentation! huggingface.co/docs/api-inference
000
Alvaro Bartolome @alvarobartt.com · 19/11/2024
🔥 Finally, if you are willing to get started quickly and experiment with LLMs feel free to give the recently released Inference Playground a try! huggingface.co/playground
100
Alvaro Bartolome @alvarobartt.com · 19/11/2024
👨‍💻 Alternatively, you can also use the Serverless Inference API programmatically via cURL, the huggingface_hub Python SDK, the openai SDK for chat completion, and much more! Find all the alternatives at huggingface.co/docs/api-inference
100
Alvaro Bartolome @alvarobartt.com · 19/11/2024
🔎 Now let's explore some of the different alternatives to run inference via the Serverless API! The most straightforward one is via the Hugging Face Hub available on the model card of the Serverless API supported models!
100
Alvaro Bartolome @alvarobartt.com · 19/11/2024
🔒 Before going on, you will first need to generate a Hugging Face fine-grained token with access to the Serverless API, as the requests need to be authenticated so keep the token safe and avoid exposing it!
100
Alvaro Bartolome @alvarobartt.com · 19/11/2024
❄️ Additionally, there are a bunch of models (around 1000) that are "Cold", meaning that those are not loaded in the Serverless API, but can be loaded when sending a request to them, also meaning that the first request may take a while until the model is loaded!
100
Alvaro Bartolome @alvarobartt.com · 19/11/2024
You may be wondering how do you know what models out of the over a million publicly available on the Hub can be used via the Serverless API? Well, we have the "Warm" tag that indicates that a model is loaded in the Serverless API and ready to be used 🔥
100
Alvaro Bartolome @alvarobartt.com · 19/11/2024
💡 Did you know that you can use over 13700 public open models and adapters on the @huggingface Hub for FREE? You just need a free account on the Hugging Face Hub (you can also subscribe to PRO to increase the requests per hour) More details on the thread 🧵
220
Alvaro Bartolome @alvarobartt.com · 19/11/2024
oh great, just updated mine, thanks for sharing!
000
Reposted by Alvaro Bartolome
Daniël de Kok @danieldk.eu · 18/11/2024
Bluesky pro-tip: you can set your domain as a handle: bsky.social/about/blog/4... Bonus: you can keep your handle if you ever want to move to another server in the future.
bsky.social
How to set your domain as your handle - Bluesky
Using a domain as your handle helps with account identity, verification, and portability. Here's how to set your domain as your handle.
241
Alvaro Bartolome @alvarobartt.com · 19/11/2024
awesome, still exploring this new realm 🤣
000
Alvaro Bartolome @alvarobartt.com · 19/11/2024
my first visit here was 18 seconds ago :/
000
Alvaro Bartolome @alvarobartt.com · 19/11/2024
where the nvim-nerds at in here?
120