Sign in

merve

@merve.bsky.social
8.6K followers 675 following 242 posts

proud mediterrenean 🧿 open-sourceress at hugging face 🤗 multimodality, zero-shot vision, vision language models, transformers

PostsRepliesMedia
merve @merve.bsky.social · 11/05/2025
llama.cpp has vision language model support now! ❤️‍🔥 get started with sota VLMs (gemma 3, Qwen2.5VL, InternVL3 & more) and serve them wherever you want 🤩 learn more github.com/ggml-org/lla... 📖
2465
merve @merve.bsky.social · 02/05/2025
If you want to ✨ speed-up & harden ✨ your RAG pipelines, use visual document retrieval models ⬇️ We have shipped a how-to guide for VDR models in Hugging Face transformers 🤗📖 huggingface.co/docs/transfo...
3283
merve @merve.bsky.social · 15/04/2025
Why do people sleep on DSE multimodal retrieval models? 👀 They're just like ColPali, but highly scalable, fast and you can even make them more efficient with binarization or matryoshka with little degradation 🪆⚡️ I collected some here huggingface.co/collections/...
1121
merve @merve.bsky.social · 15/04/2025
I'm so hooked on @hf.co Inference Providers (specifically Qwen2.5-VL-72B) for multimodal agentic workflows with smolagents 🥹 get started ⤵️ > filter models provided by different providers > test them through widget or Python/JS/cURL
0102
merve @merve.bsky.social · 14/04/2025
my weekly summary on what's released in open AI is up on @hf.co huggingface.co/posts/merve/... collection is here huggingface.co/collections/...
0181
merve @merve.bsky.social · 14/04/2025
fan-favorite open-source PDF rendering model OlmOCR goes faster and more efficient ⚡️ RolmOCR-7B follows same recipe with OlmOCR, builds on Qwen2.5VL with training set modifications and improves accuracy & performance 🤝 huggingface.co/reducto/Rolm...
0170
merve @merve.bsky.social · 12/04/2025
Hello friends 👋🏼 If visit Turkey this summer, know that millions of Turkish people are doing a boycott, once a week not buying anything and rest of the week only buying necessities if you have plans, here's a post that summarizes where you should buy stuff from www.instagram.com/share/BADrkS...
instagram.com
Login • Instagram
Welcome back to Instagram. Sign in to check out what your friends, family & interests have been capturing & sharing around the world.
0281
Reposted by merve
merve @merve.bsky.social · 09/04/2025
SmolVLM paper is out and it's packed with great findings on training a good smol vision LM! Andi summarized them below, give it a read if you want to see more insights 🤠
0294
merve @merve.bsky.social · 11/04/2025
DO NOT SLEEP ON THIS MODEL Kimi-VL-A3B-Thinking is the first ever capable open-source reasoning VLM with MIT license ❤️ > it has only 2.8B activated params 👏 > it's agentic 🔥 works on GUIs > surpasses gpt-4o I've put it to test (see below ⤵️) huggingface.co/spaces/moons...
1302
merve @merve.bsky.social · 11/04/2025
InternVL3 is out 💥 > 7 ckpts with various sizes (1B to 78B) > Built on InternViT encoder and Qwen2.5VL decoder, improves on Qwen2.5VL > Can do reasoning, document tasks, extending to tool use and agentic capabilities 🤖 > easily use with Hugging Face transformers 🤗 huggingface.co/collections/...
0122
Reposted by merve
Simon Willison @simonwillison.net · 09/04/2025
Model Context Protocol has prompt injection security problems simonwillison.net/2025/Apr/9/m...
simonwillison.net
Model Context Protocol has prompt injection security problems
As more people start hacking around with implementations of MCP (the Model Context Protocol, a new standard for making tools available to LLM-powered systems) the security implications of tools built ...
911721
Reposted by merve
jsulz @jsulz.com · 09/04/2025
Xet infra now backs 1000s of repos on @hf.co , which means we get to put on our researcher hats and peer into the bytes 👀 🤓 Xet clients chunk files (~64KB) and skip uploads of duplicate content, but what if those chunks are already in _another_ repo? We skip those too.
huggingface.co
From Chunks to Blocks: Accelerating Uploads and Downloads on the Hub
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1164
merve @merve.bsky.social · 09/04/2025
SmolVLM paper is out and it's packed with great findings on training a good smol vision LM! Andi summarized them below, give it a read if you want to see more insights 🤠
0294
merve @merve.bsky.social · 06/04/2025
X'in politikaları sebebiyle işimle alakalı post'ları burada da paylaşıyor olacağım, takip edebilirsiniz 😊
1301
merve @merve.bsky.social · 06/03/2025
icymi I shipped a tutorial on fine-tuning vision language models on videos ⏯️ learn how to fine-tune SmolVLM2 on Video Feedback dataset 📖 github.com/merveenoyan/...
github.com
smol-vision/Fine_tune_SmolVLM2_on_Video.ipynb at main · merveenoyan/smol-vision
Recipes for shrinking, optimizing, customizing cutting edge vision models. 💜 - merveenoyan/smol-vision
1323
merve @merve.bsky.social · 26/02/2025
All the multimodal document retrieval models (ColPali, DSE et al) are now under visual document retrieval at @hf.co 📝🤗 take your favorite VDR model out for multimodal RAG 🤝
0190
Reposted by merve
Andi @andimara.bsky.social · 23/01/2025
Smol but mighty: • 256M delivers 80% of the performance of our 2.2B model. • 500M hits 90%. Both beat our SOTA 80B model from 17 months ago! 🎉 Efficiency 🤝 Performance Explore the collection here: huggingface.co/collections/... Blog: huggingface.co/blog/smolervlm
1162
Reposted by merve
Andi @andimara.bsky.social · 23/01/2025
Introducing the smollest VLMs yet! 🤏 SmolVLM (256M & 500M) runs on <1GB GPU memory. Fine-tune it on your laptop and run it on your toaster. 🚀 Even the 256M model outperforms our Idefics 80B (Aug '23). How small can we go? 👀
1487
merve @merve.bsky.social · 17/01/2025
Everything that was released passed week in open AI 🤠 > Link to all models, datasets, demos huggingface.co/collections/... > Text-readable version is here huggingface.co/posts/merve/...
1323
merve @merve.bsky.social · 13/01/2025
there's a new multimodal retrieval model in town 🤠 @llamaindex.bsky.social released vdr-2b-multi-v1 > uses 70% less image tokens, yet outperforming other dse-qwen2 based models > 3x faster inference with less VRAM 💨 > shrinkable with matryoshka 🪆 huggingface.co/collections/...
1462
merve @merve.bsky.social · 10/01/2025
What a week to open the year in open ML, all the things released at @hf.co 🤠 Here's everything released, find text-readable version here huggingface.co/posts/merve/... All models are here huggingface.co/collections/...
0211
merve @merve.bsky.social · 09/01/2025
ViTPose -- best open-source pose estimation model just landed to @hf.co transformers 🕺🏻💃🏻 🔖 Model collection: huggingface.co/collections/... 🔖 Notebook on how to use: colab.research.google.com/drive/1e8fcb... 🔖 Try it here: huggingface.co/spaces/hysts...
1678
merve @merve.bsky.social · 09/01/2025
ByteDance just dropped SA2VA: a new family of vision LMs combining Qwen2VL/InternVL and SAM2 with MIT license 💗 The models are capable of tasks involving vision-language understanding and visual referrals (referring segmentation) both for images and videos ⏯️
3618
merve @merve.bsky.social · 31/12/2024
supercharge your LLM apps with smolagents 🔥 however cool your LLM is, without being agentic it can only go so far enter smolagents: a new agent library by @hf.co to make the LLM write code, do analysis and automate boring stuff! huggingface.co/blog/smolage...
thumbnail that says introducing smolagents
28516
merve @merve.bsky.social · 20/12/2024
ColPali is landed at @hf.co transformers and I have just shipped a very lean fine-tuning tutorial in smol-vision 🤠💗 QLoRA fine-tuning with 4-bit with bsz of 4 can be done with 32 GB VRAM and is very fast! ✨ github.com/merveenoyan/...
screenshot of the top of the tutorial that says "Fine-tune ColPali for Multimodal RAG"
0467
merve @merve.bsky.social · 20/12/2024
you can now stay up-to-date with big AI research labs' updates on @hf.co easily over org activity page 🥹 I have been looking forward to this feature as I felt most back to back releases are overwhelming and I tend to miss out 🤠
0221
merve @merve.bsky.social · 19/12/2024
BERT is so back 🔥 Answer AI and Lighton released ModernBERT: lightning-fast state-of-the-art BERT model with Apache 2.0 license 🥹 2x fast as debertav3 and 3x faster than nomic 💨 all models are here hf.co/collections/answerdotai/modernbert-67627ad707a4acbf33c41deb read more hf.co/blog/modernbert 📖
3736
merve @merve.bsky.social · 18/12/2024
Aya by Cohere For AI can now see! 👀 C4AI community has built Maya 8B, a new open-source multilingual VLM built on SigLIP and Aya 8B 🌱 works on 8 languages! 🗣️ The authors extend Llava dataset using Aya's translation capabilities with 558k examples! works very well ⬇️ huggingface.co/spaces/kkr51...
screenshot of model conversation turns
0304
merve @merve.bsky.social · 13/12/2024
VLMs go MoE ✨ DeepSeek AI dropped three new commercially permissive vision LMs based on SigLIP encoder and their DeepSeek-MoE decoder 🐳 the models come in 1.0B, 2.8B and 4.5B active params 🥹 models seem to catch up with state-of-the-art with less active parameters! huggingface.co/collections/...
1405
merve @merve.bsky.social · 12/12/2024
Learn how to build a complete multimodal RAG pipeline with ColQwen2 as retriever, MonoQwen2-VL as reranker, Qwen2-VL as VLM in this notebook that runs on a GPU as small as L4 🔥 huggingface.co/learn/cookbo...
screenshot of the notebook in the link
0486
merve @merve.bsky.social · 08/12/2024
This week in open-source AI was insane 🤠 A small recap🕺🏻 Text-readable version is here huggingface.co/posts/merve/... Collection to all models, datasets, demos is here huggingface.co/collections/...
2582
Reposted by merve
Linoy Tsaban 🎗️ @linoy.hf.co · 03/12/2024
'tis the season of open-source video models 🎄📹⚡️ Tencent just dropped the weights for ✨HunyuanVideo✨ - 13B parameters - competitive with closed source - code & weights released - demo coming soon 🔥
2295
merve @merve.bsky.social · 02/12/2024
this was SmolVLM fine-tuning, and you can also do this! 🤗 I made a notebook that includes all the goodies: QLoRA, gradient accumulation, gradient checkpointing with explanations on how they work 💝 below snapshot is with bsz=4 with simulated bsz=16 on L4 🤠 github.com/huggingface/...
0552
merve @merve.bsky.social · 02/12/2024
So many open-source and open releases last week! Here's a recap, find the text-readable version here huggingface.co/posts/merve/...
3794
merve @merve.bsky.social · 28/11/2024
it only takes a single CLI command to kick-off a Direct Preference Optimization fine-tuning run on SmolVLM huggingface.co/blog/smolvlm... you're welcome
1794
merve @merve.bsky.social · 28/11/2024
people on this platform will take your words out of context, twist, not mention your correction, because they just want to hate on what you work on, and insult you comfortably. I'll keep posting here about my work but will not be interacting with anyone who wants to bash on my company.
41164
merve @merve.bsky.social · 27/11/2024
muting this conversation 👋🏼 with some people there's not much winning really, nothing actionable I will keep shipping & posting about my projects, thanks everyone for the input 🤗🧡
11121
Reposted by merve
merve @merve.bsky.social · 26/11/2024
Small yet mighty! 💫 We are releasing SmolVLM: a new 2B small vision language made for on-device use, fine-tunable on consumer GPU, immensely memory efficient 🤠 We release three checkpoints under Apache 2.0: SmolVLM-Instruct, SmolVLM-Synthetic and SmolVLM-Base huggingface.co/collections/...
1115927
merve @merve.bsky.social · 27/11/2024
It's pretty sad to see the negative sentiment towards Hugging Face on this platform due to a dataset put by one of the employees. I want to write a small piece. 🧵 Hugging Face empowers everyone to use AI to create value and is against monopolization of AI it's a hosting platform above all.
2945570
merve @merve.bsky.social · 27/11/2024
The authors of ColPali trained a retrieval model based on SmolVLM 🤠 TLDR; - ColSmolVLM performs better than ColPali and DSE-Qwen2 on all English tasks - ColSmolVLM is more memory efficient than ColQwen2 💗 Find the model here huggingface.co/vidore/colsm...
4738
merve @merve.bsky.social · 27/11/2024
I had to debug idefics3/smolVLM fine-tune because the model kept on yapping 🤠 plot twist: the processor and the chat template had a different EOS token compared to what base model had, which somehow sneaked into the code 🥸
2311
merve @merve.bsky.social · 27/11/2024
I completely agree 💯
3480
merve @merve.bsky.social · 27/11/2024
my friends, this platform is likely already scraped in the hands of big corps. perhaps do not take your accumulated anger out on Daniel who has apologized & took the dataset down most of the people are taking all the anger towards big AI on him because he listened to the people, it's unfair
91058
Reposted by merve
Doron Adler @norod78.bsky.social · 27/11/2024
Cool! it can 'enumerate' objects seen in an image and output the result as valid JSON while adhering to the instructions in the prompt. Well done! 👏 Output: {'object_list': ['minion', 'lightsaber', 'power strip', 'cable', 'box', 'robot']} Input prompt in ALT, top-p: 0.2, temp: 0.1
Your task is to analyze an image and output a list of all distinct objects, people, animals, etc. that you can identify in the image. For every image the user sends, please study that image carefully and compile a json-like list of every meaningful object or entity you can discern. Each item should be on a separate line. The list should be thorough in capturing all key elements of the image, but also concise, avoiding redundancy. Minor background details can be omitted. Output your list in the following format: {"object_list":[item 1, item 2 , item 3]} The formatting of this list is important, as it is meant to be easily machine-readable for later processing. Focus only on the literal contents of the image. Do not include any subjective analysis, guesses about the context, or details that aren't clearly visible. Provide only the object list, without any other discussion or response. If no object were detected, just return: {"object_list":[]} Each item must appear only once, never repeat the same item.
162
Reposted by merve
Prince Canuma @prince-canuma.bsky.social · 26/11/2024
Now on MLX bsky.app/profile/prin...
081
merve @merve.bsky.social · 26/11/2024
Small yet mighty! 💫 We are releasing SmolVLM: a new 2B small vision language made for on-device use, fine-tunable on consumer GPU, immensely memory efficient 🤠 We release three checkpoints under Apache 2.0: SmolVLM-Instruct, SmolVLM-Synthetic and SmolVLM-Base huggingface.co/collections/...
1115927
Reposted by merve
Tom Aarsen @tomaarsen.com · 25/11/2024
✨ Jina AI just released Jina-CLIP-v2: A multimodal (images and texts) & multilingual embedding model. Details in 🧵 Model: huggingface.co/jinaai/jina-... 📈 Jina-CLIP-v2 outperforms Jina-CLIP-v1 (by 3% on text-image and text-text tasks) 🧵
27210
merve @merve.bsky.social · 25/11/2024
I looked back and found an instant in time where 7 years ago I was studying for my data science class in south of France (visited Nice for another thing) it's like foreshadowing my life because I still spend my weekends writing code lol
2480
merve @merve.bsky.social · 24/11/2024
something small but might approaches
2420
Reposted by merve
merve @merve.bsky.social · 22/11/2024
Here's a recap of everything released this week in open machine learning ❄️ All models mentioned are in this collection huggingface.co/collections/... A text-readable version of this post is found at huggingface.co/posts/merve/...
31229