Philipp Schmid @philschmid.bsky.social · 17/12/2024How we implemented test-time computing for open models to solve complex math problems like OpenAI o1. 👀 Test-time compute methods use dynamic inference strategies to have LLMs “think longer” on harder problems, e.g. difficult math problems. 2213
Philipp Schmid @philschmid.bsky.social · 10/12/2024What is better than an LLM as a Judge? Right, an Agent as a Judge! Meta created an Agent-as-a-Judge to evaluate code agents to enable intermediate feedback alongside DevAI a new benchmark of 55 realistic development tasks. Paper: huggingface.co/papers/2410....huggingface.coPaper page - Agent-as-a-Judge: Evaluate Agents with AgentsJoin the discussion on this paper page 2271
Philipp Schmid @philschmid.bsky.social · 09/12/2024A big day for AI and sad day for the EU. OpenAI releases Sora, their text-to-video model, with a dedicated UI Studio! Sora will be free for all ChatGPT Pro and Plus subscribers without additional cost. Sora will be available to later today, except if you live in the EU or UK. 🤯 251
Philipp Schmid @philschmid.bsky.social · 28/11/2024First open-weights for OpenAI-o1-like reasoning model! QwQ from the Qwen team is a 32B model that beats OpenAI O1 mini and competes w/ O1 preview and is available under Apache 2.0 on Hugging Face! 🤯 2402
Philipp Schmid @philschmid.bsky.social · 26/11/2024SmolLM can now see! 👀 Meet SmolVLM - a tiny 2B but powerful vision language model that runs on your device! Built on top of SmolLM and released under Apache 2.0. 🚀 3415
Philipp Schmid @philschmid.bsky.social · 26/11/2024How far can we push LLM optimizations? Turns out, pretty far! A new study achieves 98% accuracy recovery on key benchmarks while removing 50% of Llama 3.1 8B's parameters using pruning. Pruning strategically to remove unnecessary connections in a neural network to make it smaller and faster. 👀 1201
Philipp Schmid @philschmid.bsky.social · 25/11/2024TIL: @huggingface.bsky.social Transformers has native Tensor Parallelism support for better inference on multiple GPUs! This will enable many benefits and optimizations in the future.🚀 For now, it supports Llama. Which one would you want to see next? 3232
Philipp Schmid @philschmid.bsky.social · 25/11/2024Created a visual for how function calling works. Wdyt? 🤔 6242
Philipp Schmid @philschmid.bsky.social · 25/11/2024Does Structured Outputs hurt LLM performance? 🤔 The paper "Let Me Speak Freely" paper claimed that it does, but new experiments by @dottxtai.bsky.social (team behind outlines) show it doesn’t if you do it correctly! 👀 3213
Philipp Schmid @philschmid.bsky.social · 24/11/2024What is the latest in open-source post-training? Allen AI released Tülu last week, which includes models, all of the data, training recipes, code, infrastructure, and evaluation framework. Here are my insights! 👀 2141
Philipp Schmid @philschmid.bsky.social · 23/11/2024Open Source Post Training is going strong! In last 2 weeks, we got data or recipes released for OpenCoder, SmolLM-2, Orca Agent Instruct, and Tülu 3. Read it, learn, and iterate: 1345
Philipp Schmid @philschmid.bsky.social · 22/11/2024# 2024-11-22 SQLite is all you need! Big sqlite-vec update! 🚀 sqlite-vec is a plugin to support Vector Search in SQLite or LibSQL databases. v0.1.6 now allows storing non-vector data in vec0 virtual tables, enabling metadata conditioning and filtering! 🤯 2174
Philipp Schmid @philschmid.bsky.social · 22/11/2024Add your BSKY 🦋 to your @huggingface.bsky.social profile! 2152
Philipp Schmid @philschmid.bsky.social · 22/11/2024New small hybrid model from NVIDIA has been announced! Hymba is a 1.5B hybrid Mamba x Attention Model that outperforms other small LLMs like Meta 3.2 or SmolLM v2 being trained on only 1.5T Tokens. 🤯 1100
Philipp Schmid @philschmid.bsky.social · 21/11/2024🚀 Biggest open text dataset release of the year! SmolTalk: a 1M sample synthetic dataset used to train SmolLM v2 is here! Available under Apache 2.0, it combines newly generated datasets + publicly available ones. Here’s what you need to know 🧵👇 2181
Philipp Schmid @philschmid.bsky.social · 21/11/2024With the preview of deepseek R1 and results equal to OpenAI o1-preview, you might want to take look at "Stream of Search". R1, "thoughts" are streamed, no MCTS is used during inference. They must have baked the "search" and "backtracking" directly into the model. huggingface.co/papers/2404....huggingface.coPaper page - Stream of Search (SoS): Learning to Search in LanguageJoin the discussion on this paper page 051
Philipp Schmid @philschmid.bsky.social · 20/11/2024Mindblowing! 🤯 New reasoning model preview from Deepseek that matches OpenAI o1! 🐳 DeepSeek-R1-Lite-Preview is now live to test! 🧠 > o1-preview-level performance on AIME & MATH benchmarks. > Access to CoT and transparent thought process in real-time. > Open-source models & API coming soon! 4281
Philipp Schmid @philschmid.bsky.social · 20/11/2024Sage Attention the next Flash Attention? 🤔 > 3x speed up over Flash Attention2, maintaining 99% performance > INT4/8 for Q and K matrices, and FP8/16 for P and V + smoothing methods for Q and V > Drop-in replacement of torch scaled_dot_product_attention > SageAttention 2 code to be released soon 2110
Philipp Schmid @philschmid.bsky.social · 19/11/2024The Microsoft Ignite conference starts today. Here is a summary of all the new AI announcements around Azure, OpenAI, Github and more! 👀 260
Philipp Schmid @philschmid.bsky.social · 19/11/2024Hello, my name is Philipp. I am a Technical Lead at @huggingface.bsky.social, leading our partnerships with AWS, Google, Azure, or NVIDIA. 🧑🏻💻 I post about the AI News, Open Models, Interesting AI Paper Summaries, blog posts, and guides! My is blog at www.philschmid.de Make sure to follow! 🤗philschmid.dePhilschmidPersonal Blog of Philipp Schmid Technical Lead and LLM at Hugging Face. Learn how to use the latest AI and Cloud Technologies from fine-tuning LLMs with RLHF to deploying them in production. 1243