Sign in

Philipp Schmid

@philschmid.bsky.social
2.9K followers 323 following 75 posts

Tech Lead and LLMs at @huggingface 👨🏻‍💻 🤗 AWS ML Hero 🦸🏻 | Cloud & ML enthusiast | 📍Nuremberg | 🇩🇪 philschmid.de

PostsRepliesMedia
Philipp Schmid @philschmid.bsky.social · 17/12/2024
Code and methods open source in a new library ,“learn and search” Blog: huggingface.co/spaces/Huggi... Learn and Search Repo: github.com/huggingface/...
huggingface.co
Scaling test-time compute - a Hugging Face Space by HuggingFaceH4
Discover amazing ML apps made by the community
1101
Philipp Schmid @philschmid.bsky.social · 17/12/2024
- Introduce DVTS, a new method of performance on larger compute budgets by maintaining solution diversity - Using compute-optimal scaling, a Llama 3 3B outperforms 70B (22x larger) on mathematical reasoning tasks
160
Philipp Schmid @philschmid.bsky.social · 17/12/2024
- Process Reward Models (PRMs) played a crucial role in the search process by evaluating intermediate solution steps - Different search strategies work better for different problem difficulties - beam search for harder problems, Best-of-N for simpler ones
160
Philipp Schmid @philschmid.bsky.social · 17/12/2024
- Test-time compute scaling offers an alternative to training larger models by allowing smaller models to "think longer" - Explored Best-of-N sampling, beam search, and Diverse Verifier Tree Search (DVTS) - Llama 3 1B achieved 55% accuracy on the MATH benchmark using optimal search strategies
110
Philipp Schmid @philschmid.bsky.social · 17/12/2024
By scaling test-time compute, smaller models can match or even surpass the performance of larger models. Llama 3.2 3B can outperform Llama 3.1 70B on MATH-500!🤯
121
Philipp Schmid @philschmid.bsky.social · 17/12/2024
How we implemented test-time computing for open models to solve complex math problems like OpenAI o1. 👀 Test-time compute methods use dynamic inference strategies to have LLMs “think longer” on harder problems, e.g. difficult math problems.
2213
Philipp Schmid @philschmid.bsky.social · 10/12/2024
- 🛠️ Cuts down costs to ~2.29% and time to ~2.36% of human evaluation - 💰 Costs $30 vs $1,297 for human evaluation - ⚡ Reduced time to 118.43 minutes vs 86.5 hours - 🧑‍⚖️ LLM achieved a 60-70% alignment rate to humans - 🥇 Agent achieved a 90% alignment rate to humans huggingface.co/datasets/DEV...
huggingface.co
DEVAI-benchmark/DEVAI · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
020
Philipp Schmid @philschmid.bsky.social · 10/12/2024
The Agent-as-a-Judge is a graph-based agent with tools to locate, read, retrieve, and evaluate files and information for a code project to evaluate the results of other agents by comparing its judgments to human evaluations (alignment rate, judge shift). Github: github.com/metauto-ai/a...
120
Philipp Schmid @philschmid.bsky.social · 10/12/2024
What is better than an LLM as a Judge? Right, an Agent as a Judge! Meta created an Agent-as-a-Judge to evaluate code agents to enable intermediate feedback alongside DevAI a new benchmark of 55 realistic development tasks. Paper: huggingface.co/papers/2410....
huggingface.co
Paper page - Agent-as-a-Judge: Evaluate Agents with Agents
Join the discussion on this paper page
2271
Philipp Schmid @philschmid.bsky.social · 09/12/2024
Sora UI: sora.com Kudos to OpenAI for shipping this! The UI/UX looks really thorough! 🚢
sora.com
Sora
Transform text and images into immersive videos. Animate stories, visualize ideas, and bring your concepts to life.
011
Philipp Schmid @philschmid.bsky.social · 09/12/2024
OpenAI trained a new Turbo model to make it easier and faster to use. With "storyboards" users get a CapCut/Tiktok/Reel-like text-to-video editor, that can be used to edit and create new short-form content! Social media will be flooded.🌊
100
Philipp Schmid @philschmid.bsky.social · 09/12/2024
A big day for AI and sad day for the EU. OpenAI releases Sora, their text-to-video model, with a dedicated UI Studio! Sora will be free for all ChatGPT Pro and Plus subscribers without additional cost. Sora will be available to later today, except if you live in the EU or UK. 🤯
251
Philipp Schmid @philschmid.bsky.social · 28/11/2024
Blog: qwenlm.github.io/blog/qwq-32b... Model: huggingface.co/Qwen/QwQ-32B... Demo: huggingface.co/spaces/Qwen/...
qwenlm.github.io
QwQ: Reflect Deeply on the Boundaries of the Unknown
GITHUB HUGGING FACE MODELSCOPE DEMO DISCORD Note: This is the pronunciation of QwQ: /kwju:/ , similar to the word “quill”. What does it mean to think, to question, to understand? These are the deep wa...
120
Philipp Schmid @philschmid.bsky.social · 28/11/2024
- ⚠️ notable limitations including language mixing, recursive reasoning loops, and safety considerations - 😍 Released under Apache 2.0 on Hugging Face - 👀 Full “reasoning” (CoT) available in the demo
110
Philipp Schmid @philschmid.bsky.social · 28/11/2024
- 👨‍🔬 QwQ-32B-Preview is an experimental research - 🔧 32.5B parameters and 32,768 context length - 📊 65.2% on GPQA, 50.0% on AIME, 90.6% on MATH-500, and 50.0% on LiveCodeBench
100
Philipp Schmid @philschmid.bsky.social · 28/11/2024
First open-weights for OpenAI-o1-like reasoning model! QwQ from the Qwen team is a 32B model that beats OpenAI O1 mini and competes w/ O1 preview and is available under Apache 2.0 on Hugging Face! 🤯
2402
Philipp Schmid @philschmid.bsky.social · 26/11/2024
Models: huggingface.co/HuggingFaceT... Blog: huggingface.co/blog/smolvlm
huggingface.co
HuggingFaceTB/SmolVLM-Instruct · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
010
Philipp Schmid @philschmid.bsky.social · 26/11/2024
🎥 Surprising video capabilities with 27.14% on CinePile 🔓 Released under Apache 2.0 on @huggingface.bsky.social 📱 Can run efficiently on laptops and edge devices
110
Philipp Schmid @philschmid.bsky.social · 26/11/2024
🚀 Smallest SOTA vision language model at only 2B parameters 🛠️ Released 3 variants with Base, Synthetic, and Instruct 💾 Requires only 5GB GPU RAM and achieves 38.8% on MMMU, 81.6% on DocVQA ⚡ 3.3-4.5x faster prefill and 7.5-16x faster generation vs Qwen2-VL
110
Philipp Schmid @philschmid.bsky.social · 26/11/2024
SmolLM can now see! 👀 Meet SmolVLM - a tiny 2B but powerful vision language model that runs on your device! Built on top of SmolLM and released under Apache 2.0. 🚀
3415
Philipp Schmid @philschmid.bsky.social · 26/11/2024
Blog: neuralmagic.com/blog/24-spar... Pruning is not a new technique, but it was much harder to achieve good results and maintain performance across tasks compared to quantization. Let's see if Neural Magic can change that.
neuralmagic.com
2:4 Sparse Llama: Smaller Models for Efficient GPU Inference
Discover Sparse Llama: A 50% pruned, GPU-optimized Llama 3.1 model with 2:4 sparsity, enabling faster, cost-effective inference without sacrificing accuracy.
100
Philipp Schmid @philschmid.bsky.social · 26/11/2024
- 📈 Full recovery on fine-tuning tasks (GSM8K, Evol-CodeAlpaca, Ultrachat-200K) - ⚡ 1.4-2.1x better multi-query throughput - 🌱 Pruned using 13B tokens training, 26 hours on 32 H100s - 🔧 Optimized for NVIDIA Ampere GPUs and newer
100
Philipp Schmid @philschmid.bsky.social · 26/11/2024
- 🔄 98.4% original accuracy on on Open LLM Leaderboard v1 with 50% less parameters using 2:4 sparsity pattern - 🚀 30% higher throughput and 1.8x lower latency with up to 5.0x when combined with quantization - 💻 Works with 4-bit quantization (GPTQ) and Sparse-Marlin kernels
100
Philipp Schmid @philschmid.bsky.social · 26/11/2024
How far can we push LLM optimizations? Turns out, pretty far! A new study achieves 98% accuracy recovery on key benchmarks while removing 50% of Llama 3.1 8B's parameters using pruning. Pruning strategically to remove unnecessary connections in a neural network to make it smaller and faster. 👀
1201
Philipp Schmid @philschmid.bsky.social · 25/11/2024
TIL: @huggingface.bsky.social Transformers has native Tensor Parallelism support for better inference on multiple GPUs! This will enable many benefits and optimizations in the future.🚀 For now, it supports Llama. Which one would you want to see next?
3232
Philipp Schmid @philschmid.bsky.social · 25/11/2024
Created a visual for how function calling works. Wdyt? 🤔
6242
Philipp Schmid @philschmid.bsky.social · 25/11/2024
Blog: blog.dottxt.co/say-what-you... No-structured outputs can actually improve LLM performance when implemented correctly.
010
Philipp Schmid @philschmid.bsky.social · 25/11/2024
🎯 JSON generation reached 77% accuracy vs the paper's reported <10% 🔮 Examples in prompts should match the exact format expected in the actual tasks 🧰 Structured generation works best when implemented as "running our response parser as a generator"
110
Philipp Schmid @philschmid.bsky.social · 25/11/2024
🛠️ Key success criteria is to align your prompt, parser, and generator - it's not just about using JSON mode 📌 JSON generation requires careful prompt design, including specifying the desired schema. 📝 Good prompts should mimic information for human to understand the task and expected response format
120
Philipp Schmid @philschmid.bsky.social · 25/11/2024
📈 The "Let Me Speak Freely" poor results came from weak prompts and wrong use of structured prompting 📊 Structured outputs outperform unstructured on the test GSM8K: 0.78 vs 0.77, Last Letter: 0.77 vs 0.73, Shuffle Object: 0.44 vs 0.41
120
Philipp Schmid @philschmid.bsky.social · 25/11/2024
Does Structured Outputs hurt LLM performance? 🤔 The paper "Let Me Speak Freely" paper claimed that it does, but new experiments by @dottxtai.bsky.social (team behind outlines) show it doesn’t if you do it correctly! 👀
3213
Philipp Schmid @philschmid.bsky.social · 24/11/2024
Paper: allenai.org/papers/tulu-...
allenai.org
010
Philipp Schmid @philschmid.bsky.social · 24/11/2024
- ⚙️ DPO performs best with batch size 32, while SFT with 128 - 💻 PPO-based training is 14x more compute-intensive than DPO but effective for verifiable tasks - ✅ Binary Rewards (True/False) yield better results than Reward Models on verifiable tasks.
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
- 📈 Scaling up unique prompts in preference datasets leads to better performance - 🔄 Combining on-policy and off-policy preference data yields best results
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
- 🔍 N-gram matching for decontamination yielded the most useful results - 🌍 Ablations and evaluations show the importance of real-world diverse prompts - 📝 Ran experiments with different chat template formats - 🤖 Used LLM-as-a-judge to generate synthetic preference data
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
Insights: - 📦 Released models, all of the data, training recipes, code, infrastructure, and evaluation framework. - 📊 Good evaluation is a must to find the best data mix, first tries will not be the best. - ⚖️ Used Length-normalized DPO for Preference Tuning
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
4️⃣ Reinforcement Learning on Verifiable Rewards (RLVR): RL (PPO) based optimization on specific skills like math and instruction following verifiable rewards, .e.g. math
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
3️⃣ Preference Tuning (DPO): Optimize models using Direct Preference Optimization with a mix of on-policy (SFT vs. other models) and off-policy data.
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
2️⃣ Supervised Finetuning (SFT): Train initial models on a carefully curated mix, iteratively refining the mix and aggressively decontaminating against evaluation datasets.
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
1️⃣ Data Curation: Collect a diverse mix of public datasets and synthetic data using persona-driven methods, focusing on core skills like reasoning, coding, and safety.
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
Training Pipeline: 0️⃣ Goals and Evaluation: identify skills to improve (e.g., reasoning, math, coding, safety, precise instruction following, knowledge recall, etc.) and build an evaluation framework
100
Philipp Schmid @philschmid.bsky.social · 24/11/2024
What is the latest in open-source post-training? Allen AI released Tülu last week, which includes models, all of the data, training recipes, code, infrastructure, and evaluation framework. Here are my insights! 👀
2141
Philipp Schmid @philschmid.bsky.social · 23/11/2024
Will post summaries over the next couple of weeks of the insights i found! Open Source Post-Training feels like early 2023 again! 🚀 Lets keep building together! 🤗
010
Philipp Schmid @philschmid.bsky.social · 23/11/2024
Orca Instruct: huggingface.co/datasets/mic...
huggingface.co
microsoft/orca-agentinstruct-1M-v1 · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
140
Philipp Schmid @philschmid.bsky.social · 23/11/2024
SmolLMv2: github.com/huggingface/...
github.com
GitHub - huggingface/smollm
Contribute to huggingface/smollm development by creating an account on GitHub.
110
Philipp Schmid @philschmid.bsky.social · 23/11/2024
Tülu 3: allenai.org/blog/tulu-3-...
allenai.org
Tülu 3: The next era in open post-training | Ai2
A technical deep-dive into Tülu 3, with the model "recipe", data, and more.
100
Philipp Schmid @philschmid.bsky.social · 23/11/2024
Open Coder: opencoder-llm.github.io
opencoder-llm.github.io
OpenCoder: Top-Tier Open Code Large Language Models
OpenCoder: The Open Cookbook For Top-Tier Code Large Language Models
100
Philipp Schmid @philschmid.bsky.social · 23/11/2024
Open Source Post Training is going strong! In last 2 weeks, we got data or recipes released for OpenCoder, SmolLM-2, Orca Agent Instruct, and Tülu 3. Read it, learn, and iterate:
1345
Philipp Schmid @philschmid.bsky.social · 22/11/2024
- 📚 Auxiliary columns, prefixed with a '+', allow storage of unindexed, SELECT-only metadata without requiring separate joins. - 🔜 Improve Quantization support with `float16`, `float8`, "smarter" binary quantization alexgarcia.xyz/blog/2024/sq...
alexgarcia.xyz
sqlite-vec now supports metadata columns and filtering
Metadata, partition key, and auxiliary column support in sqlite-vec
010
Philipp Schmid @philschmid.bsky.social · 22/11/2024
- 💡 Store metadata like user_id or created_at fields directly within vec0 virtual tables. - 🔍 Metadata columns can be used in WHERE clauses of KNN queries for filtering results based on non-vector data. - 🛠️ Introducing Partition Keys to shard the vector index and speed up queries
110