Sign in

muchenli.bsky.social

@muchenli.bsky.social
2 followers 2 following 5 posts
PostsRepliesMedia
muchenli.bsky.social @muchenli.bsky.social · 02/02/2025
(3/3) We introduce LiveAoPSBench, an evolving benchmark that circumvents data contamination and offers a clearer view of real model capabilities. Interstingly, We test DeepSeek R1 and QwQ models on LiveAoPSBench-2024 and observe a noticeable performance decline over time.
020
muchenli.bsky.social @muchenli.bsky.social · 02/02/2025
(2/3) We show the effectiveness of our curated instruction tuning set by conducting supervised fine-tuning.
110
muchenli.bsky.social @muchenli.bsky.social · 02/02/2025
(1/3) Pushing forward the edge of large-scale reasoning demands high quality data on truly challenging questions. In this paper, we tap into the rich trove of online math forum discussions—transforming them into over 600k curated QA pairs to supercharge LLM reasoning.
120
muchenli.bsky.social @muchenli.bsky.social · 02/02/2025
Paper: arxiv.org/abs/2501.14275
Code & Data: github.com/DSL-Lab/aops Project Website: livemathbench.github.io Evaluation Leaderboard: livemathbench.github.io/leaderboard.... AlphaXiv(for QA): www.alphaxiv.org/abs/2501.14275
arxiv.org
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
Advances in Large Language Models (LLMs) have sparked interest in their ability to solve Olympiad-level math problems. However, the training and evaluation of these models are constrained by the limit...
120
muchenli.bsky.social @muchenli.bsky.social · 02/02/2025
🚀 Training a Large Reasoning Model, but high-quality data is scarce? Check out our paper: “Leveraging Online Olympiad-Level Math Problems for LLM Training & Contamination-Resistant Evaluation.” 📖 TL;DR: 🔹 LLM-powered high-quality data collection 🤖 🔹 647K Math QA pairs 📊 🔹 A Live Math Benchmark ⏳
arxiv.org
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
Advances in Large Language Models (LLMs) have sparked interest in their ability to solve Olympiad-level math problems. However, the training and evaluation of these models are constrained by the limit...
121