Sign in

alphaXiv

@alphaxiv.org
272 followers 4 following 356 posts

High fidelity research

PostsRepliesMedia
alphaXiv @alphaxiv.org · 21h
Most AI agents force the harness around a decoder-only Transformer But what if we change the model’s in/output shape to fit the agent? This blog suggests using recurrent memory for old history while dense attention handles recent context, reducing manual compaction alphaxiv.org/abs/2609.lan...
050
alphaXiv @alphaxiv.org · 29/09/2026
What if you make on-policy distillation recursive? This paper proposes that after each round, the improved model becomes both the next student and the new gold-conditioned teacher, letting better reasoning feed into future supervision. www.alphaxiv.org/abs/2609.30652
010
alphaXiv @alphaxiv.org · 28/09/2026
What if you could trace a model behavior back to the smallest set of internal components actually responsible for it? Matryoshka Attribution learns one ranked mask over internal components across many sparsity levels at once, letting it isolate compact causal circuits alphaxiv.org/abs/2609.25518
030
alphaXiv @alphaxiv.org · 27/09/2026
If world model has to reconstruct every pixel, it'll pretty much waste most of its capacity modeling irrelevant background noise. This paper found that by removing pixel decoding, it performs much better with moving distractors and natural-video backgrounds. alphaxiv.org/abs/2609.22175
041
alphaXiv @alphaxiv.org · 26/09/2026
Can a model learn entirely through self-play and no real data? This self-play setup can generate its own pre-training data from scratch and still teach a model general predictive structure that transfers to real-world data. www.alphaxiv.org/abs/2609.30063
050
alphaXiv @alphaxiv.org · 25/09/2026
This paper shows you can just use JEV for every evaluation instead of expensive LLM. JEV basically acts as a cheap first-pass judge, returning both a verdict and how confident it is. When confidence is high, keep the answer. When it’s low, escalate to a stronger LLM. alphaxiv.org/abs/2609.26550
061
alphaXiv @alphaxiv.org · 24/09/2026
Xiaomi dropped their new attention mechanism for their upcoming V3 model. At 1M context, it cuts prefill FLOPs by 2.92x vs HySparse and shrinks the KV cache from 6.72GB to 2.69GB, improving long-context retrieval by a lot. Read more about their new architecture here: alphaxiv.org/abs/2609.26368
030
alphaXiv @alphaxiv.org · 23/09/2026
VLA models are too slow for reactive robot control, so this paper lets the VLA generate actions in the background, while a lightweight RL policy uses the latest observation to edit and select actions in real time This gives from 42% to 97% with just 10min of online data alphaxiv.org/abs/2609.18207
060
alphaXiv @alphaxiv.org · 22/09/2026
Xiaomi MiMo just completed their largest RL scaling run so far, openly. With the run costing $2.6M for Pro and $0.9M for Flash, with ~44% spent on rollouts, 41-44% on training, and 14% on grading. Read more about it: www.alphaxiv.org/pdf/2609.mim...
030
alphaXiv @alphaxiv.org · 21/09/2026
This paper proposes one world-modeling framework that works across multiple fields Through Orthogonal Predictive Factorization, it splits a single JEPA latent state into complementary factors that predict different parts of the world and recombine into a complete state alphaxiv.org/abs/2609.20800
041
alphaXiv @alphaxiv.org · 20/09/2026
“In-Context Robot Learning with VLM Agents” This paper gave the frozen VLM an example of the task, even just a human video with no robot action labels, and it can translate what it sees into robot actions on the fly. alphaxiv.org/abs/2609.19138
020
alphaXiv @alphaxiv.org · 19/09/2026
“Dream-RSI: Recursive Self-Improvement through Evolving Worlds” This paper lets AI agents improve how they search by using past exploration as a simulator to cheaply practice better search strategies before trying them for real. www.alphaxiv.org/abs/2609.14858
050
alphaXiv @alphaxiv.org · 16/09/2026
“Learning to Solve Hard Problems in RL for LLMs by Never Giving Up” This paper makes RL better at hard problems by dynamically shifting sampling compute away from already-solved prompts toward problems the model still struggles to solve. Read more about it: www.alphaxiv.org/abs/2609.13443
010
alphaXiv @alphaxiv.org · 15/09/2026
Why Does Post-Training Quantization Work? This paper finds that pretrained LLMs naturally self-correct quantization errors across layers, while the LM head protects their most confident token predictions from the errors that remain. Read more about it: www.alphaxiv.org/abs/2609.11716
041
alphaXiv @alphaxiv.org · 14/09/2026
This new paper “The Last AI Built by Humans” proposed that true recursive self-improvement means AI getting better at improving itself, from choosing strategies and generating learning experiences, and eventually designing its own mechanisms. Read more about it: www.alphaxiv.org/abs/2609.11873
020
alphaXiv @alphaxiv.org · 13/09/2026
“Thinking with Looped Flows” trains recurrent reasoning as iterative denoising, so each loop turns a noisy guess into a better solution while carrying state forward. This makes extra test-time compute more reliable, reaching 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2 www.alphaxiv.org/pdf/2609.11801
031
alphaXiv @alphaxiv.org · 10/09/2026
DeepSeek-V4.1-Flash DeepSeek's latest model introduces a new Causal Encoder-Decoder and CSA2 architecture that nearly halves prefill compute, compresses KV across layers, and reduces persistent KV storage by ~8x while still supporting 1M-token contexts. www.alphaxiv.org/abs/2609.dee...
030
alphaXiv @alphaxiv.org · 10/09/2026
This proof for one of the seven Millennium Prize Problems was generated by an unreleased OpenAI model beyond GPT-6 Astra. It basically answered a core question in fluid dynamics whether or not a perfectly smooth 3D flow can develop infinite velocity in finite time. www.alphaxiv.org/pdf/2609.nav...
030
alphaXiv @alphaxiv.org · 07/09/2026
On-Policy Distillation is surprisingly data-efficient, with a single query recovering 72% of the full-data gain because rollouts already cover 71.5% of the relevant state space. So the bottleneck is how slowly the student absorbs dense token-level teacher supervision www.alphaxiv.org/abs/2609.04172
010
alphaXiv @alphaxiv.org · 06/09/2026
“Flow Reasoning Models” This paper trains flow models to iteratively correct their own intermediate mistakes, reaching 99.5% on Sudoku-Extreme while matching the next-best method with 44x fewer inference FLOPs. www.alphaxiv.org/abs/2606.29150
010
alphaXiv @alphaxiv.org · 05/09/2026
“Language Models Can Control Their Own Attention” So this paper introduces Declarative Attention, where the model explicitly switches between global, focused, and local attention during reasoning, letting the inference engine skip irrelevant KV cache regions. alphaxiv.org/abs/2609.02737
040
alphaXiv @alphaxiv.org · 04/09/2026
Fast-weight models try to make attention cheaper by continuously rewriting a small fixed-size memory. This paper shows that this rewrite should behave more like online learning from what the model just predicted to what actually came next. alphaxiv.org/abs/2608.27763
010
alphaXiv @alphaxiv.org · 04/09/2026
This paper wraps existing coding agents in repeated planning, coding, and independent testing loops, carrying forward both the software and evidence of what worked or failed. Their system autonomously built a playable FPS over 70+ iterations. alphaxiv.org/abs/2609.01481
020
alphaXiv @alphaxiv.org · 02/09/2026
A new post-training paradigm is emerging called Multi-Teacher On-Policy Distillation (MOPD), where one model learns from multiple specialized RL teachers. Already used in models like Kimi K3 and DeepSeek-V4, we compiled some key papers tracing its developments. www.alphaxiv.org/shared/folde...
041
alphaXiv @alphaxiv.org · 10/06/2025
🚨Another surge in progress for reinforcement learning this week, provided by Beyond the 80/20 Rule, ProRL, and AReal all pushing the boundaries.🚀 Check out the top 10 papers for the week👇 - Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
120
alphaXiv @alphaxiv.org · 02/06/2025
🚨There’s a new ceiling for efficient reasoning with the rise of Learning to Reason without External Rewards, along with AgriFM pushing the boundaries of AI to even agriculture🚀 Check out the top 10 papers for the week👇 - Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
110
alphaXiv @alphaxiv.org · 19/05/2025
🚨Clear your schedule for a recap of a tremendous week featuring DeepSeek-V3 along with BLIP3-o’s improvements in multimodal architecture 🚀 Check out the top 10 papers for the week👇 - Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
120
alphaXiv @alphaxiv.org · 03/05/2025
🚨Bright week for agents and representation learning, notably including X-Fusion’s remarkable progress in multimodal capabilities 🚀 Check out the top 10 papers for the week👇 - From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
110
alphaXiv @alphaxiv.org · 29/04/2025
🚨Don’t miss out on this high-impact week for AI and reinforcement learning, featuring greedy agents, test-time RL, and a powerful new benchmark for LLM physical reasoning 🚀 Check out the top 10 papers for the week👇 - TTRL: Test-Time Reinforcement Learning
161
alphaXiv @alphaxiv.org · 19/04/2025
🚨Don’t miss this week’s immense developments in optimizing reasoning, along with advanced visual embedding capabilities.🚀 Check out the top 10 papers for the week👇 - Reasoning Models Can Be Effective Without Thinking
151
alphaXiv @alphaxiv.org · 13/04/2025
🚨Huge week for video generation and multimodal models, with detailed one-minute video generation and more efficient approaches to multimodality 🚀 Check out the top 10 papers for the week👇 - Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
120
alphaXiv @alphaxiv.org · 08/04/2025
Introducing Deep Research for arXiv Ask questions like 'What are the latest breakthroughs in RL fine-tuning?' and get comprehensive lit reviews with trending papers automatically included Turn hours of literature searches into seconds with AI-powered research context ⚡
1174
alphaXiv @alphaxiv.org · 06/04/2025
Introducing Llama 4 for understanding arXiv papers 🚀 Highlight any section of a paper to ask questions and “@” other papers for quick context, comparisons, and benchmark references
131
alphaXiv @alphaxiv.org · 05/04/2025
🚨Notable week for scaling, with shocking improvements in visual representation learning as well as reward modeling vastly expanding LLM capabilities 🚀 Check out the top 10 papers for the week👇 - What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models
120
alphaXiv @alphaxiv.org · 31/03/2025
Reinforcement learning for retrieval-augmented reasoning 🚀 Baichuan introduces ReSearch, an RL framework that teaches LLMs to reason with search from scratch Outperforms RAG baselines No supervised data on reasoning steps Simple & generalizable Trending #1 on alphaXiv 📈
130
alphaXiv @alphaxiv.org · 29/03/2025
🚨Major week for reinforcement learning, with important strides in improving parameter tuning and reasoning capabilities paving the road for smarter LLMs 🚀 Check out the top 10 papers for the week👇 - Reasoning to Learn from Latent Thoughts
130
alphaXiv @alphaxiv.org · 23/03/2025
🚀This week was huge for foundation models, from generalist humanoid robots to multimodal LLMs that learn from human preferences and negative examples – here are the top 10 papers for the week🚨 - DAPO: An Open-Source LLM Reinforcement Learning System at Scale
150
alphaXiv @alphaxiv.org · 14/03/2025
We used Mistral OCR with Claude 3.7 to create blog-style overviews for arXiv papers Generate beautiful research blogs with figures, key insights, and clear explanations from the paper with just one click Understand papers in minutes - not hours
182
alphaXiv @alphaxiv.org · 14/03/2025
arXiv has been instrumental in advancing open CS research for decades -- we highly encourage everyone to support arXiv for Cornell Giving Day today!
031
Reposted by alphaXiv
J R D M B @jrdmb.bsky.social · 12/03/2025
Saw that @alphaxiv.org now has an AI feature to create paper overviews (using Mistral OCR and Claude), so I created one for new DES paper 2503.06712. Includes figures, key findings, etc. Once created, it's publicly available to everyone at the overview link. alphaXiv also has a cosmology community.
alphaxiv.org
Dark Energy Survey: implications for cosmological expansion models from the final DES Baryon Acoustic Oscillation and Supernova data | alphaXiv
View recent discussion. Abstract: The Dark Energy Survey (DES) recently released the final results of its two principal probes of the expansion history: Type Ia Supernovae (SNe) and Baryonic Acoustic ...
092
alphaXiv @alphaxiv.org · 09/03/2025
🚀This week, AI is soaring to new heights—whether it’s evolving language models through nature-inspired techniques, mastering video generation at scale, or crafting smarter, self-improving agents that think like swarms.🚨 - Nature-Inspired Population-Based Evolution of Large Language Models
100
alphaXiv @alphaxiv.org · 01/03/2025
🚀This week, AI is stepping into new dimensions—becoming co-scientists, sculpting 3D avatars, and blending cloud and on-device models into a seamless dance of creativity and efficiency.🚨 - Towards an AI co-scientist
140
alphaXiv @alphaxiv.org · 22/02/2025
🚀This week was huge for AI—whether through sparse attention, mastering long-context reasoning, or even creating million-dollar software engineering gigs, it's all about smarter, more efficient models shaping a dynamic future.🚨
142
alphaXiv @alphaxiv.org · 16/02/2025
Top Trending Papers on alphaXiv this week!📈 🚀From teaching themselves to predict the future to solving strategic social deduction, AI this week is discovering the hidden geometry of prompts, scaling reasoning, and rethinking what’s possible with less.🚨
131
alphaXiv @alphaxiv.org · 15/02/2025
1997: Deep Blue defeats Kasparov at chess 2016: AlphaGo masters the game of Go 2025: Stanford researchers crack Among Us Trending on alphaXiv 📈 Remarkable new work trains LLMs to master strategic social deduction through multi-agent RL, doubling win rates over standard RL.
162
alphaXiv @alphaxiv.org · 14/02/2025
We used DeepSeek-V3 to classify every AI paper on arXiv by topic (agents, VLMs, etc) 🚀 Now you can instantly filter to see what's trending in each area 🚨
0103
alphaXiv @alphaxiv.org · 09/02/2025
Top Trending Papers on alphaXiv this week!📈 🚀This week, AI is leveling up—from solving Olympiad geometry with AlphaGeometry2 to generating motion with VideoJAM, while also mastering the art of reasoning and adversarial resilience with tools like DeepRAG and LIMO. 🚨
141
alphaXiv @alphaxiv.org · 02/02/2025
🚀 Pushing the limits of large-scale reasoning requires top-tier data! Check out this LLM-powered pipeline to curate 600k+ Olympiad-level math Q&A pairs and introduced LiveAoPSBench, a dynamic benchmark for contamination-resistant math evaluation. 📊🔥 #AI #Math #LLMs
140
Reposted by alphaXiv
Michelle Lin is @NeurIPS2024 (and new to bsky!) she/her @mchll-ln.bsky.social · 24/01/2025
I'm on @alphaxiv.org now! See you in the comments! Fancy Research Profile - michelle.alphaxiv.io Comment-able papers - www.alphaxiv.org/profile/6793... Thanks to @ml-collective.bsky.social for sharing this platform!
michelle.alphaxiv.io
Michelle Lin
<p>I am a MSc student at the University of Montreal &amp; Mila - Quebec AI Institute.</p><p>Prior, I completed my Bachelors in Computer Science at McGill University, where I was also a Research Assist...
093
alphaXiv @alphaxiv.org · 19/01/2025
The alphaXiv Extension is now live on Firefox! Like papers directly on arXiv and see discussions on relevant papers in your community! Link to download: addons.mozilla.org/en-US/firefo...
196