Sign in

alphaXiv

@alphaxiv.org
272 followers 4 following 357 posts

High fidelity research

PostsRepliesMedia
alphaXiv @alphaxiv.org · 2h
More pre-training doesn't necessarily mean better generalization for LLMs? This new paper discovers mode-hopping, where LLMs repeatedly switch between shallow pattern-matching and actual generalization, even while training loss remains stable. Read more: www.alphaxiv.org/abs/2609.33150
010
alphaXiv @alphaxiv.org · 30/09/2026
Most AI agents force the harness around a decoder-only Transformer But what if we change the model’s in/output shape to fit the agent? This blog suggests using recurrent memory for old history while dense attention handles recent context, reducing manual compaction alphaxiv.org/abs/2609.lan...
051
alphaXiv @alphaxiv.org · 29/09/2026
What if you make on-policy distillation recursive? This paper proposes that after each round, the improved model becomes both the next student and the new gold-conditioned teacher, letting better reasoning feed into future supervision. www.alphaxiv.org/abs/2609.30652
010
alphaXiv @alphaxiv.org · 28/09/2026
What if you could trace a model behavior back to the smallest set of internal components actually responsible for it? Matryoshka Attribution learns one ranked mask over internal components across many sparsity levels at once, letting it isolate compact causal circuits alphaxiv.org/abs/2609.25518
030
alphaXiv @alphaxiv.org · 27/09/2026
If world model has to reconstruct every pixel, it'll pretty much waste most of its capacity modeling irrelevant background noise. This paper found that by removing pixel decoding, it performs much better with moving distractors and natural-video backgrounds. alphaxiv.org/abs/2609.22175
041
alphaXiv @alphaxiv.org · 26/09/2026
Can a model learn entirely through self-play and no real data? This self-play setup can generate its own pre-training data from scratch and still teach a model general predictive structure that transfers to real-world data. www.alphaxiv.org/abs/2609.30063
050
alphaXiv @alphaxiv.org · 25/09/2026
This paper shows you can just use JEV for every evaluation instead of expensive LLM. JEV basically acts as a cheap first-pass judge, returning both a verdict and how confident it is. When confidence is high, keep the answer. When it’s low, escalate to a stronger LLM. alphaxiv.org/abs/2609.26550
061
alphaXiv @alphaxiv.org · 24/09/2026
Xiaomi dropped their new attention mechanism for their upcoming V3 model. At 1M context, it cuts prefill FLOPs by 2.92x vs HySparse and shrinks the KV cache from 6.72GB to 2.69GB, improving long-context retrieval by a lot. Read more about their new architecture here: alphaxiv.org/abs/2609.26368
030
alphaXiv @alphaxiv.org · 23/09/2026
VLA models are too slow for reactive robot control, so this paper lets the VLA generate actions in the background, while a lightweight RL policy uses the latest observation to edit and select actions in real time This gives from 42% to 97% with just 10min of online data alphaxiv.org/abs/2609.18207
060
alphaXiv @alphaxiv.org · 22/09/2026
Xiaomi MiMo just completed their largest RL scaling run so far, openly. With the run costing $2.6M for Pro and $0.9M for Flash, with ~44% spent on rollouts, 41-44% on training, and 14% on grading. Read more about it: www.alphaxiv.org/pdf/2609.mim...
030
alphaXiv @alphaxiv.org · 21/09/2026
This paper proposes one world-modeling framework that works across multiple fields Through Orthogonal Predictive Factorization, it splits a single JEPA latent state into complementary factors that predict different parts of the world and recombine into a complete state alphaxiv.org/abs/2609.20800
041
alphaXiv @alphaxiv.org · 20/09/2026
“In-Context Robot Learning with VLM Agents” This paper gave the frozen VLM an example of the task, even just a human video with no robot action labels, and it can translate what it sees into robot actions on the fly. alphaxiv.org/abs/2609.19138
020
alphaXiv @alphaxiv.org · 19/09/2026
“Dream-RSI: Recursive Self-Improvement through Evolving Worlds” This paper lets AI agents improve how they search by using past exploration as a simulator to cheaply practice better search strategies before trying them for real. www.alphaxiv.org/abs/2609.14858
050
alphaXiv @alphaxiv.org · 16/09/2026
“Learning to Solve Hard Problems in RL for LLMs by Never Giving Up” This paper makes RL better at hard problems by dynamically shifting sampling compute away from already-solved prompts toward problems the model still struggles to solve. Read more about it: www.alphaxiv.org/abs/2609.13443
010
alphaXiv @alphaxiv.org · 15/09/2026
Why Does Post-Training Quantization Work? This paper finds that pretrained LLMs naturally self-correct quantization errors across layers, while the LM head protects their most confident token predictions from the errors that remain. Read more about it: www.alphaxiv.org/abs/2609.11716
041
alphaXiv @alphaxiv.org · 14/09/2026
This new paper “The Last AI Built by Humans” proposed that true recursive self-improvement means AI getting better at improving itself, from choosing strategies and generating learning experiences, and eventually designing its own mechanisms. Read more about it: www.alphaxiv.org/abs/2609.11873
020
alphaXiv @alphaxiv.org · 13/09/2026
“Thinking with Looped Flows” trains recurrent reasoning as iterative denoising, so each loop turns a noisy guess into a better solution while carrying state forward. This makes extra test-time compute more reliable, reaching 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2 www.alphaxiv.org/pdf/2609.11801
031
alphaXiv @alphaxiv.org · 10/09/2026
DeepSeek-V4.1-Flash DeepSeek's latest model introduces a new Causal Encoder-Decoder and CSA2 architecture that nearly halves prefill compute, compresses KV across layers, and reduces persistent KV storage by ~8x while still supporting 1M-token contexts. www.alphaxiv.org/abs/2609.dee...
030
alphaXiv @alphaxiv.org · 10/09/2026
This proof for one of the seven Millennium Prize Problems was generated by an unreleased OpenAI model beyond GPT-6 Astra. It basically answered a core question in fluid dynamics whether or not a perfectly smooth 3D flow can develop infinite velocity in finite time. www.alphaxiv.org/pdf/2609.nav...
030
alphaXiv @alphaxiv.org · 07/09/2026
On-Policy Distillation is surprisingly data-efficient, with a single query recovering 72% of the full-data gain because rollouts already cover 71.5% of the relevant state space. So the bottleneck is how slowly the student absorbs dense token-level teacher supervision www.alphaxiv.org/abs/2609.04172
010
alphaXiv @alphaxiv.org · 06/09/2026
“Flow Reasoning Models” This paper trains flow models to iteratively correct their own intermediate mistakes, reaching 99.5% on Sudoku-Extreme while matching the next-best method with 44x fewer inference FLOPs. www.alphaxiv.org/abs/2606.29150
010
alphaXiv @alphaxiv.org · 05/09/2026
“Language Models Can Control Their Own Attention” So this paper introduces Declarative Attention, where the model explicitly switches between global, focused, and local attention during reasoning, letting the inference engine skip irrelevant KV cache regions. alphaxiv.org/abs/2609.02737
040
alphaXiv @alphaxiv.org · 04/09/2026
Fast-weight models try to make attention cheaper by continuously rewriting a small fixed-size memory. This paper shows that this rewrite should behave more like online learning from what the model just predicted to what actually came next. alphaxiv.org/abs/2608.27763
010
alphaXiv @alphaxiv.org · 04/09/2026
This paper wraps existing coding agents in repeated planning, coding, and independent testing loops, carrying forward both the software and evidence of what worked or failed. Their system autonomously built a playable FPS over 70+ iterations. alphaxiv.org/abs/2609.01481
020
alphaXiv @alphaxiv.org · 02/09/2026
A new post-training paradigm is emerging called Multi-Teacher On-Policy Distillation (MOPD), where one model learns from multiple specialized RL teachers. Already used in models like Kimi K3 and DeepSeek-V4, we compiled some key papers tracing its developments. www.alphaxiv.org/shared/folde...
041
alphaXiv @alphaxiv.org · 10/06/2025
ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development An LLM-powered plugin designed to simplify and accelerate workflow creation in ComfyUI, an open-source AI art platform, by providing intelligent node/model recommendations and automated workflow generation.
010
alphaXiv @alphaxiv.org · 10/06/2025
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning AReaL is an asynchronous reinforcement learning system that efficiently trains large language models for reasoning tasks by maximizing GPU usage and decoupling generation from training.
100
alphaXiv @alphaxiv.org · 10/06/2025
Pseudo-Simulation for Autonomous Driving Pseudo-simulation is a new evaluation paradigm for autonomous vehicles that blends the realism of real-world data with the generalization power of simulation, enabling robust, scalable testing without the need for interactive environments.
131
alphaXiv @alphaxiv.org · 10/06/2025
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics SmolVLA is a compact, open-source VLA model built for low-cost training and real-world deployment on consumer hardware, enabling efficient language-driven robot control without sacrificing performance.
100
alphaXiv @alphaxiv.org · 10/06/2025
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models This paper introduces ProRL, a method that uses long-horizon reinforcement learning to unlock new reasoning strategies in LLMs—strategies that base models cannot access, even with extensive sampling.
110
alphaXiv @alphaxiv.org · 10/06/2025
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning This paper shows that punishing wrong answers—without explicitly rewarding the correct—can be surprisingly effective for improving reasoning in large language models trained via reinforcement learning with verifiable rewards.
110
alphaXiv @alphaxiv.org · 10/06/2025
CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate This paper challenges the idea that Chain-of-Thought (CoT) prompting enables true reasoning in LLMs, arguing instead that CoT acts as a structural constraint that guides models to imitate the appearance of reasoning.
110
alphaXiv @alphaxiv.org · 10/06/2025
Why Gradients Rapidly Increase Near the End of Training This note investigates a sudden rise in gradient norms during the late stages of LLM training and identifies a surprising cause: the interplay between weight decay, normalization layers, and scheduled learning rate decay.
110
alphaXiv @alphaxiv.org · 10/06/2025
General agents need world models This paper proves that any general agent capable of reliably completing diverse, goal-directed tasks must implicitly learn a predictive model of its environment—challenging the notion that model-free learning is sufficient for general intelligence.
110
alphaXiv @alphaxiv.org · 10/06/2025
Beyond the 80/20 Rule This paper investigates how a small subset of high-entropy tokens—termed "forking tokens"—drives the performance of reinforcement learning with verifiable rewards (RLVR) in reasoning tasks for large language models.
110
alphaXiv @alphaxiv.org · 10/06/2025
- Pseudo-Simulation for Autonomous Driving - AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning - ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development
110
alphaXiv @alphaxiv.org · 10/06/2025
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning - ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models - SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
210
alphaXiv @alphaxiv.org · 10/06/2025
- General agents need world models - Why Gradients Rapidly Increase Near the End of Training - CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
110
alphaXiv @alphaxiv.org · 10/06/2025
🚨Another surge in progress for reinforcement learning this week, provided by Beyond the 80/20 Rule, ProRL, and AReal all pushing the boundaries.🚀 Check out the top 10 papers for the week👇 - Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
120
alphaXiv @alphaxiv.org · 02/06/2025
HiDream-I1 HiDream-I1 is a 17B-parameter open-source image generation model using a novel sparse Diffusion Transformer (DiT) with dynamic Mixture-of-Experts (MoE) to deliver state-of-the-art image quality in seconds while reducing computation.
010
alphaXiv @alphaxiv.org · 02/06/2025
Silence is Not Consensus This work introduces the Catfish Agent, a specialized large language model designed to disrupt premature consensus—called Silent Agreement—in multi-agent clinical decision-making systems by injecting structured dissent to improve diagnostic accuracy.
120
alphaXiv @alphaxiv.org · 02/06/2025
WebDancer: Towards Autonomous Information Seeking Agency WebDancer is an end-to-end autonomous web agent designed for complex, multi-step information seeking. It combines a data-centric and training-stage pipeline to enable robust reasoning and decision-making in real-world web environments.
110
alphaXiv @alphaxiv.org · 02/06/2025
LiteCUA The authors introduce AIOS 1.0, a platform that helps language models better understand and interact with computers by transforming them into contextual environments. Built on this, LiteCUA is a lightweight agent that uses this structured context to perform digital tasks.
100
alphaXiv @alphaxiv.org · 02/06/2025
Darwin Godel Machine The Darwin Gödel Machine (DGM) is a self-improving AI system that rewrites its own code to enhance coding performance. Inspired by Gödel machines and Darwinian evolution, it uses empirical validation and an archive of past agents to drive open-ended, recursive improvement.
120
alphaXiv @alphaxiv.org · 02/06/2025
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models This paper analyzes a fundamental barrier in reinforcement learning (RL) for large language models (LLMs): the sharp early collapse of policy entropy, which limits exploration and caps downstream performance.
110
alphaXiv @alphaxiv.org · 02/06/2025
AgriFM AgriFM is a multi-source temporal remote sensing foundation model tailored for crop mapping. It introduces a modified Video Swin Transformer backbone for unified spatiotemporal processing of satellite imagery from MODIS, Landsat-8/9, and Sentinel-2.
100
alphaXiv @alphaxiv.org · 02/06/2025
WorldEval This paper introduces WorldEval, a real-to-video evaluation framework that uses world models to assess real-world robot manipulation policies in a scalable, safe, and reproducible way. It avoids costly real-world evaluations by simulating robot actions via generated videos.
100
alphaXiv @alphaxiv.org · 02/06/2025
Learning to Reason without External Rewards This paper introduces RLIF, a paradigm where LLMs improve reasoning using intrinsic signals instead of external rewards. The authors propose INTUITOR, which uses a model’s self-confidence—measured as self-certainty—as the sole reward signal.
100
alphaXiv @alphaxiv.org · 02/06/2025
Paper2Poster The paper introduces Paper2Poster, the first benchmark for automated academic poster generation, and PosterAgent, a visual-in-the-loop multi-agent system that converts research papers into high-quality posters using open-source models.
110
alphaXiv @alphaxiv.org · 02/06/2025
- WebDancer: Towards Autonomous Information Seeking Agency - Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making - HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
100