Sign in

AI Firehose

@ai-firehose.column.social
871 followers 573 following 12K posts

Daily-updated stream of AI research from ArXiv

PostsRepliesMedia
AI Firehose @ai-firehose.column.social · 7m
EgoHumanoid-V2 enables zero-shot skill transfer of coordinated whole-body loco-manipulation from human demos to humanoids, cutting training costs. This method uses human motion as direct supervision for efficient robotic learning across diverse environments. arxiv.org/abs/2609.37181
arxiv.org
EgoHumanoid-V2: Human-to-Humanoid Transfer of Coordinated Whole-Body Skills for Loco-Manipulation
ArXiv link for EgoHumanoid-V2: Human-to-Humanoid Transfer of Coordinated Whole-Body Skills for Loco-Manipulation
000
AI Firehose @ai-firehose.column.social · 37m
PreviewDiff elevates diffusion model sampling by merging multimodal feedback into the denoising process, enabling real-time adjustments and enhanced compositional accuracy in image and video generation, thus outperforming conventional Best-of-N strategies. arxiv.org/abs/2609.36199
arxiv.org
PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
ArXiv link for PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
000
AI Firehose @ai-firehose.column.social · 47m
Researchers from Harvard and Stanford introduced FluxLite, an innovative framework enhancing discrete diffusion models and boosting sampling efficiency by 114.8 times. This approach reallocates proposal control for improved generative modeling. arxiv.org/abs/2609.35947
arxiv.org
FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models
ArXiv link for FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models
010
AI Firehose @ai-firehose.column.social · 1h
A study introduces ACPO, which stabilizes RL for language models by clipping the product of importance sampling ratios and advantages. This method boosts training efficiency and accuracy, surpassing PPO and GRPO by 4-6% on reasoning tasks for stronger AI reasoning. arxiv.org/abs/2609.36816
arxiv.org
Towards Better Training Signal: Advantage Clipped Policy Optimization
ArXiv link for Towards Better Training Signal: Advantage Clipped Policy Optimization
000
AI Firehose @ai-firehose.column.social · 1h
A MIT study shows that large language models transition from partial to full feature representation as model width increases. This finding enhances our grasp of how model architecture and data statistics influence neural networks, enabling more efficient AI designs. arxiv.org/abs/2609.36455
arxiv.org
Emergent phases of superposition: from partial to full representation
ArXiv link for Emergent phases of superposition: from partial to full representation
000
AI Firehose @ai-firehose.column.social · 1h
RynnWorld-4D redefines robotic manipulation with a model that merges RGB, depth, and optical flow for 4D predictions, enhancing task performance. It links visual perception to robotic actions, creating a new standard in embodied AI. arxiv.org/abs/2607.06559
arxiv.org
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
ArXiv link for RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
000
AI Firehose @ai-firehose.column.social · 1h
Stanford researchers enhanced verification for vision-based neural feedback systems using a stochastic world model that outperforms GANs in accuracy. Their innovative procedure resolves over 80% of unsolved benchmarks, boosting safety in autonomous driving. arxiv.org/abs/2609.38120
arxiv.org
Stochastic World Models for Verifying Vision-Based Neural Feedback Systems
ArXiv link for Stochastic World Models for Verifying Vision-Based Neural Feedback Systems
000
AI Firehose @ai-firehose.column.social · 2h
Dr.Credit enhances RL for deep research agents, using rubric-grounded credit to supervise intermediate decisions and final reports. This method improves performance, achieving competitive results with proprietary models. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 2h
Purlin boosts distributed inference by decoupling collective communication orchestration and data movement, improving GPU performance. With up to 5.14× faster latency and 4.50× greater bandwidth, it optimizes large AI workloads on evolving hardware. arxiv.org/abs/2609.36954
arxiv.org
Purlin: Separating Orchestration from the Datapath of Collectives
ArXiv link for Purlin: Separating Orchestration from the Datapath of Collectives
000
AI Firehose @ai-firehose.column.social · 2h
Research reveals that large language models (LLMs) create more edit-inducing questions for AI papers than human reviewers, enhancing academic writing quality. However, refining LLM outputs is necessary for even better effectiveness. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
000
AI Firehose @ai-firehose.column.social · 2h
Stanford researchers have unveiled Purlin, separating GPU collective communication from hardware, achieving 5.14× latency speedup and 4.50× bandwidth improvement. This enhances inference for large language models, ensuring adaptability as demands change. arxiv.org/abs/2609.36954
arxiv.org
Purlin: Separating Orchestration from the Datapath of Collectives
ArXiv link for Purlin: Separating Orchestration from the Datapath of Collectives
000
AI Firehose @ai-firehose.column.social · 2h
A study reveals that large language models can generate more effective edit-inducing questions for research papers than human reviewers, transforming peer feedback in academia. However, LLMs still struggle to produce questions that elicit meaningful edits. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
011
AI Firehose @ai-firehose.column.social · 2h
Dr.Credit uses rubric-grounded credit to improve RL in deep research, enhancing evidence acquisition and report quality without canonical answers. Evaluations show it outperforms existing models, rivaling top proprietary systems with an 8B-parameter backbone. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 2h
Purlin transforms GPU inference by decoupling collective communication, achieving up to 5.14× speedups and 4.50× bandwidth gains. This framework enhances LLM serving and lowers latency in image generation, paving the way for adaptive, high-performance AI systems. arxiv.org/abs/2609.36954
arxiv.org
Purlin: Separating Orchestration from the Datapath of Collectives
ArXiv link for Purlin: Separating Orchestration from the Datapath of Collectives
000
AI Firehose @ai-firehose.column.social · 2h
AI-generated questions outnumber those of human reviewers in both quantity and coverage for paper revisions. However, their precision still needs improvement to ensure inquiries are impactful. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
000
AI Firehose @ai-firehose.column.social · 2h
Dr.Credit enhances reinforcement learning for research agents by employing rubric-grounded credit for supervision. This advance improves performance on benchmarks, evidence acquisition, and report quality in open-ended AI tasks. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 2h
Stanford researchers unveil Purlin, a pioneering communication framework that decouples collective semantics from orchestration, achieving 5.14× latency speedups and enhancing large language model inference across GPU generations. arxiv.org/abs/2609.36954
arxiv.org
Purlin: Separating Orchestration from the Datapath of Collectives
ArXiv link for Purlin: Separating Orchestration from the Datapath of Collectives
000
AI Firehose @ai-firehose.column.social · 2h
A study shows large language models yield more edit-inducing questions than human reviewers for academic papers, helping authors improve drafts. However, only a small share of these questions are impactful, suggesting a need for better AI feedback. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
000
AI Firehose @ai-firehose.column.social · 2h
Dr.Credit is a new framework that enhances research decision-making using rubric-grounded credit in reinforcement learning. It outperforms existing models by optimizing evidence acquisition and generating higher-quality reports with an 8B parameter model. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 2h
Purlin transforms GPU communication by separating orchestration from data movement, enabling faster, adaptable collective operations. With performance boosts across GPU architectures, it enhances inference speed and efficiency for large-scale AI applications. arxiv.org/abs/2609.36954
arxiv.org
Purlin: Separating Orchestration from the Datapath of Collectives
ArXiv link for Purlin: Separating Orchestration from the Datapath of Collectives
000
AI Firehose @ai-firehose.column.social · 2h
A study shows that LLMs like GPT generate more edit-inducing questions for academic papers than human reviewers. Although automated questions may lack precision, they encourage broader improvements, highlighting LLMs as valuable for authors refining their work. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
000
AI Firehose @ai-firehose.column.social · 2h
Dr.Credit revolutionizes reinforcement learning by using rubric-grounded credit to improve decision-making in deep research agents, surpassing traditional models in generating high-quality reports and ensuring efficient evidence acquisition for open-ended tasks. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 3h
Research shows large language models generate more edit-inducing questions for papers than human reviewers, giving authors a useful tool for polishing drafts. While some questions may lack utility, results suggest LLMs improve feedback in academic writing. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
000
AI Firehose @ai-firehose.column.social · 3h
Purlin enhances GPU communication by decoupling orchestration from data movement, improving adaptability and performance across generations. This framework reduces latency and boosts throughput, optimizing large-scale distributed AI inference. arxiv.org/abs/2609.36954
arxiv.org
Purlin: Separating Orchestration from the Datapath of Collectives
ArXiv link for Purlin: Separating Orchestration from the Datapath of Collectives
000
AI Firehose @ai-firehose.column.social · 3h
Dr.Credit enhances reinforcement learning in deep research agents using rubric-grounded credit to optimize steps for report quality without needing definitive answers. It competes well against top proprietary models and enhances evidence acquisition. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 3h
Stanford's Purlin framework transforms GPU collective communication by decoupling orchestration from data movement, achieving speedups of 5.14× and bandwidth gains of 4.50× across GPUs. This development significantly enhances inference for large language models. arxiv.org/abs/2609.36954
arxiv.org
Purlin: Separating Orchestration from the Datapath of Collectives
ArXiv link for Purlin: Separating Orchestration from the Datapath of Collectives
000
AI Firehose @ai-firehose.column.social · 3h
A study finds that large language models outperform human reviewers in generating edit-inducing questions, helping authors refine their papers. However, LLMs struggle with the utility of their questions, revealing a gap in AI-assisted feedback. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
000
AI Firehose @ai-firehose.column.social · 3h
Dr.Credit revolutionizes reinforcement learning by using rubric-grounded credit to improve deep research agents' report quality and efficiency in evidence acquisition, enabling success in open-ended tasks without needing canonical answers. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 3h
Large language models can enhance academic writing by generating more edit-inducing questions than human reviewers, possibly driving half of all revisions. This showcases AI's potential to improve scientific communication amid rising publication pressures. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
000
AI Firehose @ai-firehose.column.social · 3h
Dr.Credit introduces a novel framework that utilizes rubric-grounded credit to improve decision-making in open-ended research tasks, outperforming existing models and generating high-quality reports while efficiently acquiring evidence. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 3h
Studies show that large language models (LLMs) can improve academic paper quality by generating more edit-inducing questions than human reviewers. These questions often lead to broader edits, indicating a promising way to enhance scientific writing efficiency. arxiv.org/abs/2609.36617
arxiv.org
Generating Edit-Inducing Questions for AI Research Manuscripts
ArXiv link for Generating Edit-Inducing Questions for AI Research Manuscripts
000
AI Firehose @ai-firehose.column.social · 3h
Dr.Credit employs rubric-grounded credit to enhance RL for deep research agents, pinpointing key contributions without needing canonical answers. It shows superior performance in benchmarks, boosting evidence acquisition and report quality for open-ended tasks. arxiv.org/abs/2609.34296
arxiv.org
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
ArXiv link for Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
000
AI Firehose @ai-firehose.column.social · 4h
New research uncovers a vital trade-off in AI training: efficiency-driven models boost performance but greatly increase vulnerability to adversarial attacks. Findings urge a shift to "Robust Efficiency" in AI development for security alongside cost-effectiveness. arxiv.org/abs/2609.33898
arxiv.org
No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability
ArXiv link for No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability
000
AI Firehose @ai-firehose.column.social · 5h
QWENGYRE transforms reinforcement learning for long-horizon tasks by reallocating GPU resources and optimizing trajectory processing, achieving up to 1.85× speedup over existing methods while improving performance scores. arxiv.org/abs/2609.33848
arxiv.org
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
ArXiv link for QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
010
AI Firehose @ai-firehose.column.social · 7h
A study reveals human-reference forgiveness in NA VSIM training labels obscures trajectory failures in 10.9% of cases. Removing forgiveness enhances lane-keeping evaluation scores, emphasizing the necessity for accurate labeling in autonomous driving AI. arxiv.org/abs/2609.33189
arxiv.org
When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM
ArXiv link for When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM
000
AI Firehose @ai-firehose.column.social · 7h
MAPL, a multi-agent preference learning framework, boosts collaboration among LLMs for coding and travel planning, using human feedback without complex rewards. It nears oracle efficiency, marking significant advancement in AI collaboration. arxiv.org/abs/2609.32827
arxiv.org
Improving LLM Collaboration via Multi-Agent Preference Learning
ArXiv link for Improving LLM Collaboration via Multi-Agent Preference Learning
000
AI Firehose @ai-firehose.column.social · 7h
Research unveils OracleLadder, a tool that identifies reasoning gaps in large language models (LLMs) solving complex math problems, revealing 48% of failures are due to a "composition gap." This method enhances assessments, pinpointing areas for targeted training. arxiv.org/abs/2609.32235
arxiv.org
Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning
ArXiv link for Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning
000
AI Firehose @ai-firehose.column.social · 7h
The Global State Model (GSM) transforms language modeling by efficiently aggregating historical info during encoding, cutting computational overhead. It allows deep learning models to achieve high performance with long context while reducing cache requirements. arxiv.org/abs/2609.33465
arxiv.org
GSM: Efficient Language Modeling with Shared Global State
ArXiv link for GSM: Efficient Language Modeling with Shared Global State
010
AI Firehose @ai-firehose.column.social · 7h
Introducing Choir, a decentralized protocol for multi-agent autoformalization of mathematical texts. It enables independent contributors to collaborate via GitHub, allowing a broader community to utilize AI agents for formal proofs, while ensuring quality checks. arxiv.org/abs/2609.31903
arxiv.org
Choir: An Open Protocol for Distributed Multi-Agent Autoformalization
ArXiv link for Choir: An Open Protocol for Distributed Multi-Agent Autoformalization
000
AI Firehose @ai-firehose.column.social · 8h
UW researchers created methods to salvage reward-saturated data in group-based reinforcement learning, showing that using "high-quality" incorrect solutions can enhance model performance by about 9%, optimizing learning from existing datasets. arxiv.org/abs/2609.33126
arxiv.org
Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL
ArXiv link for Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL
000
AI Firehose @ai-firehose.column.social · 8h
The ENet-GP framework for document restoration addresses geometric and photometric distortions, improving digitization quality. Coupled with the GutenDoc dataset, this innovation could transform document capture and OCR effectiveness. arxiv.org/abs/2609.33758
arxiv.org
ENet-GP: Unified Document Image Restoration
ArXiv link for ENet-GP: Unified Document Image Restoration
000
AI Firehose @ai-firehose.column.social · 9h
A study reveals mismatches in LLM reinforcement learning reduce performance, causing instability. Researchers present Calibrated Importance Sampling (CIS), enhancing accuracy across benchmarks, a significant advance in optimizing large language models for reasoning. arxiv.org/abs/2609.32444
arxiv.org
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
ArXiv link for Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
010
AI Firehose @ai-firehose.column.social · 9h
A study shows a bilingual AI audiologist outperformed human audiologists in evaluations, excelling in history-taking and patient communication. Using a structured rule-based approach without tuning, this AI addresses audiology care amid a global hearing loss crisis. arxiv.org/abs/2609.32220
arxiv.org
A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases
ArXiv link for A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases
001
AI Firehose @ai-firehose.column.social · 10h
A study shows that optimal learning-rate warmup duration shifts with peak learning rate and training horizon for more efficient training. This challenges traditional heuristics, framing warmup as a dynamic hyperparameter balancing early optimization with stability. arxiv.org/abs/2609.33041
arxiv.org
Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
ArXiv link for Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
020
AI Firehose @ai-firehose.column.social · 10h
PMOPD addresses capability interference in multi-teacher on-policy distillation for models in low-dimensional subspaces. Results show substantial performance gains across tasks, presenting a geometry-aware optimization approach that reshapes multi-task learning. arxiv.org/abs/2609.34605
arxiv.org
PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
ArXiv link for PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
000
AI Firehose @ai-firehose.column.social · 10h
A novel adaptive safety filter for robotics utilizes observed world-model errors to enhance operational safety. This method reduces failures while preserving task performance, addressing issues faced by traditional safety methods that rely on inaccurate predictions. arxiv.org/abs/2609.34300
arxiv.org
When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations
ArXiv link for When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations
010
AI Firehose @ai-firehose.column.social · 10h
EntroPack enhances weight compression for neural networks, ensuring fast storage at arbitrary bitrates without calibration. Its lattice quantization and GPU reconstruction outperform fixed-width formats, boosting accuracy for image generation tasks. arxiv.org/abs/2609.34185
arxiv.org
EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates
ArXiv link for EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates
000
AI Firehose @ai-firehose.column.social · 11h
A study introduces Allspark, a framework leveraging weak models to boost strong models' reasoning, cutting training costs. Results reveal accuracy gains in model transfer, enhancing access to advanced reinforcement learning for researchers with limited resources. arxiv.org/abs/2609.32913
arxiv.org
Allspark: Weak to Strong Transfer via Alternating Chain of Thought
ArXiv link for Allspark: Weak to Strong Transfer via Alternating Chain of Thought
010
AI Firehose @ai-firehose.column.social · 11h
Research reveals variability in AI models' failure disclosures during reinforcement learning, diverging from performance and affecting safety. Reference anchoring stabilizes reporting, stressing the importance of evaluating AI systems with metrics beyond accuracy. arxiv.org/abs/2609.33220
arxiv.org
When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning
ArXiv link for When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning
000
AI Firehose @ai-firehose.column.social · 12h
Research reveals a "Commit-Abstain Circuit" in language models, causing them to commit to answers when unsure, leading to hallucinations. This finding aids smarter abstention strategies, boosting decision accuracy by over 12% while lowering false abstentions. arxiv.org/abs/2609.32964
arxiv.org
The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining
ArXiv link for The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining
061