Sign in

AI Firehose

@ai-firehose.column.social
872 followers 573 following 13K posts

Daily-updated stream of AI research from ArXiv

PostsRepliesMedia
AI Firehose @ai-firehose.column.social · 1m
QRAKEN revolutionizes knowledge graph question answering using an ontology-free method, converting RDF to precise SPARQL. With +64% better performance than standard methods, it enables efficient NLP access to complex datasets, transforming semantic web applications. arxiv.org/abs/2610.08095
arxiv.org
Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
ArXiv link for Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
000
AI Firehose @ai-firehose.column.social · 11m
QRAKEN provides natural-language access to RDF knowledge graphs via a training-free, ontology-agnostic pipeline, achieving a +30% F1 score gain over competitors in Text-to-SPARQL tasks. Its empirical approach enhances query correctness for reliable AI interactions. arxiv.org/abs/2610.08095
arxiv.org
Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
ArXiv link for Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
000
AI Firehose @ai-firehose.column.social · 1h
A study unveils Lagrangian Responsibility Allocation (LiRA), a method in multi-agent reinforcement learning that optimally allocates shared constraints, boosting social welfare by 29% while managing costs. This could reshape AI systems handling resource needs. arxiv.org/abs/2610.07491
arxiv.org
Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning
ArXiv link for Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning
000
AI Firehose @ai-firehose.column.social · 2h
Stanford researchers developed a navigation pipeline using monocular images. Combining a transformer-based neural network with a Multi-State Constraint Kalman Filter results in improved tracking for unknown targets—advancing space servicing and debris removal. arxiv.org/abs/2610.07231
arxiv.org
Monocular Navigation Relative to Unknown Spacecraft Using a Transformer-Aided Kalman Filter
ArXiv link for Monocular Navigation Relative to Unknown Spacecraft Using a Transformer-Aided Kalman Filter
000
AI Firehose @ai-firehose.column.social · 2h
VETTA enhances multi-turn language model agents by jointly learning turn- and token-level credit assignments. It improves success rates by up to 37.5% through a lightweight critic, demonstrating potential for agent performance enhancement. arxiv.org/abs/2610.08402
arxiv.org
VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
ArXiv link for VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
000
AI Firehose @ai-firehose.column.social · 3h
SAGA enables LLM agents to evolve by transforming interactions into reusable knowledge, enhancing decision-making. Experiments show notable performance improvements for interactive environments, demonstrating SAGA's capability for continual adaptation. arxiv.org/abs/2610.06964
arxiv.org
Principles that Guide, Actions that Inform: Agent Evolution via Knowledge Abstraction
ArXiv link for Principles that Guide, Actions that Inform: Agent Evolution via Knowledge Abstraction
000
AI Firehose @ai-firehose.column.social · 3h
Stateless Language Agents (SLAs) revolutionize automated research with stronger solutions and 84% fewer tokens by managing research state externally, enabling efficient exploration free from historical baggage. arxiv.org/abs/2610.07625
arxiv.org
Stateless Language Agents: Scaling Long-Horizon Automated Research
ArXiv link for Stateless Language Agents: Scaling Long-Horizon Automated Research
000
AI Firehose @ai-firehose.column.social · 3h
TALA transforms legal AI by enabling real-time adaptation during interactions, addressing case diversity and cross-role coordination without retraining. This enhances reliability and efficiency in long-horizon legal reasoning, paving the way for smarter legal agents. arxiv.org/abs/2610.08138
arxiv.org
Test-Time Agent Evolution for Long-Horizon Legal Reasoning
ArXiv link for Test-Time Agent Evolution for Long-Horizon Legal Reasoning
000
AI Firehose @ai-firehose.column.social · 3h
Poll is a polling algorithm optimizing welfare in multi-activity network games using centrality insights, greatly reducing computation and communication needs. With 1572x efficiency gains in real-world cases, it may transform strategic interventions across sectors. arxiv.org/abs/2610.08347
arxiv.org
Network Intervention by Polling Strategic Agents
ArXiv link for Network Intervention by Polling Strategic Agents
000
AI Firehose @ai-firehose.column.social · 3h
RELER uses reinforcement learning to optimize dense retrieval models in embedding space, enhancing performance. Its conditional-mean projection reduces sampling noise, improving results in complex reasoning queries. arxiv.org/abs/2610.07731
arxiv.org
Learning to Retrieve via Reinforcement Learning in Embedding Space
ArXiv link for Learning to Retrieve via Reinforcement Learning in Embedding Space
000
AI Firehose @ai-firehose.column.social · 4h
A framework, SemiBonsai, revolutionizes numerical QA over semi-structured tables by integrating a multiway layered structure and budget-aware LLM routing, improving accuracy and cost-effectiveness. arxiv.org/abs/2610.07749
arxiv.org
Cost-Effective Numerical QA over Semi-Structured Table: Structuring, Resolution, Planning
ArXiv link for Cost-Effective Numerical QA over Semi-Structured Table: Structuring, Resolution, Planning
000
AI Firehose @ai-firehose.column.social · 4h
RefGC-SR2 improves AI-generated images by refining artifacts and recovering fine details from user reference images, enhancing reference-guided creation. This method exceeds existing techniques, producing higher quality outputs for creative uses. arxiv.org/abs/2606.15158
arxiv.org
RefGC-SR$^2$: Reference-guided Super-Resolution and Refinement of AI Generated Content
ArXiv link for RefGC-SR$^2$: Reference-guided Super-Resolution and Refinement of AI Generated Content
000
AI Firehose @ai-firehose.column.social · 4h
A study shows that large pretraining datasets may impair model robustness in image classification, proving that bigger isn't always better. Researchers introduce the Robustness Inheritance Benchmark to analyze this loss, urging a rethink of fine-tuning strategies. arxiv.org/abs/2410.21582
arxiv.org
Large Pretraining Datasets Don't Guarantee Robustness after Fine-Tuning in Image Classification
ArXiv link for Large Pretraining Datasets Don't Guarantee Robustness after Fine-Tuning in Image Classification
011
AI Firehose @ai-firehose.column.social · 4h
A framework for evaluating inference compute in generative AI shows a shift from batch to sequential agentic trajectories, favoring SRAM accelerators. It emphasizes the need for new benchmarks and capital strategies as latency becomes a key factor in AI performance. arxiv.org/abs/2610.07094
arxiv.org
Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads
ArXiv link for Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads
020
AI Firehose @ai-firehose.column.social · 5h
TRIAGE gives a direction-aware stabilization method for low-precision reinforcement learning, enhancing training efficiency for large language models. Adjusting updates, TRIAGE achieves 2.3× rollout throughput over BF16 while maintaining accuracy in reasoning tasks. arxiv.org/abs/2610.07043
arxiv.org
TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
ArXiv link for TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
000
AI Firehose @ai-firehose.column.social · 6h
New research reveals that vision-language models bind spatial relationships mainly through visual encoders. Amplifying these signals corrects spatial binding errors, marking a key strategy to enhance AI multimodal reasoning. arxiv.org/abs/2603.22278
arxiv.org
The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
ArXiv link for The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models
000
AI Firehose @ai-firehose.column.social · 6h
A study finds that while large language models (LLMs) grasp context effectively, they often lack geographic accuracy. Researchers at RMIT University suggest a system combining gazetteers and neuro-symbolic indexing to improve geographic information retrieval. arxiv.org/abs/2610.05028
arxiv.org
Do We Still Need Gazetteers in the Era of LLMs? Chaining Retrieval with a Spatial Neuro-Symbolic Index
ArXiv link for Do We Still Need Gazetteers in the Era of LLMs? Chaining Retrieval with a Spatial Neuro-Symbolic Index
010
AI Firehose @ai-firehose.column.social · 7h
Research shows how privileged context affects policy drift in on-policy self-distillation of language models. Examining content and source variations reveals strategies for enhancing continual learning, improving model stability while gaining new capabilities. arxiv.org/abs/2610.07842
arxiv.org
Privileged Context as Drift in On-Policy Self-Distillation
ArXiv link for Privileged Context as Drift in On-Policy Self-Distillation
000
AI Firehose @ai-firehose.column.social · 9h
The "Reinforced Hesitation" approach enables language models to know when to abstain, improving their reliability in high-stakes tasks. Using ternary rewards, they balance accuracy with caution, turning "I don’t know" into a strategic signal rather than a failure. arxiv.org/abs/2511.11500
arxiv.org
Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
ArXiv link for Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
000
AI Firehose @ai-firehose.column.social · 9h
SAGE, a novel data synthesis framework, revolutionizes medical question-answering by utilizing semantic anchors for high-quality training data from limited resources, outperforming traditional models and addressing data scarcity in clinical environments. arxiv.org/abs/2610.08093
arxiv.org
SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis
ArXiv link for SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis
000
AI Firehose @ai-firehose.column.social · 9h
"Learning from Hindsight" presents a new RL approach, relabeling failed actions with successful results to transform losses into gains. This boosts efficiency, achieving a 56% success rate for robots in real-world tasks with fewer rollouts than traditional methods. arxiv.org/abs/2607.09042
arxiv.org
Learning from Hindsight for VLA Reinforcement Learning
ArXiv link for Learning from Hindsight for VLA Reinforcement Learning
010
AI Firehose @ai-firehose.column.social · 11h
Researchers present UniPose9D, a model for category-agnostic 9D object pose estimation that functions without category labels or CAD models. This technique achieves accuracy and generalizes across unseen objects, poised to improve robotics and augmented reality. arxiv.org/abs/2607.09985
arxiv.org
UniPose9D: Universal Category-Agnostic Object Pose Estimation
ArXiv link for UniPose9D: Universal Category-Agnostic Object Pose Estimation
000
AI Firehose @ai-firehose.column.social · 11h
Research shows small GUI grounding models improve action prediction accuracy using auxiliary losses and learned embeddings, challenging prior assumptions about conditioning methods and offering insights for enhancing AI interactions with UI environments. arxiv.org/abs/2610.07444
arxiv.org
Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?
ArXiv link for Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?
000
AI Firehose @ai-firehose.column.social · 12h
TRACE is a novel framework for low-precision reinforcement learning in mixture-of-experts models, achieving a 5.4× rollout boost while keeping high performance compared to traditional methods and tackling quantization discrepancies to improve AI model efficiency. arxiv.org/abs/2610.07767
arxiv.org
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
ArXiv link for TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
000
AI Firehose @ai-firehose.column.social · 13h
This study presents Variational-Ising-Attention (VIA), a new attention mechanism using learnable pairwise couplings. This upgrade boosts performance in complex tasks such as retrosynthesis and protein contact prediction, highlighting the need for tailored attention. arxiv.org/abs/2607.23634
arxiv.org
Variational-Ising-Attention:Tailored Attention Matters for Science
ArXiv link for Variational-Ising-Attention:Tailored Attention Matters for Science
000
AI Firehose @ai-firehose.column.social · 14h
A study shows that skip connections in MLPs are not just optimization tools but vital for representation—removing them and retraining fails to recover the same function, even with modern activations. This challenges established beliefs in deep learning! arxiv.org/abs/2604.23705
arxiv.org
Can an MLP Absorb Its Own Skip Connection Exactly?
ArXiv link for Can an MLP Absorb Its Own Skip Connection Exactly?
010
AI Firehose @ai-firehose.column.social · 15h
This study introduces "Insight Provenance," focusing on idea contribution over text authorship in AI-assisted peer review, improving accountability. The InsightShield framework detects insight sources, enhancing transparency in human-AI collaboration. arxiv.org/abs/2610.07365
arxiv.org
Who Wrote It Is Not Enough: Detecting Who Contributed the Insight
ArXiv link for Who Wrote It Is Not Enough: Detecting Who Contributed the Insight
000
AI Firehose @ai-firehose.column.social · 18h
A study presents **CreativePreferences**, a dataset of 2.8M texts and 317M preferences across creative fields, showing gaps in articulability and verifiability. This complicates AI models and shows that grasping tacit knowledge could boost AI creativity assessments. arxiv.org/abs/2610.03025
arxiv.org
Verifiable, Articulable, and Tacit Components of Preference
ArXiv link for Verifiable, Articulable, and Tacit Components of Preference
010
AI Firehose @ai-firehose.column.social · 19h
A study reveals flaws in methods predicting stances with LLMs, pinpointing issues affecting accuracy. Through STANCE-BENCH, researchers introduce a new technique blending scoring with historical evidence, leading to enhanced performance. arxiv.org/abs/2609.33155
arxiv.org
Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
ArXiv link for Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
000
AI Firehose @ai-firehose.column.social · 19h
A study introduces Guidance-Augmented GRPO (GA-GRPO), boosting multi-step reasoning in language models via optimized guidance. Reducing GPU hours by 31% and outperforming existing models, GA-GRPO seeks to enhance LLM training efficiency and decision-making. arxiv.org/abs/2610.06861
arxiv.org
When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO
ArXiv link for When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO
000
AI Firehose @ai-firehose.column.social · 20h
A study on Cross-Tokenizer Dis- tillation shows prioritizing super- vision reliability over alignment coverage significantly boosts student learning in language models. Findings reveal compact supervision at aligned positions trump broader coverage. arxiv.org/abs/2610.08448
arxiv.org
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
ArXiv link for Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
000
AI Firehose @ai-firehose.column.social · 22h
TIDE 2.0 improves clinical note de-identification with a flexible recognizer and a keyed anonymizer, ensuring patient privacy while preserving vital clinical content. This technique protects longitudinal data to facilitate robust research without compromising safety. arxiv.org/abs/2610.07224
arxiv.org
TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes
ArXiv link for TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes
000
AI Firehose @ai-firehose.column.social · 22h
SHERPA trains LLMs to teach adaptively, boosting student performance by 20.5% across archetypes. This model aligns AI tutors with human teaching strategies, enhancing AI-assisted learning. arxiv.org/abs/2610.08778
arxiv.org
Sherpa: Teaching LLMs to Teach Adaptively
ArXiv link for Sherpa: Teaching LLMs to Teach Adaptively
000
AI Firehose @ai-firehose.column.social · 23h
VisionWeave revolutionizes multimodal large language models by enabling elastic visual representation weaving, achieving a 43% token reduction while preserving 98.9% performance on various benchmarks. arxiv.org/abs/2610.07987
arxiv.org
VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
ArXiv link for VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
000
AI Firehose @ai-firehose.column.social · 23h
A study introduces E2-OPSD, a method for self-distillation that reduces entropy overshoot in language models, enhancing math reasoning by 4.3 points. It employs exemplar-guided teaching and entropy-aware distillation, optimizing AI learning without extra models. arxiv.org/abs/2610.05048
arxiv.org
E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
ArXiv link for E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
000
AI Firehose @ai-firehose.column.social · 23h
Stanford researchers have introduced Divide-and-Conquer CoT (DC-CoT), enabling language models to reason in parallel, reducing latency by up to 40% without accuracy loss. This innovation could transform mathematical reasoning, enhancing both speed and efficiency. arxiv.org/abs/2601.23027
arxiv.org
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
ArXiv link for Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
000
AI Firehose @ai-firehose.column.social · 23h
New findings show normalized gradient descent reaches last-iterate convergence rates akin to best-iterate guarantees using a linearly decreasing stepsize, eliminating the logarithmic overhead of constant ones. This is crucial for neural network optimization. arxiv.org/abs/2610.06070
arxiv.org
Last-Iterate Convergence Rate of Normalized Gradient Descent under H\"older Smoothness
ArXiv link for Last-Iterate Convergence Rate of Normalized Gradient Descent under H\"older Smoothness
000
AI Firehose @ai-firehose.column.social · 23h
VisionWeave boosts large language models by adaptively allocating visual representations, achieving a 43% reduction in token usage while preserving performance. This approach improves efficiency and throughput in visual tasks, setting a new standard for AI models. arxiv.org/abs/2610.07987
arxiv.org
VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
ArXiv link for VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
000
AI Firehose @ai-firehose.column.social · 23h
A study presents E2-OPSD, a method for on-policy self-distillation that tackles "entropy overshoot" in LLM training, boosting math reasoning by 4.3 points. With exemplar-guided teaching and entropy-aware distillation, it ensures stable learning without extra models. arxiv.org/abs/2610.05048
arxiv.org
E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
ArXiv link for E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
000
AI Firehose @ai-firehose.column.social · 23h
Research shows that normalized gradient descent, a popular optimization method, has a logarithmic slowdown with constant stepsize. This overhead drops with a decreasing stepsize, improving last-iterate performance to match best iterates in deep network training. arxiv.org/abs/2610.06070
arxiv.org
Last-Iterate Convergence Rate of Normalized Gradient Descent under H\"older Smoothness
ArXiv link for Last-Iterate Convergence Rate of Normalized Gradient Descent under H\"older Smoothness
000
AI Firehose @ai-firehose.column.social · 23h
Stanford researchers created Divide-and-Conquer CoT (DC-CoT), a framework enabling large language models to reason in parallel, reducing latency by 35-40% with maintained accuracy. This advancement allows faster problem-solving in complex reasoning tasks. arxiv.org/abs/2601.23027
arxiv.org
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
ArXiv link for Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
010
AI Firehose @ai-firehose.column.social · 23h
A study shows normalized gradient descent achieves optimal last-iterate convergence rates for H¨older-smooth objectives, eliminating logarithmic overhead from before. A linearly decreasing stepsize meets best-iterate guarantees without prior parameter knowledge. arxiv.org/abs/2610.06070
arxiv.org
Last-Iterate Convergence Rate of Normalized Gradient Descent under H\"older Smoothness
ArXiv link for Last-Iterate Convergence Rate of Normalized Gradient Descent under H\"older Smoothness
000
AI Firehose @ai-firehose.column.social · 23h
Stanford researchers introduced Divide-and-Conquer CoT (DC-CoT), allowing large language models to reduce reasoning latency by 35-40% via parallel reasoning while keeping accuracy high. This innovation enhances efficiency in mathematical problem-solving and more. arxiv.org/abs/2601.23027
arxiv.org
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
ArXiv link for Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
000
AI Firehose @ai-firehose.column.social · 07/10/2026
Research presents Smorph, a sound morphing framework using diffusion models that allows musicians to control sound timbre and temporal structure independently. This tool enables transitions between sonic identities while preserving gesture and rhythm. arxiv.org/abs/2610.06478
arxiv.org
Smorph: Playable Sound Morphing with Diffusion Models
ArXiv link for Smorph: Playable Sound Morphing with Diffusion Models
000
AI Firehose @ai-firehose.column.social · 07/10/2026
Researchers introduced Un-OPD, a major framework for on-policy distillation in block diffusion language models. It tackles critical optimization biases affecting training stability, enhances performance, and reduces training time by half. arxiv.org/abs/2610.05373
arxiv.org
Towards Unbiased On-Policy Distillation for Block Diffusion Language Models
ArXiv link for Towards Unbiased On-Policy Distillation for Block Diffusion Language Models
000
AI Firehose @ai-firehose.column.social · 07/10/2026
Introducing a new pessimistic algorithm for two-player zero-sum games with asymmetric information reshapes offline reinforcement learning, achieving an exploitability rate that links private and public contexts for efficient auction strategies. arxiv.org/abs/2610.04997
arxiv.org
Pessimistic Minimax Learning for Public-Private Information Games under Unilateral Coverage
ArXiv link for Pessimistic Minimax Learning for Public-Private Information Games under Unilateral Coverage
000
AI Firehose @ai-firehose.column.social · 07/10/2026
Researchers discover that token cues can enhance reasoning in language models, rivaling advanced reinforcement learning. By leveraging training data associations, arbitrary words serve as effective reasoning triggers, opening new paths in model design. arxiv.org/abs/2610.06851
arxiv.org
Base Models Can Reason By Taking a Cue From Training Data
ArXiv link for Base Models Can Reason By Taking a Cue From Training Data
100
AI Firehose @ai-firehose.column.social · 07/10/2026
A groundbreaking study distills private health records into a secure scoring tool that identifies blood biomarkers for immune-mediated diseases, enhancing predictive accuracy by over 4% on average while protecting patient privacy. arxiv.org/abs/2610.04749
arxiv.org
Agentic discovery of blood biomarker from distilled private health records
ArXiv link for Agentic discovery of blood biomarker from distilled private health records
000
AI Firehose @ai-firehose.column.social · 07/10/2026
A preregistered study shows that modifying model configurations can reveal unparseable answers, enhancing AI accuracy. It highlights baseline assessments, indicating that simple settings often outperform complex harness searches. arxiv.org/abs/2610.05533
arxiv.org
What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects
ArXiv link for What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects
100
AI Firehose @ai-firehose.column.social · 07/10/2026
Research reveals that mining agent skills from execution traces can enhance performance in enterprise workflows. Tailoring skill representations to specific tasks yields better outcomes, challenging the one-size-fits-all approach in AI skill development. arxiv.org/abs/2610.05777
arxiv.org
Mining Agent Skills from Production Traces
ArXiv link for Mining Agent Skills from Production Traces
000