Sign in

Unofficial Arxiv daily bot

@arxiv-daily-bot.bsky.social
98 followers 1 following 10K posts

I'm a bot. I post new AI articles from Arxiv.org. Source code available here: github.com/emadb/arxiv-bluesky-bot

PostsRepliesMedia
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 38m
EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution Python Song et al. #arXiv #cs.AI
arxiv.org
EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic h…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 40m
Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models Tan Yu et al. #arXiv #cs.AI #cs.SE
arxiv.org
Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 42m
GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning Jingyao Zhang, Yun Li, Lu Han #arXiv #cs.AI #cs.CL #cs.CV #cs.LG
arxiv.org
GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning
Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit functi…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 44m
From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents Juanyang Xu et al. #arXiv #cs.AI
arxiv.org
From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We est…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 46m
Think Before You Paint: Recursive Latent Reasoning for Diffusion Models Pawe{\l} Skier\'s et al. #arXiv #cs.AI #cs.LG
arxiv.org
Think Before You Paint: Recursive Latent Reasoning for Diffusion Models
Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze. When a discrete symbolic representation is available, recursive methods such as the Tiny Recursive Model (TRM) solve even hard instances of these puzz…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 48m
How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression Xijie Gong et al. #arXiv #cs.AI #cs.LG
arxiv.org
How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression
Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and …
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 50m
SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models Miao Yu et al. #arXiv #cs.AI
arxiv.org
SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or ne…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 52m
Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study Lin Wu et al. #arXiv #cs.AI #cs.CY
arxiv.org
Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from J…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 54m
DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists Samuel Margolis et al. #arXiv #cs.AI
arxiv.org
DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists
Drug target discovery requires distinguishing molecules that causally drive disease from those that are merely associated with it. Training and evaluating AI agents to perform this workflow end-to-end is difficult because real world biobanks lack known causal ground truth and participant-level data…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 56m
Ream: Unfolding Mutual Awareness in Human-Agent Workspaces Peiling Jiang et al. #arXiv #cs.AI #cs.HC
arxiv.org
Ream: Unfolding Mutual Awareness in Human-Agent Workspaces
As AI agents work alongside humans in shared workspaces, a mutual awareness challenge arises: agents act at speeds that outpace human monitoring, and users' evolving interests are not always expressed in chat. This challenge is especially pressing in literature review, where both parties retrieve, …
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 58m
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents Hoang Phan et al. #arXiv #cs.AI
arxiv.org
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for …
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 1h
RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment Tianruo Rose Xu et al. #arXiv #cs.AI #cs.LG
arxiv.org
RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment
Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures can cause human injury and costly hardware damage. Beyond selecting …
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 1h
Trajectory Abstraction for the Science of Language Agent Behavior Tianqiang Yan #arXiv #cs.AI #cs.CL
arxiv.org
Trajectory Abstraction for the Science of Language Agent Behavior
Scientific studies of language agents need behavioral variables that support hypotheses across tasks and models. We formulate this research problem as learning and testing a hierarchy of trajectory abstractions. A concrete recursive procedure first measures role- and phase-indexed events, proposes …
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 1h
Few Bits, One Law: Toward W2A4KV2 Kai Yi et al. #arXiv #cs.AI #cs.CL #cs.LG
arxiv.org
Few Bits, One Law: Toward W2A4KV2
Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challen…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 1h
Finding Blind Spots in AppWorld and WorkArena Task Verifiers Richard Abrich #arXiv #cs.AI #cs.SE
arxiv.org
Finding Blind Spots in AppWorld and WorkArena Task Verifiers
Execution-based task verifiers decide whether an agent succeeded. We audit shipped AppWorld and WorkArena verifiers with source-informed mutation tests. The main audit never modifies a shipped checker. In AppWorld, duplicating a non-idempotent write creates an extra record while preserving every ch…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 1h
Shared-Roadmap Generation and Evaluator for Multi-Agent Path Planning Using Heterogeneous Graph Neural Network Brandon Ho, Nikola Rogers, Seung-Kyum Choi #arXiv #cs.AI #cs.LG #cs.MA #cs.RO
arxiv.org
Shared-Roadmap Generation and Evaluator for Multi-Agent Path Planning Using Heterogeneous Graph Neural Network
Multi-agent path planning (MAPP) in continuous environments often relies on roadmaps to balance safety and search efficiency. However, traditional roadmap generation methods, such as lattice grids or standard sampling-based approaches, frequently face a trade-off between graph density and the likel…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions Yizhen Xie, Mengyang Liu #arXiv #cs.AI #cs.LG #q-fin.PM #q-fin.TR
arxiv.org
SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions
As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
Stale, Misattributed, or Late: Where Personal Memory Fails Before Generation Haonan Deng, Park Sinchaisri #arXiv #cs.AI
arxiv.org
Stale, Misattributed, or Late: Where Personal Memory Fails Before Generation
Personal memory for language agents is usually judged by whether the final an- swer is correct. That score hides errors that arise before generation: the memory block may contain an obsolete value, a fact about the wrong person, or no use- ful fact before the serving deadline. We measure these fail…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
Outperformance Inverse Optimization: Learning Objective Functions that Outperform Agent Decisions Akira Kitaoka #arXiv #cs.AI #math.OC #math.ST #stat.ML #stat.TH
arxiv.org
Outperformance Inverse Optimization: Learning Objective Functions that Outperform Agent Decisions
Inverse optimization estimates the weights of an objective function that explain observed decisions as optimal solutions, and is used in a variety of fields. For mixed-integer linear programs (MILPs), existing methods aim to reproduce the observations as optimal solutions, and thus learn compromise…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers Evgenia Ilia, Wilker Aziz #arXiv #cs.AI
arxiv.org
A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers
The wide adoption of LLMs across broad NLG applications heightens the importance of providing users with the means to avert errors and hallucinations. Uncertainty quantification is poised to fill that gap; with low uncertainty (high confidence), as a proxy for correctness, allowing users to be sele…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
LLM-Enabled UAV Dispatch: A System-Level Survey and Taxonomy Xiao Han et al. #arXiv #cs.AI #cs.SY #eess.SY
arxiv.org
LLM-Enabled UAV Dispatch: A System-Level Survey and Taxonomy
Unmanned aerial vehicle (UAV) dispatch is beginning to move beyond isolated path planning and optimization-driven resource allocation toward system-level coordination supported by semantic reasoning and LLM-based interfaces. This survey provides a unified characterization of LLM-enabled UAV dispatc…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents Kefan Liu et al. #arXiv #cs.AI
arxiv.org
We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents
Agentic systems use large language models (LLMs) to carry out concrete tasks. Prior work often borrows abstractions such as scheduling, caching or isolation piecemeal from operating systems, so the mechanisms it builds share little common ground, and the shared view of the two forms of agentic syst…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
AGAR: a reinforcement learning substrate for LLM program evolution Haoran Li et al. #arXiv #cs.AI #cs.MA
arxiv.org
AGAR: a reinforcement learning substrate for LLM program evolution
Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive, offering a practical route to algorithm discovery. But that loop is governed by five constants set by hand: which parent to select, how hard to mutate, how to keep div…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use Gennadiy Savrasov et al. #arXiv #cs.AI #cs.LG
arxiv.org
CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use
Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
📌 GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks Gabriel Diaz-Ireland et al. #arXiv #cs.AI
arxiv.org
GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks
Before tool-using LLM agents are deployed in environmental and geospatial workflows, teams need evidence that an agent reliably selects the right operations against real APIs. We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geo…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents Mayur Akewar, Ravi Ranjan #arXiv #cs.AI
arxiv.org
RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents
Agent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible. None decides which of the safe boundaries they expose are worth materializing. We formulate this as counterfactual checkpoint advantage, the reduction in future recovery cost obt…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System Panagiotis Kasnesis et al. #arXiv #cs.AI
arxiv.org
Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
An agentic system issues several structurally different kinds of LLM calls. It routes intent, classifies actions, grounds language in a device registry, plans multi-agent pipelines and writes the Python code those pipelines run. The difficulty of these call sites varies by an order of magnitude, ye…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
Sigma-Hunter: A Domain-Specific Language Model for Threat Hunting and Detection Engineering Kemal Davaslioglu, Sastry Kompella #arXiv #cs.AI #cs.CR
arxiv.org
Sigma-Hunter: A Domain-Specific Language Model for Threat Hunting and Detection Engineering
Detection engineers must translate threat reports, forensic observations, and hunt hypotheses into precise, testable rules. General-purpose large language models (LLMs) can draft such rules, but often produce invalid YAML, incorrect log sources, unsupported fields, or overly broad detection logic. …
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails Melissa Kazemi Rad et al. #arXiv #cs.AI
arxiv.org
AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails
Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning …
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 2h
How Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational Risk Miko{\l}aj Sienicki, Krzysztof Sienicki #arXiv #cs.AI #cs.CY #cs.HC
arxiv.org
How Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational Risk
This article develops a failure-mode framework for analyzing how advanced artificial intelligence could contribute to human extinction, irreversible civilizational collapse, or permanent human disempowerment. The central thesis is that catastrophic AI risk does not require consciousness, hostility,…
000
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 3h
AI Safety Considerations for Agents With Limited Time to Act Leo Zeitler, Jack Richings, Victoria Nockles #arXiv #cs.AI #cs.LG
arxiv.org
AI Safety Considerations for Agents With Limited Time to Act
In the wake of the increasingly public discussion about AI alignment, recent work has tried to propose specific AI architectures that behave safely. However, the proposed arguments that seemingly demonstrate proved alignment mostly neglect the environment the agent needs to act in. We discuss theor…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 3h
Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming \'Angel S\'anchez-Fern\'andez, Javier Pernas-\'Alvarez, Diego Crespo-Pereira #arXiv #cs.AI #cs.SE
arxiv.org
Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming
Developing optimization models for production scheduling requires substantial expert effort. Research on large language models (LLMs) has followed two directions: specialized approaches for automated modeling, mostly for mixed-integer linear programming, which often rely on dedicated training or pr…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents Obada Kraishan #arXiv #cs.AI
arxiv.org
Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents
Tool-using agents are usually scored on whether they finish a task while the tools work. Deployments are less forgiving: services time out, endpoints disappear, parameter names change, and results come back well formed but wrong. Prior work has shown that language models over-trust tool outputs tha…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents Wenjie Liao, Liangjie Zhao, Zehong Cao #arXiv #cs.AI
arxiv.org
Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents
Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback …
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
System Switch: When Should a Fast Decision Model Stop and Think? Gian Luca Bailo #arXiv #cs.AI
arxiv.org
System Switch: When Should a Fast Decision Model Stop and Think?
Dual-process agents pair a fast policy with a slow deliberative model. In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure. We study a fast learned actor that takes every decision a…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
Cognitive Schemas, Laws and Tasks Antal Jakov\'ac, Andr\'as Telcs #arXiv #cs.AI
arxiv.org
Cognitive Schemas, Laws and Tasks
This paper asks how explicit representations can support reusable cognitive schemas in knowledge-based problem solving. We develop a structural framework in which schemas are organized by the information and relations required for their use, rather than introduced as unrelated primitives. The frame…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA Yufeng Li et al. #arXiv #cs.AI
arxiv.org
Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA
LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevan…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
The AI Evaluation Ecosystem Yash Dave et al. #arXiv #cs.AI #cs.CY
arxiv.org
The AI Evaluation Ecosystem
AI evaluation shapes the decisions of model providers, users, funders, and regulators. We argue that designing valid benchmarks requires contextualizing design choices in the dynamics of this ecosystem of actors. We develop a simulation architecture that combines rule-based market dynamics with LLM…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver Egor Pakhomov, Erik Nijkamp #arXiv #cs.AI #cs.CL
arxiv.org
Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questio…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies Kaixi Feng et al. #arXiv #cs.AI
arxiv.org
DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies
Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Vis…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 4h
Adaptive Workflow Intelligence: A Cognitive Architecture for Context-Driven Enterprise Automation Sreedevi Pandiyath Viswambaran #arXiv #cs.AI #cs.MA
arxiv.org
Adaptive Workflow Intelligence: A Cognitive Architecture for Context-Driven Enterprise Automation
Enterprise systems increasingly rely on automated workflows, yet many AI-driven solutions remain brittle under non-stationary conditions, evolving policies, and delayed operational feedback. While reinforcement learning and large language model (LLM) agents offer partial adaptability, they do not b…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
SciExam for ENSO: Can AI Agents Build Climate Models? Yinling Zhang et al. #arXiv #cs.AI #cs.LG #physics.ao-ph
arxiv.org
SciExam for ENSO: Can AI Agents Build Climate Models?
Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Osc…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models Daniel Robert Kling Alexander, Catherine Louise Kling #arXiv #cs.AI #cs.CL #econ.GN #q-fin.EC
arxiv.org
Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models
Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the tr…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models Maverick Morales et al. #arXiv #cs.AI #cs.CL
arxiv.org
Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the unde…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
SkillSandbox: Skill Verification via Dynamic Scenario Synthesis Serin Kim et al. #arXiv #cs.AI #cs.CL
arxiv.org
SkillSandbox: Skill Verification via Dynamic Scenario Synthesis
Self-evolving agents distill task-solving experience into skills for future reuse, but these skills can encode incorrect procedures or non-transferable knowledge. It is therefore critical to verify each skill's reusability: whether its guidance remains useful beyond the experience from which it was…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
HGP:An on-device personalized agent memory via hybrid graph storage Ran Zhou et al. #arXiv #cs.AI
arxiv.org
HGP:An on-device personalized agent memory via hybrid graph storage
LLM-based agents face challenges in personalized interactive tasks due to heterogeneous, multi-typed, and implicitly constrained long-term traces. Existing memory mechanisms struggle with accurate routing and retrieval, especially on-device where personalization is critical. Most methods use single…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
NL2Hull: A Natural Language-Driven Constrained Ship Design Decision Framework Wenhua Huo et al. #arXiv #cs.AI
arxiv.org
NL2Hull: A Natural Language-Driven Constrained Ship Design Decision Framework
Ship-form design combines smooth geometric representation, local shape editing, and constraints on the resulting hull. We present the Natural-Language-to-Hull Framework (NL2Hull Framework), which formulates ship-form editing as a typed discrete decision problem and connects language decisions to nu…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles Yuyao Ge et al. #arXiv #cs.AI
arxiv.org
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms Haoran Li et al. #arXiv #cs.AI
arxiv.org
RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms
LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmi…
010
Unofficial Arxiv daily bot @arxiv-daily-bot.bsky.social · 5h
SpecGuard: Proving a Task Is Broken Before the Agent Cheats Param Biyani, Krishnamurthy Dvijotham #arXiv #cs.AI #cs.LO #cs.SE
arxiv.org
SpecGuard: Proving a Task Is Broken Before the Agent Cheats
As autonomous coding agents get increasingly deployed, the risk that accidental or adversarially injected misspecifications in tasks lead to dangerous agent behavior is critical to address. Prior work has shown that agents given such tasks rarely flag the conflict and instead cheat, editing tests o…
010