Sign in

Salesforce AI Research

@sfresearch.bsky.social
167 followers 57 following 215 posts

We advance state-of-the-art #AI techniques paving the path for innovative products at @Salesforce.com. Focus areas: #AIAgents, #EnterpriseAI, #EGI, and #TrustedAI.

PostsRepliesMedia
Salesforce AI Research @sfresearch.bsky.social · 01/10/2026
(5/5) Fractured Chain-of-Thought Reasoning: examines how disruptions in chain-of-thought processes affect reasoning in language models. arxiv.org/abs/2505.12992 #COLM2026
000
Salesforce AI Research @sfresearch.bsky.social · 01/10/2026
(4/5) RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation: grounds agent benchmarking in realistic user simulation. arxiv.org/abs/2605.20204 #COLM2026
100
Salesforce AI Research @sfresearch.bsky.social · 01/10/2026
(3/5) InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation: introduces a scalable framework for simulating personalities grounded in real interviews. arxiv.org/abs/2602.20294 #COLM2026
100
Salesforce AI Research @sfresearch.bsky.social · 01/10/2026
(2/5) MTA-Agent: An Open Recipe for Multimodal Deep Search Agents: presents an open recipe for building multimodal deep search agents. arxiv.org/abs/2604.06376 #COLM2026
100
Salesforce AI Research @sfresearch.bsky.social · 01/10/2026
(1/5) We are pleased to announce our participation in COLM 2026, the Third Annual Conference on Language Modeling, at the Hilton Union Square in San Francisco, October 6–9. Our researchers will present 4 accepted papers. Full list below ⬇️ #COLM2026
100
Salesforce AI Research @sfresearch.bsky.social · 08/09/2026
(3/3) Evidence-Backed Video QA asks video models to show their work: an answer paired with precise spatio-temporal evidence, tracked masks, not just text. We introduce ST-Evidence, plus a 160k-scale training set. arxiv.org/abs/2607.11862 #FutureOfAI #EnterpriseAI #ECCV2026
arxiv.org
Evidence-Backed Video Question Answering
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainabi...
021
Salesforce AI Research @sfresearch.bsky.social · 08/09/2026
(2/3) DreamHouse asks whether vision-language models can build the real world, not just picture it. Our new benchmark grounds physical generative reasoning in timber-frame construction. arxiv.org/abs/2603.24866 #ECCV2026
arxiv.org
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
The physical world is not merely visual; it is governed by rigorous structural and procedural constraints. Yet, the evaluation of vision-language models (VLMs) remains heavily skewed toward perceptual...
110
Salesforce AI Research @sfresearch.bsky.social · 08/09/2026
🧵 (1/3) We're heading to #ECCV2026 in Malmö, Sweden. 🇸🇪 This week, our team will present two new papers spanning physical generative reasoning and video question answering. More below ⬇️
121
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(13/13) UserRL: Training Interactive User-Centric Agent via Reinforcement Learning: a unified RL framework with eight gym environments for training agents to better assist users in multi-turn interactions. arxiv.org/abs/2509.19736 #EMNLP2026
arxiv.org
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
Reinforcement learning (RL) has shown promise in training agentic models that move beyond static benchmarks to engage in dynamic, multi-turn interactions. Yet, the ultimate value of such agents lies i...
010
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(12/13) UserBench: An Interactive Gym Environment for User-Centric Agents: a gym environment testing whether LLM agents can uncover and align with evolving user preferences, not just complete tasks. arxiv.org/abs/2507.22034 #EMNLP2026
arxiv.org
UserBench: An Interactive Gym Environment for User-Centric Agents
Large Language Models (LLMs)-based agents have made impressive progress in reasoning and tool use, enabling them to solve complex tasks. However, their ability to proactively collaborate with users, e...
100
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(11/13) 🎲 Agentic Uncertainty Quantification: presents methods for quantifying uncertainty in multi-step agentic reasoning. arxiv.org/abs/2601.15703 Authors: Jiaxin Zhang, Prafulla Kumar Choubey, Kung-Hsiang Huang, Caiming Xiong, Chien-Sheng Wu #EMNLP2026
arxiv.org
Agentic Uncertainty Quantification
Although AI agents have demonstrated impressive capabilities in long-horizon reasoning, their reliability is severely hampered by the ``Spiral of Hallucination,'' where early epistemic errors propagat...
100
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(10/13) 🌱 OpenSkill: Open-World Self-Evolution for LLM Agents: introduces a framework for LLM agents to self-evolve skills in open-world settings. arxiv.org/abs/2606.06741 Authors: Yan, Song, Zhang, Liang, Zhang, Dai, He, Yu, Xu, Li, Sun #EMNLP2026
arxiv.org
OpenSkill: Open-World Self-Evolution for LLM Agents
Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful trajectories, or verifier signals. Real open-world ...
110
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(9/13) 📄 UniDoc-Bench: A Unified Benchmark for Document-Centric Multimodal RAG (Findings): a unified benchmark for retrieval-augmented generation over multimodal documents. arxiv.org/abs/2510.03663 Authors: Xiangyu Peng, Can Qin, Zeyuan Chen, Ran Xu, Caiming Xiong, Chien-Sheng Wu #EMNLP2026
arxiv.org
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
Multimodal retrieval-augmented Generation (MM-RAG) is a key approach for applying large language models (LLMs) and agents to real-world knowledge bases, yet current evaluations are fragmented -- focus...
110
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(8/13) Seeing is Believing? New MMPersuade benchmark: multimodal inputs make VLMs more persuadable than text alone, with visual cues often bypassing safety defenses. Needs image-aware guardrails. 🛡️ arxiv.org/abs/2510.22768 Authors: Qiu, Venkit, Huang, Zhang, Peng, Wu #EMNLP2026
arxiv.org
Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion
As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored the dynamics of Agent-to-Agent (A2A) persuasion, the r...
120
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(7/13) ✅ TEX: Test-Time Scaling Testing Agents via Execution-based Cross-Validation (Findings): proposes execution-based cross-validation to scale testing agents at inference time. sforce.co/4dhzzkL Authors: Srijan Bansal, Akash Gokul, Yingbo Zhou, Semih Yavuz #EMNLP2026
sforce.co
TEX: Test-Time Scaling Testing Agents via Execution-based Cross-Validation
The era of software engineering agents is underway. Benchmarks and real-world usage (e.g., tools like Cursor and Claude Code) illustrate that LLMs can be incredibly effective at writing code for real-...
100
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(6/13) ⚖️ Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity: studies calibration and bias failures when multimodal LLMs judge culturally ambiguous content. arxiv.org/abs/2606.20676 Authors: Lee, Sharma, Park, Venkit, Kim, Chia, Vlachos, Joty #EMNLP2026
arxiv.org
Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity
MLLM-as-a-Judge is conventionally validated by agreement with human annotations, but this metric is undefined when the human pool is culturally heterogeneous. We introduce VOIR DIRE, a multimodal benc...
100
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(5/13) 🐛 VIBEPASS: Can Vibe Coders Really Pass the Vibe Check? An Empirical Analysis of Fault-Targeted Reasoning in Frontier LLMs (Findings): examines how well frontier LLMs reason about and localize code faults. arxiv.org/abs/2603.15921 Authors: Bansal, Fangkai, Zhou, Xu, Joty, Yavuz #EMNLP2026
arxiv.org
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
As Large Language Models shift the programming toward human-guided ''vibe coding'', agentic coding tools increasingly rely on models to self-diagnose and repair their own subtle faults -- a capability...
110
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(4/13) VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?: evaluates whether vision-language models can revise visualization code from multimodal feedback. arxiv.org/abs/2608.10408 Authors: Rahman, Azimlu, Rahman, Laskar, Bhuiyan, Joty, Prince #EMNLP2026
arxiv.org
VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative...
100
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(3/13) 📊 DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?: a benchmark testing whether agents can complete real, end-to-end data science workflows. arxiv.org/abs/2608.10366 Authors: Rahman, Islam, Mahbub, Laskar, Joty, Prince #EMNLP2026
arxiv.org
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, te...
100
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(2/13) 🤖 Harnessing LLM Agents with Skill Programs: introduces reusable skill programs that let LLM agents plan and execute complex tasks more reliably. arxiv.org/abs/2605.17734 Authors: Hongjun Liu, Yifei Ming, Shafiq Joty, Chen Zhao #EMNLP2026
arxiv.org
Harnessing LLM Agents with Skill Programs
Equipping LLM agents with reusable skills derived from past experience has become a popular and successful approach for tackling complex and long-horizon tasks. However, such lessons are often encoded...
100
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
(1/13) 🎉 We are pleased to announce our participation in #EMNLP2026, the Conference on Empirical Methods in Natural Language Processing, Budapest, Hungary, October 24–29. 🇭🇺 Our researchers will present 12 accepted papers spanning agentic reasoning, evaluation, and multimodal AI. Full list below ⬇️
100
Salesforce AI Research @sfresearch.bsky.social · 02/09/2026
Traditional monitoring tells you what happened last week; #OperationalIntelligence tells you what's happening now and why. Our EVP & Chief Scientist Silvio Savarese breaks down the shift to continuous, proactive systems that surface what matters before it becomes a problem. sforce.co/4cquhTH
sforce.co
Operational Intelligence: Turning Enterprise Data Into Enterprise Decisions with AI
Silvio Savarese on operational intelligence: the AI layer that makes enterprise signal legible, so leaders can act before the moment passes.
110
Salesforce AI Research @sfresearch.bsky.social · 10/07/2026
What happens when AI agents start negotiating with each other, and who's responsible when they do? Silvio Savarese and Sabastian Niles break down the New AI Trust Architecture: 5 requirements for safe, accountable agent-to-agent communication. sforce.co/4gwgT3p #AgenticAI #AItrust #AIgovernance
salesforce.com
The New AI Trust Architecture: 5 Requirements for Agent-to-Agent Communication
5 requirements for trustworthy AI agent-to-agent communication, from Salesforce's Chief Scientist and Chief Legal Officer. Standards, identity, and accountability, explained.
010
Salesforce AI Research @sfresearch.bsky.social · 09/06/2026
(5/5) Live now in the model cards for First Name Match, Account Match, and TextEval, on the Salesforce Trust site: sforce.co/4eaIwMu #ResponsibleAI #Sustainability #FutureOfAI
sforce.co
Salesforce-owned Model Cards - Salesforce Compliance Site
Browse documents within the category: Salesforce-owned Model Cards.
030
Salesforce AI Research @sfresearch.bsky.social · 09/06/2026
(4/5) For customers, that means models with similar performance but very different carbon footprints become easy to compare, so sustainability can factor into which model they choose.
100
Salesforce AI Research @sfresearch.bsky.social · 09/06/2026
(3/5) The disclosures estimate energy use and carbon emissions across three lifecycle stages: pre-training, post-training, and inference, using the AI Energy Score methodology, an industry framework Salesforce helped develop. 🔋
100
Salesforce AI Research @sfresearch.bsky.social · 09/06/2026
(2/5) @salesforce.com AI Research worked with the Impact team to build these estimates into the standard model evaluation workflow, so a model's environmental footprint is measured on the same footing as performance and risk.
100
Salesforce AI Research @sfresearch.bsky.social · 09/06/2026
(1/5) Model cards are nutrition labels for AI. Now they include environmental impact. 🌱 @Salesforce.com is adding standardized energy + carbon metrics: sforce.co/4umu8qm
sforce.co
Measuring AI’s Environmental Impact: How We’re Operationalizing Transparency Through Model Cards
Today, Salesforce is expanding its AI model cards with standardized environmental impact metrics. This update helps customers better understand the energy
100
Salesforce AI Research @sfresearch.bsky.social · 03/06/2026
(8/8) AuthorS: Ye Liu, Srijan Bansal, Bo Pang, Yang Li, Zeyu Leo Liu, Yifei Ming, Zixuan Ke, Shafiq Joty, and Semih Yavuz. 📝 Blog: sforce.co/4dAjQOu
sforce.co
Can Language Models Remember What They Learn?
Post-training methods (RLVR, On-policy distillation) are Episode-local Language models are getting better at learning from feedback during post-training. In reinforcement learning with verifiable rewa...
020
Salesforce AI Research @sfresearch.bsky.social · 03/06/2026
(7/8) PMD doesn't give models a permanent notebook. It lets them use one while learning, absorb the useful lessons into their weights, and move on. Every training step contains signals about what works and what fails. Most methods throw it away. 📄 Paper: sforce.co/4dXlVTE
sforce.co
Procedural Memory Distillation
Online Reflection for Self-Improving Language Models
120
Salesforce AI Research @sfresearch.bsky.social · 03/06/2026
(6/8) Memory transfers across scales. Memory distilled from Qwen3-8B improves models from 1.7B to 32B, and more retrieved memory means better performance. As rollout budgets grow, PMD keeps improving while SDPO saturates.
110
Salesforce AI Research @sfresearch.bsky.social · 03/06/2026
(5/8) The full loop matters. ❌ Freeze policy, evolve memory: model never internalizes it ❌ Freeze memory, evolve policy: memory goes stale ✅ Co-evolve both: consistent gains Experience becomes memory. Memory updates the learner. Break either link, the cycle weakens.
110
Salesforce AI Research @sfresearch.bsky.social · 03/06/2026
(4/8) Against the strongest self-distillation baseline (SDPO): Qwen3-8B 📐 SciKnowEval: 74.4 → 77.2 💻 LiveCodeBench: 47.9 → 51.7 OLMo3-7B 📐 Science: 69.5 → 73.3 💻 Code: 45.0 → 51.1 Same backbone. Only difference: memory.
110
Salesforce AI Research @sfresearch.bsky.social · 03/06/2026
(3/8) The loop: attempt → verify → store → reflect → train a memory-conditioned self-teacher → repeat. Policy and memory co-evolve. Every iteration, the teacher gets smarter, built from the model's own history.
110
Salesforce AI Research @sfresearch.bsky.social · 03/06/2026
(2/8) PMD organizes experience into 3 levels: 📝 Experience: trajectories, wins, losses, feedback 💡 Insight: strategies from wins, lessons from failures 🧠 Behavior: cross-task patterns distilled into reusable skills Memory is used only in training. No retrieval at inference.
110
Salesforce AI Research @sfresearch.bsky.social · 03/06/2026
(1/8) Can Language Models Remember What They Learn? LLMs learn from feedback. But most post-training is amnesiac: rollout → reward → update → forget. What if you keep the signal? Procedural Memory Distillation (PMD): learning from experience, not just feedback. 🧵
110
Salesforce AI Research @sfresearch.bsky.social · 02/06/2026
6th Multimodal Algorithmic Reasoning Workshop at #CVPR2026 is Thursday (6/4), 8:55 AM–12:30 PM MDT, Room 601, Colorado Convention Center. 👥 Keynote by Juan Carlos Niebles @jcniebles.bsky.social, organized by Honglu Zhou @hongluzhou.bsky.social 📆 sforce.co/4ueOT7j
011
Salesforce AI Research @sfresearch.bsky.social · 28/05/2026
Can Language Models Remember What They Learn? Introducing Procedural Memory Distillation (PMD): sforce.co/4dAjQOu PMD turns model attempts into reusable training memory, conditions a self-teacher on it, and distills the guidance into the student's weights.
sforce.co
Can Language Models Remember What They Learn?
Post-training methods (RLVR, On-policy distillation) are Episode-local Language models are getting better at learning from feedback during post-training. In reinforcement learning with verifiable rewa...
020
Salesforce AI Research @sfresearch.bsky.social · 26/05/2026
5/5 Authors: The broader AI takeaway: showing a capability in a model's reasoning trace does not mean it can actually deploy it causally in sequential interaction. Romain Cosentino, Sarath Shekkizhar, Adam Earle, Silvio Savarese #FutureOfAI #EnterpriseAI #AgenticAI
010
Salesforce AI Research @sfresearch.bsky.social · 26/05/2026
4/5 Because agents fail to leverage the underlying utility structure, negotiations devolve into anchor-following. And forcing agents to articulate explicit give/ask trade plans before each offer didn't improve efficiency. The missing capability is reciprocal multi-turn strategy.
110
Salesforce AI Research @sfresearch.bsky.social · 26/05/2026
3/5 This creates a new failure mode: more information can actually help the other side. When sellers were given the buyer's preferences, they didn't extract more value. Instead, buyer utility rose while seller utility fell. Sellers accommodated without extracting compensating gains.
110
Salesforce AI Research @sfresearch.bsky.social · 26/05/2026
2/5 The problem isn't reading the room. Informed agents rapidly form accurate early beliefs about the opponent's priorities. The breakdown happens downstream: they fail to convert this social understanding into multi-turn strategic execution.
110
Salesforce AI Research @sfresearch.bsky.social · 26/05/2026
📣 1/5 Counterparty Modeling is Not Strategy: The Limits of LLM Negotiators sforce.co/3RTTU7Q New research tests whether LLM agents can turn knowledge of a partner's preferences into better deals. The finding? LLMs can know what the other side wants without bargaining strategically on it.
sforce.co
Counterparty Modeling is Not Strategy: The Limits of LLM Negotiators
Negotiation requires more than inferring what the other side wants: it requires using that information to make advantageous offers and counteroffers over multiple turns. We study whether large languag...
112
Salesforce AI Research @sfresearch.bsky.social · 22/05/2026
5/5 Binary rewards leave information on the table. Teaching models to interpret why they failed, not just that they failed, unlocks a complementary learning signal. Paper: sforce.co/4uv2f0k Authors: Yang Li, Erik Nijkamp, Semih Yavuz, Shafiq Joty
sforce.co
Learning from Language Feedback via Variational Policy Distillation
Reinforcement learning from verifiable rewards (RLVR) suffers from sparse outcome signals, creating severe exploration bottlenecks on complex reasoning tasks. Recent on-policy self-distillation method...
010
Salesforce AI Research @sfresearch.bsky.social · 22/05/2026
4/5 Tested on 3 model families (Qwen3-4B/8B, Llama-3.1-8B) across code generation (LiveCodeBench) and scientific reasoning (SciKnowEval): ✅ Consistent gains over GRPO and self-distillation baselines ✅ Stable training where prior methods collapse ✅ Best gains on domains with rich error signals
110
Salesforce AI Research @sfresearch.bsky.social · 22/05/2026
3/5 VPD frames this as variational EM with a co-evolving teacher and student sharing one network: ➡️ E-step: train the model to interpret feedback (compiler errors, critiques, self-correction) ➡️ M-step: distill that understanding into the student on its rollouts No extra model, no extra GPU memory.
100
Salesforce AI Research @sfresearch.bsky.social · 22/05/2026
2/5 When a coding solution fails 1 test out of 50, RLVR treats it the same as random garbage. The compiler error tells you exactly what went wrong, but the reward is just 0. On hard problems, nearly all rollouts fail → zero gradient signal → the model stops improving.
100
Salesforce AI Research @sfresearch.bsky.social · 22/05/2026
1/5 RLVR trains LLMs with pass/fail rewards — but every near-miss rollout is wasted. What if models could actually *learn* from their mistakes? New paper: "Learning from Language Feedback via Variational Policy Distillation" Read: sforce.co/4uv2f0k 🧵👇
111
Salesforce AI Research @sfresearch.bsky.social · 20/05/2026
When AI Becomes Invisible: The Rise of Ambient Intelligence Silvio Savarese on the shift from AI as a tool you go to, to a presence already there: always-on, aware, adaptive, and anticipatory. sforce.co/3Pwo68s #FutureOfAI #EnterpriseAI #AmbientIntelligence
sforce.co
When AI Becomes Invisible: The Rise of Ambient Intelligence
A perspective on the future of enterprise AI that is ambient: context-aware, always on, and invisible to users.
000
Salesforce AI Research @sfresearch.bsky.social · 18/05/2026
Nvidia has adopted our xLAM function calling dataset in NeMo Gym, their new library for building RL training environments for LLM agents. Check out our function calling dataset: sforce.co/4uZkL0P And see how it works in Nemo Gym: bit.ly/4tuu0pg #AgenticAI #ReinforcementLearning
010