Sign in

Adrian Chan

@gravity7.bsky.social
799 followers 618 following 407 posts

Bridging IxD, UX, & Gen AI design & theory. Ex Deloitte Digital CX. Stanford '88 IR. Edinburgh, Berlin, SF. Philosophy, Psych, Sociology, Film, Cycling, Guitar, Photog. Linkedin: adrianchan. Web: gravity7.com. Insta, X, medium: @gravity7

PostsRepliesMedia
Adrian Chan @gravity7.bsky.social · 31/03/2026
A recursively self-improving AI: its input is its own output. Its training data is its own generation. Its evaluation is its own judgment. Nothing from outside enters. This isn't exploration. This is confirmation. The most sophisticated echo chamber ever built.
020
Adrian Chan @gravity7.bsky.social · 09/06/2025
Those #LLM reward models like sycophancy even more than you do! Researchers find preferences for verbosity, listicles, vagueness, and jargon even higher among LLM-based reward models (synthetic data) than among us humans. #AI #AIalignment arxiv.org/abs/2506.05339
arxiv.org
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
Language models serve as proxies for human preference judgements in alignment and evaluation, yet they exhibit systematic miscalibration, prioritizing superficial patterns over substantive qualities. ...
040
Adrian Chan @gravity7.bsky.social · 08/06/2025
Everybody talking about the "new" apple paper might find this MLST interview with @rao2z.bsky.social interesting. "Reasoning" and "inner thoughts" of LLMs were exposed as self-mumblings and fumblings long ago. #LLMs #AI www.youtube.com/watch?v=y1Wn...
youtube.com
Do you think that ChatGPT can reason?
YouTube video by Machine Learning Street Talk
050
Adrian Chan @gravity7.bsky.social · 21/05/2025
This is interesting, published yesterday. CoT type reasoning shifts attention away from instruction tokens. Paper proposes "constraint attention" to keep models attentive to instructions when doing CoT. #AI #LLM www.arxiv.org/abs/2505.11423
arxiv.org
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
Reasoning-enhanced large language models (RLLMs), whether explicitly trained for reasoning or prompted via chain-of-thought (CoT), have achieved state-of-the-art performance on many complex reasoning ...
030
Adrian Chan @gravity7.bsky.social · 16/05/2025
"What's the best way to think about this?" #LLM research produces encyclopedia of reasoning strategies, allowing models to select the best way to reason through problems. arxiv.org/abs/2505.10185
arxiv.org
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
Long chain-of-thought (CoT) is an essential ingredient in effective usage of modern large language models, but our understanding of the reasoning strategies underlying these capabilities remains limit...
041
Adrian Chan @gravity7.bsky.social · 14/05/2025
Clarifying questions w #LLMs increase user satisfaction when users can see the point of answering them. Specific questions beat generic ones. But I wonder if this changes when #agents are personal assistants, & are more personal & more aware. #UX #AI #Design arxiv.org/abs/2402.01934
arxiv.org
Clarifying the Path to User Satisfaction: An Investigation into Clarification Usefulness
Clarifying questions are an integral component of modern information retrieval systems, directly impacting user satisfaction and overall system performance. Poorly formulated questions can lead to use...
000
Adrian Chan @gravity7.bsky.social · 14/05/2025
Interesting - could #LLMs in search capture context missed when googling? "backtracing ... retrieve the cause of the query from a corpus. ... targets the information need of content creators who wish to improve their content in light of questions from information seekers." arxiv.org/abs/2403.03956
arxiv.org
Backtracing: Retrieving the Cause of the Query
Many online content portals allow users to ask questions to supplement their understanding (e.g., of lectures). While information retrieval (IR) systems may provide answers for such user queries, they...
000
Adrian Chan @gravity7.bsky.social · 14/05/2025
@tedunderwood.me In case you haven't seen this paper, you might find interesting. Researchers extract style vectors (incl from Shakespeare) and apply to an LLM internal layers instead of training on original texts. Generations can then be "steered" to a desired style. arxiv.org/abs/2402.01618
arxiv.org
Style Vectors for Steering Generative Large Language Model
This research explores strategies for steering the output of large language models (LLMs) towards specific styles, such as sentiment, emotion, or writing style, by adding style vectors to the activati...
110
Adrian Chan @gravity7.bsky.social · 14/05/2025
"LLMs tend to (1) generate overly verbose responses, leading them to (2) propose final solutions prematurely in conversation, (3) make incorrect assumptions about underspecified details, and (4) rely too heavily on previous (incorrect) answer attempts." arxiv.org/abs/2505.06120
arxiv.org
LLMs Get Lost In Multi-Turn Conversation
Large Language Models (LLMs) are conversational interfaces. As such, LLMs have the potential to assist their users not only when they can fully specify the task at hand, but also to help them define, ...
100
Adrian Chan @gravity7.bsky.social · 14/05/2025
"LLMs ... recognize graph-structured data... However... we found that even when the topological connection information was randomly shuffled, it had almost no effect on the LLMs’ performance... LLMs did not effectively utilize the correct connectivity information." www.arxiv.org/abs/2505.02130
arxiv.org
Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
Attention mechanisms are critical to the success of large language models (LLMs), driving significant advancements in multiple fields. However, for graph-structured data, which requires emphasis on to...
000
Adrian Chan @gravity7.bsky.social · 12/05/2025
Let's dose an LLM and study its hallucinations! LLMs were fed "blended" prompts, impossible conceptual combinations, meant to elicit hallucinations. Models did not trip, but instead tried to reason their way through their responses. arxiv.org/abs/2505.00557
arxiv.org
Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
Hallucinations in large language models (LLMs) present a growing challenge across real-world applications, from healthcare to law, where factual reliability is essential. Despite advances in alignment...
110
Adrian Chan @gravity7.bsky.social · 04/05/2025
An unsurprisingly sharp deep dive into sycophancy & RLHF. Whilst this episode has its tech explanations, the social interaction aspects are unsolved & will rise as models are personalized, as mirroring is an effective design feature, & we're not good at distinguishing affirmation from agreement.
110
Adrian Chan @gravity7.bsky.social · 04/05/2025
Analogous to honing in on the solution to a problem, this #LLM research draws conclusions from the most common "reasoning trace" & compares to final answer. Overthinking by reasoning models can be a waste of tokens - but which tokens? #AI www.arxiv.org/abs/2504.20708
arxiv.org
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
Large Language Models (LLMs) leverage step-by-step reasoning to solve complex problems. Standard evaluation practice involves generating a complete reasoning trace and assessing the correctness of the...
010
Adrian Chan @gravity7.bsky.social · 03/05/2025
"Opportunistic sycophancy?" Anthropic's discovery of an influence bot network shows how conversation can be weaponized. This isn't the Cambridge Analytica-styled user preference targeting, but language/conversation instead. #AIethics www.anthropic.com/news/detecti...
anthropic.com
Detecting and Countering Malicious Uses of Claude
Detecting and Countering Malicious Uses of Claude
122
Adrian Chan @gravity7.bsky.social · 03/05/2025
Wld b interesting to apply this resrch to multi-turn convos w AI - AI can't anticipate convo as we do, as it can't anticipate "where we are going" - thus the temporal anticipation on AI's side of convo will lack look ahead phrases.
000
Adrian Chan @gravity7.bsky.social · 02/05/2025
This prompted the idea of using mechanistic probing on a model trained on period literature to conduct a digital Foucauldian examination of latent sociocultural features (as evidenced by latent layer features, activitations, circuits etc). ;-) Period models - Black Mirror Hotel Reverie episode?
010
Adrian Chan @gravity7.bsky.social · 02/05/2025
One potential but not primary benefit to personalizing LLMs to user preferences (memory & sycophancy) might be to create more signal for model verifiers & reasoning - but is there a risk then we get individualized AI-generated filter bubbles? Cognitive biases mirrored back to us?
040
Adrian Chan @gravity7.bsky.social · 01/05/2025
Sublime reflections on AI & consciousness here. If an LLM reads Deleuze's Logic of Sense, does Alice still become both smaller & larger at the same time? (No, because AI has neither temporality nor a "sense" for it) #philosophy #LLMs covidianaesthetics.substack.com/p/from-apoph...
covidianaesthetics.substack.com
From Apophasis to Apophenia. A Response to Murray Shanahan
Inmachination #03
000
Reposted by Adrian Chan
Nicole Hennig @nic221.bsky.social · 01/05/2025
When ChatGPT Broke an Entire Field: An Oral History | Quanta Magazine www.quantamagazine.org/when-chatgpt… #AI #history (very interesting)
Text Shot: R. THOMAS MCCOY (assistant professor, department of linguistics, Yale University): During that summer, I vividly remember members of the research team I was on asking, “Should we look into these transformers?” and everyone concluding, “No, they’re clearly just a flash in the pan.”
362
Adrian Chan @gravity7.bsky.social · 25/04/2025
Sleep-time compute - offline reasoning used to speculate on the user's reasoning, and generate and anticipate future queries. And if used by a group, even offer a kind of "collective intelligence" by speculating on the group or team's aggregate queries. #LLM #AI www.alphaxiv.org/abs/2504.13171
alphaxiv.org
Sleep-time Compute: Beyond Inference Scaling at Test-time | alphaXiv
View 1 comments: Really like this idea. I've long thought it strange that LLMs reason on their own reasons but not also on the reasoning of the user. Which is what we do when we communicate — take the...
020
Adrian Chan @gravity7.bsky.social · 25/04/2025
This missive from Dario is worth the read (and mech interp on #AI is truly fascinating). Among features/concepts found in #LLMs: "genres of music that express discontent." I'm reminded of Borges Chinese Encyclopedia of Animals www.darioamodei.com/post/the-urg...
Borges' Chinese Encyclopedia
010
Adrian Chan @gravity7.bsky.social · 23/04/2025
Loved this piece, and reminded of Anthony Giddens' view of tech as "disembedding mechanism" - how soon and how varied will be AI's disembedding of intellect, work, roles, jobs, tasks from contemporary society?
010
Adrian Chan @gravity7.bsky.social · 23/04/2025
Does this suggest a future of tiny reasoning models each with a different "cognitive" architecture (or reasoning methods and applications)?
020
Adrian Chan @gravity7.bsky.social · 21/04/2025
"The Model is the Message. Models don’t talk—they reshape the conditions of thought. Where once the medium shaped the message, now the model shapes the mind." That's what ChatGPT gave me after a lengthy conversation about semiotics, McLuhan, and #LLMs
010
Adrian Chan @gravity7.bsky.social · 19/04/2025
Don't use summarizers for the papers by @rao2z.bsky.social because the reasoning traces therein are, unlike the LRMs & LLMs under investigation, substantively meaningful, semantically well-ordered, and stylistically compelling and engaging! #AI #LLMs #CoT arxiv.org/abs/2504.09762
arxiv.org
(How) Do reasoning models reason?
We will provide a broad unifying perspective on the recent breed of Large Reasoning Models (LRMs) such as OpenAI o1 and DeepSeek R1, including their promise, sources of power, misconceptions and limit...
182
Reposted by Adrian Chan
Dare Obasanjo @carnage4life.bsky.social · 02/04/2025
Another chapter in AI biting the hand that feeds it: Wikipedia’s bandwidth surged 50% since January thanks to AI crawlers. Unlike search engines, they send no traffic back so no new users, no new donors. Just rising costs and a shrinking audience. A raw deal for a cornerstone of the free web.
diff.wikimedia.org
How crawlers impact the operations of the Wikimedia projects
Since the beginning of 2024, the demand for the content created by the Wikimedia volunteer community – especially for the 144 million images, videos, and other files on Wikimedia Commons – has grow…
919953
Adrian Chan @gravity7.bsky.social · 28/03/2025
Fascinating paper on semantic entropy in graph reasoning models w implications that #genAI reasoning continuously explores novel connections. Very interesting implications for #LLMs in research. Not an easy read but you can use Gemini to help if you read it here: www.alphaxiv.org/abs/2503.18852
alphaxiv.org
Self-Organizing Graph Reasoning Evolves into a Critical State for Continuous Discovery Through Structural-Semantic Dynamics | alphaXiv
View 1 comments: Fascinating paper - this suggests to me that working with reasoning models in the future will entail learning how to think differently. To think in ways that can leverage the explorat...
020
Adrian Chan @gravity7.bsky.social · 27/03/2025
I disagree w the premise of this paper that text is a "culmination of a thought process" - what do others think?
131
Reposted by Adrian Chan
Sung Kim @sungkim.bsky.social · 27/03/2025
Microsoft canceling dara center leases, then this. China built hundreds of AI data centers to catch the AI boom. Now many stand unused. www.technologyreview.com/2025/03/26/1...
technologyreview.com
China built hundreds of AI data centers to catch the AI boom. Now many stand unused.
The country poured billions into AI infrastructure, but the data center gold rush is unraveling as speculative investments collide with weak demand and DeepSeek shifts AI trends.
3224
Adrian Chan @gravity7.bsky.social · 18/03/2025
Adolescence. #Fiilmsky
010
Adrian Chan @gravity7.bsky.social · 15/03/2025
I used to say that to learn how a social technology works, and what it does, turn it off. This seems now to be a principle mobilized against government also.
100
Adrian Chan @gravity7.bsky.social · 14/03/2025
Beautiful work. Now words for animal sounds are also communicated verbally and in writing between people - I wonder how much impact this also has on representing animal sounds (woof in English, wau in German, both use w but w is pronounced as a v in German)
010
Reposted by Adrian Chan
Nathan Lambert @natolambert.bsky.social · 10/03/2025
The most accessible article I’ve written on understanding how very rapid improvements to models have been happing with “post-training.” The upshot of it is that scaling still matters and the new forms of RL training make us deeply appreciate the value of a strong base model.
buff.ly
Elicitation, the simplest way to understand post-training
An F1 analogy to help understand fast improvements in post-training on top of slow improvements in scaling.
1225
Adrian Chan @gravity7.bsky.social · 13/03/2025
There should be a taxonomy of hallucinations, as all are not equal. Factual inaccuracies, inventions, errors of logic, alignment/sycophancy-based, etc. We know the technical causes of hallucinations, but they don't result in meaning-based (interpretation) descriptions of linguistic fallacies.
120
Reposted by Adrian Chan
Maria Antoniak @mariaa.bsky.social · 28/02/2025
Required reading If you ever use models to brainstorm. Really nice study design: Domain expert researchers search for plagiarized papers of LLM- systems that generate research plans. Many successfully find papers, verified by authors of those papers. But very hard to detect, see screenshot.
Figure 2 from the paper showing a one-to-one mapping between methods generated by a model and methods from a real paper. The mapping is exact by the language used is very different.
15112
Adrian Chan @gravity7.bsky.social · 24/02/2025
The shift from "task automation" to "capability augmentation" vis-a-vis #GenAI has great implications for design - org design, product design, #UX, #IxD, #LLM design. Because it becomes about information > knowledge > choices > recommendations > decisions > communication.
020
Adrian Chan @gravity7.bsky.social · 23/02/2025
Enshittification now seems such a quaint and minor irritant...
000
Adrian Chan @gravity7.bsky.social · 20/02/2025
Compare this view of an LLM diffusion model generating its response to the "reasoning" we see in conventional LLMs. This view illustrates the degree to which seeing an AI "think" step by step sustains an illusion that it's actually thinking. Really it's just choosing its words carefully. #LLM #AI
https://www.alphaxiv.org/abs/2502.09992
020
Adrian Chan @gravity7.bsky.social · 19/02/2025
Interesting challenge to CoT: "implicit learning ... outperforms both explicit learning and chain-of-thought prompting... findings ... call into question the benefits of widely-used corrective rationales to aid LLMs in learning from mistakes." #LLM #ML #AI www.alphaxiv.org/abs/2502.08550
alphaxiv.org
LLMs can implicitly learn from mistakes in-context | alphaXiv
View recent discussion. Abstract: Learning from mistakes is a fundamental feature of human intelligence. Previous work has shown that Large Language Models (LLMs) can also learn from incorrect answers...
010
Adrian Chan @gravity7.bsky.social · 19/02/2025
Multi-agent reasoning is overrated, according to this paper: www.alphaxiv.org/abs/2502.08788 #LLM #ML #AI
alphaxiv.org
If Multi-Agent Debate is the Answer, What is the Question? | alphaXiv
View 4 comments: With regards to the Multi-Persona (MP) design, is there a common trend across the results of the approach in which a certain agent (between the figurative "angel" and "devil" agent) i...
010
Adrian Chan @gravity7.bsky.social · 18/02/2025
If you’re not part of the solution you’re part of the precipitate – DFW
020
Adrian Chan @gravity7.bsky.social · 06/02/2025
To think we believed deep fakes were the greatest threat to civility.
000
Reposted by Adrian Chan
Zeke Hausfather @zekehausfather.com · 05/02/2025
Effectively paralyzing all new solar and wind development – even on private land – seems pretty contrary to conservative deregulation and free market values and non-sensical given our declared "energy emergency". This is bad: heatmap.news/plus/th...
heatmap.news
Trump Has Paralyzed Renewables Permitting, Leaked Memo Reveals
The American Clean Power Association wrote to its members about federal guidance that has been “widely variable and changing quickly.”
12348129
Adrian Chan @gravity7.bsky.social · 05/02/2025
A grand proliferation of #AI & #LLM thinking strategies at inference is underway. Here's Chain of Action of Thought (COAT), which could also be "trial and error," or "internalized search," or, perhaps, just brainstorming. #ML #CoT www.alphaxiv.org/abs/2502.02508
alphaxiv.org
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search | alphaXiv
View recent discussion. Abstract: Large language models (LLMs) have demonstrated remarkable reasoning capabilities across diverse domains. Recent studies have shown that increasing test-time computati...
000
Adrian Chan @gravity7.bsky.social · 05/02/2025
"significant decline in accuracy as problem complexity grows—a phenomenon we term the “curse of complexity....” persists even with larger models and increased inference-time computation, suggesting inherent constraints in current #LLM reasoning capabilities." #AI #ML www.arxiv.org/abs/2502.01100
arxiv.org
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
We investigate the logical reasoning capabilities of large language models (LLMs) and their scalability in complex non-monotonic reasoning. To this end, we introduce ZebraLogic, a comprehensive evalua...
000
Reposted by Adrian Chan
Mor Naaman @informor.bsky.social · 04/02/2025
You are not going to get a more authoritative technical view of the DeepSeek story than Sasha's.
0102
Reposted by Adrian Chan
AI Firehose @ai-firehose.column.social · 05/02/2025
A groundbreaking study presents MODS, a multi-LLM framework enhancing debatable query-focused summarization by treating documents as discussion speakers. This method boosts summary balance and coverage by up to 58%, facilitating unbiased exploration of controversies. arxiv.org/abs/2502.00322
arxiv.org
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
ArXiv link for MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
001
Adrian Chan @gravity7.bsky.social · 03/02/2025
Intermediate reasoning tokens are not evaluated, @rao2z.bsky.social reminds us. In human reasoning, our reasons do "add up" to our conclusions & decisions. The anthropomorphism of exposed "thoughts" is more deceptive BUT also more explanatory. #LLM #AI #ML x.com/rao2z/status...
x.com
x.com
010
Adrian Chan @gravity7.bsky.social · 02/02/2025
Are the reasonings of LLMs trust-building for AI? Hmm. Let's think this through. Well, if the reasons are solid, then.... I really think #LLM reasoning is #UX and even #IxD if it's used to prompt the user w questions etc. In a chat agent, it's "inner thoughts." ...
110