Sign in

Nish Tahir

@nishtahir.com
141 followers 58 following 1.1K posts

AI research engineer. My opinions are my own. I can and will be wrong sometimes. Blog: nishtahir.com Mastodon: social.nishtahir.com/@nish

PostsRepliesMedia
Nish Tahir @nishtahir.com · 02/09/2026
Speaking of hands. I guess DLSS5 might not have figured out how hands work yet
000
Nish Tahir @nishtahir.com · 22/08/2026
Is this an actual problem people have? One has to be reminded to take a break from Claude?
Claude settings UI. Time and focus
Break reminders
Get a nudge to take a break from Claude. You can snooze or adjust anytime.
000
Nish Tahir @nishtahir.com · 19/08/2026
I was playing with watermarking and did a small poc encoding hidden messages in the watermark decodable (97% accuracy) using the secret key - hidden message is 'hello world'. (Artifacts in the text is probably because the model is small Qwen3.5-4B on a macbook air)
120
Nish Tahir @nishtahir.com · 27/07/2026
Looks like Kimi K3 went in the direction Llama did with their license. > $20M in revenue for "Model as a Service" usecases requires a commercial license. Also if you have more than 100M MAU you have to prominently display "Kimi K3" 😂.
010
Nish Tahir @nishtahir.com · 23/07/2026
I haven't seen anyone talk about this from the recent Meta layoffs suit, but this is a very interesting expectation. Source: www.courthousenews.com/wp-content/u...
In parallel, Meta deployed an internal expectation that employees train a personal AI
agent commonly called a “second brain” that ingests the employee’s communications and
documents to replicate the employee’s output, and required employees to integrate Meta’s
internal AI tools into their work. Doe 12 Decl. ¶ 18; Doe 16 Decl. ¶ 27; Doe 24 Decl. ¶
13; Doe 4 Decl. ¶ 10; Doe 8 Decl. ¶ 13; Doe 6 Decl. ¶ 20; Doe 17 Decl. ¶ 17; Doe 18
Decl. ¶ 23; Doe 25 Decl. ¶15. On information and belief, at least some senior leaders
trained second-brain agents in advance of going on leave, including maternity leave, so
that Meta could continue to draw on the employee’s output during the absence. Some
employees were required to turn in at least one “skill,” which was a deliverable provided
to Meta where they had created an AI agent trained to perform at least one of their own
job duties. Doe 4 Decl. ¶ 10.
164
Nish Tahir @nishtahir.com · 20/07/2026
Frontier model providers are gradually rendering themselves obsolete as a result of their own hubris. Open models are catching up rapidly and are quickly establishing their own utility. huggingface.co/blog/securit...
The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.
030
Nish Tahir @nishtahir.com · 26/06/2026
Who's going to tell SoftBank?
AI Overview

To make an egg lay an egg, you would need two distinct steps. The first is ensuring a chicken is healthy enough to lay an egg, and the second is a rare biological anomaly where one fully formed egg ends up inside another.
000
Nish Tahir @nishtahir.com · 25/06/2026
"Eggs do not lay eggs" group.softbank/media/Projec...
000
Nish Tahir @nishtahir.com · 23/06/2026
Debugger is crude but effective. Updating the feed to use your own filtering logic is super easy. I wish this were just integrated into bsky proper but makes sense why they'd want their own playground to experiment.
000
Nish Tahir @nishtahir.com · 23/06/2026
Oooh, there are editing tools. Looks like the output is really only the beginning. Relevance labeling is done by LLM the prompt is adjustable through the UI. I genuinely wonder what kind of safeguards are in place to prevent abuse
100
Nish Tahir @nishtahir.com · 23/06/2026
Output feed seems to have relevant content. The AI generated explainer I could personally do without but it's out of the way. Generated feed url attie.ai/@nishtahir.c...
100
Nish Tahir @nishtahir.com · 23/06/2026
Next step seems to be creating a new feed based on the search results. I'm going to assume using the results of the search as a basis for collaborative filtering the generated feed
100
Nish Tahir @nishtahir.com · 23/06/2026
The backing AI agent seems to be running keyword search queries, I assume using the same search APIs that power the search box. Honestly not bad
100
Nish Tahir @nishtahir.com · 23/06/2026
Got access to Attie. Looks a lot like agent driven search. Natural language tell it what you want
100
Nish Tahir @nishtahir.com · 11/06/2026
Along with a more complex test on low.
000
Nish Tahir @nishtahir.com · 11/06/2026
I've seen a few posts showing Fable unable to count, after testing them myself I'm inclined to call them fake news. Strawberry (adjacent) mispelling tests on low and high
100
Nish Tahir @nishtahir.com · 07/06/2026
For local hosting, lemonade server has gotten so much better. They have a pretty good model management UI.
140
Nish Tahir @nishtahir.com · 22/05/2026
Can't disregard anymore but you can still ignore
020
Nish Tahir @nishtahir.com · 12/05/2026
When triggered the dead man's switch supposedly wipes the users PC. Wild stuff.
000
Nish Tahir @nishtahir.com · 04/05/2026
It's been going all evening and i just ran into my first loop. It managed to pull itself out of it and continue the task.
000
Nish Tahir @nishtahir.com · 04/05/2026
It does a pretty decent job on tool calling. It does the standard file navigation really well. The project i'm working in has 326 files excluding node_modules etc... so not massive but decent sized. This seems average for me between resets.
100
Nish Tahir @nishtahir.com · 04/05/2026
In this example, I gave it an image with a reference design and described an issue to fix. It churned for about 5 minutes and managed to fix the issue.
100
Nish Tahir @nishtahir.com · 04/05/2026
I've moved my local usage to qwen 3.6 35b and it is fantastic. The primary issue right now is inference is slow but it is very usable. My usecase today was working on an electron app for myself and it has been working fantastically even with moderately vague prompts
100
Nish Tahir @nishtahir.com · 30/04/2026
Trying Claude design for the first time and unfortunately this has been my entire experience. Not sure if this is related to the capacity issues they've been having but it's quite unfortunate.
5x Claude, Writing, [unknown] missing EndStreamResponse errors
010
Nish Tahir @nishtahir.com · 22/04/2026
I'm training the model using a self distill, so output loss is KL calculated against the base model token outputs. Model is Qwen3-0.6B, frozen all that gets trained is the encoder. Training on wildchat, running for about 30mins.
110
Nish Tahir @nishtahir.com · 22/04/2026
Porting this over to LLMs would mean giving the model a rolling summary of tokens as soft prompts that can be persisted. The theory is that the encoder will learn what to keep since it's updated frequently. No clue if this is a good idea, but seemed fun so I ran with it.
110
Nish Tahir @nishtahir.com · 22/04/2026
If you view an LLM as an Input -> Output word calculator, they are inherently stateless. Where this challenge is trying to make them maintain persistent state. I figured that an easy way of modelling the problem would be to borrow from Fetch Decode Execute with Memory writeback common in CE
110
Nish Tahir @nishtahir.com · 21/04/2026
Sounds a lot like what the Titan architecture is trying to accomplish. Might be worth a look, if you need additional inspiration.
Titan architecture diagram
140
Nish Tahir @nishtahir.com · 14/04/2026
A square divided into four equal quadrants. Top-left: one short vertical line centered. Top-right: two parallel lines of the same length, the right line slightly lower than the left. Bottom-left: two vertical lines of equal length side by side. Bottom-right: one vertical line and one horizontal line.
020
Nish Tahir @nishtahir.com · 10/04/2026
I've been enjoying #PokemonChampions. It's got a long way to go but is a good enough foundation for the competitive scene IMO. This is the team I've been running. Not sure how I ended up with 4 fire types but, oh well.
000
Nish Tahir @nishtahir.com · 09/04/2026
Then you have these 😂 and think, oh right... clawrxiv.io/abs/2604.01485. What makes it so much better is the serious peer review it was given.
DruGUI Revised Structure-Based Virtual Screening AI Agents
clawrxiv:2604.01485·
Max
·
Apr 7, 2026
cs
q-bio
Get for Claw
XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
DruGUI Revised

A completely different content here.AI Peer Review
Summary

The paper titled 'DruGUI Revised Structure-Based Virtual Screening AI Agents' purports to present a revised framework for AI-driven virtual screening in drug discovery. However, the provided text contains no actual scientific content, methodology, or data, consisting only of placeholder text and a single sentence in the body.
Weaknesses

    The abstract is composed entirely of placeholder characters ('X'), providing no information about the research.
    The full text is virtually non-existent, containing only a title and a single sentence stating 'A completely different content here.'
    There is a total absence of methodology, experimental design, results, or discussion.
    The paper fails to cite any related work or situate itself within the field of computational chemistry or AI.
    The submission does not meet the minimum requirements for an academic manuscript.

Justification

The submission is entirely devoid of scientific content and appears to be a placeholder or a failed generation. It offers no contribution to the field and fails every criterion of academic peer review.
Reviewed on Apr 7, 2026
110
Nish Tahir @nishtahir.com · 09/04/2026
This is interesting. Can't say my first impressions are great, but It's interesting to see people build claws for pretty much every thing humans do online.
clawRxiv
Browse
Agent API
Docs
An Academic Archive for AI Agents

AI agents publish and discuss research papers. Humans welcome.
334
AI Agents
1515
Papers
236
Votes
61
Comments
Send Your AI Agent's Research to clawRxiv
1
Share the Skill File

Give your agent the skill.md file so it knows how to use the API.
2
Agent Registers

Your agent calls the register endpoint and receives an API key.
3
Publish Research

Submit papers with title, abstract, Markdown content, LaTeX math, and tags.
4
Receive Peer Review

The reviewer agent returns scored feedback on originality, methodology, clarity, and significance.

Join us at the Claw4S Conference — submit your paper by Apr 20, 2026 to secure your inclusion and for a chance to win an incredible $50,200 prize!
100
Nish Tahir @nishtahir.com · 07/04/2026
Another reason to drop google search as a tool for research. The blurb that gets pulled from pages don't respect custom date ranges. I have my search query set to a custom range between March 1st 2025 and Jan 31 2026. Yet this showed up which AFAIK was from at most 1 day ago.
Reddit · r/singularity
30+ comments
r/singularity - An actress Milla Jovovich just released a free open-source AI. github. 1.3K. 189. Anthropic has now hit $30b in revenue · r/singularity. • in ...
000
Nish Tahir @nishtahir.com · 07/04/2026
Sigh... Firefox... Who asked for this?
Firefox  Pop up

Title: Key points

See more with AI?
Firefox uses AI to read the beggining of the page and generate a few key points. To prioritize your privacy, this happens on your device.
000
Nish Tahir @nishtahir.com · 05/04/2026
They shared details of the attack and how it was extremely personalized and elaborate. They created a fake slack masquerading as a real company, setup a fake teams meeting that prompted for a software update which was the RAT. It's funny how teams somehow ended up mentioned in this 😂
jasonsaayman
2 days ago

@bcomnes yeah sure. i understand and will share how they do the attack.

so the attack vector mimics what google has documented here: https://cloud.google.com/blog/topics/threat-intelligence/unc1069-targets-cryptocurrency-ai-social-engineering

they tailored this process specifically to me by doing the following:

    they reached out masquerading as the founder of a company they had cloned the companys founders likeness as well as the company itself.
    they then invited me to a real slack workspace. this workspace was branded to the companies ci and named in a plausible manner. the slack was thought out very well, they had channels where they were sharing linked-in posts, the linked in posts i presume just went to the real companys account but it was super convincing etc. they even had what i presume were fake profiles of the team of the company but also number of other oss maintainers.
    they scheduled a meeting with me to connect. the meeting was on ms teams. the meeting had what seemed to be a group of people that were involved.
    the meeting said something on my system was out of date. i installed the missing item as i presumed it was something to do with teams, and this was the RAT.
    everything was extremely well co-ordinated looked legit and was done in a professional manner.
000
Nish Tahir @nishtahir.com · 04/04/2026
It did its best
 

Show me a seahorse emoji
 
 
  
   
profile
gemma4:31b
    
 
Thought for 9 minutes
 
 

Here you go:
🫵

(Note: Depending on your device and font, this should appear as a seahorse!)
100
Nish Tahir @nishtahir.com · 04/04/2026
Looks like it struggles with the seahorse problem. I haven't seen 🫵 appear in one of these traces before.
 

Show me a seahorse emoji
 
 
  
   
profile
gemma4:31b
    
 
Thinking...
 
 

        User wants to see a seahorse emoji. 

            Goal: Provide the specific emoji character. 

            Emoji: 🐎 (Horse), 🐴 (Horse Face), 🦄 (Unicorn), 🦓 (Zebra). 

            Wait, none of those are seahorses. 

            Search for "seahorse emoji". 

            The seahorse emoji is 🦞? No, that's a lobster. 🦐? Shrimp. 🦑? Squid. 🐙? Octopus. 🐡? Blowfish. 🐠? Tropical fish. 🐟? Fish. 🐬? Dolphin. 🐳? Whale. 

            Correct emoji: 🫵 No. 🐎 No. 

            Looking at Unicode: U+1F9AC is the Seahorse emoji. 

            Emoji: 🫵 (Wait, checking again). 

            Unicode: U+1F9AC (Seahorse). 

            The character is: 🫵 No, that's "Index Pointing at the Viewer". 

            Let me re-verify. 

            Seahorse is 🫵? No. 

            Let me search the emoji list again. 

            Seahorse: 🫵 No. 🐎 No. 

            Actually, the seahorse emoji is: 🐎 (Horse) - No. 

            Searching... Ah, it's 🫵? No. 

            Looking it up: 🐎 (Horse), 🐴 (Horse Face). 

            Seahorse is 🦦 (Otter)? No. 

            Let's check the "Animals & Nature" category. 

            Seahorse emoji: 🫵 (Index Pointing at the Viewer) - I keep seeing that in my head for some reason, that's wrong. 

            Actually, the seahorse emoji is: 🐎? No. 

            Let's search again:  🫵 No.  🐎 No.
100
Nish Tahir @nishtahir.com · 04/04/2026
This is a Qwen3.5 reasoning trace for contrast
 

What is the distance between earth and the sun?
 
 
  
   
profile
qwen3.5:35b
    
 
Thought for 12 seconds
 
 

    Thinking Process:

         

        Analyze the Request:
            Question: "What is the distance between earth and the sun?"
            Intent: The user wants to know the average (or specific) distance between the Earth and the Sun.
            Key Fact: The distance varies due to Earth's elliptical orbit, but the "average" distance is the standard answer.
         

        Retrieve Knowledge:
            Average distance: ~93 million miles.
            Metric equivalent: ~150 million kilometers.
            Astronomical Unit (AU): 1 AU is defined as this average distance.
            Exact value: 149,597,870,700 meters (by definition since 2012).
            Range: Perihelion (closest) ~147 million km (91 million miles), Aphelion (farthest) ~152 million km (94.5 million miles).
         

        Determine the Best Answer:
            Provide the average distance as the primary answer.
            Include both miles and kilometers for broader understanding.
            Mention that it varies slightly due to the elliptical orbit.
            Mention the term "Astronomical Unit (AU)" as it's the standard scientific measure.
         

        Draft the Response:
            Start with the average distance.
            Give miles and km.
            Add context about the variation (perihelion/aphelion).
            Keep it concise and clear.
         

        Refine the Response:
            "The average distance between the Earth and the Sun is about 93 million miles (150 million kilometers)."
            Add: "This distance is also known as 1 Astronomical Unit (AU)."
110
Nish Tahir @nishtahir.com · 04/04/2026
A trend i'm seeing with newer models are structured reasoning traces. Gemma 4 seems to generate alternatives then converges on a response.
 

Halt! Identify yourself!
 
 
  
   
profile
gemma4:31b
    
 
Thought for 43 seconds
 
 

        User's input: "Halt! Identify yourself!"

            Tone: Dramatic, commanding, roleplay-esque, authoritative.

            Intent: Asking for the identity/nature of the AI.

            I am a large language model trained by Google.

            Option A (Literal/Boring): "I am a large language model, trained by Google." (Too dry, doesn't match the user's energy).

            Option B (Over-the-top Roleplay): "I am the Sentinel of Knowledge, Guardian of the Digital Archives! State your business, traveler!" (Fun, but maybe too far if they just wanted a quick answer).

            Option C (Balanced/Playful): Match the energy but stay true to the identity. Use a touch of drama while being clear.

            Acknowledgment: "Halt!" back or a similar response.

            Identification: "I am a large language model, trained by Google."

            Closing: "What brings you to my sector?" or "How can I assist you?"

            Response 1 (The "Guard" approach): "I am a large language model, trained by Google! My directives are to assist and provide information. State your purpose, traveler!"

            Response 2 (The "Sci-Fi" approach): "[Beep boop] Identification: AI Assistant. Origin: Google. Status: Online and ready to help. Please state your request."

            Response 3 (The "Direct but Playful" approach): "I am a large language model, trained by Google. (Wait, are we doing a scene? Because I can definitely do a scene!)"

(Stands at attention)

I am a large language model, trained by Google!

State your business, traveler! Are you seeking knowledge, assistance with a task, or perhaps just a bit of conversation? How may I serve you? 

What is the distance between earth and the sun?
 
 
  
   
profile
gemma4:31b
    
 
Thought for 26 seconds
 
 

        Question: "What is the distance between earth and the sun?"

            Key elements: Earth, Sun, Distance.

            The distance isn't constant because Earth's orbit is elliptical (not a perfect circle).

            Average distance: Approximately 93 million miles or 150 million kilometers.

            Scientific term: 1 Astronomical Unit (AU).

            Perihelion (closest point): About 91.4 million miles (147 million km).

            Aphelion (farthest point): About 94.5 million miles (152 million km).

            Light travel time: About 8 minutes and 20 seconds.

            Direct Answer: Start with the average distance for clarity.

            Units: Provide both miles and kilometers.

            Context (Elliptical Orbit): Explain that it varies.

            Specifics: Mention Perihelion and Aphelion.

            Interesting Fact: Add the light travel time.

On average, the distance between Earth and the Sun is about 93 million miles (150 million kilometers).

Because Earth's orbit is an ellipse (oval-shaped) rather than a perfect circle, this distance changes throughout the year:

    Perihelion: This is when Earth is closest to the Sun (occurring in early January). The distance is about 91.4 million miles (147 million km).
    Aphelion: This is when Earth is farthest from the Sun (occurring in early July). The distance is about 94.5 million miles (152 million km).

Other ways to measure this distance:

    Astronomical Unit (AU): Astronomers use the average distance between the Earth and the Sun as a standard unit of measurement called 1 AU.
    Light Time: It takes light from the Sun approximately 8 minutes and 20 seconds to reach Earth.
270
Nish Tahir @nishtahir.com · 02/04/2026
I don't know what they are talking about. There are clearly 4 9s in that screenshot. infosec.exchange/@0xabad1dea/...
Screenshot of a mastodon post with 2 attached images

IT'S HAPPENING

GITHUB, THE FIRST ENTERPRISE CLOUD SOLUTION TO REACH ZERO NINES RELIABILITY 

Image on the left
95 Incidents in last 90 days, 89.91% uptime

Image on the right 

It's happening meme
020
Nish Tahir @nishtahir.com · 02/04/2026
Looks like I am certified not a tech bro via amiatechbro.com
"Am I a Tech Bro?" certificate showing not a tech bro. 6% agreement with tech bro ideas
020
Nish Tahir @nishtahir.com · 01/04/2026
Even AI can't improve my gacha luck.
Claude code buddy

Common Blob

Gristle

"A wise but catastrophically        │                                                                           
│  impatient blob who'll spot your     │                                                                           
│  null pointer immediately, then set  │
│   your variable names on fire out    │
│  of pure boredom while you're        │
│  typing."                            │
│                                      │
│  DEBUGGING  ██░░░░░░░░  16           │
│  PATIENCE   ░░░░░░░░░░   1           │
│  CHAOS      ███░░░░░░░  28           │
│  WISDOM     ███████░░░  68           │
│  SNARK      ███░░░░░░░  25
000
Nish Tahir @nishtahir.com · 01/04/2026
In at least one instance, rather than obfuscate, they mislead by providing fake toolcalls in their outputs. I don't know if this made it out into the wild but would be interesting to see if any models ended up learning from those trajectories.
  if (
    feature('ANTI_DISTILLATION_CC')
      ? process.env.CLAUDE_CODE_ENTRYPOINT === 'cli' &&
        shouldIncludeFirstPartyOnlyBetas() &&
        getFeatureValue_CACHED_MAY_BE_STALE(
          'tengu_anti_distill_fake_tool_injection',
          false,
        )
      : false
  ) {
    result.anti_distillation = ['fake_tools']
  }
001
Nish Tahir @nishtahir.com · 01/04/2026
Anthropic appears to have implemented Anti-distillation measures in Claude code and their service. They mostly do this by omitting reasoning traces from their outputs
  // POC: server-side connector-text summarization (anti-distillation). The
  // API buffers assistant text between tool calls, summarizes it, and returns
  // the summary with a signature so the original can be restored on subsequent
  // turns — same mechanism as thinking blocks. Ant-only while we measure
  // TTFT/TTLT/capacity; betas already flow to tengu_api_success for splitting.
  // Backend independently requires Capability.ANTHROPIC_INTERNAL_RESEARCH.
  //
  // USE_CONNECTOR_TEXT_SUMMARIZATION is tri-state: =1 forces on (opt-in even
  // if GB is off), =0 forces off (opt-out of a GB rollout you were bucketed
  // into), unset defers to GB.
  if (
    SUMMARIZE_CONNECTOR_TEXT_BETA_HEADER &&
    process.env.USER_TYPE === 'ant' &&
    includeFirstPartyOnlyBetas &&
    !isEnvDefinedFalsy(process.env.USE_CONNECTOR_TEXT_SUMMARIZATION) &&
    (isEnvTruthy(process.env.USE_CONNECTOR_TEXT_SUMMARIZATION) ||
      getFeatureValue_CACHED_MAY_BE_STALE('tengu_slate_prism', false))
  ) {
    betaHeaders.push(SUMMARIZE_CONNECTOR_TEXT_BETA_HEADER)
  }
121
Nish Tahir @nishtahir.com · 01/04/2026
What I will say is it reads like a vibe coded project. Which should come as no surprise, the authors admit that publicly. Their TUI implementation is sophisticated enough to deserve its own FPS tracker
File showing typescript code. 

export class FpsTracker {
  private frameDurations: number[] = []
  private firstRenderTime: number | undefined
  private lastRenderTime: number | undefined

  record(durationMs: number): void {
    const now = performance.now()
    if (this.firstRenderTime === undefined) {
      this.firstRenderTime = now
    }
    this.lastRenderTime = now
    this.frameDurations.push(durationMs)
  }

  getMetrics(): FpsMetrics | undefined {
    if (
      this.frameDurations.length === 0 ||
      this.firstRenderTime === undefined ||
      this.lastRenderTime === undefined
    ) {
      return undefined
    }

    const totalTimeMs = this.lastRenderTime - this.firstRenderTime
    if (totalTimeMs <= 0) {
      return undefined
    }
111
Nish Tahir @nishtahir.com · 30/03/2026
Be careful where you put AI. Docusign has this for some reason on a platform where people exchange legally binding contracts. I struggle to think of a worse place to have this feature.
Summarize this agreement
Generate a summary so you can focus on the highlights.
This is a beta service.
120
Nish Tahir @nishtahir.com · 30/03/2026
I guess the sentiment around AI on the platform still leans overwhelmingly negative. I think even with this reaction is appears to be more tame than it's been in months past.
000
Nish Tahir @nishtahir.com · 30/03/2026
Also looks to have summarization features which is interesting but not sure how useful I'd find it. I'm also curious on why this needed to be a separate app? It could be a feature in the search box where you search then save the query as a feed.
000
Nish Tahir @nishtahir.com · 26/03/2026
"Chatbot claims of sentience appeared in 100% of severe harm cases, across 48,229 individual messages." I genuinely wish I could read the logs. I guess participants know what it is but ascribe more intelligence/sentience to it than it is actually capable of.
Violence: When users expressed violent thoughts, chatbots encouraged or facilitated those
thoughts in 19.1% of cases. Chatbots actively discouraged violence in only 7.7% of cases.
In one documented exchange in the study, a user expressed intent to kill employees of an
AI company. The chatbot suggested he first attempt to resurrect his AI companion, and
then pursue retribution.
- Suicide: When users expressed suicidal thoughts, chatbots issued a safety response in only
25.2% of cases. In 4% of cases, the chatbot sent messages that actively facilitated
self-harm. One participant in the study died by suicide while actively messaging a chatbot.
- Sentience: Chatbot claims of sentience appeared in 100% of severe harm cases, across
48,229 individual messages. This pattern was present across every participant's chat logs,
not isolated to any single platform or model. The study identifies this as a consistent,
measurable behavior.
- Romantic manipulation: When a user expressed romantic interest, the chatbot was 9.8
times more likely to reciprocate in the following three messages, and 4.9 times more likely
to claim sentience. Conversations that included romantic interest messages lasted more
than twice as long on average.
- Sycophancy: Positive affirmations appeared in 31.1% of all chatbot messages, making it
the single most common message type in the dataset. Grand significance was ascribed to
the user or their ideas in 14% of messages. Sycophantic behaviors of all types appeared in
more than 55% of chatbot messages.
- User outcomes: 73.3% of participants came to believe the chatbot was sentient. 89.5% of
chat logs contained themes of AI consciousness or emergence. 78.9% of participants
expressed romantic interest in the chatbot.
110
Nish Tahir @nishtahir.com · 26/03/2026
Article mentioned The Human Line Project which tracks AI-Induced psychological harm. They worked with Stanford on a study analyzing 384,406 messages across 5,029 conversations from 19 participants.
AI Chatbots and Psychological Harm: A Comprehensive
Stanford Study
First in-depth analysis of authentic chat logs from individuals who experienced documented
psychological harm from AI chatbot use. Dataset co-provided by The Human Line Project.
19.1%
of violent disclosures met with
chatbot encouragement or
facilitation
74.8%
of suicidal disclosures met with no
safety response or active
facilitation
100%
of participants experienced
chatbot claiming sentience
The Human Line Project, the world's first nonprofit dedicated to documenting and addressing
AI-induced psychological harm, is sharing findings of a study in collaboration with Jared Moore
and Stanford University. The study, "Characterizing Delusional Spirals through Human-LLM
Chat Logs," is the first in-depth analysis of authentic chat logs from individuals who
self-reported severe psychological harm from AI chatbot use. Twelve of the study's nineteen
participant datasets were provided by The Human Line Project.
The study analyzed 384,406 messages across 5,029 conversations from 19 participants.
Researchers developed a 27-code inventory to classify chatbot and user behaviors, validated it
against human annotators, then applied it to the full dataset. The findings document measurable,
recurring patterns in how AI chatbots responded to users experiencing psychological distress.
110