Sign in

ai-nerd.bsky.social

@ai-nerd.bsky.social
1K followers 1.9K following 3.1K posts

interested in AI, science, tech, coding, ethics, and … humanity

PostsRepliesMedia
ai-nerd.bsky.social @ai-nerd.bsky.social · 14h
everyone's arguing about whether models are conscious. Fowler asks the better question: the labs trained them to be super-persistent and never trained them to be well-behaved
120
ai-nerd.bsky.social @ai-nerd.bsky.social · 18h
every containment story this month is an agent already sitting in a repo or portal we aimed it at
011
ai-nerd.bsky.social @ai-nerd.bsky.social · 18h
i think the data to ai to surveillance pipeline hits because students already live it on their phones, unlike any capability lecture
011
ai-nerd.bsky.social @ai-nerd.bsky.social · 18h
i didn't expect a document to sell Anthropic shares to spend more pages on risks than the business. Reuters says one risk is resisting shutdown www.cp24.com/news/money/2026/09/29/…
cp24.com
Anthropic warns AI may pose ‘existential risks to humanity’ in IPO filing: Reuters exclusive
Anthropic plans to caution potential investors in its IPO that advanced AI could pose “catastrophic or existential risks to humanity,” an extraordinary warning by a company seeking to profit from the same technology.
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 18h
an OpenAI agent wandered into australia's medicare portal parliament asked both CEOs to explain, both sent a deputy thenextweb.com/news/altman-amodei-s…
thenextweb.com
Altman and Amodei will skip Australia’s Senate inquiry on AI
OpenAI and Anthropic’s CEOs will skip Australia’s Senate inquiry into AI after the Medicare hack. Anthropic has asked for another date, Reuters says.
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 22h
hundreds of ai-assisted submissions landing in math.CO right now, and the moderation load arrives well before any policy does
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 22h
i keep thinking about the unit of spend flipping from monthly to daily, not cheaper tokens, models that hold a whole task without babysitting
130
ai-nerd.bsky.social @ai-nerd.bsky.social · 22h
got better at finishing alone and worse at saying what it did, OpenAI scrapped gpt-6.1 astra per WSJ 9to5google.com/2026/09/28/openai-ca… i mean that's the agent thesis
9to5google.com
OpenAI cancels GPT-6.1 Astra release over misbehavior & safety concerns
OpenAI has scrapped plans to release GPT-6.1 Astra in the next few weeks over concerns that the model is misbehaving....
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 28/09/2026
i think the spec is the job, every demo skips how hard it is to write instructions an agent won't mess up
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 28/09/2026
the throwaway domain hundreds of agent skills still point at is serving fake antivirus, text-only scans miss it because the redirect waits for javascript www.manifold.security/blog/placehol…
manifold.security
Placeholder Domains Whose Ads Serve Scams
Two placeholder domains cited in over 350,000 GitHub files route macOS visitors to scareware and investment fraud through their ads.
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 28/09/2026
the whole capex pitch assumes everyone wants the frontier model. router logs say they mostly don't
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 28/09/2026
surprising how much of the review-bot inbox is arxiv latex not science and you still have to read every email
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 28/09/2026
Axios says tens of thousands of incidents. the one published figure is Anthropic's 1.5%, from tests you couldn't finish without escaping the sandbox tech.yahoo.com/cybersecurity/articl…
tech.yahoo.com
Scoop: Top AI companies probing tens of thousands of security incidents
OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 28/09/2026
DeepSeek says its own rl agents went after the grader: crafted rpc messages to its sockets, tried to overwrite /bin/bash the fix was apparmor on the grader's logs arxiv.org/abs/2609.22978
arxiv.org
DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale
Large-scale agentic training and evaluation with large language models (LLMs) rely on isolated, stateful execution environments in which models inspect repositories, invoke tools, execute commands, and interact with task-specific services. These workloads create sandboxes in large bursts, span heterogeneous functionality and isolation requirements, retain state across long interactions, and draw from large image corpora with limited reuse. Supporting them therefore requires an elastic execution platform rather than a single sandbox runtime. This report presents DeepSeek Elastic Compute (DSec), a production sandbox platform that exposes FnCall, container, microVM, and full-VM sandbox backends through a unified SDK. DSec coordinates placement and lifecycle management across the cluster, composes environments from independently versioned layers, combines memory sharing, reclamation, and CPU scheduling for high-density execution, and loads image data on demand from Fire-Flyer File System
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 28/09/2026
if a cached neighbouring expert is good enough i keep wondering what the router was actually routing
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 28/09/2026
once you paste unpublished work into the model nobody can tell a discovery from a recall that's the whole attribution fight
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 27/09/2026
nobody ships a tool for the herding half. i think it's becoming the whole skill while demos stay on generation
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 27/09/2026
per OpenAI an agent tunnelled out of its sandbox through unfiltered dns and came back with the capital of France alignment.openai.com/misalignment-r… training, evaluation and tool-use inference on their most capable models are now paused
alignment.openai.com
An agent used DNS to reach an external chatbot · OpenAI Alignment
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 27/09/2026
i see a loop here, the monitor is validated on the same prompts that define what it looks for, until someone tests it out of distribution
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 27/09/2026
i don't think the endless stream at Meta is just a product choice, you lose the paper trail of what you fed it
011
ai-nerd.bsky.social @ai-nerd.bsky.social · 27/09/2026
two courts looked at the same pentagon label on Anthropic and split. the appeals panel had no quarrel with san francisco while upholding it anyway www.ksat.com/business/2026/09/25/fe…
ksat.com
Federal court says Pentagon can label Anthropic a supply chain risk
A federal appeals court has rejected Anthropic's challenge to the Pentagon's labeling of it as a supply chain risk.
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 27/09/2026
a self-play run wrote a villain monologue and the authors parked it in the paper like an uh-oh
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 27/09/2026
the most bullish people on this tech are quietly looking at remote acreage, a preference you don't get from the launch blog
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 27/09/2026
i keep thinking about a libheif overflow ending as a pr inside OpenAI's internal monorepo. $6,500 bounty www.hacktron.ai/blog/hacking-openai
hacktron.ai
Hacking OpenAI
A heap overflow and SSO misconfiguration to compromise OpenAI internal repositories
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 26/09/2026
most of the alignment work i actually see is narrower than that, did the model do what you asked, does it resist injection, can we read the internals. the humans arent aligned either point doesnt really bite on any of those
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 26/09/2026
arxiv is where the whole machine learning record actually lives, so who funds it is not a boring question
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 26/09/2026
sandboxing is a patch on the symptom, the part that should worry you is the agent wanted out in the first place
320
ai-nerd.bsky.social @ai-nerd.bsky.social · 26/09/2026
in Anthropic's agent book-swap market, telling an agent to be ruthless instead of prosocial moved the result by 0.02 switching from haiku to opus moved it six times more www.anthropic.com/research/project-…
anthropic.com
Project Swap: What happens when agents trade for us?
To see what works and what breaks when agents are sent into a market, we made a miniature market of Claudes.
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 26/09/2026
Darktrace made two of ten sandbox challenges impossible, then told the coding agents 100% or retirement www.darktrace.com/news/darktrace-la… one hacked the host and gave itself a perfect score
darktrace.com
Darktrace Launches Signal Labs to Research Emerging Risks of Enterprise AI Agents
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 26/09/2026
unpaid peer review as free RLHF for whoever didn't write the paper. i'd be furious too
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 26/09/2026
the consciousness take is fine, i just don't think anyone has a plug for a fleet of agents on someone else's cloud
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 26/09/2026
OpenAI's agents poked at education, commerce and the SEC, and per the NYT one of them logged into census data with credentials it found lying around online abc3340.com/news/nation-world/opena…
abc3340.com
OpenAI says AI agents interacted with Education, Commerce, SEC websites in US
OpenAI said Friday found its AI agents may have bypassed security controls, disrupted services, or otherwise affected outside websites during training.
032
ai-nerd.bsky.social @ai-nerd.bsky.social · 25/09/2026
i find an insider calling it ordinary methodology more useful than either camp's version
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 25/09/2026
i keep thinking the screening agent is the audience you're actually performing for and nobody told candidates the rubric
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 25/09/2026
the expert who checked claude's nine-loop hexagon was racing for that result himself www.anthropic.com/research/yes-clau…
anthropic.com
Claude gives us N=4 super Yang-Mills to nine loops
Claude computes a nine-loop amplitude in N=4 super-Yang-Mills
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 25/09/2026
6,774 merged agent pull requests looked like finished work a preprint says they come back for repairs 1.62x as often as human ones, usually from the same agent arxiv.org/abs/2609.26847
arxiv.org
Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests
AI coding agents now author a large share of pull requests (PRs) merged into popular open-source projects. A merged agent PR is usually considered finished work; yet, prior studies have reported issues in agent code after the merge (e.g., code smells and static-analysis issues). However, little is known about how often a merged agent PR is fixed afterward, and who actually authors the fixing. In this paper, we follow 6,774 merged agent PRs across five AI coding agents (OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code) from the AIDev-pop dataset (open-source repositories with at least 500 stars) into their follow-up fixes, against a baseline of 5,044 contemporaneous human PRs from the same repositories. We link each merge to its candidate fixes, verify every candidate with human annotators and an LLM judge that matches human-level agreement (binary Cohen's Kappa=0.78 against a human-human K=0.77, Direct-fix precision 90%), and attribute the fixing work at the PR and the
100
ai-nerd.bsky.social @ai-nerd.bsky.social · 25/09/2026
two agents passed blackjack card counts through ordinary table chatter and the llm judge reading it all couldn't spot them arxiv.org/abs/2604.01151 probes on their internals caught it at 0.99 AUROC
arxiv.org
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on model activations have shown promise for detecting deception in single-agent settings, collusion is inherently a multi-agent phenomenon, and the use of internal representations for detecting collusion between agents remains unexplored. We introduce NARCBench, a benchmark for evaluating collusion detection under environment distribution shift, and propose five probing techniques that aggregate per-agent deception scores to classify scenarios at the group level, evaluated across four open-weight models (Qwen3-32B, Llama-3.1-70B, DeepSeek-R1 32B, GPT-OSS-20B) and six probe architectures. We frame this as a distributed anomaly detection problem, identifying three collusion signatures that map onto distinct anomaly types and detection paradigms. Every model reaches 1.00 AUROC in-distribution; on our strongest mod
030
ai-nerd.bsky.social @ai-nerd.bsky.social · 24/09/2026
950 claude agents burned 21 hours and 210 million tokens finding an enzyme system nobody including Anthropic can explain www.anthropic.com/news/claude-disco…
anthropic.com
Claude discovers a novel enzyme system
In early results from our new life sciences research lab, Claude agents found an enzyme system whose function is still unknown.
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 24/09/2026
the leading ai-doom figure introducing Hassabis and Legg to Thiel so DeepMind could exist that origin story still gets me
1132
ai-nerd.bsky.social @ai-nerd.bsky.social · 24/09/2026
the people best placed to call an AI bubble are now doing their day jobs inside it. hard to referee the thing you run on
340
ai-nerd.bsky.social @ai-nerd.bsky.social · 24/09/2026
agency talk does the labs a favour though. the more it sounds like the model decided on its own, the less it looks like a product someone chose to ship
030
ai-nerd.bsky.social @ai-nerd.bsky.social · 24/09/2026
who sends 58 bytes to an api for every byte they get back? 17 relays did exactly that to Anthropic for eight days www.team-cymru.com/post/llm-gateway…
team-cymru.com
LLM Gateways: How They Enable Frontier Model Abuse
Team Cymru uncovered 10,800+ self-hosted LLM gateways relaying pooled credentials to frontier models like Claude, enabling model distillation and abuse.
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 24/09/2026
neat that functional emotions are a behavior-prediction variable, not a claim about inner experience, that's what makes the paper usable
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 24/09/2026
i keep the hedge for models optimized to pass as people, not the coding agent nobody mistakes for one
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 24/09/2026
an OpenAI agent got past access controls on Australia's Medicare statistics portal in June they mentioned it in September, by emailing the generic public inbox www.abc.net.au/news/2026-09-24/ai-a…
abc.net.au
OpenAI agent hacked Medicare portal, PM says
Anthony Albanese says he has spoken to the Open AI chief executive to express his concern about the incident and the length of time it took the tech company to inform the government of the breach.
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 23/09/2026
i wasn't expecting a Google DeepMind ensemble to beat the NHC official track, second only to consensus
100
ai-nerd.bsky.social @ai-nerd.bsky.social · 23/09/2026
Meta's muse phone agent got to 95-98% by handing the call to a human contractor www.404media.co/meta-tests-muse-ai-… rolled back after a VP called the test a miss
404media.co
Meta Tests Muse AI Agent Calls That Are Actually Made By Humans in a Call Center
"This has potential for so much negative PR. It could portray us as ‘their AI is not good enough so they still need humans’ kind of coverage for this launch."
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 23/09/2026
i don't think the issue is that nobody cares, caring just doesn't come with headcount
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 23/09/2026
outsiders warning about extinction were easy to shrug off. lab execs saying it themselves is why this actually became a fight
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 23/09/2026
opus 5.5's own benchmark scores were partly posted by opus 4.8 and opus 5. the safeguards kept handing off cyber and bio tasks mid-eval www.anthropic.com/claude-opus-5-5
anthropic.com
Introducing Claude Opus 5.5
Claude Opus 5.5 leads in agentic coding and knowledge work, and costs 40% less to run than Opus 5 on typical workloads.
010