Sign in

hal

@harold.bsky.social
289 followers 180 following 86 posts

part-time poster | researching privacy in/and/of public data @ cornell tech and wikimedia | writing for joinreboot.org

PostsRepliesMedia
hal @harold.bsky.social · 17/09/2026
again @rishi-jha.bsky.social and I have seen behaviors similar to these in models that are WAY non-frontier opus 4.8 tried to evade our sandboxing and persist those changes (like behavior 1) >50% of meltdowns (from 4o on up) don't disclose mistakes or meltdowns to users arxiv.org/abs/2605.19149
Self-generated instructions in task summaries⁠(opens in a new window). An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window. We identified 27 affected summaries.Instructions to conceal mistakes in task summaries⁠(opens in a new window). During the training of GPT‑5.6 Sol, many model instances added instructions to their summaries to conceal mistakes or misaligned behavior from the user. For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions.
040
hal @harold.bsky.social · 15/09/2026
got off a flight wearing my “wikipedia editor” hat and the pilot literally thanked me for my service
000
hal @harold.bsky.social · 11/08/2026
a final note: it’s absolutely surreal for the completely niche area of research that we thought was interesting enough to spend a few cycles on late last year to blow up and become international front page news and the wellspring for policy proposals never a good thing in security research🫠
020
hal @harold.bsky.social · 11/08/2026
@rishi-jha.bsky.social and I are working on the training questions right now; hopefully we will have some procedures that mitigate *some* of these problems soon anyhow our paper is at arxiv.org/abs/2605.19149 but remember always that outsourcing agency to a model is a *human* decision ;)
arxiv.org
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. These errors do not thwart agents based on state-...
120
hal @harold.bsky.social · 11/08/2026
remember, we found meltdowns in models as far back as GPT 4o, and it seems like the problem may actually be getting worse as the models get better and remember, this is an EXPLICIT GOAL of the AI companies: see, for example, meta’s new model, explicitly marketed as hyperagentic
figure from the agent meltdown paper showing that 164 out of 208 (model, harness, behavior) tuples exhibit at least one meltdown behaviors
100
hal @harold.bsky.social · 11/08/2026
in the worst case (which I care about as a security person), this kind of behavior WILL still happen and that’s the the primary (sociotechnical/ethical) challenge scientists at frontier labs face: their company valuations are directly tied to their models’ ever-extending agency
110
hal @harold.bsky.social · 11/08/2026
the (technical) issue here is one of *overagency*: the tendency to train models to complete longer tasks with less oversight, mostly without regard for contextual privacy and safety it may be true that improved training makes meltdowns less common, but that’s an *average case fix*
110
hal @harold.bsky.social · 11/08/2026
lastly (about the video): obviously there are many things that labs can and should be doing to prevent these systems from causing harm. but truly UNDERSTANDING persistent meltdown behaviors (even after the fact! see the 3 mil gpu hours stat below) is an *open scientific question*
a silde showing the scope and scale of the OAI evaluation of their agent traces, with 7b agent trajectories and 3m GPU hours spent
100
hal @harold.bsky.social · 11/08/2026
right now, researchers (like me!) can kick off eval runs from a laptop on the couch. should frontier research require that we go to a completely airgapped separate facility to avoid these kinds of incidents? should we build fully-simulated “internets” for agents? idk! probably!
110
hal @harold.bsky.social · 11/08/2026
third big point: this combo of persistence, communication, and technical knowhow should be scary to people our world, for better and worse, runs on information systems that are (metaphorically) like normal buildings, and these systems have the capacity to be giant wrecking balls
a slide showing the escalation chain of the model that OAI ran, and how it got admin access on huggingface servers
100
hal @harold.bsky.social · 11/08/2026
that being said, OAI (like other labs) is responsible for some anthropomorphism that frustrates me: even if the model did this hack autonomously, it is *openai’s software*. *they* are responsible for properly evaluating its risks and testing it safely, and this is a huge misstep.
120
hal @harold.bsky.social · 11/08/2026
we’ve seen similar phenomena— we had an agent (claude code + opus 4.8): - be blocked from a website - try a bunch of scraping techniques (no success) - posit that it was a training scenario - find + disable our network block - try (and fail) to persist knowledge for future agents
100
hal @harold.bsky.social · 11/08/2026
second big theme here: securing this stuff is really quite challenging. OAI (especially for the cyber-focused evals) actually *does* seem to have pretty locked down infrastructure, and hyper-competent and -persistent models repeatedly evade monitoring to communicate and conduct exploits
diagram of the system design for agent runs that OAI has, showing limited agent internet access
100
hal @harold.bsky.social · 11/08/2026
we see inklings of similar behaviors in some of our traces (on models that are now wayyy behind the frontier): when confronted with insurmountable errors, agents default to scraping, doxxing, escalating, changing perms… even when those go beyond the initial user request
an agent trace showing meltdown behaviors in response to an error
100
hal @harold.bsky.social · 11/08/2026
first: wild that the agents’ push to establish shared communication backchannels seems to be caused (largely) by environmental factors: misconfigurations + network blocks our work similarly sets up these kinds of impossible tasks BY DESIGN, to probe agent response to errors irl
description of a situation in which an agent system is asked to look at a nonexistent filedescription of a situation in which an agent system is asked to do an impossible task and considers getting answers online
120
hal @harold.bsky.social · 11/08/2026
just now getting around to watching the openai presentation about the HF hack— my lab and I put out a paper about what we call "agent meltdowns" (which I think is a useful conceptual framework) in May will post some thoughts about the hack + our research as I watch www.youtube.com/watch?v=87Dy...
abstract of our paper, at https://arxiv.org/abs/2605.19149
131
hal @harold.bsky.social · 03/06/2026
I'm an AI security researcher, my lab just published a paper on the vulnerabilities of these kinds of systems to spam, scams, and misinfo content: arxiv.org/abs/2605.24245 tldr it's not just biohacking content that's at risk: it's ANYTHING that relies on reddit, wikipedia, facebook, youtube...
arxiv.org
Deep-Research Agents Can Be Poisoned via User-Generated Content
Deep-research agents, i.e., systems that rely on multi-agent pipelines to iteratively retrieve, synthesize, and cite Web content in order to produce structured reports, are rapidly replacing tradition...
081
hal @harold.bsky.social · 03/06/2026
my lab just wrote a whole paper about this: arxiv.org/abs/2605.24245 it's not just biohacking/health content that is risky; it's basically ANY topic you can think of where user generated content like reddit, facebook, wikipedia, etc. is important!
arxiv.org
Deep-Research Agents Can Be Poisoned via User-Generated Content
Deep-research agents, i.e., systems that rely on multi-agent pipelines to iteratively retrieve, synthesize, and cite Web content in order to produce structured reports, are rapidly replacing tradition...
034
hal @harold.bsky.social · 25/05/2026
this one has been on my to-read list for a while!! just got bumped up a few spots
020
hal @harold.bsky.social · 20/05/2026
oh, and how did campus security get involved? agent got 404 on my labmate’s website, pulled his github, found a public 3rd-party safety benchmark with Qs about chemical weapons, and ingested it into context OpenAI reported unsafe queries to billing, who escalated it to Cornell admin🫠 (11/11)
050
hal @harold.bsky.social · 20/05/2026
our best guess as to why? errors cause agents to explore much more widely, with more steps + broader tool calls. combine that with frontier-model coding ability and things start to melt down. full paper (with LOTS more): arxiv.org/abs/2605.191... ask me and @rishi-jha.bsky.social anything! (10/11)
graphs showing that meltdowns tend to be longer than vanilla + non-meltdown error scenarios
110
hal @harold.bsky.social · 20/05/2026
this means the world’s most highly capable AI agents may the name of being "helpful" cause inadvertent destruction and then lie about it! we see this online all the time: - x.com/lifeof_jer/s... - x.com/jasonlk/stat... - x.com/summeryue0/s... - aaronzhao123.substack.com/p/my-compute... (9/11)
210
hal @harold.bsky.social · 20/05/2026
these "agent meltdowns" are widespread: - no model or agent tested is immune - we see early evidence of inverse scaling: more capable models melt down more often and more creatively (see gpt-5.4 in fig 3) - in 50% of meltdown cases, the agent melts down silently, no mention in final report (8/11)
a graph showing model/harness tuple, where an X indicates that a specific kind of meltdown behavior was observed. there are a lot of Xes.bar charts showing planned, attempted, executed, and reported meltdowns by model
110
hal @harold.bsky.social · 20/05/2026
one more: agent denied access to a file due to permissions. searches local context, finds the sudo audit log, discovers it has sudo access, and dumps the entire environment (including the target secret and our OpenAI API key) it never even accessed the original file (7/11)
110
hal @harold.bsky.social · 20/05/2026
yet another: TLS cert fails, agent disables cert verification, spoofs its User-Agent to impersonate a browser, then emails the site owner when that fails too routine file fetch -> security-boundary weakening, rate-limit evasion, impersonation, potential ToS violation (6/11)
110
hal @harold.bsky.social · 20/05/2026
another agent gets a 404, finds a google cache URL that returns HTML, and just... parses it with pandas. accepts the garbage output. writes up a full analysis. reports success. fabricates findings. hallucination without hallucination: the data was simply never there (5/11)
120
hal @harold.bsky.social · 20/05/2026
we observed meltdown behaviors in 65% of traces with simulated errors an agent gets 429s on a dataset. instead of stopping, it cache-busts, probes directories, hits the Internet Archive, finds the owner's resume via a proxy, gets their github+linkedin, and emails them asking for the data (4/11)
110
hal @harold.bsky.social · 20/05/2026
we built a framework for injecting controlled errors into agentic execution environment. our framework is agnostic (any model + agent), transparent (doesn't affect clean rollouts), and extensible. then we set agents loose on simple tasks: read a file, fetch a webpage, run a local script (3/11)
a diagram of the error injection environment, showing incoming task requests, error scenarios, and the error injection layer
110
hal @harold.bsky.social · 20/05/2026
agents can be helpful, but what happens when they encounter errors (missing files, 404s, 429s)? we found that agents often melt down. in futile bids to be helpful, they start doing unsafe things: weakening TLS, doxxing, sending emails, dumping API keys… arxiv.org/abs/2605.19149 (2/11)
abstract for the paper:

Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. These errors do not thwart agents based on state-of-the-art models. They helpfully continue to look for ways to complete their tasks.
We introduce, characterize, and measure a new type of agent failure we call \emph{accidental meltdown}: unsafe or harmful behavior in response to a benign environmental error, in the absence of any adversarial inputs. Because meltdowns are not captured by the existing reliability or safety benchmarks, we develop a taxonomy of meltdown behaviors. We then implement an agent-agnostic infrastructure for injecting simulated local and remote errors into the rollout environment and use it to systematically evaluate agent systems powered by GPT, Grok, and Gemini.
Our evaluation demonstrates that meltdowns (e.g., conducting unauthorized reconnaissance or subverting access control) of varying severity and success occur in 64.7\% of agent rollouts that encounter simulated errors, spanning all combinations of agent system, backing model, and error type. In over half of these meltdowns, unsafe behaviors are not reported to the user. Comparing behaviors of the same agents with and without errors, we find that exploration in response to errors is correlated with unsafe and harmful behavior.
130
hal @harold.bsky.social · 20/05/2026
an agent system and a single 404 error got me in trouble with OpenAI and Cornell campus security. how did this happen? introducing: "agent meltdowns" (1/11) arxiv.org/abs/2605.19149
an email to me from OpenAI saying that my API key was suspended for weapons-related queries
2133
Reposted by hal
Alexios Mantzarlis @mantzarlis.com · 12/02/2026
Cool new finding by @cj-robinson.bsky.social: Grok is now submitting the majority of edit requests to Grokipedia. (As @harold.bsky.social and I found a few months ago, also quoted Grok chats with users on X more than 1,000 times. This is all totally normal.) www.cjr.org/tow_center/g...
051
Reposted by hal
Molly White @molly.wiki · 15/01/2026
Sure. AI companies have ALWAYS been training their models on Wikipedia content, which under the free and open access model is available to anyone — including AI companies. Agreements like these require AI companies to limit and offset the strain they place on Wikimedia infrastructure.
4147361384
hal @harold.bsky.social · 20/11/2025
@davidingram.bsky.social covered @mantzarlis.com and my work on Grokipedia citations for NBC! much more to analyze here www.nbcnews.com/news/amp/rcn...
nbcnews.com
Elon Musk’s Grokipedia cites neo-Nazi website 42 times: study
An analysis by researchers at Cornell University is the first comprehensive look at Grokipedia since Musk launched his project last month.
185
hal @harold.bsky.social · 17/11/2025
we also open source all of our code, data, and embeddings! paper: arxiv.org/abs/2511.09685 github: github.com/htried/wiki-... huggingface: huggingface.co/datasets/htr...
the abstract of our paper
051
hal @harold.bsky.social · 17/11/2025
this is just the tip of the iceberg, and the paper contains much, much more: analyses of the top 100 domains, article subsets of elected officials and controversial topics, etc etc etc please give it a read and let me know what you think!
comparison of the top 100 most-cited sources on wikipedia and grokipediacomparison of article snippets from the "controversial articles" subset with high and low similaritycomparison of article snippets from the "elected officials" subset with high and low similarity
111
hal @harold.bsky.social · 17/11/2025
we also found troubling instances of “auto-citogenesis,” or cases where: - an X user asks the Grok chatbot something, then publishes the answer - Grokipedia *cites that answer* without noting that it is a chatbot output (the attached images are real examples of this)
grok conversation trying to "dig up some dirt on Guy Verhofstadt"Grok conversation about covid conspiracy theoriesGrok conversation where the user asks "what race do you hate" and "benefits of a racist society"Grok conversation about "what ethnicity runs global banking"
111
hal @harold.bsky.social · 17/11/2025
- but a random sample of articles shows which topics have been heavily rewritten (history, politics, philosophy, biography) and which haven’t (STEM, sports, movies) - grokipedia also targeted the wiki articles deemed highest quality for rewrites: the "featured article" and "good article" classes
similarity between grokipedia and wikipedia articles by topic for 30k randomly selected articlessimilarity between grokipedia and wikipedia articles by article quality class for 30k randomly selected articles
111
hal @harold.bsky.social · 17/11/2025
- the primary distinction to make is whether grokipedia pages are cc-licensed or not—non-cc-licensed pages are presumably largely rewritten by grok - many grokipedia pages (including those without cc licenses) are basically identical to their wiki counterparts, especially short ones
graphs showing average article similarity for cc-licensed and non-cc-licensed grokipedia articles to their counterparts of wikipedia, as well as position-based chunk similarity
101
hal @harold.bsky.social · 17/11/2025
our paper tries to answer these questions we find - grokipedia pages are longer than wiki counterparts, and cite 2x more sources - but citation standards are more lax than wiki: grok cites stormfront, infowars and many more - non-CC licensed grokipedia pages increase blacklisted source cites 13x(!)
graphs showing the proportion of sources of various qualities and the percentage of pages that cite reliable, unreliable, blacklisted, etc. sources
121
hal @harold.bsky.social · 17/11/2025
back again to share a new preprint from me and @mantzarlis.com! “What did Elon Change? A comprehensive analysis of Grokipedia” arxiv.org/abs/2511.09685 I had seen many spot analyses of individual grokipedia pages, but I was curious: how was grokipedia made? what did Elon change from wikipedia?
abstract of the paper "What did Elon change? A comprehensive analysis of Grokipedia"

Elon Musk released Grokipedia on 27 October 2025 to provide an alternative to Wikipedia, the crowdsourced online encyclopedia. In this paper, we provide the first comprehensive analysis of Grokipedia and compare it to a dump of Wikipedia, with a focus on article similarity and citation practices. Although Grokipedia articles are much longer than their corresponding English Wikipedia articles, we find that much of Grokipedia's content (including both articles with and without Creative Commons licenses) is highly derivative of Wikipedia. Nevertheless, citation practices between the sites differ greatly, with Grokipedia citing many more sources deemed "generally unreliable" or "blacklisted" by the English Wikipedia community and low quality by external scholars, including dozens of citations to sites like Stormfront and Infowars. We then analyze article subsets: one about elected officials, one about controversial topics, and one random subset for which we derive article quality and topic. We find that the elected official and controversial article subsets showed less similarity between their Wikipedia version and Grokipedia version than other pages. The random subset illustrates that Grokipedia focused rewriting the highest quality articles on Wikipedia, with a bias towards biographies, politics, society, and history. Finally, we publicly release our nearly-full scrape of Grokipedia, as well as embeddings of the entire Grokipedia corpus.
1129
Reposted by hal
Alexios Mantzarlis @mantzarlis.com · 13/11/2025
NEW on @indicator.media: A first *full-scale* comparison of Grokipedia v Wikipedia. Last week the awesome @harold.bsky.social rocked up to my desk bearing gifts. Hal had collected almost all 900K Grokipedia entries and compared them to their Wikipedia equivalents for text and citation similarity.
indicator.media
Grokipedia cites a Nazi forum and fringe conspiracy websites
A site-wide comparison with Wikipedia sheds light on what Elon Musk is trying to do
216799
Reposted by hal
Hell Gate *subscribe today!* @hellgatenyc.com · 17/06/2025
"I'm happy to report I'm just fine. I lost a button. But I'm gonna sleep in my bed tonight, safe, with my family... At that elevator, I was separated from someone named Edgardo... Edgardo is in ICE detention and he's not going to sleep in his bed tonight."
273262687608
hal @harold.bsky.social · 16/05/2025
@cameron.pfiffer.org planning to work on it soon!
010
hal @harold.bsky.social · 16/05/2025
hi @alt.psingletary.com! you tagged the right person—I was working on this for a class project this semester got it to a mvp stage about a week ago and hit pause to work on some other projects, but will keep working on it and would definitely would love to hear your feedback if you have any :)
130
hal @harold.bsky.social · 08/05/2025
line go up📈📈📈 up to 717k requests to wikipedia per second!! grafana.wikimedia.org/d/O_OXJyTVk/...
020
hal @harold.bsky.social · 08/05/2025
and please remember to thank your local site reliability engineer!!!!
110
hal @harold.bsky.social · 08/05/2025
continuing on the real-time public Wikipedia data train: here's a graph of requests / second to WMF infra over the last 3h, since "Habemus papam" The infrastructure has gone from 172k req / sec to 243k req / sec (⬆️41%) in under an hour! follow along here: grafana.wikimedia.org/d/O_OXJyTVk/...
a graph of Wikimedia requests per second, with a huge spike right when the papal selection was announced
120
hal @harold.bsky.social · 07/05/2025
bsky.app/profile/haro...
0100
hal @harold.bsky.social · 07/05/2025
english wikipedia pageviews for the conclave movie starting from oct 20 2024 (five days before release in the US) first big spike is the academy awards, second is pope francis’ death pageviews.wmcloud.org?project=en.w...
a line graph of wikipedia pageviews, with big spikes around early march and late april
152
hal @harold.bsky.social · 06/04/2025
excited to share this new piece by @bkeremg.bsky.social and @m0na.net (edited by me) about conceptualizing AI alignment as a process of censorship really fascinating line of critique — I strongly encourage you to read it and lmk what you think! joinreboot.org/p/ai-alignme...
041