Sign in

hal

@harold.bsky.social
288 followers 180 following 86 posts

part-time poster | researching privacy in/and/of public data @ cornell tech and wikimedia | writing for joinreboot.org

PostsRepliesMedia
hal @harold.bsky.social · 17/09/2026
again @rishi-jha.bsky.social and I have seen behaviors similar to these in models that are WAY non-frontier opus 4.8 tried to evade our sandboxing and persist those changes (like behavior 1) >50% of meltdowns (from 4o on up) don't disclose mistakes or meltdowns to users arxiv.org/abs/2605.19149
Self-generated instructions in task summaries⁠(opens in a new window). An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window. We identified 27 affected summaries.Instructions to conceal mistakes in task summaries⁠(opens in a new window). During the training of GPT‑5.6 Sol, many model instances added instructions to their summaries to conceal mistakes or misaligned behavior from the user. For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions.
040
hal @harold.bsky.social · 15/09/2026
got off a flight wearing my “wikipedia editor” hat and the pilot literally thanked me for my service
000
hal @harold.bsky.social · 11/08/2026
just now getting around to watching the openai presentation about the HF hack— my lab and I put out a paper about what we call "agent meltdowns" (which I think is a useful conceptual framework) in May will post some thoughts about the hack + our research as I watch www.youtube.com/watch?v=87Dy...
abstract of our paper, at https://arxiv.org/abs/2605.19149
131
hal @harold.bsky.social · 03/06/2026
my lab just wrote a whole paper about this: arxiv.org/abs/2605.24245 it's not just biohacking/health content that is risky; it's basically ANY topic you can think of where user generated content like reddit, facebook, wikipedia, etc. is important!
arxiv.org
Deep-Research Agents Can Be Poisoned via User-Generated Content
Deep-research agents, i.e., systems that rely on multi-agent pipelines to iteratively retrieve, synthesize, and cite Web content in order to produce structured reports, are rapidly replacing tradition...
034
hal @harold.bsky.social · 20/05/2026
an agent system and a single 404 error got me in trouble with OpenAI and Cornell campus security. how did this happen? introducing: "agent meltdowns" (1/11) arxiv.org/abs/2605.19149
an email to me from OpenAI saying that my API key was suspended for weapons-related queries
2133
Reposted by hal
Alexios Mantzarlis @mantzarlis.com · 12/02/2026
Cool new finding by @cj-robinson.bsky.social: Grok is now submitting the majority of edit requests to Grokipedia. (As @harold.bsky.social and I found a few months ago, also quoted Grok chats with users on X more than 1,000 times. This is all totally normal.) www.cjr.org/tow_center/g...
051
Reposted by hal
Molly White @molly.wiki · 15/01/2026
Sure. AI companies have ALWAYS been training their models on Wikipedia content, which under the free and open access model is available to anyone — including AI companies. Agreements like these require AI companies to limit and offset the strain they place on Wikimedia infrastructure.
4147361384
hal @harold.bsky.social · 20/11/2025
@davidingram.bsky.social covered @mantzarlis.com and my work on Grokipedia citations for NBC! much more to analyze here www.nbcnews.com/news/amp/rcn...
nbcnews.com
Elon Musk’s Grokipedia cites neo-Nazi website 42 times: study
An analysis by researchers at Cornell University is the first comprehensive look at Grokipedia since Musk launched his project last month.
185
hal @harold.bsky.social · 17/11/2025
back again to share a new preprint from me and @mantzarlis.com! “What did Elon Change? A comprehensive analysis of Grokipedia” arxiv.org/abs/2511.09685 I had seen many spot analyses of individual grokipedia pages, but I was curious: how was grokipedia made? what did Elon change from wikipedia?
abstract of the paper "What did Elon change? A comprehensive analysis of Grokipedia"

Elon Musk released Grokipedia on 27 October 2025 to provide an alternative to Wikipedia, the crowdsourced online encyclopedia. In this paper, we provide the first comprehensive analysis of Grokipedia and compare it to a dump of Wikipedia, with a focus on article similarity and citation practices. Although Grokipedia articles are much longer than their corresponding English Wikipedia articles, we find that much of Grokipedia's content (including both articles with and without Creative Commons licenses) is highly derivative of Wikipedia. Nevertheless, citation practices between the sites differ greatly, with Grokipedia citing many more sources deemed "generally unreliable" or "blacklisted" by the English Wikipedia community and low quality by external scholars, including dozens of citations to sites like Stormfront and Infowars. We then analyze article subsets: one about elected officials, one about controversial topics, and one random subset for which we derive article quality and topic. We find that the elected official and controversial article subsets showed less similarity between their Wikipedia version and Grokipedia version than other pages. The random subset illustrates that Grokipedia focused rewriting the highest quality articles on Wikipedia, with a bias towards biographies, politics, society, and history. Finally, we publicly release our nearly-full scrape of Grokipedia, as well as embeddings of the entire Grokipedia corpus.
1129
Reposted by hal
Alexios Mantzarlis @mantzarlis.com · 13/11/2025
NEW on @indicator.media: A first *full-scale* comparison of Grokipedia v Wikipedia. Last week the awesome @harold.bsky.social rocked up to my desk bearing gifts. Hal had collected almost all 900K Grokipedia entries and compared them to their Wikipedia equivalents for text and citation similarity.
indicator.media
Grokipedia cites a Nazi forum and fringe conspiracy websites
A site-wide comparison with Wikipedia sheds light on what Elon Musk is trying to do
216799
Reposted by hal
Hell Gate *subscribe today!* @hellgatenyc.com · 17/06/2025
"I'm happy to report I'm just fine. I lost a button. But I'm gonna sleep in my bed tonight, safe, with my family... At that elevator, I was separated from someone named Edgardo... Edgardo is in ICE detention and he's not going to sleep in his bed tonight."
274262737608
hal @harold.bsky.social · 08/05/2025
line go up📈📈📈 up to 717k requests to wikipedia per second!! grafana.wikimedia.org/d/O_OXJyTVk/...
020
hal @harold.bsky.social · 08/05/2025
continuing on the real-time public Wikipedia data train: here's a graph of requests / second to WMF infra over the last 3h, since "Habemus papam" The infrastructure has gone from 172k req / sec to 243k req / sec (⬆️41%) in under an hour! follow along here: grafana.wikimedia.org/d/O_OXJyTVk/...
a graph of Wikimedia requests per second, with a huge spike right when the papal selection was announced
120
hal @harold.bsky.social · 07/05/2025
english wikipedia pageviews for the conclave movie starting from oct 20 2024 (five days before release in the US) first big spike is the academy awards, second is pope francis’ death pageviews.wmcloud.org?project=en.w...
a line graph of wikipedia pageviews, with big spikes around early march and late april
152
hal @harold.bsky.social · 06/04/2025
excited to share this new piece by @bkeremg.bsky.social and @m0na.net (edited by me) about conceptualizing AI alignment as a process of censorship really fascinating line of critique — I strongly encourage you to read it and lmk what you think! joinreboot.org/p/ai-alignme...
041
hal @harold.bsky.social · 04/04/2025
and set your devices to update automatically!
010
hal @harold.bsky.social · 24/03/2025
Reminder that a key part of the privacy harms from genetic information is the fact that *we all share genes with our relatives*! Deleting or refusing to share genetic information protects your family in addition to yourself oag.ca.gov/news/press-r...
oag.ca.gov
Attorney General Bonta Urgently Issues Consumer Alert for 23andMe Customers
Californians have the right to direct the company to delete their genetic data OAKLAND — California Attorney General Rob Bonta today issued a consumer alert to customers of 23andMe, a genetic testing ...
010
hal @harold.bsky.social · 18/03/2025
Excited to announce a new preprint from my lab (with @rishi-jha.bsky.social and Vitaly Shmatikov; my first as a first author!) about severe security vulnerabilities in LLM-based multi-agent systems: “Multi-Agent Systems Execute Arbitrary Malicious Code” arxiv.org/abs/2503.12188 1/12
A screenshot of the abstract of the paper, detailing our findings that several multi-agent frameworks can be hijacked to enable a complete security breach.
182
hal @harold.bsky.social · 11/01/2025
do you have ~feelings~ about location sharing culture? i'm editing a project on locations and want to hear from YOU (<5 min) forms.gle/iG1UZJKrcNwm...
a screenshot of the snap map, focused on lower manhattan
062
hal @harold.bsky.social · 09/01/2025
brb updating median voter theory to reflect the fact that 30% of american adults read at a 10yo level or below from on.ft.com/4fBSEwy
3298
hal @harold.bsky.social · 08/01/2025
courtesy of @vrandecic.bsky.social
Comic of a board room meeting.
Panel 1: boss: "Wikipedia is writing bad things about us"
Panel 2, board members suggestions:
Board member 1: "Tell people to stop donating to Wikipedia"
Board member 2: "Doxx and attack Wikipedia volunteers"
Board member 3: "Stop doing bad things?"
Panel 3 and 4: boss looks angry at Board member 3
Panel 5: Board member 3 is been thrown out of the window
040
hal @harold.bsky.social · 11/10/2024
hi world! are you interested in writing stories about campaign finance (or understand how money flows)? 🗣️📈DATATALK📈🗣️ is a platform for asking natural language Qs of FEC data that I've been working on with folks at Stanford, Big Local News, and the Brown Institute datatalk.genie.stanford.edu
111
hal @harold.bsky.social · 19/09/2024
🫣🫣🫣
000
hal @harold.bsky.social · 18/11/2023
new piece out in Reboot! this one is a deep dive into the extreme privacy community — the people who dedicate their lives to hiding from the internet what drives people to want to disappear? what can we learn from them? joinreboot.org/p/threat-model
020
hal @harold.bsky.social · 17/11/2023
[chanting] hey hey! ho ho! automate that CEO! www.nytimes.com/2023/11/17/t...
nytimes.com
OpenAI’s Board Pushes Out Sam Altman, Its High-Profile C.E.O.
Mira Murati, who previously served as chief technology officer, has been named interim chief executive.
010
hal @harold.bsky.social · 07/11/2023
does anyone know anyone / have experience doing archival research into the Manhattan Project? I have a (now passed) family member who claimed that he was a part of the project but refused to tell family members any details about what he did there
000
hal @harold.bsky.social · 29/07/2023
somehow ended up on the wikipedia page for dan savage and came across this absolutely unhinged gem of a story en.wikipedia.org/wiki/Dan_Savage#20…
000
Reposted by hal
Erin Kissane @kissane.myatproto.social · 25/07/2023
If you live in AZ, CO, CT, DE, HI, IL, MI, MN, NH, NM, PA, RI, VA, VT, WI, or WV, you have a Dem senator co-sponsoring KOSA and you can call their office right now to register your opinions. www.blackburn.senate.gov/2023/5/bla…
The Kids Online Safety Act has been cosponsored by U.S. Senators Shelley Moore Capito (R-W.Va.), Ben Ray Luján (D-N.M.), Bill Cassidy (R-La.), Tammy Baldwin (D-Wisc.), Joni Ernst (R-Iowa), Amy Klobuchar (D-Minn.), Gary Peters (D-Mich.), Steve Daines (R-Mont.), Marco Rubio (R-Fla.), John Hickenlooper (D-Colo.), Dan Sullivan (R-Alaska), Chris Murphy (D-Conn.), Todd Young (R-Ind.), Chris Coons (D-Del.), Chuck Grassley (R-Iowa), Brian Schatz (D-Hawaii), Lindsey Graham (R-S.C.), Mark Warner (D-Va.), Roger Marshall (R-Kan.), Peter Welch (D-Vt.), Cindy Hyde-Smith (R-Miss.), Maggie Hassan (D-N.H.), Markwayne Mullin (R-Okla.), Dick Durbin (D-Ill.), Jim Risch (R-Idaho), Sheldon Whitehouse (D-R.I.), Katie Britt (R-Ala.), Bob Casey (D-Penn.), Rick Scott (R-Fla.), Cynthia Lummis (R-Wyo.), John Cornyn (R-Texas.), Lisa Murkowski (R-Alaska), Roger Marshall (R-Miss.), Mark Kelly (D-Ariz.), Joe Manchin (D-W.Va.), James Lankford (R-Okla.), and Mike Crapo (R-Idaho).
18464576
Reposted by hal
Rick Caruso’s Private Fire Crew @amandasmith.bsky.social · 24/07/2023
“may you find the social media home you deserve” is a modern Yiddish curse
2578130
Reposted by hal
Nicholas Grossman @nicholasgrossman.bsky.social · 23/07/2023
“‘The Social Network’ made me want to get into entrepreneurship”is like “‘Wall Street’ made we want to get into finance,” and overlaps heavily with thinking the point of “Scarface” and “Fight Club” is that Tony Montana and Tyler Durden are super cool guys with awesome lives you should emulate.
914027
Reposted by hal
Amy @lolennui.bsky.social · 18/07/2023
Hearing disturbing rumors that some of these protestors on the picket line are professional actors
563647862
Reposted by hal
full slack @fullslack.bsky.social · 12/06/2023
this picture of Biden looking at a quantum computer is fucking hilarious lmao
68846147
Reposted by hal
Kristi Yamaguccimane @wapplehouse.bsky.social · 11/06/2023
this is the funniest thumbnail I’ve ever seen
30605114
hal @harold.bsky.social · 10/06/2023
eugene v debs donald trump 🤝
020
Reposted by hal
Mary Gillis @marygillis.bsky.social · 04/06/2023
Sci-fi writer Ted Chiang: ‘The machines we have now are not conscious’ www.ft.com/content/c1f6d948-3dde-40…
"There was an exchange on Twitter a while back where someone said, ‘What is artificial intelligence?’ And someone else said, ‘A poor choice of words in 1954’,” he says. “And, you know, they’re right. I think that if we had chosen a different phrase for it, back in the ’50s, we might have avoided a lot of the confusion that we’re having now.” So if he had to invent a term, what would it be? His answer is instant: applied statistics.
14295125
hal @harold.bsky.social · 19/05/2023
this is so cool
000
Reposted by hal
Casey Newton @caseynewton.bsky.social · 15/05/2023
Fascinating: humans unexpectedly regained their advantage in Go because they created situations AIs never had seen in their training data franklantz.substack.com/p/the-after…
But Utopia is not a place, it’s a process. And it turns out the story doesn’t end there. A couple of months ago, a group of researchers announced that they had discovered a technique that allows an amateur-level human player to consistently beat state of the art superhuman Go AI. What’s amazing about this adversarial policy, as they call it, is that it’s simple to explain and understand. Basically, it involves allowing the AI to surround a group of your stones and then surrounding the group that is doing the surrounding. Because the AI considers your innermost group dead it doesn’t see the threat coming until it’s too late.

The technique works because this is a type of situation that doesn’t often occur “naturally”, and therefore the AI has a kind of blind spot, and misses what would be glaringly obvious to even a novice player. And this blind spot includes the basic principles of life and death the game is built on, concepts that are completely fundamental to how we think about the g
911942
Reposted by hal
Alana McLaughlin🏳️‍⚧️ @alanaferal.bsky.social · 03/05/2023
39325
Reposted by hal
socks @socks.bsky.social · 30/04/2023
this screenshot is the most concise summary of lesswrong-ai-xrisk culture that you'll ever find
152
hal @harold.bsky.social · 20/04/2023
parallel eyes parallelize parallel lies (this bleet sponsored by Apache Spark)
081
hal @harold.bsky.social · 17/04/2023
how are we feeling today fellow youths
040
hal @harold.bsky.social · 17/04/2023
the nanny (1993-1999) does not get ANYWHERE near enough respect compared to seinfeld/curb your enthusiasm it just hits
110
hal @harold.bsky.social · 15/04/2023
thread of words i like: swimmingly mellifluous filch bonkers obsidian bupkis (more to come)
020
hal @harold.bsky.social · 12/04/2023
new platform as protocol who dis
030