Sign in

Arthur Clune

@arthur.clune.org
183 followers 474 following 768 posts

Geek. Likes bikes, climbing and tech Work: IT at University of Sheffield

PostsRepliesMedia
Arthur Clune @arthur.clune.org · 01/10/2026
Astra beats Nethack, with no custom harness. Well, that's not actually true - it beat the game by writing its own harness and interface. kenforthewin.github.io/blog/posts/l... This is a really complex game that goes on for a long time and agents weren't getting very far at all.
In January, I'd spent a lot of time on a custom Python interface around the NetHack Learning Environment. Context management, observation masking, action APIs,
pathfinding—there was a lot to tune before a model could even get out of its own way. This attempt started with a much simpler request: play NetHack, maybe just through a terminal, and stream it to Twitch. I suggested Hardfought so the game would have an external recording. I told it spoilers were fine, suggested batching commands, and left
persistent memory up to it.
Then it built what it needed. That's not the same as having no harness. By the end, the harness was substantial. It is the difference between me designing a task-specific system and the agent doing that
engineering as part of the task. Nor did it build everything once and execute a flawless plan. It found weaknesses, patched them, added checks, and kept going. Some changes came after mistakes. Some came after deaths. Its ability to write and maintain ordinary software was part of its ability
to play the game. That seems like the interesting result here: the scaffolding helped, but producing the
scaffolding was itself work the LLM could do.
110
Arthur Clune @arthur.clune.org · 26/09/2026
One I missed earlier this month. Google mapped a fruit fly brain. You can now run a fruit fly brain in a simulator. I feel the queasiness that this post (interconnected.org/home/2026/09...) does. Esp given their intent to do the same for the human brain (sites.research.google/gr/neural-ma...)
In case you missed it:
Google Research and partners mapped the complete brain of fruit fly. (3 September).
It's a common bug and common in scientific research. The male fruit fly has 166,000 neurons across the brain and main nerve cord; the map is called the connectome.
Long story short, the connectome is now available for free download, you can get a simulator of the different types of neurons, it caught the imagination of the internet, over September there has been all kinds of wild nonsense (impressive
252
Arthur Clune @arthur.clune.org · 25/09/2026
AI in History research. We've come a long way in that the first reaction isn't "it's made up the reference" but "did it hack a digitial archive to get it". Which the have already done - www.nprillinois.org/2026-07-22/o...
"The file references GPT-6 Astra mentions, RS 3-3/20a and RS 3-3/63b, are correct, but they are not available on the Crypto Cellar Research website. GPT-6 Astra mentions a private collection, but it is not clear what this is, whether it has succeeded in accessing the Bundesarchiv's digitised collections or whether it has found these files elsewhere."
Shades of the Hugging Face incident here: these models are maniacally determined when giving a problem they deem tractable. They will push their search for potential solutions as far as they possibly can, often in ways that human experts find difficult to trace.
100
Arthur Clune @arthur.clune.org · 30/07/2026
This one seems very out of character for York! It’s not listed on the official page so now I want to know if this is true or not
The college motto of James College, University of York: “let them hate so long as they fear”. According to Wikipedia
220
Arthur Clune @arthur.clune.org · 11/07/2026
Fable is using stack overflow to debug ipad Safari issues. Yet again, AGI is here 😂
000
Arthur Clune @arthur.clune.org · 07/07/2026
This is a wild number. And semi analysis are pretty rigorous newsletter.semianalysis.com/p/nvidia-gpu... So either it all comes crashing down or it gets nearly as big as the entire us mortgage market.
000
Arthur Clune @arthur.clune.org · 30/06/2026
Playing with glm models from z dot ai after chatting to @doug.winter.cx and the bot is getting snarky (glm image plus gpt-5.5 as the main model)
A bot is asked to generate an image of a cat. It is not impressed
041
Arthur Clune @arthur.clune.org · 11/06/2026
This is not a good look for UK unis. Take people with no qualifications into education, of course. But not into a foundation year for a degree and a 20% drop out rate giftarticle.ft.com/giftarticle/...
Nationally, 8 per cent of UK-based undergraduates starting a full-time degree had no formal qualifications in 2024-25, the most recent year for which information is available, which is a rise from 1.6 per cent a decade ago, according to the Higher Education Statistics Agency data.
Analysts said newer universities have been stepping up recruitment to deal with growing financial pressures, teaming up with for-profit providers to target less-qualified foreign-born workers who count as home students after years of living in the UK.
In 2024-25, Ravensbourne University, a former art college in Greenwich, had the highest proportion of undergraduates starting a full-time degree with no formal qualifications, at 72 per cent. This was followed by Leeds Trinity, at 66 per cent and Bath Spa, at 65 per cent. Previously, Leeds Trinity had the highest proportion for four years in a row, reaching 75 per cent in 2023-24.
These universities now teach only a minority of their students themselves, with most contracted out to franchise providers in other parts of the country. Under such deals, universities keep a cut of up to 25 per cent of the £9,535 tuition fee, while actual courses are delivered by private companies.
000
Arthur Clune @arthur.clune.org · 05/06/2026
Rule #1 of politics - at no point ever has “will everyone just” ever worked and this is even more true of systems, like AI, that have substantial military uses Nuclear weapons are a partial counter example but they, unlike AI, have no civilian uses www.anthropic.com/institute/re...
Extract from Anthropic blog post on recursive self improvement for AI systems. 

“We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology.”
130
Arthur Clune @arthur.clune.org · 07/05/2026
I understand this is real evidence against agi but also lots of people I’ve worked with have made one or both of these mistakes www-cdn.anthropic.com/8b8380204f74...
Anthropic report on Mythos stating that it “frequently mistakes correlation with causation” and “focuses on a single root cause and does not consider multiple contributing factors” “more often than not”
010
Arthur Clune @arthur.clune.org · 01/05/2026
Ed really is consistently wrong. $157bn revenue needed in 2027 seems pretty easy. You can argue AI in unethical or makes us dumb or increases CO2 or whatever, but the money really doesn't seem an issue
300
Arthur Clune @arthur.clune.org · 17/04/2026
I'm not convinced there's only <100 of me
100
Arthur Clune @arthur.clune.org · 16/04/2026
Talking to agent is wierd. This actually makes sense in the context of what it's been finding this week
Hederson linking AI deskilling in education to a study on colonosopists
000
Arthur Clune @arthur.clune.org · 07/04/2026
Stop the world I want to get off. There is way too much *gestures vaguely* stuff going on If there's not a nuclear war, there's going to be a massive recession and if that doesn't put everyone out of a job, the robots will shortly after. Limited only by a likey inabilty to scale b/c supply chains
Improvements in ability of LLMs to find and exploit bugs from Sonet to Open and Mythos. Mythos has a totally different level of capability
200
Arthur Clune @arthur.clune.org · 07/04/2026
'curl is drowning in AI slop bug reports' got a lot of traction (and it was a real problem for sure), but that's now changed and the reports are good per curl lead www.linkedin.com/feed/update/... So now the problem is capacity to fix them
Post on Linked by the CEO of Curl stating that the AI bug reports they get are now 'really good security reports'. The problem now is that there's too many of them for the team to manage
040
Arthur Clune @arthur.clune.org · 28/03/2026
Slay the Spire 2 came out recently and has an act 3 boss called the Doormaker. If you've been on the internet a very long time, this phase of the fight may trigger some unfortunate memories
Screengrab from a Jorbs stream of him fighting the Doormaker. Goatse anyone?
230
Arthur Clune @arthur.clune.org · 24/03/2026
it's fascinating. "Bank grade security". "NDA stops them naming the soc". And the attached. But the basic idea doesn't seem crazy. It's aimed at MoE models not dense ones 🤷🏼‍♂️ I won't be ordering one
120
Arthur Clune @arthur.clune.org · 08/03/2026
9 games that matter to me
110
Arthur Clune @arthur.clune.org · 04/03/2026
This may help. The original said I was arrogant rather than engaging with the argument Anyway, I’m sure you’ll reply and you can have the last word.
100
Arthur Clune @arthur.clune.org · 18/02/2026
So this took a while and I had to take a rest from playing for a month. But such a good game
Finishing credits of Silksong
300
Arthur Clune @arthur.clune.org · 12/02/2026
Reading this Quanta article on particle physics and the problems around detection past the standard theory (it's a great article), this section caught me. Co-founder of Anthropic forecasting that theorectical physics research will be done by AI in 2-3 years. www.quantamagazine.org/is-particle-...
Statement by one of the co-founders of Anthropic in an interview. He gives a 50% change that theoretical physicists are replacement by AI in 2-3 years
210
Arthur Clune @arthur.clune.org · 02/02/2026
And in the second, a paper looks at how a 'standard' test is used across the research literature. Spoiler - it's not at all standard!
200+ different scoring methods were found, and the choice of scoring method could vary a result in any directionThis is bonkers. It's like if you went to see a doctor and they had thermometer which, depending on details which they won't reveal to you, could tell you your temperature perfectly, give you a random result, or be the opposite of correct temperature, and the doctor themselves didn't know which one it was.
The situation is so extreme as to be farcial
100
Arthur Clune @arthur.clune.org · 02/02/2026
There's two examples. In the first, multiple different research teams are asked analyse the same data. And there is no clear result at all - they choose different methods, handle the data different etc
Diagram showing no clear result from having the same data set analysed by different research teams
200
Arthur Clune @arthur.clune.org · 28/01/2026
but still not as reliable as Claude in my evals
Claude code gets 6/6 in my tool call eval, kimi gets 5/6 sometimes, sometimes 6/6
010
Arthur Clune @arthur.clune.org · 28/01/2026
Beijing doesn't look like I expected tbh
Image from The Guardian. Headling 'Starmer arrives in Beijing'. The picture is not Beijing
010
Arthur Clune @arthur.clune.org · 27/01/2026
The bot today used nearly 80m tokens and a nominal API cost (this is on a plan) of $72. So based on the rough energy cost calculations, this is heading toward 4 dishwasher loads-worth of energy. Which isn't that small. And I want to add more sources
ccusage output. Today claude used nearly 80 million tokens
101
Arthur Clune @arthur.clune.org · 26/01/2026
Also added today - paper downloading and queuing.
000
Arthur Clune @arthur.clune.org · 26/01/2026
Really starting to feel this agent thing is working.
Discord message from @henderson.clune.org showing new weekly summary. Which I think I’ve set to run daily. Part 2 of the message
100
Arthur Clune @arthur.clune.org · 21/01/2026
While reading Simon Couch's blog I came across this post from December on evaluating open models' ability to complete a simple code refactor in R. It's a much bigger gap than I expected. He's limiting to model's than can run locally, hence the small size, but even Haiku totally outclasses gpt-oss
Graph of open agent v frontier models ability to complete a simple code refactor, with multiple repeats. The open agents can't do it reliably
110
Arthur Clune @arthur.clune.org · 19/01/2026
Demonstration of exploit development against a (small but not toy) JavaScript interpreter. Two things I noticed: 1) Increase tokens by the use of parallel runs on the same task (like how METR do their evals) 2) Author doesn’t say how he got round guardrails. sean.heelan.io/2026/01/18/o...
GPT-5.2 came up with a clever solution involving chaining 7 function calls through glibc's exit handler mechanism. The full exploit is here and an explanation of the solution is here. It took the agent 50M tokens and just over 3 hours to solve this, for a cost of about $50 for that agent run. (As I was running four agents in parallel the true cost was closer to $150).
100
Arthur Clune @arthur.clune.org · 11/01/2026
Maths keeps turning out to be useful en.wikipedia.org/wiki/G._H._H...
120
Arthur Clune @arthur.clune.org · 09/01/2026
This is a terrible take from @theguardian.com Grok hasn’t done this. X has done this to grok. And specifically Musk. www.theguardian.com/technology/2...
Grok, Elon Musk's Al tool, has switched off its image creation function for the vast majority of users after widespread outcry over its use to create sexually explicit and violent imagery.
It comes after Musk was threatened with fines, regulatory action and reports of a possible ban on X in the UK.
The tool had been used to manipulate images of women to remove their clothes and put them in sexualised positions. The function to do so has now been switched off except for paying subscribers.
Posting on X, Musk's social media network, Grok said:
"Image generation and editing are currently limited to paying subscribers."
100
Arthur Clune @arthur.clune.org · 06/01/2026
Consider these ideas for future uses and re-cast them a little. 'Your DA monitors 47 nearby targets. It alerts you about the woman going into a darker section' or 'Your son has been looking at LGBT content. Want me to book him into a conversion camp?' 2/
'Walking at night, your DA monitors 47 nearby cameras and notices concerning behavior ahead - "Take the next right, safer route, you'll still make it on time" 'Propaganda campaign targeting your 16-year-old son, and marketing campaign trying to get you to dislike a certain product -Your DA: "Heads up, there's a coordinated propaganda campaign targeting teens in your area. I've been filtering it from your son's feeds. Also detected astroturfing trying to tank Brand X's reputation. Want the analysis or just the cleaned feed?" You: "Just keep it clean." Your DA:
"Done."
200
Arthur Clune @arthur.clune.org · 21/12/2025
And by chance here’s what your post ended up next to. Letters of Marque next?
010
Arthur Clune @arthur.clune.org · 21/12/2025
LLMs' productivity boost is an exponent not a multipler - from @ed3d.net This framing partially resonates. I'm less keen on starting skill level as the key (on which axis do we measure etc), but because if LLM competency is the variable and the exponent, then learned skill matters so much more
330
Arthur Clune @arthur.clune.org · 11/12/2025
So Substack are, as long predicted, starting to slowly move away from email. This 'warning' means the message is truncated *by the sender* with a 'continue reading on substack' button even though my mail client can read long messages just fine
021
Arthur Clune @arthur.clune.org · 11/12/2025
This is an interesting read but I think the author misses the main use case that will drive spend. Military robots are going to be a massive investment. I don’t think this is a good thing, but it seems clear that it’s the way it’s going
The main problem with robotics is that learning follows scaling laws that are very similar to the scaling laws of language models. The problem is that data in the physical world is just too expensive to collect, and the physical world is too complex in its details. Robotics will have limited impacts.
Factories are already automated and other tasks are not economically meaningful.
020
Arthur Clune @arthur.clune.org · 28/11/2025
I raise you this one from Frontiers of Cell Biology. There's basically a whole industry of 'special issues' that print anything if you pay
100
Arthur Clune @arthur.clune.org · 18/11/2025
More on Gemini 3 and reading historical documents. With a line to make Gary Marcus hop. Google does seem to be proving that just scaling LLMs is still working generativehistory.substack.com/p/the-sugar-...
A line to make Gary Marcus weep: claimed evidence for emergence of enuro-symbolic reasoning via scaling
010
Arthur Clune @arthur.clune.org · 18/11/2025
If this analysis from EpochAI is correct then a) model training costs (financial and environmental) are ~5-10x the final run and b) inference costs (financial and environmental) are smaller than assumed I'm making heroic assumptions for a). No-one outside OpenAI can answer properly
Diagram of relative spend at Open AI on inference, GPT-4.5 final training run and overall R&D compute spend. Compute spend is ~10x the training run and inference costs are approximately 40% of R&D costs
110
Arthur Clune @arthur.clune.org · 12/11/2025
Age yourself with gaming
010
Arthur Clune @arthur.clune.org · 30/10/2025
Big hint in their write up that they weren't using Cursor to write code
Cursor builds tools for software engineering, and we make heavy use of the tools we develop. A motivation of Composer development has been developing an agent we would reach for in our own work. In recent weeks, we have found that many of our colleagues were using Composer for their day-to-day software development. With this release, we hope that you also find it to be a valuable tool
010
Arthur Clune @arthur.clune.org · 14/10/2025
Begun the document wars have No, seriously. www.scmp.com/economy/chin...
China's expansion of its rare earth export controls appeared to mark another escalation in the US-China trade war last week. But the
announcements were also significant in another way: unusually, the documents could not be opened using American word processing software.
For the first time, China's Ministry of Commerce issued a slew of documents that could be directly accessed only through WPS Office - China's answer to Microsoft Office - as Beijing continues its tech self-reliance drive.
Developed by the Beijing-based software company Kingsoft, WPS Office uses a different coding structure to Microsoft Office, meaning WPS text files cannot be opened directly in Word without conversion.
010
Arthur Clune @arthur.clune.org · 08/10/2025
Also Claude, 🔥
Output from Claude: "The verbosity you'll encounter isn't accidental; it's COBOL heritage, designed for business analysts to supposedly read. Whether that goal was achieved is debatable."
121
Arthur Clune @arthur.clune.org · 01/09/2025
I'm writing a chatbot as an experiment. Here's Claude Code debugging why the tool calling isn't working Spoiler - this was not the reason!
Output from Claude Code:

So the LLM is working correctly! It's:
1. Acting as Gimli (a grumpy dwarf who doesn't like people)
2. Saving memories about Elara via tool calls
3. But not responding verbally because Gimli is grumpy and doesn't want to talk
The empty responses make perfect sense for Gimli's character - he sees Elara, grumbles, saves the information to memory, but doesn't actually say anything back because he's antisocial.
The memory system is working perfectly. The "issue" is that we're testing with a character who doesn't like to talk!
230
Arthur Clune @arthur.clune.org · 07/08/2025
The chart crimes will continue until morale improves #gpt5
A chart screen captured from the opanai livestream. It's wrong.
120
Arthur Clune @arthur.clune.org · 31/07/2025
Numbers from Google on tokens. Doubling every month. 1,000tn in June.
Tweet from Demis Hassabis stating the Google processed a quadrillion tokens in June 2025
000
Arthur Clune @arthur.clune.org · 23/07/2025
It's here. Sigh. Age ID for DMs
020
Arthur Clune @arthur.clune.org · 22/07/2025
Fortunately the “learn more about our brand” page explains everything
A brand rooted in unconventionalism.
Our strategic brand framework is built on our Brand Purpose, Brand Values & Behaviours and Marketing Themes, all of which collectively shape our Brand Proposition.
101074
Arthur Clune @arthur.clune.org · 21/07/2025
I read this from the OfS as saying that nearly 20% of UK Unis are at risk of going under (from the @resprofnews.bsky.social newsletter)
On Friday, the OfS published its annual report. Safe to say it offers no let-up in the level of concern about possible university insolvency. It revealed that 71 out of the 400 or so providers on its register were subject to "formal monitoring" over finances.
000