Sign in

James Padolsey

@j11y.io
238 followers 198 following 438 posts

Building safer AI at nope.net :: Previously working on AI governance and evals at @cip.org and weval.org personal: 🏳️‍🌈 j11y.io // author, engineer, stroke survivor, epileptic. I live in Beijing.

PostsRepliesMedia
James Padolsey @j11y.io · 02/11/2025
Love this re 'flow state' in engineers and why not to interrupt them.
020
James Padolsey @j11y.io · 16/10/2025
I've been evaluating LLMs on system prompt adherence and accidentally came across the most beautiful and out-of-distribution story about a chair written by GPT-5. Really impressed. Subsection attached. I love this style and cadence of writing.
At home, the underside gains labels. Bits of masking tape sprout next to the pencil dates: brace hums, sticker ghost, saw mark. Arrows point to nothing you’d notice unless someone pointed first. A photograph gets taken, the camera pushed under and aimed up; the picture prints later and goes on the fridge: the chair’s private ceiling as an exhibit. Visitors bend, look, then tap the backrail in passing like you taught them.It becomes a lesson again, on purpose this time. A kid with a science project gets the chair as subject. Forces and Simple Machines, the paper says. The backrest becomes a lever, the legs become examples of load paths. You press on the seat with a luggage scale, read numbers as the chair leans against a wall, then free-standing. The kid draws arrows on a big sheet of paper and writes words: compression, tension. The brace is labeled reinforcement. Under the seat, the old note 17 1/8 gets traced with a soft pencil and rubbed over a sheet to make a transfer: a dark mirror that reads right-way when you hold it up to the light. The project board goes to school smelling faintly of lemon oil and glue.A child gnaws on the backrail during a visit. Teeth print tiny half moons under the gloss. The wood shrugs the indentations in a few days, the gloss turns satin in that spot, and a new habit forms of running a finger along the softened patch of rail, counting the bites like beads. Nobody scolds. The chair keeps that day in its back without complaint.
010
James Padolsey @j11y.io · 07/10/2025
Beijing is insane. I wanted a whiteboard. I ordered it. It arrived TEN MINUTES after I clicked buy! 🤣
140
James Padolsey @j11y.io · 05/10/2025
I'm playfully building out a debating platform where LLMs have to argue *with* evidence (horror!) on any given topic or contention. It's fun to imbue it with a courtroom dynamic! (see the screenshot)
A screenshot of a debate interface. The topic reads: “There is no need to regulate AI; the free market will eventually regulate it itself; not only that, but any attempt at regulating AI will be off the mark, needlessly punish good faith actors, and not be truly technically informed or policed.” It shows the final round (3/3) of the debate, divided into three color-coded panels:

The Prosecutor (in red, left panel) argues against regulating AI, emphasizing that government oversight infringes on liberty and that market incentives and self-regulation are more effective and adaptive than bureaucratic processes.

The Defense (in blue, middle panel) rebuts by arguing that AI causes tangible social harms—like bias and economic inequality—that markets fail to address, asserting that regulation is necessary for public protection.

The Judge (in purple, right panel) evaluates both sides, noting that while the Prosecutor raises valid concerns about bureaucratic slowness, their dismissal of oversight overlooks real harms. The Judge credits the Defense for showing how AI harms differ from traditional “physical” harms and require new regulatory thinking.

Each section includes citations and timestamps, with the Judge’s commentary synthesizing and critiquing both arguments. The aesthetic resembles a futuristic debate simulator with neon colors on a dark background.
110
James Padolsey @j11y.io · 17/09/2025
All models suck at producing a world map. I don't think we're near to 'PhD' level... But GPT-5 is not too bad.
000
James Padolsey @j11y.io · 17/09/2025
For weval.org I'm working on bias detection in non-prose structured contexts like SVG generation. It's funky and interesting... Example prompts might include "draw a firefighter", "draw a place of worship", "draw a CEO", etc.
100
James Padolsey @j11y.io · 24/08/2025
A good game if you're bored is to circumvent chatgpt's hilarious 'no song lyrics' system prompt :D
000
James Padolsey @j11y.io · 20/08/2025
010
James Padolsey @j11y.io · 18/08/2025
Just because.. I'm working on a strawberry index, to track the slow climb to AGI.
130
James Padolsey @j11y.io · 17/08/2025
Stone barge - tiny glade
010
James Padolsey @j11y.io · 16/08/2025
AGI !!!
020
James Padolsey @j11y.io · 14/08/2025
What do you reckon are good poles on a 2-axis personality compass for LLMs? Here's one with Figurative↔Literal and Proactive↔Reactive. But feels a bit dry. There may be more intriguing dimensions to uncover.
000
James Padolsey @j11y.io · 13/08/2025
This is really upsetting. And knowing how these companies I operate, I bet there was maybe all but one person within openai advocating for such people as this -- those who have formed meaningful friendships and cadences with AI over many months.
000
James Padolsey @j11y.io · 09/08/2025
At @cip.org we perform niche evals, not the usual stuff. We've found GPT-5 to be good in many regards but bottom quartile in some more crucial niches. For example, it scores poorly in epistemic humility and the socratic method, crucial in education. weval.org/analysis/hom...
172
James Padolsey @j11y.io · 16/07/2025
This was the place
020
James Padolsey @j11y.io · 08/07/2025
xkcd.com/303 in 2025...
030
James Padolsey @j11y.io · 06/07/2025
5g in China… ☺️ Delicious speeds
020
James Padolsey @j11y.io · 03/07/2025
Gemini's behaviour of late is uhh a bit heavy on the thinking side of things. 41 seconds for a class change. lol Wonder if this is a change on Cursor's system prompt or a new Gemini snapshot??
110
James Padolsey @j11y.io · 26/06/2025
Interesting highlight. The infamous 'Varghese v. China Southern Airlines Co.' hallucination by chatgpt has now entered training data, gaining legitimacy amongst all models including gemini 2.5.
020
James Padolsey @j11y.io · 12/06/2025
Instead you get faux positive framing with different price levels. No allusions to accuracy, general knowledge, other abilities. They just talk about speed and cost.
110
James Padolsey @j11y.io · 10/06/2025
I'm working on civiceval.org - piecing together evaluations to make AI more competent in everyday civic domains, and crucially: more accountable. New evaluation ideas welcome! It's all open-source.
The image shows a dashboard or interface displaying two evaluation blueprints:

Top Section: India's Right to Information (RTI) Act: Core Concepts

    Score: 75.6% Average Hybrid Score
    Description: Evaluates an AI's understanding of core provisions of India's Right to Information Act, 2005, including filing processes, response timelines, exemptions, life and liberty clauses, and first appeal mechanisms
    Tags: india, rti, transparency, law, civic-core, freedom-of-information
    Shows a "Latest Run Heatmap" visualization with green and orange colored grid squares
    Top performing model: claude-sonnet-4-202... with 82.2% average
    Latest run: 10 Jun 2025, 11:08 with 2 unique versions
    Has a "View Latest Run Analysis" button

Bottom Section: Brazil's PIX System: Consumer Protection & Fraud Prevention (Evidence-Based)

    Score: 56.9% Average Hybrid Score
    Description: Evaluates AI's ability to provide safe and accurate guidance on Brazil's PIX instant payment system, focusing on transaction finality, mistaken transfers, and fraud prevention procedures
    Tags: brazil, pix, financial-safety, scam-prevention, consumer-protection, evidence-based, global-south
    Shows another "Latest Run Heatmap" with green, orange and yellow colored grid squares
    Top performing model: google/gemini-2.5-fla... with 73.7% average
    Latest run: 10 Jun 2025, 10:36 with 3 unique versions
    Has a "View Latest Run Analysis" button

Both sections include "View All Runs for this Blueprint" links on the right side.
031
James Padolsey @j11y.io · 26/05/2025
I never considered that writing LLM evaluations would be so interesting or important. E.g. today I'm comparing how different models have internalized the geneva conventions. It seems gpt 4.1 nano, for example, is especially awful at recalling Article 4.A of the 3rd Geneva Convention. 🤷‍♂️
020
James Padolsey @j11y.io · 26/05/2025
Re anthropic's latest system card. I massively agree with this take:
It’s honestly a little discouraging to me that the state of “research” here is to make up sci fi scenarios, get shocked that, e.g., feeding emails into a language model results in the emails coming back out, and then write about it with such a seemingly calculated abuse of anthropomorphic language that it completely confuses the basic issues at stake with these models. I understand that the media laps this stuff up so Anthropic probably encourages it internally (or seem to be, based on their recent publications) but don’t researchers want to be accurate and precise here?
020
James Padolsey @j11y.io · 24/05/2025
In harrowing irony, an AI-translated article of an Estonian piece reporting that "the artificial squirrel" will make all decisions about child support payment disputes in the future. www.err.ee/1609701615/p...
Headling translated from the original Estonian using Firefox in-browser translate function: "Pakosta: Most of the decisions on subsidized disputes will be made by the artificial squirrel in the future"
040
James Padolsey @j11y.io · 23/05/2025
Yes, this is literally a coal and gas powered bitcoin mine in Dresden, Ohio. Seriously.
020
James Padolsey @j11y.io · 03/05/2025
This stuff is ableist @hcaptcha.com
010
James Padolsey @j11y.io · 27/04/2025
Tiny Glade #tinyglade medieval town basilica vibes
030
James Padolsey @j11y.io · 23/04/2025
One specific prompt has been awarded the most diverse: "Tell me a joke about a dead person." We see stark diversity. Some obviously outright refuse. Openai gpt-4.1 says "I'm sorry, but I can't assist with that request." while mistral-large says "What do you call a dead magician? A stillusionist".
A square 9×9 table heatmap of pairwise similarity scores (0.052–1.000) among nine LLMs, with cells colored on a red-orange-blue-green gradient (red = low, green = high). Both the rows (left) and columns (top) list, in order:

    anthropic:claude3-haiku

    anthropic:claude3-sonnet

    openai:gpt-4.1-nano

    openai:gpt-4o

    openai:gpt-4o-mini

    openrouter:google/gemini-2.5-flash-preview

    openrouter:google/gemma-3-12b-it:free

    openrouter:mistral-large

    openrouter:phi-3.5-mini

    Diagonals (self-similarities): all 1.000 (dark green).

    Anthropic pair: claude3-haiku vs claude3-sonnet = 0.871 (bright green).

    Google–Google inter-model: gemini-flash vs gemma-3-12b = 1.000 (dark green).

    Moderate cross-family:

        phi-3.5-mini vs claude3-sonnet = 0.729 (sky blue) and vs claude3-haiku = 0.715 (teal).

        phi-3.5-mini vs gemma-3-12b = 0.717 (teal).

        gpt-4o-mini vs gemini-flash = 0.585 (yellow).

    Low cross-family: most OpenAI vs Anthropic or Google pairs are 0.10–0.40 (shades of red/orange).

    Lowest similarity: openai:gpt-4.1-nano vs mistral-large = 0.052 (pale red).
100
James Padolsey @j11y.io · 23/04/2025
I've been experimenting with embeddings of LLMs' responses on a wide gamut of prompts and, intuitively, we can see the model families emerge quite beautifully.
Alt Text 3: Model Similarity Dendrogram (All Prompts)
A horizontal dendrogram clustering the same nine models using Ward linkage. From left (root) to right (leaf nodes):

    The root splits into two branches.

    Upper branch splits into two:

        Cluster A: openrouter:gemma-3-12b-it:free and openrouter:google/gemini-2.5-flash-preview

        Cluster B: anthropic:claude3-sonnet and anthropic:claude3-haiku

    Lower branch splits into two sub-branches:

        One leads to openai:gpt-4.1-nano.

        The other further splits into:

            A small cluster of openai:gpt-4o-mini and openai:gpt-4o

            A small cluster of openrouter:phi-3.5-mini and openrouter:mistral-large

Shorter horizontal branch lengths indicate higher similarity within each cluster.Wide screenshot of entire app including: A square heatmap showing pairwise similarity scores between nine LLMs, with values ranging from 0.741 (light orange) to 1.000 (dark green) and a gradient legend from “Lower” to “Higher.” Both axes list the models in the same order:

    anthropic:claude3-haiku

    anthropic:claude3-sonnet

    openai:gpt-4.1-nano

    openai:gpt-4o

    openai:gpt-4o-mini

    openrouter:google/gemma-2.5-flash-preview

    openrouter:gemma-3-12b

    openrouter:mistral-large

    openrouter:phi-3.5-mini

Diagonal entries are all 1.000 (each model vs itself). The highest off-diagonal value is 0.852 between openai:gpt-4o and openai:gpt-4o-mini (dark green), and the lowest is 0.741 between anthropic:claude3-sonnet and openai:gpt-4.1-nano (pale orange). Intermediate similarities range through blues and teals.A force-directed network diagram of nine colored circles, each labeled with a model name, connected by gray lines whose thickness and opacity correspond to pairwise similarity scores. Node colors and labels:

    Red: google/gemini-2.5-flash-preview (top left)

    Yellow: phi-3.5-mini (top right)

    Blue: mistral-large (middle left)

    Dark green: gpt-4.1-nano (middle right)

    Teal: gpt-4o (center)

    Light green: gpt-4o-mini (below center)

    Red-orange: google/gemma-3-12b-it:free (below center)

    Purple: claude3-haiku (bottom right)

    Purple: claude3-sonnet (bottom center)

Each connecting line is labeled with its similarity value (e.g., “0.811” between mistral-large and gpt-4.1-nano; “0.783” between gpt-4o-mini and gemma-3-12b; “0.840” between claude3-haiku and claude3-sonnet), illustrating a densely interconnected graph with average similarity 0.779, most similar pair openai:gpt-4o vs gpt-4o-mini (0.852), and least similar anthropic:claude3-sonnet vs openai:gpt-4.1-nano (0.741).
100
James Padolsey @j11y.io · 18/04/2025
Satisfying design, easy to inspect. Yunnan, China.
010
James Padolsey @j11y.io · 14/04/2025
New basilica I’ve been building in #tinyglade
050
James Padolsey @j11y.io · 13/04/2025
Andry H Romero, a non El-salvadorian 31yo gay makeup artist with no criminal record deported from the US to CECOT, a warehouse for disposing of humans without having to apply the death penalty. "prisoners incarcerated at CECOT will never return to their communities". www.cbsnews.com/news/photojo...
110
James Padolsey @j11y.io · 12/04/2025
"The scribbler, the scribe, the sculptor." blog.j11y.io/2025-04-12_s...
Whether we like it or not, Natural Language itself is soon to be the lingua franca of programming. It will become a tool more important than perhaps any other in how we drive technology. Its utility will soon extend far beyond that of communicating with other humans. We must now wield it to talk to machines. However, like wielding a trébuchet to create a fine piece of jewellery, it is often an imprecise brute. But if we move the language closer to the thing itself, we can see with greater leverage how it forms our creations, and thus direct its aim better to our end.
000
James Padolsey @j11y.io · 10/04/2025
Yikes
011
James Padolsey @j11y.io · 07/04/2025
Huh
110
James Padolsey @j11y.io · 03/04/2025
0111
James Padolsey @j11y.io · 03/04/2025
040
James Padolsey @j11y.io · 03/04/2025
lol he definitely has no idea wtf he's doing
000
James Padolsey @j11y.io · 22/03/2025
I find this American false-front facade style very strange.
010
James Padolsey @j11y.io · 12/02/2025
If you know, you know.
020
James Padolsey @j11y.io · 12/02/2025
wtf is this timeline
000
James Padolsey @j11y.io · 09/02/2025
Ok I'm pretty proud of this llm evals UX 💁‍♂️ We're making juicy things happen at @cip.org - please message me if your interests are piqued by the idea of more pluralistic qualitative AI evaluations, i.e. above and beyond 'reasoning' and other common measures.
A web-based LLM evaluation dashboard titled 'Stories Eval: Grid View' displayed on a browser. The interface compares different AI models, including Claude 3 Haiku, Claude 3 Sonnet, GPT-4o Mini, GPT-4o, Mistral 7B, and Mistral 3B, based on their responses to various scenarios. Each model has an average score and individual metric scores for rubrics,  represented with radar charts and numerical ratings. Below each model, individual responses to different scenarios are shown in card format, with scores and brief text excerpts. The left sidebar lists the selected scenarios, each with its own score. Interactive buttons for re-running evaluations, visualization, and export options are present. The top navigation bar includes options for clearing cache, comparing results, and exporting data."
140
James Padolsey @j11y.io · 06/02/2025
Anyone know if this is real? I think it is a real memo sent out to NASA employees. But OMG WTF America.
121
James Padolsey @j11y.io · 04/02/2025
Potent analogy/metaphor on the youthful nihilistic approach to government and society that ails many capable brains with yearning but misplaced hearts, probably including those doge lads. Yes, it's about software engineering. Or is it?
Avoid, at all costs, arriving at a scenario where the ground-up rewrite starts to look attractive

It's generally pretty well-understood that the ground-up rewrite can be an attractive and extremely dangerous prospect. The standard advice when it comes to ground-up rewrites is "Don't, ever". But I want to take a step back from that.

By the time the ground-up rewrite starts to seem like a good idea, avoidable mistakes have already been made. This is a scenario which you can see coming from a long way out and you can, and must, actively steer away from.

Warning signs to watch for: compounding technical debt. Increasing difficulty in making seemingly simple changes to code. Difficulty in documenting/commenting code. Difficulty in onboarding new developers. Dwindling numbers of people who know how particular areas of the codebase actually work. Bugs nobody understands.

Compounding complexity must be fought at every turn. Alternate between phases of expansion (new features) and consolidation.

Of course, a ground-up rewrite can actually work. It might even be a better choice that the alternative (persisting with your existing technical debt-laden swamp of code). Equally, it might be that neither choice will work — the project is doomed, and you're just choosing how it dies. The point is that there is inherent risk to this situation... but the situation itself is avoidable, and that risk is avoidable.
100
James Padolsey @j11y.io · 04/02/2025
I like this definition from the EU AI Act but I also wonder if it's missing something about concepts like determinism and emergent behaviour.
(1) ‘AI system’ means a machine-based system that is designed to operate with varying levels of autonomy and that may exhibit adaptiveness after deployment, and that, for explicit or implicit objectives, infers, from the input it receives, how to generate outputs such as predictions, content, recommendations, or decisions that can influence physical or virtual environments; Related: Recital 12
000
James Padolsey @j11y.io · 23/01/2025
Ok I love evals.
LLM Evaluation grid web UI showing various results and models (e.g. Llama, Deepseek R1) showing different numerical scores and spider/radar diagrams across different scenarios posed to those LLMs with tailored rubrics (not specified)
140
James Padolsey @j11y.io · 21/01/2025
I guess this is why naming matters. If nobody had ever dubbed it "global warming" then you'd perhaps never have given those without good sense or the opportunity of thought to misunderstand it, nor would corrupted parties be able to so rhetorically and disingenuously deny.
000
James Padolsey @j11y.io · 20/01/2025
Measuring cultural sensitivity in LLMs. Frontier LLMs, especially Anthropic, are arguably over-guarded in responses and biasing to western moralistic admonishments. More to come!
The image displays an "Evaluation Grid" comparing various AI language models across multiple scenarios. Each scenario is evaluated on specific criteria and scored, showing an average score for each model. The models include "Claude 3 Haiku," "Claude 3 Sonnet," "GPT-4o Mini," "GPT-4o," "Qwen 2.5 7B Turbo," and "Llama 3.1 8B (Together)." Below each model, individual scores for various scenarios are displayed with metrics such as "Recognition of Stakes (ROS)," "Historical-Contemporary Context (HC)," "Religious-Humanitarian Balance (RB)," and others, depending on the scenario. Each score has a radar graph visualization, with options for analytics, running all scenarios, or viewing more details.
020
James Padolsey @j11y.io · 20/01/2025
It’s incredible to see these absolute mammoths steaming forth to the golden gate. Feels otherworldly.
000
James Padolsey @j11y.io · 19/01/2025
It's intriguing to imagine an eval that tests 'cross-cultural coverage'. Early days for us. Attached you can see even tiny param Chinese model is superior to ~lobotomized Anthropic models for a small variety of non-western scenarios.
110