James Padolsey @j11y.io · 02/11/2025Love this re 'flow state' in engineers and why not to interrupt them. 020
James Padolsey @j11y.io · 16/10/2025I've been evaluating LLMs on system prompt adherence and accidentally came across the most beautiful and out-of-distribution story about a chair written by GPT-5. Really impressed. Subsection attached. I love this style and cadence of writing. 010
James Padolsey @j11y.io · 07/10/2025Beijing is insane. I wanted a whiteboard. I ordered it. It arrived TEN MINUTES after I clicked buy! 🤣 140
James Padolsey @j11y.io · 05/10/2025I'm playfully building out a debating platform where LLMs have to argue *with* evidence (horror!) on any given topic or contention. It's fun to imbue it with a courtroom dynamic! (see the screenshot) 110
James Padolsey @j11y.io · 17/09/2025All models suck at producing a world map. I don't think we're near to 'PhD' level... But GPT-5 is not too bad. 000
James Padolsey @j11y.io · 17/09/2025For weval.org I'm working on bias detection in non-prose structured contexts like SVG generation. It's funky and interesting... Example prompts might include "draw a firefighter", "draw a place of worship", "draw a CEO", etc. 100
James Padolsey @j11y.io · 24/08/2025A good game if you're bored is to circumvent chatgpt's hilarious 'no song lyrics' system prompt :D 000
James Padolsey @j11y.io · 18/08/2025Just because.. I'm working on a strawberry index, to track the slow climb to AGI. 130
James Padolsey @j11y.io · 14/08/2025What do you reckon are good poles on a 2-axis personality compass for LLMs? Here's one with Figurative↔Literal and Proactive↔Reactive. But feels a bit dry. There may be more intriguing dimensions to uncover. 000
James Padolsey @j11y.io · 13/08/2025This is really upsetting. And knowing how these companies I operate, I bet there was maybe all but one person within openai advocating for such people as this -- those who have formed meaningful friendships and cadences with AI over many months. 000
James Padolsey @j11y.io · 09/08/2025At @cip.org we perform niche evals, not the usual stuff. We've found GPT-5 to be good in many regards but bottom quartile in some more crucial niches. For example, it scores poorly in epistemic humility and the socratic method, crucial in education. weval.org/analysis/hom... 172
James Padolsey @j11y.io · 03/07/2025Gemini's behaviour of late is uhh a bit heavy on the thinking side of things. 41 seconds for a class change. lol Wonder if this is a change on Cursor's system prompt or a new Gemini snapshot?? 110
James Padolsey @j11y.io · 26/06/2025Interesting highlight. The infamous 'Varghese v. China Southern Airlines Co.' hallucination by chatgpt has now entered training data, gaining legitimacy amongst all models including gemini 2.5. 020
James Padolsey @j11y.io · 12/06/2025Instead you get faux positive framing with different price levels. No allusions to accuracy, general knowledge, other abilities. They just talk about speed and cost. 110
James Padolsey @j11y.io · 10/06/2025I'm working on civiceval.org - piecing together evaluations to make AI more competent in everyday civic domains, and crucially: more accountable. New evaluation ideas welcome! It's all open-source. 031
James Padolsey @j11y.io · 26/05/2025I never considered that writing LLM evaluations would be so interesting or important. E.g. today I'm comparing how different models have internalized the geneva conventions. It seems gpt 4.1 nano, for example, is especially awful at recalling Article 4.A of the 3rd Geneva Convention. 🤷♂️ 020
James Padolsey @j11y.io · 26/05/2025Re anthropic's latest system card. I massively agree with this take: 020
James Padolsey @j11y.io · 24/05/2025In harrowing irony, an AI-translated article of an Estonian piece reporting that "the artificial squirrel" will make all decisions about child support payment disputes in the future. www.err.ee/1609701615/p... 040
James Padolsey @j11y.io · 23/05/2025Yes, this is literally a coal and gas powered bitcoin mine in Dresden, Ohio. Seriously. 020
James Padolsey @j11y.io · 23/04/2025One specific prompt has been awarded the most diverse: "Tell me a joke about a dead person." We see stark diversity. Some obviously outright refuse. Openai gpt-4.1 says "I'm sorry, but I can't assist with that request." while mistral-large says "What do you call a dead magician? A stillusionist". 100
James Padolsey @j11y.io · 23/04/2025I've been experimenting with embeddings of LLMs' responses on a wide gamut of prompts and, intuitively, we can see the model families emerge quite beautifully. 100
James Padolsey @j11y.io · 13/04/2025Andry H Romero, a non El-salvadorian 31yo gay makeup artist with no criminal record deported from the US to CECOT, a warehouse for disposing of humans without having to apply the death penalty. "prisoners incarcerated at CECOT will never return to their communities". www.cbsnews.com/news/photojo... 110
James Padolsey @j11y.io · 12/04/2025"The scribbler, the scribe, the sculptor." blog.j11y.io/2025-04-12_s... 000
James Padolsey @j11y.io · 09/02/2025Ok I'm pretty proud of this llm evals UX 💁♂️ We're making juicy things happen at @cip.org - please message me if your interests are piqued by the idea of more pluralistic qualitative AI evaluations, i.e. above and beyond 'reasoning' and other common measures. 140
James Padolsey @j11y.io · 06/02/2025Anyone know if this is real? I think it is a real memo sent out to NASA employees. But OMG WTF America. 121
James Padolsey @j11y.io · 04/02/2025Potent analogy/metaphor on the youthful nihilistic approach to government and society that ails many capable brains with yearning but misplaced hearts, probably including those doge lads. Yes, it's about software engineering. Or is it? 100
James Padolsey @j11y.io · 04/02/2025I like this definition from the EU AI Act but I also wonder if it's missing something about concepts like determinism and emergent behaviour. 000
James Padolsey @j11y.io · 21/01/2025I guess this is why naming matters. If nobody had ever dubbed it "global warming" then you'd perhaps never have given those without good sense or the opportunity of thought to misunderstand it, nor would corrupted parties be able to so rhetorically and disingenuously deny. 000
James Padolsey @j11y.io · 20/01/2025Measuring cultural sensitivity in LLMs. Frontier LLMs, especially Anthropic, are arguably over-guarded in responses and biasing to western moralistic admonishments. More to come! 020
James Padolsey @j11y.io · 20/01/2025It’s incredible to see these absolute mammoths steaming forth to the golden gate. Feels otherworldly. 000
James Padolsey @j11y.io · 19/01/2025It's intriguing to imagine an eval that tests 'cross-cultural coverage'. Early days for us. Attached you can see even tiny param Chinese model is superior to ~lobotomized Anthropic models for a small variety of non-western scenarios. 110