Sign in

Lex

@notesbylex.com
89 followers 11 following 77 posts

Senior MLE at @Canva. Full-stack developer. Talks about Software dev, LLMs, Agentic Reasoning, Obsidian, note-taking, dog photos.

PostsRepliesMedia
Lex @notesbylex.com · 26/09/2026
In this article, I put Claude Opus 5.5 to the test to see if it could find any new leads in the mystery of Satoshi Nakamoto. notesbylex.com/can-claude-...
notesbylex.com
Can Claude Opus 5.5 find any new leads on Satoshi Nakamoto?
Satoshi Nakamoto I'm not a huge crypto guy, but I do think Bitcoin is a fascinating project, both in the technical implementation and the origin story. Like a lot of people, I've had a lot of...
000
Lex @notesbylex.com · 24/09/2026
"Jev is not fundamentally a language model that learned to emit cleaner JSON. It is a preference model promoted into a software interface". Great article. di-zhang-llm.github.io/blog/what-i...
di-zhang-llm.github.io
What Is RLCD? The Secret Behind Jev
RLCD is a calibrated, schema-conditioned extension of pairwise reward modeling: Bradley–Terry becomes Plackett–Luce, and the reward model becomes Jev’s typed decision interface.
000
Lex @notesbylex.com · 23/09/2026
Opus 5.5 - posted a few notes and a demo on my blog. Major difference from previous Claude versions is that thinking cannot be disabled, as usual substantial improvement on benchmarks. notesbylex.com/claude-opus...
notesbylex.com
Claude Opus 5.5
Claude Opus 5.5 is Anthropic's first model in the Claude 5.5 family, released on 22 September 2026. It improves on Opus 5 in coding, computer use and knowledge work, while reducing the price...
000
Lex @notesbylex.com · 21/09/2026
Jev is a cool project, but it's kinda surprising how enthusiastic people are about classification models again.
000
Lex @notesbylex.com · 12/09/2026
After about a year of wearing a smart ring, I've decided to throw in the towel. A short essay about the downsides to smart rings: notesbylex.com/giving-up-o...
notesbylex.com
Giving Up on Smart Rings
After about a year of wearing a smart ring, I've decided to throw in the towel. I did really like my Oura Ring. Tracking my sleep and steps, amongst other things, has been really helpful for...
101
Lex @notesbylex.com · 06/08/2026
I wrote some notes and a code examples about Stateless MCP servers notesbylex.com/what-is-sta...
notesbylex.com
What is Stateless MCP?
Stateless MCP is the new stateless protocol core for MCP introduced in the 2026-07-28 Specification. The original MCP specification described a handshake protocol that started with an...
010
Lex @notesbylex.com · 05/08/2026
LLMs are still the wrong tool for tabular data "no amount of prompt polishing aimed at for- mat, precision, or batching should be expected to fix high-dimensional in-context prediction, and purpose-built tabular models remain the right tool" arxiv.org/abs/2608.02412
arxiv.org
Why Large Language Models Fail at Tabular Prediction
Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads:...
000
Lex @notesbylex.com · 05/08/2026
Stateless MCP: "MCP is transforming from a bidirectional stateful protocol into a request/response stateless protocol" blog.modelcontextprotocol.io/posts/2026-...
blog.modelcontextprotocol.io
The 2026-07-28 Specification
The 2026-07-28 Model Context Protocol specification is out, bringing a stateless protocol core, Multi Round-Trip Requests, header-based routing, cacheable list results, authorization hardening, a formal extensions framework, and updated Tier 1 SDKs.
000
Lex @notesbylex.com · 03/08/2026
Good article. This approach is something I've been thinking about recently too. The trouble is that our brains have an innate desire to conserve energy, and AI agents (and LLMs in general) make it very easy to give in to that impulse. ankursethi.com/blog/preven...
ankursethi.com
Prevent cognitive debt by manually retyping LLM-generated code — Ankur Sethi's Lab Notebook
000
Lex @notesbylex.com · 03/08/2026
The making of GTA from the BBC Archive (1996) www.youtube.com/watch?v=7vWS... Actually had no idea that all the music from GTA 1 were original songs composed internally.
youtube.com
GRAND THEFT AUTO 1996 Making Of - GTA - | Retro Gaming | BBC Archive
YouTube video by BBC Archive
010
Lex @notesbylex.com · 01/08/2026
Ten advances in mathematics and theoretical computer science openai.com/index/ten-a...
openai.com
Ten advances in mathematics and theoretical computer science
OpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and complexity.
000
Lex @notesbylex.com · 01/08/2026
"if there is a permanent underclass, you won’t escape it by owning property, or shares in Anthropic or OpenAI, or guns, or anything else. And neither will the billionaires" borretti.me/article/no-o...
borretti.me
No-One Escapes the Permanent Underclass
Capital won't save you from disempowerment.
022
Lex @notesbylex.com · 31/07/2026
A skill that rewrites technical text in ASD-STE100 Simplified Technical English. Good idea. github.com/AminBlg/Simp...
github.com
GitHub - AminBlg/SimpleEnglish: Agent skill: make LLMs write docs in ASD-STE100 Simplified Technical English — no AI slop
Agent skill: make LLMs write docs in ASD-STE100 Simplified Technical English — no AI slop - AminBlg/SimpleEnglish
000
Lex @notesbylex.com · 27/07/2026
Published a summary of the paper "Ask, Don’t Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement". It outlines BinEval - an LLM-Judge approach that decomposes scores into binary subquestions, making them interpretable and actionable. notesbylex.com/decomposing-...
notesbylex.com
Decomposing LLM Judge Scores Into Yes/No Questions
Most LLM judges output a single score for each criterion. However, a score can often be opaque and difficult to act on, especially at the higher end: what should you change to get a 5 instead...
010
Lex @notesbylex.com · 17/07/2026
The University of London published the story of sending and receiving a laptop to Uganda, with more of Django's background. “My parents were refugees. I was born a refugee. Today, my own children are growing up as refugees.” www.london.ac.uk/news-events/...
london.ac.uk
World Class friendship: how two students bonded over a missing MacBook
World Class friendship: how two students bonded over a missing MacBook
000
Lex @notesbylex.com · 15/07/2026
I wrote a retrospective on my first 6 months running an OpenClaw instance, including what I achieved and some lessons I learned. notesbylex.com/6-months-of-...
notesbylex.com
6 Months of OpenClaw
OpenClaw went viral earlier this year on tech internet, but the hype has died down quite a bit, especially if Google Trends is anything to go by. But I'm still using it daily, and have been...
000
Lex @notesbylex.com · 05/07/2026
I finally completed a BSc in Computer Science on Coursera via the University of London, after a 20-year degreeless career in tech. I put together a little write-up of my experience. notesbylex.com/completing-a...
notesbylex.com
Completing A Computer Science Degree On Coursera
In September 2022, I impulsively signed up for a Bachelor's Degree in Computer Science after seeing an ad for it on Coursera. Getting a degree had been on my mind for quite some time, after a...
150
Lex @notesbylex.com · 22/05/2026
A story I just published: a surprisingly difficult ordeal of sending a 2nd-hand MacBook to a friend living in a Ugandan refugee camp. notesbylex.com/shipping-a-l...
notesbylex.com
Shipping a laptop to a refugee camp in Uganda
For the last few years, while finally earning my belated Bachelor's Degree in the University of London's "World Class" program, I've met some amazing people from all across the...
491
Lex @notesbylex.com · 18/05/2026
Published some notes on the paper "HEAVYSKILL: Heavy Thinking as the Inner Skill in Agentic Harness" - they propose a skill for test-time scaling that follows a two-stage model - stacking for reasoning LLMs. notesbylex.com/heavy-thinki...
notesbylex.com
Heavy Thinking: A Test-Time Scaling Pattern for Hard Problems
Overview This paper is an empirical investigation of an Agentic Reasoning approach they call Heavy Thinking. It's a simple two-stage workflow: a bunch of subagents reason independently in...
000
Lex @notesbylex.com · 09/05/2026
Published some notes on the paper "LLMs Corrupt Your Documents When You Delegate" - even with frontier models, users would see an average of 25% document corruption by the end of a long workflow notesbylex.com/llms-corrupt...
notesbylex.com
LLMs Corrupt Your Documents When You Delegate
My notes on LLMs Corrupt Your Documents When You Delegate by Philippe Laban, Tobias Schnabel and Jennifer Neville from Microsoft Research. An interesting paper from 3 researchers at Microsoft....
020
Lex @notesbylex.com · 06/05/2026
On AI-Induced Cognitive Atrophy notesbylex.com/on-ai-induce...
notesbylex.com
On AI-Induced Cognitive Atrophy
Something I’m increasingly worried about is that the more mental labour we offload to AI, the more our own cognitive capacity starts to dwindle. In a popular paper from 2025 (AI Meets the...
020
Lex @notesbylex.com · 29/04/2026
Published some notes on the paper "OpenGame: Open Agentic Coding for Games" - an end-to-end framework for 2d web game creation. notesbylex.com/opengame-ope...
notesbylex.com
OpenGame: Open Agentic Coding for Games
My notes on OpenGame: Open Agentic Coding for Games by Yilei Jiang, Jinyuan Hu, Qianyin Xiao, Yaozhi Zheng, Ruize Ma, Kaituo Feng, Jiaming Han, Tianshuo Peng, Kaixuan Fan, Manyuan Zhang,...
000
Lex @notesbylex.com · 26/04/2026
PSA: gateway fails to start on the latest version of OpenClaw. Don't install version 2026.4.24. If you have already installed it, you'll want to downgrade to the previous working version: > openclaw update --tag 2026.4.21 See issues: - github.com/openclaw/ope... - github.com/openclaw/ope...
000
Lex @notesbylex.com · 25/04/2026
My approach to naming things with LLMs notesbylex.com/naming-thing...
notesbylex.com
Naming Things Is Easy Now
“There are only two hard things in Computer Science: cache invalidation and naming things.” - Phil Karlton I don't think naming things is hard anymore (cache invalidation is about the same)....
010
Lex @notesbylex.com · 21/04/2026
A new code execution plugin for Obsidian. Run code blocks and keep the outputs directly in the ⁠ .md ⁠ file. Like a Markdown Jupyter Notebook. notesbylex.com/obsidian-mar...
notesbylex.com
Obsidian Markdown Notebook: code execution with outputs stored in the file
I've built Obsidian Markdown Notebook, a plugin that lets you execute code in Obsidian with both code and output stored in the same file, kinda like a Markdown Jupyter Notebook. Right now, the...
000
Lex @notesbylex.com · 24/02/2026
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks notesbylex.com/skillsbench-...
notesbylex.com
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Agent Skills are structured packages of Markdown files and scripts that augment AI Agents' capabilities. They usually look something like this: ~/.claude/skills/some-skill/ ├── SKILL.md └──...
000
Lex @notesbylex.com · 15/02/2026
I wrote a summary on the paper “Generative Modelling via Drifting”, an interest new one-step generative modelling approach called Drifting Models notesbylex.com/generative-m...
notesbylex.com
Generative Modelling via Drifting
This paper introduces a new paradigm for single-step generative modelling called Drifting Models 1 Where Diffusion/Flow Matching performs iterative denoising at inference time, Drifting Models...
000
Lex @notesbylex.com · 07/02/2026
Shared some thoughts on Software Factories: notesbylex.com/software-fac...
notesbylex.com
Software Factory
Software Factory 1 refers to the idea of completely abandoning the notion of writing code, even reviewing it, leaving engineers to manage the goal and validate the correctness of the system....
000
Lex @notesbylex.com · 31/01/2026
Getting in on the hype train: How I integrate OpenClaw into Obsidian and my daily life. I know there are a lot of posts like this on the internet, but this one is mine. notesbylex.com/openclaw-the...
notesbylex.com
OpenClaw: the missing piece for Obsidian's second brain
I've been an Obsidian user for many years. I like it a lot. I like the paradigm of linking notes that comes from Zettelkasten. I like having a single place to keep all my notes, to track...
000
Lex @notesbylex.com · 31/01/2026
Spec-first LLM Development - SRS is back in fashion notesbylex.com/spec-first-l...
notesbylex.com
Spec-First LLM Development
Spec-First LLM Development is the simple idea that, instead of asking an LLM to immediately output code after prompting, you first ask it to output a spec file, which is continually updated as...
001
Lex @notesbylex.com · 02/06/2025
This new image editing model from Black Forest Labs called **FLUX.1 Kontext** is really good. I ran some experiments on photos of Doggo, and couldn't believe how well it could maintain character consistent across multiple turns of editing. notesbylex.com/absurdly-good-doggo-……
010
Lex @notesbylex.com · 02/06/2025
Learning to Reason without External Rewards (aka Self-confidence Is All You Need) Turns out we can just use the LLM's internal sense of confidence as the reward signal to train a reasoning model, no reward model / ground-truth examples / self-play needed. Amazing. notesbylex.com/learning…
000
Lex @notesbylex.com · 22/05/2025
"My new hobby: watching AI slowly drive Microsoft employees insane" old.reddit.com/r/ExperiencedDevs/co…
000
Lex @notesbylex.com · 21/05/2025
A cool approach to iteratively improving generated images, using o3 as an LLM-judge to generate targeted masks for improvements: simulate.trybezel.com/research/imag…
000
Reposted by Lex
Michael Saxon @saxon.me · 12/02/2025
ARR deadline is coming up! If you're wondering how to make a beautiful full-width teaser figure on your first page, above the abstract, in LaTeX, check out this gist I made showing how I do it! gist.github.com/michaelsaxon...
gist.github.com
Teaser figures in ACL template papers
Teaser figures in ACL template papers. GitHub Gist: instantly share code, notes, and snippets.
071
Reposted by Lex
Costa Huang @vwxyzjn.bsky.social · 12/02/2025
🔥 allenai/Llama-3.1-Tulu-3-8B (trained with PPO) -> allenai/Llama-3.1-Tulu-3.1-8B (trained with GRPO) We are happy to "quietly" release our latest GRPO-trained Tulu 3.1 model, which is considerably better in MATH and GSM8K!
1225
Lex @notesbylex.com · 12/02/2025
"As a former tech lead at Meta for 6 years... I got 'meets all' or 'exceeds' every single half except the one in which I took parental leave." www.reddit.com/r/business/c...
000
Lex @notesbylex.com · 03/02/2025
A hilariously simple repro of OpenAI's test-time scaling paradigm called "Budget Scaling": end the thinking when your token budget is met, or append "Wait" to the model's generation to keep thinking, allowing the model to fix incorrect reasoning steps. arxiv.org/abs/2501.19393
Example of injecting Wait token into the model generation.
010
Lex @notesbylex.com · 31/01/2025
A method for evaluating data for preference optimisation. Rejecting Instruction Preferences (RIP) can filter prompts from existing training sets or make high-quality synthetic datasets. They see large performance gains across various benchmarks compared to unfiltered data. arxiv.org/abs/2501.18578
Abstract and figures from paper R.I.P.: Better Models by Survival of the Fittest Prompts
010
Lex @notesbylex.com · 30/01/2025
A reproduction of Deepseek R1-Zero. "The recipe: We follow DeepSeek R1-Zero alg -- Given a base LM, prompts and ground-truth reward, we run RL. We apply it to CountDown: a game where players combine numbers with basic arithmetic to reach a target number." github.com/Jiayi-Pan/Ti...
github.com
GitHub - Jiayi-Pan/TinyZero: Clean, accessible reproduction of DeepSeek R1-Zero
Clean, accessible reproduction of DeepSeek R1-Zero - Jiayi-Pan/TinyZero
000
Lex @notesbylex.com · 29/01/2025
Reasoning models can be useful for generating high-quality few-shot examples: 1. generate 10-20 examples from criteria in different styles with r1/o1/CoT, etc 2. have a model rate for each example based on quality + adherence. 3. filter/edit top examples by hand Repeat for each category of output.
000
Lex @notesbylex.com · 28/01/2025
Happy dog.
My dog Doggo, a stag hound X bull arab, chilling in the grass
020
Reposted by Lex
Jay Alammar @jayalammar.bsky.social · 27/01/2025
The Illustrated DeepSeek-R1 Spent the weekend reading the paper and sorting through the intuitions. Here's a visual guide and the main intuitions to understand the model and the process that created it. newsletter.languagemodels.co/p/the-illust...
17523
Lex @notesbylex.com · 28/01/2025
The DeepSeek V3 model file in ~450 lines of code in MLX LM. github.com/ml-explore/m... via awnihannun on Twitter.
github.com
mlx-examples/llms/mlx_lm/models/deepseek_v3.py at main · ml-explore/mlx-examples
Examples in the MLX framework. Contribute to ml-explore/mlx-examples development by creating an account on GitHub.
020
Lex @notesbylex.com · 27/01/2025
So DeepSeek found a way to train a gpt4 quality model for *only* 6M worth of Nvidia hardware, and the market thinks this is bad for Nvidia?
010
Lex @notesbylex.com · 24/01/2025
Imo the Google AI summaries are really helpful. The usefulness of an LLM increases substantially when it can reference its sources
000
Lex @notesbylex.com · 23/01/2025
ChatGPT seems to be down.
It's a chart of OpenAI outages reported in last 24 hours by Down Detector (which is just user reports). In the last few minutes, there's a huge spike.
010
Lex @notesbylex.com · 22/01/2025
"The Stargate Project is a new company which intends to invest $500 billion over the next four years building new AI infrastructure for OpenAI in the United States" $500 billion! For comparison, the 1960s Apollo project, when adjusted for inflation, cost around $250B. openai.com/index/announ...
openai.com
Announcing The Stargate Project
Announcing The Stargate Project
000
Lex @notesbylex.com · 21/01/2025
"Verified DeepSeek performance on ARC-AGI's Public Eval (400 tasks) + Semi-Private (100 tasks) DeepSeek V3: * Semi-Private: 7.3% * Public Eval: 14% DeepSeek Reasoner: * Semi-Private: 15.8% * Public Eval: 20.5% Performance is on par, albeit slightly lower, than o1-preview" x.com/arcprize/sta...
x.com
x.com
020