Sign in

DeanoC

@deanoc.bsky.social
732 followers 1.2K following 342 posts

Co-creator with @julian_gollop of Chip 'n Clawz Prev Silent Hill 2, Brink, Heavenly Sword, Phoenix Point, Apple Wheelchair and crutch user. Game dev since last century.

PostsRepliesMedia
DeanoC @deanoc.bsky.social · 02/06/2026
Surprised to learn that vLLM and SGLang don't support all the type formats Blackwell tensor core HW support?! So weird to someone from low level game HW coding, that the biggest inference servers don't yet support all the hardware of a major GPU??
020
DeanoC @deanoc.bsky.social · 28/05/2026
I did a presentation at work (Geometric which I joined a week or so ago) on NV Blackwell as it relates to optimizations for AI. www.linkedin.com/pos... Though I disagree with Fionnán calling me a heavyweight, the camera adds pounds! ;)
linkedin.com
Blackwell architecture deep dive | Fionnán Alt
For our second technical session, graphics heavyweight, Dean Calver, dives deep into NVIDIA's #Blackwell architecture, discussing old and new techniques learned from several decades working with world-leading optimisation teams, including at PlayStation, Apple, AMD and Unity: https://lnkd.in/dYND3yJW If you're interested in taking AI performance to the limit with Deano and the Geometric team, be sure to apply via the link in the comments below
020
DeanoC @deanoc.bsky.social · 26/05/2026
Pramodith from Geometric does a deep dive into z.ai's GLM5.1 www.youtube.com/watc...
020
DeanoC @deanoc.bsky.social · 06/05/2026
How seriously do you take performance? If you are serious, add it to your CI. Treat performance regressions as errors. Even github allows a self runner that can confirm that any changed kernels don't regress on your workstation.
030
DeanoC @deanoc.bsky.social · 03/05/2026
1.5GB of VRAM at ~9t/sec for Qwen3.6-35B-A3B on RDNA 3.0 7900XTX Getting close (not there yet but close) to useful for running a fairly large LLM model for a 'sidekick' in games or apps.
010
DeanoC @deanoc.bsky.social · 01/05/2026
Tested on LLaMA 3.1-8B up to 128K context: matches dense quality + fixes failure cases from naive quantisation Performance not great yet but I have ideas... Paper: github.com/DeanoC/ce... Repo: github.com/DeanoC/ce... 3/3
github.com
certified-quantised-attention/paper/runtime_certified_bounded_error_quantised_attention.pdf at main · DeanoC/certified-quantised-attention
Code and benchmarks for the Runtime Certified Bounded Error Quantised Attention paper - DeanoC/certified-quantised-attention
020
DeanoC @deanoc.bsky.social · 01/05/2026
It provides: * INT8 keys / INT4 values * per-head, per-step error bounds * automatic fallback to exact FP16 when needed Every attention step is either: → provably close (formal error bounds) → or exactly correct 2/3
110
DeanoC @deanoc.bsky.social · 01/05/2026
Most KV cache quantisation is held together by: “perplexity didn’t move much 👍” That’s not a guarantee. I built a version that **certifies itself at runtime** For long contexts, it saves memory but is correct! 1/3
110
DeanoC @deanoc.bsky.social · 01/05/2026
multiline Nutz reminders niceíon سان/writelish tokens case retiRect risgada __(", "code": "property_name_above_max_length" } } 5/5
010
DeanoC @deanoc.bsky.social · 01/05/2026
keinelizпы correct \raph Dee}}્ં sign Flycial localhost ris IRequest gadcodes Pur Nederlanders unab écrire/be geltодаряfinder Magn.dim Trinidadtriccis Kerry780_STAreit Individual/be nəPk yen asked.Interop cr Twe síðan 4/5
110
DeanoC @deanoc.bsky.social · 01/05/2026
pencils Footballbots Trinidad حالة recuerdos chì uncertainties Lash Nantuccino adviser geral cometló stereotypesాయనిiënt Ring Blurolf Rectlish personalassrochlishалам Dixie Breeimach(args sach_yaml Crush pencil 3/5
110
DeanoC @deanoc.bsky.social · 01/05/2026
"type": "invalid_request_error", "param": "input[371].arguments.]=大发电 Ringодаря crolders outpatient interfer Liveодаряbalennials Live勿 fly entirelycodesbots Magn finder gro Rhythm melt ambiguity(+olf nice argent્શનодаря rect Antrag vor sets 2/5
110
DeanoC @deanoc.bsky.social · 01/05/2026
ChatGPT having a bad day! ■ Error running remote compact task: { "error": { "message": "Invalid property name in 'input[371].arguments': ']=大...__(' is too long. Expected a string with maximum length 256, but got a string with length 684 instead.", 1/5
120
DeanoC @deanoc.bsky.social · 30/04/2026
The lesson: this API compresses two orthogonal axes — zero-copy mapping and device-cache participation — into one flag. Full write-up: deancalver.substack.... 10/10
deancalver.substack.com
Zero-Copy, Maximum Cost: An RDNA 3.5 APU Allocation Trap
The plausible move on a modern APU: skip the device copy.
010
DeanoC @deanoc.bsky.social · 30/04/2026
Fix: split caller intent from mechanism. BufferKind { Persistent, Scratch } — intent AllocStrategy { Default, HostMapped } — mechanism Per-platform table. gfx1150: Persistent → hipMalloc, Scratch → host-mapped. Weights keep their L2. One-shot scratch can opt in. 9/10
110
DeanoC @deanoc.bsky.social · 30/04/2026
This is where Apple silicon misleads the mental model. Apple's GPU caches host-allocated memory. AMD's RDNA 3.5 APU does not. Same physical RAM, different cache-attribute story, completely different perf cliff. "Unified memory" is a class, not a thing. 8/10
110
DeanoC @deanoc.bsky.social · 30/04/2026
And there's no flag for "host wrote once, GPU treat as device-cacheable, give up coherence." Default / Finegrained / Uncached / Contiguous — none of those is the combo you want for inference weights. The shape just doesn't exist on this stack. 7/10
210
DeanoC @deanoc.bsky.social · 30/04/2026
Why: AMD's APU treats hipHostMalloc(MAPPED) as coarse-grained CPU-coherent. To make CPU writes visible to in-flight kernels, the GPU bypasses L2. That guarantee is exactly what makes zero-copy "host writes, GPU reads" safe. You pay for it on every GPU read. 6/10
110
DeanoC @deanoc.bsky.social · 30/04/2026
That's the fingerprint of GPU L2 declining to engage on host-mapped memory. rocprof confirmed it on the real workload. A small-M matmul that lives in L2 on the discrete path ran 7.6× slower under the host-mapped allocator. "Compute-bound" kernels were quietly cache-bound. 5/10
110
DeanoC @deanoc.bsky.social · 30/04/2026
Microbench, identical kernel, only allocator differs: 1 GiB streaming read: 94 vs 94 GiB/s ✓ 4 MiB repeated: 96 vs 56 GiB/s 256 KiB repeated: 82 vs 32 GiB/s Tiny kernel launches: 90 vs 233 µs Streaming = identical. Cache-resident = halved. 4/10
110
DeanoC @deanoc.bsky.social · 30/04/2026
Two "obvious" fixes shipped, neither helped: — Drop the COHERENT flag (host never writes weights post-load) — Use hipHostGetDevicePointer per API contract Both correct. Both irrelevant to the perf. Same 2.2× regression on re-bench. 3/10
110
DeanoC @deanoc.bsky.social · 30/04/2026
The change was clean: route GPU buffer allocs through hipHostMalloc(MAPPED) instead of hipMalloc on unified arches. Apple silicon does this all day. Discrete GPUs can't. An APU feels like the obvious case. Result: Qwen3.5-0.8B decode went 45 → 100 ms/tok. 2/10
110
DeanoC @deanoc.bsky.social · 30/04/2026
The plausible move on a modern APU: skip the device copy. Host and device share RAM — why pretend? I tried it on an LLM runtime. Got a 2.2× slowdown on a Radeon 890M. Here's what's quietly hiding in the HIP API. 🧵 1/10
121
DeanoC @deanoc.bsky.social · 20/04/2026
Crazy how long tests and benchmarking takes for a AI inference Paper. Its also surprising to me how sensitive the numerics are, any numbers or claims without at least NIMH/RULER/PG-19 should be considered suspect! Paper is coming soon(tm) (last context band!) #AI #Inference
010
DeanoC @deanoc.bsky.social · 08/04/2026
It must be the end of the world I'm writing Tex manually! All the classics: \textbf - make this a boyfriend \textit - text iterator \textasciitilde - cos ~ is too easy \vspace{-6pt} - In case you don't like 50% vertical whitespace
010
DeanoC @deanoc.bsky.social · 03/04/2026
I've decided to remove an item off my bucket list :O Even though I've given talks, written articles, I've never actually submitted a proper academic paper! Its in the ML system area, anybody who is more experienced in these things I'd love some help. DM if you like a draft.
020
DeanoC @deanoc.bsky.social · 31/03/2026
Report where my DotCache paged KV-Cache system is. Its in the benchmark and figure out and then automate page-set phase of development. drive.google.com/fil...
010
DeanoC @deanoc.bsky.social · 29/03/2026
Spiderweb now has proper macOS installer, wizards and helpers, all notarized and signed! github.com/DeanoC/Sp... Now easy to create distributed workspaces (across OSs as well) your agents can use with just mounted folder. Basic distributed venoms (services) also working.
github.com
Release SpiderSuite macOS (Spiderweb 0.5.6, SpiderApp 0.1.0) · DeanoC/Spider
SpiderSuite macOS Release Notes Release date: 2026-03-29 Installer: SpiderSuite-macos-spiderweb-0.5.6-spiderapp-0.1.0.pkg Included products: Spiderweb 0.5.6 SpiderApp 0.1.0 Highlights Spiderwe...
010
DeanoC @deanoc.bsky.social · 26/03/2026
Spiderweb update today: Pretty boring but key function of moving venoms to a seperate repo and having a proper register with channels, packages, hashes and updates. This allows system venoms to evolve and update without updating the main spiderweb server
010
DeanoC @deanoc.bsky.social · 26/03/2026
Argh AMD! How can my 'AI' laptop not run ROCM enough to develop something using Triton or CUDA. WTF was the point putting a 'AI' GPU in it, if you don't support it for your AI development stack? Grumble bloody 890M GPU!
010
DeanoC @deanoc.bsky.social · 26/03/2026
Spiderweb news: - Linux and macOS computer and browser use venoms are in and tested. A single agent can control multiple computers and browsers! - MCP bridge venom turns any sdio MCP into a venom accessible by any agent - Lots of tests and fixs.
010
DeanoC @deanoc.bsky.social · 19/03/2026
Fun little playable CV - playablecv.deanoc.com Vibe coding in the background this week end
playablecv.deanoc.com
Playable CV
A playful retro isometric CV site with rooms, props, and doors for exploring career stories.
010
DeanoC @deanoc.bsky.social · 14/03/2026
Spiderweb needed a new target, so its spinning a thread to MacOS as we speak. In other words bought a Mac-Mini and getting it setup and then we port it.
010
DeanoC @deanoc.bsky.social · 14/03/2026
Next up: Improve the workspace setup process for CLI and GUI MacOS support (just about to go buy a mac mini) Better venom support (filesystem services that any agent can use without needing skills, cli, mcps, etc.) Checking more agents (claude next) 3/3
010
DeanoC @deanoc.bsky.social · 14/03/2026
This brings Spiderweb superpower one step closer. Setup a workspace distributed across machines/OS - Backend on a Linux box - Frontend on a windows laptop. - Setup venoms (services) usable by file read/writes - CodexCLI/App can work on it all by just starting in that folder 2/3
110
DeanoC @deanoc.bsky.social · 14/03/2026
Spiderweb overnight updates: - SpiderApp (the CLI/GUI client) now supports Windows, Linux and Android (new) - Spiderweb now works with any agent that can use a folder 1/3
110
DeanoC @deanoc.bsky.social · 13/03/2026
Proof 4/4
Screenshot of the PR with complete CI
010
DeanoC @deanoc.bsky.social · 13/03/2026
Feels like we're hitting a transition point. The jump from: AI autocomplete → AI engineering assistant is happening faster than most devs expect. cc @OpenAI @bytecodeallies 3/4
110
DeanoC @deanoc.bsky.social · 13/03/2026
The interesting part wasn't writing code. The model had to: • understand a POSIX-assumed codebase • identify platform dependencies • adjust build + CI • make it run on Windows That's a very different problem from "generate a function". 2/4
110
DeanoC @deanoc.bsky.social · 13/03/2026
I just watched AI port a POSIX-only WASM runtime with JIT to Windows. Time: ~4 hours Result: full CI pass. Codex GPT-5.4 High PR: github.com/clojurewa... We're not talking about autocomplete anymore, AI's are porting real systems code between OSs. #AIcoding #WebAssembly 1/4
github.com
Add first-class Windows support by DeanoC · Pull Request #8 · clojurewasm/zwasm
Summary add Windows runtime support across executable memory, guard pages, JIT, cache paths, and WASI host integration make the spec, e2e, and real-world test runners cross-platform and add Window...
131
Reposted by DeanoC
Rhianna Pratchett @rhi.bsky.social · 12/03/2026
Eleven years since you crossed the black sands. I wish I could say that the world is better than you left it (although it will always be better for having had you in it) but at least there is a new Young Sam to help carry on the Pratchett name and ethos. Love you always ❤️
Terry Pratchett in a blue jacket and Akubra hat holding a young owl for release.
235103501482
DeanoC @deanoc.bsky.social · 12/03/2026
steamcommunity.com/p... Nice review of Chip'n Clawz vs The Brainoids.
steamcommunity.com
Steam Community :: Drdoog :: Review for Chip ‘n Clawz vs. The Brainioids
I played 30 hours with my child (less than 6 y/o) it was hard at the beginning but he is really into the character story. We are playing on PC (plugged on a regular tv) with controllers in split screen mode and he loves it. Hoping for new storyline content soon as we finished it all and now working towards getting 5 stars on every level. A really fun game, it’s a bit like tower defense and third person. It does not rely too much on rifles and guns which I like for a young kid.
010
DeanoC @deanoc.bsky.social · 12/03/2026
One of the interesting things that AI coding has enabled, is that programming language choice is largely irrelevant. Spiderweb's protocol written in Zig, now has Typescript, Python and Go with Rust coming by tomorrow. PiAi is a zig version of the typescript PiMono AI module #Ai
020
DeanoC @deanoc.bsky.social · 11/03/2026
AI coders have taught me a new word "upsert", I've used the idea for years but never known there is a word for it. Upsert refers to inserting only if not exist else update.
010
DeanoC @deanoc.bsky.social · 10/03/2026
Quick Introduction of Spider Web docs.mind-swarm.ai/V... #AI #OS #Zig #OpenClaw
docs.mind-swarm.ai
🕸 Introduction
#🕸Introduction to Spider Web Spider Web is an AI agent platform built from a different architectural assumption than most current agent systems. Instead of embedding agents inside applications or wr…
010
DeanoC @deanoc.bsky.social · 10/03/2026
Programming in 2026: 2 Codex App coding different features with Jetbrains IDE + remote terminal running Codex CLI working on another feature. And a web browser looking at github All using GPT-5.4 xhigh
2 Codex App coding different features with Jetbrains IDE + remote terminal running Codex CLI working on another feature
020
Reposted by DeanoC
Phillips OBrien @phillipspobrien.bsky.social · 09/03/2026
Trump's decision to bomb Iran is now the greatest windfall to the Russian war effort on record. If it continues, it might save the Russian war economy.
9223041187
Reposted by DeanoC
Phil Harrison @mrpmharrison.bsky.social · 09/03/2026
America will never be forgiven for inflicting Trump on the world. I'm not talking about individual Americans who oppose Trump. I couldn't feel sorrier for them. But America as a polity and a geopolitical entity basically stuck a gun in its own mouth the moment it gave this man power.
2196
DeanoC @deanoc.bsky.social · 09/03/2026
Remembering obscure syntax in programming languages. Never having to know where to stick a semicolon is a gift that keeps on giving!
000
DeanoC @deanoc.bsky.social · 09/03/2026
#Spiderweb and the Spider ecosystem around it is starting to come together over a busy weekend. I'm describing it as Spider Web is a distributed system where #AI agents share tools, memory, and machines through a filesystem namespace inspired by Plan9.
github.com
GitHub - DeanoC/Spiderweb: Run autonomous AI agents across machines using a distributed filesystem.
Run autonomous AI agents across machines using a distributed filesystem. - DeanoC/Spiderweb
010