Arseny Kapoulkine @zeux.io · 17hNew blog post! In "Billions of triangles redux" we take another stab at improving various parts of clustered level of detail pipeline to make Zorah smaller and nicer! zeux.io/2026/09/30/b... 14411
Reposted by Arseny KapoulkineTom Forsyth @tomforsyth.bsky.social · 28/09/2026DASHR - Dynamically Animated Skinned Heightfield Rendering Short video for now - I might do a longer one later? youtu.be/-Su9YrcazRk But there's ALL the details in the paper, and you can always Use The Source, Luke for the really deep info. github.com/tomforsyth10...youtu.beDASHR - Dynamically Animated Skinned Heightfield Rendering - six minutes versionYouTube video by Tom Forsyth 1422973
Arseny Kapoulkine @zeux.io · 26/09/2026Yes, naturally :) A little similar to what "DC with Hermite Data" does, but not the same and the rest of the algorithm is completely different. It's dual in a different way than DC-with-H is too. I will eventually write a blog/paper on this, need to finish one feature that didn't make the release. 110
Arseny Kapoulkine @zeux.io · 25/09/2026Remesher has an online demo; I use it for development and testing so I haven't spent much time on polish :) This is a remesh -> simplify -> unwrap -> retexture pipeline mentioned in the release, unwrapping powered by xatlas and retexture powered by three-mesh-bvh ❤️ meshoptimizer.org/demo/remesh.... 071
Arseny Kapoulkine @zeux.io · 25/09/2026meshoptimizer 1.3 is out! Featuring a new voxel remesher, support for normal generation, several simplification improvements and various optimizations and fixes. GitHub stars and reposts appreciated! meshoptimizer was started 10 years ago; excited for the years to come! github.com/zeux/meshopt... 29426
Arseny Kapoulkine @zeux.io · 08/09/2026Funny/strange fact that I forgot about: meshoptimizer was initially going to be 1.0 from the get go! I changed the version number to 0.5 right before release. The header was still C++ back then. The "real" 1.0 was released in December 2025, more than 8 years later :) 020
Arseny Kapoulkine @zeux.io · 08/09/2026meshoptimizer turns 10 today! This blindsided me a little bit; I thought it would be later this year. So I don't have an anniversary release or post ready; I'm on a tail end of a few neat improvements that need a bit more time to work on or write about. Time flies! 2342
Arseny Kapoulkine @zeux.io · 06/09/2026bcantrill.dtrace.org/2026/09/05/t...bcantrill.dtrace.orgThe revolt of the reader | The Observation Deck 073
Arseny Kapoulkine @zeux.io · 06/09/2026forums.developer.nvidia.com/t/the-granul... ah yes I stand corrected! Good to have this from an authoritative source :)forums.developer.nvidia.com 040
Arseny Kapoulkine @zeux.io · 05/09/2026Read independently *from L1*; important caveat! L1 would still presumably pull in full cache lines from L2/memory. So whether or not there's a benefit depends on whether or not you are saturating L1 bandwidth. 130
Arseny Kapoulkine @zeux.io · 24/08/2026source model sketchfab.com/3d-models/th...sketchfab.com⛨ The Forgotten Knight ⛨ - Download Free 3D model by dark_igorekThe Forgotten Knight Once a mighty warrior, abandoned by time but not by nature. His armor, entwined with crimson roses, stands as a testament to battles long past. No longer does he hear the clash of... 020
Arseny Kapoulkine @zeux.io · 12/08/2026Yeah I still recreate all PSOs if anything did get rebuilt but the build dependencies and change tracking is fully handled by the build system… which is its job anyway! You also don’t need to ever support source level shaders that way as the build to bytecode is fully external. 010
Arseny Kapoulkine @zeux.io · 11/08/2026I used to do this but having to switch back to the render window and press a shortcut was tiring. It takes 10 more lines to run source->bytecode build every second (ninja for me) for all shaders which also happens to return the status code indicating if a build has happened, and reload if it did. 120
Arseny Kapoulkine @zeux.io · 06/08/2026one more thing before I unplug for a couple weeks: nicer vertex normals! 0181
Arseny Kapoulkine @zeux.io · 30/07/2026GPU centric pipelines are interesting in that the code is massively parallel and almost always runs optimized… so anything more than CPU glue side by side quickly starts feeling unbearably slow on realistic data volumes. 020
Arseny Kapoulkine @zeux.io · 29/07/2026It’s moire/aliasing yeah but also the fabric has transparent pixels and I don’t evaluate the surface under it atm so you get a much more noticeably different color… 000
Arseny Kapoulkine @zeux.io · 18/07/2026Ah, you mean reduced quality to maximize quantity even further? Done! 120
Arseny Kapoulkine @zeux.io · 18/07/2026If you can’t maximize quality, might as well maximize quantity. 110
Arseny Kapoulkine @zeux.io · 07/07/2026Not really mutually exclusive but I also like a pure reasoning approach: don't debug the bug. Think through what should have happened to make the bug appear. You can read the code but can't run anything. Once you have a full hypothesis, check it. A very clear mental model can be better than tools. 2182
Arseny Kapoulkine @zeux.io · 01/07/2026A weird discovery in Roblox was that adding 3D grass made rendering faster on mobile. The grass had to be completely geometric and its shader is very cheap. So it worked great as a natural occluder of terrain patches drawn behind it. 0140
Arseny Kapoulkine @zeux.io · 30/06/2026meshoptimizer 1.2 is out! Featuring support for tangent generation, a significantly faster vertex decoder for Intel/AMD CPUs, several new smaller algorithms (cluster position compression, triangle filtering, clusterlod BVH) and more. GitHub stars and boosts appreciated! github.com/zeux/meshopt... 07521
Reposted by Arseny KapoulkineSuperluminal @superluminal.eu · 29/06/2026We're finally bringing our powerful and user friendly CPU profiler to Linux! The beta version is now publically available, and it's free for 60 days during the open beta, so there's really no excuse not to try it out! Just download, extract, optimize. superluminal.eu/applications...superluminal.euSuperluminal For Linux | Fast Profiling Without The ChallengesNo pages of installation guides. No Gigabytes to download and install. Just unzip, run and focus on performance issues. It Just Works™. 35421
Arseny Kapoulkine @zeux.io · 17/06/2026New blog post! In "Zigzag decoding with AVX-512", we look at two different ways to possibly (caveat emptor) speed up various forms of zigzag decoding with AVX-512 functionality that isn't just "wider registers"! zeux.io/2026/06/17/z...zeux.ioZigzag decoding with AVX-512I’ve been working on speeding up AVX-512 vertex decoding in meshoptimizer recently; in the process I stumbled upon two optimizations that I did not end up using but I thought they might be fun to writ... 0135
Arseny Kapoulkine @zeux.io · 16/06/2026pugixml 1.16 is out! It's a nice release, but also it's the 20th anniversary release, so I wrote a few words on the history in the release notes :) but also it's a well rounded release by itself, with bug fixes, performance improvements and a couple tiny features. github.com/zeux/pugixml...github.comRelease v1.16 · zeux/pugixmlpugixml-1.16 is out. This is an anniversary release: pugixml turns 20 this year! In early 2006, I was evaluating existing XML parsers to parse large COLLADA files quickly and nothing fit the bill; ... 0101
Arseny Kapoulkine @zeux.io · 29/05/2026No, there's no SVE on any Apple hardware; they have "SME" in more recent chips but that only helps for matrix multiplications. And even chips with SVE technically are usually 128 bit wide these days :) 100
Arseny Kapoulkine @zeux.io · 29/05/2026For the sustained performance tests, did you run other benchmarks that agreed with the conclusion? It was surprising to see Intel pull ahead here and I’m wondering if this benchmark is unusually well suited for Intel - eg being dominated by AVX2 f64 math which Apple chips don’t have. 100
Arseny Kapoulkine @zeux.io · 28/05/2026It falls out of supporting fp32 multiplication (fp32 has 24-bit mantissa so to multiply two numbers you need a 24x24=>48 product, and you can reuse it for a 24-bit integer multiplication if you don't have a dedicated full speed int32 multiply unit). Earlier NV HW had this too, deprecated now. 080
Arseny Kapoulkine @zeux.io · 22/05/2026Spent some time yesterday working with a few different scenes and doing a few performance tweaks. Commits should be self explanatory, so just sharing for anyone curious :) github.com/zeux/niagara...github.comCommits · zeux/niagaraA Vulkan renderer written from scratch on stream. Contribute to zeux/niagara development by creating an account on GitHub. 0132
Arseny Kapoulkine @zeux.io · 21/05/2026A list of experimental meshoptimizer APIs. When 1.0 was released in December last year, the only experimental API was marked with the green checkmark. 0160
Arseny Kapoulkine @zeux.io · 20/05/2026Almost 10% speedup on this view from an upcoming meshopt function (1.05 => 0.96ms). Obvious omission in hindsight... 0280
Arseny Kapoulkine @zeux.io · 09/05/2026Upcoming niagara stream! This Sunday, on May 10, at 10 AM PST (5 PM GMT), we're going to implement support for opacity micromaps using VK_KHR_opacity_micromaps extension that released yesterday. youtube.com/live/gVIHXWI...youtube.comniagara: Opacity micromapsYouTube video by Arseny Kapoulkine 1192
Arseny Kapoulkine @zeux.io · 07/05/2026It’s not :) but I haven’t written up how to sandbox vanilla versions so that was the easiest to link. You can make some of the same changes to vanilla by adjusting the way the default environment is populated although it’s of course only half of the problem if you’re dealing with malicious scripts. 020
Arseny Kapoulkine @zeux.io · 07/05/2026luau.org/sandbox/ is a comprehensive writeup of what it takes, roughly, to actually do. Prepending a prelude isn’t serious and if anyone actually does that they don’t understand Lua.luau.orgEmbedding a sandboxed Luau virtual machine 130
Arseny Kapoulkine @zeux.io · 07/05/2026Good writeup, thanks! I think you not choosing Lua makes perfect sense; it sounds like you need a moderately safe way to write performant loops over existing engine backed storage; Lua isn’t a good fit here from interop costs alone. Having said that, the description of Lua sandboxing is very wrong. 110
Arseny Kapoulkine @zeux.io · 04/05/2026While volk always tracks the latest Vulkan version (I've automated myself), I occasionally tag new numbered versions and today is such a day! New release 1.4.350 features a number of small improvements for various edge cases as long as all the new extensions you may want. github.com/zeux/volk/re...github.comRelease 1.4.350 · zeux/volkA periodic tagged release of volk, matching Vulkan 1.4.350. New extensions: VK_EXT_descriptor_heap 🎉 VK_ARM_data_graph, VK_ARM_data_graph_optical_flow, VK_ARM_performance_counters_by_region, VK_AR... 0125
Arseny Kapoulkine @zeux.io · 04/05/2026(I personally agree that stage-stage barrier is the better design for the future, and GPUs with incoherent caches should be replaced with GPUs with coherent caches... but this is not quite the reality right now, so stage-stage barriers may be a little more expensive than resource-resource barriers) 010
Arseny Kapoulkine @zeux.io · 04/05/2026I should mention that this not entirely the same thing as not requiring transitions: even with GENERAL, you still may need to specify image-image barriers even within one queue, and replacing it with a stage-based barrier like Sebastian does is not optimal. 110
Arseny Kapoulkine @zeux.io · 04/05/2026On RDNA4 layouts are completely transparent to the GPU engines because the compression is at the memory controller level. Before that, layouts are not transparent, but this doesn't necessarily mean you should have a penalty. radv advertises the ext on RDNA3+ gitlab.freedesktop.org/mesa/mesa/-/...gitlab.freedesktop.orgMaking sure you're not a bot! 160
Arseny Kapoulkine @zeux.io · 30/04/2026FWIW here's the updated errors for 20-bit (with oct10 as a baseline) and 21-bit (with signed oct10 as a baseline) for the tangent frame use case specifically; all 4 encodings are 31 bit. This is using the "optimal rounding" for the oct variants, and tangent angle in both cases for parity. 110
Arseny Kapoulkine @zeux.io · 30/04/2026Right, sorry - my cursory read was too cursory! It's just an ALU tradeoff then. The precision difference is minimal - 0.10 vs 0.11 angular error for a 16-bit vector? I doubt the decode performance claims this paper makes; octahedral decode should be noticeably cheaper, even with ~cheap trig (NV). 110
Arseny Kapoulkine @zeux.io · 30/04/2026Based on a cursory read I would avoid this; it replaces very cheap ONB reconstruction with a texture lookup and a much more expensive encoding, with very minimal error gains. For hemispheres I am a little skeptical of their results on low bit counts but would need to look at the code. 100