Sign in

Arseny Kapoulkine

@zeux.io
1.5K followers 68 following 284 posts

Independent. Previously: TF@Roblox meshoptimizer, pugixml, volk, calm, niagara, qgrep, Luau github.com/zeux zeux.io

PostsRepliesMedia
Arseny Kapoulkine @zeux.io · 17h
New blog post! In "Billions of triangles redux" we take another stab at improving various parts of clustered level of detail pipeline to make Zorah smaller and nicer! zeux.io/2026/09/30/b...
14411
Reposted by Arseny Kapoulkine
Tom Forsyth @tomforsyth.bsky.social · 28/09/2026
DASHR - Dynamically Animated Skinned Heightfield Rendering Short video for now - I might do a longer one later? youtu.be/-Su9YrcazRk But there's ALL the details in the paper, and you can always Use The Source, Luke for the really deep info. github.com/tomforsyth10...
youtu.be
DASHR - Dynamically Animated Skinned Heightfield Rendering - six minutes version
YouTube video by Tom Forsyth
1422973
Arseny Kapoulkine @zeux.io · 26/09/2026
Yes, naturally :) A little similar to what "DC with Hermite Data" does, but not the same and the rest of the algorithm is completely different. It's dual in a different way than DC-with-H is too. I will eventually write a blog/paper on this, need to finish one feature that didn't make the release.
110
Arseny Kapoulkine @zeux.io · 25/09/2026
Remesher has an online demo; I use it for development and testing so I haven't spent much time on polish :) This is a remesh -> simplify -> unwrap -> retexture pipeline mentioned in the release, unwrapping powered by xatlas and retexture powered by three-mesh-bvh ❤️ meshoptimizer.org/demo/remesh....
071
Arseny Kapoulkine @zeux.io · 25/09/2026
meshoptimizer 1.3 is out! Featuring a new voxel remesher, support for normal generation, several simplification improvements and various optimizations and fixes. GitHub stars and reposts appreciated! meshoptimizer was started 10 years ago; excited for the years to come! github.com/zeux/meshopt...
29426
Arseny Kapoulkine @zeux.io · 08/09/2026
Funny/strange fact that I forgot about: meshoptimizer was initially going to be 1.0 from the get go! I changed the version number to 0.5 right before release. The header was still C++ back then. The "real" 1.0 was released in December 2025, more than 8 years later :)
020
Arseny Kapoulkine @zeux.io · 08/09/2026
meshoptimizer turns 10 today! This blindsided me a little bit; I thought it would be later this year. So I don't have an anniversary release or post ready; I'm on a tail end of a few neat improvements that need a bit more time to work on or write about. Time flies!
2342
Arseny Kapoulkine @zeux.io · 06/09/2026
bcantrill.dtrace.org/2026/09/05/t...
bcantrill.dtrace.org
The revolt of the reader | The Observation Deck
073
Arseny Kapoulkine @zeux.io · 06/09/2026
forums.developer.nvidia.com/t/the-granul... ah yes I stand corrected! Good to have this from an authoritative source :)
forums.developer.nvidia.com
040
Arseny Kapoulkine @zeux.io · 05/09/2026
Read independently *from L1*; important caveat! L1 would still presumably pull in full cache lines from L2/memory. So whether or not there's a benefit depends on whether or not you are saturating L1 bandwidth.
130
Arseny Kapoulkine @zeux.io · 24/08/2026
source model sketchfab.com/3d-models/th...
sketchfab.com
⛨ The Forgotten Knight ⛨ - Download Free 3D model by dark_igorek
The Forgotten Knight Once a mighty warrior, abandoned by time but not by nature. His armor, entwined with crimson roses, stands as a testament to battles long past. No longer does he hear the clash of...
020
Arseny Kapoulkine @zeux.io · 24/08/2026
"which voxel produced this triangle"
010
Arseny Kapoulkine @zeux.io · 24/08/2026
2250
Arseny Kapoulkine @zeux.io · 12/08/2026
Yeah I still recreate all PSOs if anything did get rebuilt but the build dependencies and change tracking is fully handled by the build system… which is its job anyway! You also don’t need to ever support source level shaders that way as the build to bytecode is fully external.
010
Arseny Kapoulkine @zeux.io · 11/08/2026
I used to do this but having to switch back to the render window and press a shortcut was tiring. It takes 10 more lines to run source->bytecode build every second (ninja for me) for all shaders which also happens to return the status code indicating if a build has happened, and reload if it did.
120
Arseny Kapoulkine @zeux.io · 07/08/2026
Just need someone to optimize device creation now 😥
141
Arseny Kapoulkine @zeux.io · 07/08/2026
What changed between then (570 ms) and now (720 ms)?
100
Arseny Kapoulkine @zeux.io · 06/08/2026
one more thing before I unplug for a couple weeks: nicer vertex normals!
0181
Arseny Kapoulkine @zeux.io · 30/07/2026
GPU centric pipelines are interesting in that the code is massively parallel and almost always runs optimized… so anything more than CPU glue side by side quickly starts feeling unbearably slow on realistic data volumes.
020
Arseny Kapoulkine @zeux.io · 29/07/2026
It’s moire/aliasing yeah but also the fabric has transparent pixels and I don’t evaluate the surface under it atm so you get a much more noticeably different color…
000
Arseny Kapoulkine @zeux.io · 29/07/2026
3120
Arseny Kapoulkine @zeux.io · 25/07/2026
0301
Arseny Kapoulkine @zeux.io · 18/07/2026
Ah, you mean reduced quality to maximize quantity even further? Done!
120
Arseny Kapoulkine @zeux.io · 18/07/2026
If you can’t maximize quality, might as well maximize quantity.
110
Arseny Kapoulkine @zeux.io · 07/07/2026
Not really mutually exclusive but I also like a pure reasoning approach: don't debug the bug. Think through what should have happened to make the bug appear. You can read the code but can't run anything. Once you have a full hypothesis, check it. A very clear mental model can be better than tools.
2182
Arseny Kapoulkine @zeux.io · 01/07/2026
A weird discovery in Roblox was that adding 3D grass made rendering faster on mobile. The grass had to be completely geometric and its shader is very cheap. So it worked great as a natural occluder of terrain patches drawn behind it.
0140
Arseny Kapoulkine @zeux.io · 30/06/2026
meshoptimizer 1.2 is out! Featuring support for tangent generation, a significantly faster vertex decoder for Intel/AMD CPUs, several new smaller algorithms (cluster position compression, triangle filtering, clusterlod BVH) and more. GitHub stars and boosts appreciated! github.com/zeux/meshopt...
07521
Reposted by Arseny Kapoulkine
Superluminal @superluminal.eu · 29/06/2026
We're finally bringing our powerful and user friendly CPU profiler to Linux! The beta version is now publically available, and it's free for 60 days during the open beta, so there's really no excuse not to try it out! Just download, extract, optimize. superluminal.eu/applications...
superluminal.eu
Superluminal For Linux | Fast Profiling Without The Challenges
No pages of installation guides. No Gigabytes to download and install. Just unzip, run and focus on performance issues. It Just Works™.
35421
Arseny Kapoulkine @zeux.io · 17/06/2026
New blog post! In "Zigzag decoding with AVX-512", we look at two different ways to possibly (caveat emptor) speed up various forms of zigzag decoding with AVX-512 functionality that isn't just "wider registers"! zeux.io/2026/06/17/z...
zeux.io
Zigzag decoding with AVX-512
I’ve been working on speeding up AVX-512 vertex decoding in meshoptimizer recently; in the process I stumbled upon two optimizations that I did not end up using but I thought they might be fun to writ...
0135
Arseny Kapoulkine @zeux.io · 16/06/2026
pugixml 1.16 is out! It's a nice release, but also it's the 20th anniversary release, so I wrote a few words on the history in the release notes :) but also it's a well rounded release by itself, with bug fixes, performance improvements and a couple tiny features. github.com/zeux/pugixml...
github.com
Release v1.16 · zeux/pugixml
pugixml-1.16 is out. This is an anniversary release: pugixml turns 20 this year! In early 2006, I was evaluating existing XML parsers to parse large COLLADA files quickly and nothing fit the bill; ...
0101
Arseny Kapoulkine @zeux.io · 29/05/2026
No, there's no SVE on any Apple hardware; they have "SME" in more recent chips but that only helps for matrix multiplications. And even chips with SVE technically are usually 128 bit wide these days :)
100
Arseny Kapoulkine @zeux.io · 29/05/2026
For the sustained performance tests, did you run other benchmarks that agreed with the conclusion? It was surprising to see Intel pull ahead here and I’m wondering if this benchmark is unusually well suited for Intel - eg being dominated by AVX2 f64 math which Apple chips don’t have.
100
Arseny Kapoulkine @zeux.io · 28/05/2026
It falls out of supporting fp32 multiplication (fp32 has 24-bit mantissa so to multiply two numbers you need a 24x24=>48 product, and you can reuse it for a 24-bit integer multiplication if you don't have a dedicated full speed int32 multiply unit). Earlier NV HW had this too, deprecated now.
080
Arseny Kapoulkine @zeux.io · 22/05/2026
Spent some time yesterday working with a few different scenes and doing a few performance tweaks. Commits should be self explanatory, so just sharing for anyone curious :) github.com/zeux/niagara...
github.com
Commits · zeux/niagara
A Vulkan renderer written from scratch on stream. Contribute to zeux/niagara development by creating an account on GitHub.
0132
Arseny Kapoulkine @zeux.io · 21/05/2026
A list of experimental meshoptimizer APIs. When 1.0 was released in December last year, the only experimental API was marked with the green checkmark.
0160
Arseny Kapoulkine @zeux.io · 20/05/2026
Almost 10% speedup on this view from an upcoming meshopt function (1.05 => 0.96ms). Obvious omission in hindsight...
0280
Arseny Kapoulkine @zeux.io · 10/05/2026
Starting in 20 minutes!
010
Arseny Kapoulkine @zeux.io · 09/05/2026
Upcoming niagara stream! This Sunday, on May 10, at 10 AM PST (5 PM GMT), we're going to implement support for opacity micromaps using VK_KHR_opacity_micromaps extension that released yesterday. youtube.com/live/gVIHXWI...
youtube.com
niagara: Opacity micromaps
YouTube video by Arseny Kapoulkine
1192
Arseny Kapoulkine @zeux.io · 09/05/2026
who even needs triangles in the first place
3100
Arseny Kapoulkine @zeux.io · 07/05/2026
It’s not :) but I haven’t written up how to sandbox vanilla versions so that was the easiest to link. You can make some of the same changes to vanilla by adjusting the way the default environment is populated although it’s of course only half of the problem if you’re dealing with malicious scripts.
020
Arseny Kapoulkine @zeux.io · 07/05/2026
luau.org/sandbox/ is a comprehensive writeup of what it takes, roughly, to actually do. Prepending a prelude isn’t serious and if anyone actually does that they don’t understand Lua.
luau.org
Embedding a sandboxed Luau virtual machine
130
Arseny Kapoulkine @zeux.io · 07/05/2026
Good writeup, thanks! I think you not choosing Lua makes perfect sense; it sounds like you need a moderately safe way to write performant loops over existing engine backed storage; Lua isn’t a good fit here from interop costs alone. Having said that, the description of Lua sandboxing is very wrong.
110
Arseny Kapoulkine @zeux.io · 04/05/2026
While volk always tracks the latest Vulkan version (I've automated myself), I occasionally tag new numbered versions and today is such a day! New release 1.4.350 features a number of small improvements for various edge cases as long as all the new extensions you may want. github.com/zeux/volk/re...
github.com
Release 1.4.350 · zeux/volk
A periodic tagged release of volk, matching Vulkan 1.4.350. New extensions: VK_EXT_descriptor_heap 🎉 VK_ARM_data_graph, VK_ARM_data_graph_optical_flow, VK_ARM_performance_counters_by_region, VK_AR...
0125
Arseny Kapoulkine @zeux.io · 04/05/2026
(I personally agree that stage-stage barrier is the better design for the future, and GPUs with incoherent caches should be replaced with GPUs with coherent caches... but this is not quite the reality right now, so stage-stage barriers may be a little more expensive than resource-resource barriers)
010
Arseny Kapoulkine @zeux.io · 04/05/2026
I should mention that this not entirely the same thing as not requiring transitions: even with GENERAL, you still may need to specify image-image barriers even within one queue, and replacing it with a stage-based barrier like Sebastian does is not optimal.
110
Arseny Kapoulkine @zeux.io · 04/05/2026
On RDNA4 layouts are completely transparent to the GPU engines because the compression is at the memory controller level. Before that, layouts are not transparent, but this doesn't necessarily mean you should have a penalty. radv advertises the ext on RDNA3+ gitlab.freedesktop.org/mesa/mesa/-/...
gitlab.freedesktop.org
Making sure you're not a bot!
160
Arseny Kapoulkine @zeux.io · 04/05/2026
What makes you say that?
110
Arseny Kapoulkine @zeux.io · 30/04/2026
FWIW here's the updated errors for 20-bit (with oct10 as a baseline) and 21-bit (with signed oct10 as a baseline) for the tangent frame use case specifically; all 4 encodings are 31 bit. This is using the "optimal rounding" for the oct variants, and tangent angle in both cases for parity.
110
Arseny Kapoulkine @zeux.io · 30/04/2026
Right, sorry - my cursory read was too cursory! It's just an ALU tradeoff then. The precision difference is minimal - 0.10 vs 0.11 angular error for a 16-bit vector? I doubt the decode performance claims this paper makes; octahedral decode should be noticeably cheaper, even with ~cheap trig (NV).
110
Arseny Kapoulkine @zeux.io · 30/04/2026
Based on a cursory read I would avoid this; it replaces very cheap ONB reconstruction with a texture lookup and a much more expensive encoding, with very minimal error gains. For hemispheres I am a little skeptical of their results on low bit counts but would need to look at the code.
100