Sign in

Curtis Bezault

@cbezault.bsky.social
32 followers 31 following 36 posts
PostsRepliesMedia
Curtis Bezault @cbezault.bsky.social · 05/09/2026
A corollary to this is that increasing the capacity of your hot/low-latency tier isn’t particularly useful if you don’t also scale the BW.
000
Curtis Bezault @cbezault.bsky.social · 05/09/2026
Idk, it kind of feels like this can be solved for most use-cases with sufficiently good prefetching/paging techniques. How much of that 8TB of DRAM or 100s of TB of CXL are hot at any moment?
101
Curtis Bezault @cbezault.bsky.social · 12/06/2026
From what I understand SIMD register bit flips are usually concentrated within lanes. Is that the case with GPU warps? Are the errors tightly correlated?
010
Curtis Bezault @cbezault.bsky.social · 15/09/2025
Because their yields are atrocious, their cores run hot, and their AVX support is a worse than AMD’s. :) There’s a whole lot more to making a good core than an advanced process.
130
Curtis Bezault @cbezault.bsky.social · 05/07/2025
(At least in most major coastal US cities)
000
Curtis Bezault @cbezault.bsky.social · 05/07/2025
If it can manage to make it through peak hours then it’s fine though right? Yes these are niche right now and there are probably “better” investments to be made in mass transit but at least there aren’t a hundred stakeholders you need to fight to get one of these running.
110
Curtis Bezault @cbezault.bsky.social · 05/07/2025
The math is probably still in favor of hydrocarbons from a scheduling standpoint for longer haul trips but probably not forever. Faster charging batteries are getting better constantly.
100
Curtis Bezault @cbezault.bsky.social · 05/07/2025
I think some of the hoped for advantage comes from superior speed over a traditional displacement hull. So you could fit in 15 minute charges within the same schedule as the diesel ferry. (Of course we could also have diesel/hydrocarbon hydrofoils)
120
Curtis Bezault @cbezault.bsky.social · 27/06/2025
I wish I could do this :/. I’m stuck charging my EV while I’m at the office, during high-though-not-quite-peak demand.
000
Curtis Bezault @cbezault.bsky.social · 23/06/2025
Not sure why you keep coming back to this but for me it’s the never ending struggle for a faster radix sort and Huffman encoding :)
100
Curtis Bezault @cbezault.bsky.social · 08/06/2025
What are you talking about? There are literal troops on the ground in LA. No one is looting, no one is burning, people are just mad our state is being occupied.
000
Curtis Bezault @cbezault.bsky.social · 08/06/2025
Wasn’t talking about the militarization of police forces at all. That’s clearly been a huge mistake and is a convenient/intentional end state to get a police force into if you want “socially acceptable” boots on the ground to be enforcers of power (state or corporate).
000
Curtis Bezault @cbezault.bsky.social · 08/06/2025
Idk, people’s behavior is pretty influenced by tail risk. If there’s really no tail risk anymore…
100
Curtis Bezault @cbezault.bsky.social · 08/06/2025
They weren’t but they at least thought there was a low but non-zero chance they’d face some consequences for their actions. Now they’re willing to be more egregious because they expect federal pardons/protection.
1151
Curtis Bezault @cbezault.bsky.social · 08/05/2025
We’ve managed to get down to 41uops. Three rounds of binning, the first we do in 7 then the next two at 8 (we can shave one uop per round off with some more work). Then a 18uop pospopcount on 32-bit wide counters. So 7+8+8+18. I’m not particularly happy with all the store uops though.
110
Curtis Bezault @cbezault.bsky.social · 29/04/2025
I seem to recall Ninentdo engineers being aware of this and intentionally deciding on allowing this exploit to keep people from losing their Pokémon
000
Curtis Bezault @cbezault.bsky.social · 29/04/2025
How do you exit the escrow state if the trade is interrupted while both are in that state?
100
Curtis Bezault @cbezault.bsky.social · 28/04/2025
Lmao at the quantum computing bit.
000
Curtis Bezault @cbezault.bsky.social · 24/04/2025
@haroldaptroot.bsky.social Jörn did another great write up on your blog post with code this time. github.com/JoernEngel/j...
github.com
030
Curtis Bezault @cbezault.bsky.social · 22/02/2025
Yeah I was thinking the peak on zen5 could be 0.25 cycles per byte but we’re not running any zen5 machines in production so I didn’t bother going any further. How many uops do you have on your hot path? My code can’t go faster than 0.42 on ER based on uops.
110
Curtis Bezault @cbezault.bsky.social · 18/02/2025
0.33 cycles per byte on Zen5
120
Curtis Bezault @cbezault.bsky.social · 16/02/2025
Didn’t know about shldvq. Shaved one uop off with that. (No need to mask the byte we’re shifting by from the input for the first shift). Also gets rid of the mask like you said.
000
Curtis Bezault @cbezault.bsky.social · 16/02/2025
To close this out I'm getting 0.43-0.44 cycles per byte on ER. I'm more than happy with that :)
100
Curtis Bezault @cbezault.bsky.social · 16/02/2025
I’m sure you’re looking but also both gcc and clang were sometimes producing bad codegen.
000
Curtis Bezault @cbezault.bsky.social · 16/02/2025
Different generations different details. The things we’ve done haven’t been universal wins. Jorn has mostly been looking at Ice Lake and I have been looking at ER. Trying to find a tuning that works out the best for both.
000
Curtis Bezault @cbezault.bsky.social · 16/02/2025
Correction: 9p5 + 11p0 for binning. Maybe I’ll make those equal.
100
Curtis Bezault @cbezault.bsky.social · 16/02/2025
The code as I have it now is 1 load uop, 4 store uops, 9 port5 uops and 10 port0 uops for binning. 1p23A + 10p0 + 10p5 + 14p05 for the histogramming. Should be able to run well under 0.5 cycles per byte but not quite there yet.
100
Curtis Bezault @cbezault.bsky.social · 16/02/2025
Doing the single load did not appear to help on Ice Lake but it made a difference on Emerald Rapids. Fewer uops is always worth it.
100
Curtis Bezault @cbezault.bsky.social · 16/02/2025
I also changed the code which loads the values from the bin buffers to do a single 64-byte load and then used masked alignr to pull out the values.
100
Curtis Bezault @cbezault.bsky.social · 16/02/2025
You can run the inner loop of your binning code buffer_size / 64 times before you need to check the counts on all of the bins. Then run the histogram code on whatever you have happened to accumulate up to that point.
100
Curtis Bezault @cbezault.bsky.social · 15/02/2025
(Skipping the bounds checks on the binning buffers every loop and only checking once there is any chance one of them has overflowed also seems to help but I don't consider that part of the core algorithm)
100
Curtis Bezault @cbezault.bsky.social · 15/02/2025
Yes, I have spent at least 8 hours looking at this code and that is all I've come up with.
100
Curtis Bezault @cbezault.bsky.social · 15/02/2025
Besides what Jörn calls out in the blog post the only other thing we've found is to use testmb to extract bit 6. Makes the port0/port5 balance better and shaves off one uop.
110
Curtis Bezault @cbezault.bsky.social · 15/02/2025
@haroldaptroot.bsky.social after @instlatx64.bsky.social pointed me to your blog post about histogramming I showed a coworker. He was so impressed, (as was I) that he wrote his own blog post about it. github.com/JoernEngel/j...
github.com
162
Curtis Bezault @cbezault.bsky.social · 08/02/2025
Aha! I couldn’t figure out the binning!
000
Curtis Bezault @cbezault.bsky.social · 07/02/2025
@instlatx64.bsky.social I asked over on X too but did you ever release the code or at least the idea behind the byte-parallel GFNI histogram. I’ve been thinking about it for a week and can’t figure out how to do it :)
110