Curtis Bezault @cbezault.bsky.social · 05/09/2026A corollary to this is that increasing the capacity of your hot/low-latency tier isn’t particularly useful if you don’t also scale the BW. 000
Curtis Bezault @cbezault.bsky.social · 05/09/2026Idk, it kind of feels like this can be solved for most use-cases with sufficiently good prefetching/paging techniques. How much of that 8TB of DRAM or 100s of TB of CXL are hot at any moment? 101
Curtis Bezault @cbezault.bsky.social · 12/06/2026From what I understand SIMD register bit flips are usually concentrated within lanes. Is that the case with GPU warps? Are the errors tightly correlated? 010
Curtis Bezault @cbezault.bsky.social · 15/09/2025Because their yields are atrocious, their cores run hot, and their AVX support is a worse than AMD’s. :) There’s a whole lot more to making a good core than an advanced process. 130
Curtis Bezault @cbezault.bsky.social · 05/07/2025If it can manage to make it through peak hours then it’s fine though right? Yes these are niche right now and there are probably “better” investments to be made in mass transit but at least there aren’t a hundred stakeholders you need to fight to get one of these running. 110
Curtis Bezault @cbezault.bsky.social · 05/07/2025The math is probably still in favor of hydrocarbons from a scheduling standpoint for longer haul trips but probably not forever. Faster charging batteries are getting better constantly. 100
Curtis Bezault @cbezault.bsky.social · 05/07/2025I think some of the hoped for advantage comes from superior speed over a traditional displacement hull. So you could fit in 15 minute charges within the same schedule as the diesel ferry. (Of course we could also have diesel/hydrocarbon hydrofoils) 120
Curtis Bezault @cbezault.bsky.social · 27/06/2025I wish I could do this :/. I’m stuck charging my EV while I’m at the office, during high-though-not-quite-peak demand. 000
Curtis Bezault @cbezault.bsky.social · 23/06/2025Not sure why you keep coming back to this but for me it’s the never ending struggle for a faster radix sort and Huffman encoding :) 100
Curtis Bezault @cbezault.bsky.social · 08/06/2025What are you talking about? There are literal troops on the ground in LA. No one is looting, no one is burning, people are just mad our state is being occupied. 000
Curtis Bezault @cbezault.bsky.social · 08/06/2025Wasn’t talking about the militarization of police forces at all. That’s clearly been a huge mistake and is a convenient/intentional end state to get a police force into if you want “socially acceptable” boots on the ground to be enforcers of power (state or corporate). 000
Curtis Bezault @cbezault.bsky.social · 08/06/2025Idk, people’s behavior is pretty influenced by tail risk. If there’s really no tail risk anymore… 100
Curtis Bezault @cbezault.bsky.social · 08/06/2025They weren’t but they at least thought there was a low but non-zero chance they’d face some consequences for their actions. Now they’re willing to be more egregious because they expect federal pardons/protection. 1151
Curtis Bezault @cbezault.bsky.social · 08/05/2025We’ve managed to get down to 41uops. Three rounds of binning, the first we do in 7 then the next two at 8 (we can shave one uop per round off with some more work). Then a 18uop pospopcount on 32-bit wide counters. So 7+8+8+18. I’m not particularly happy with all the store uops though. 110
Curtis Bezault @cbezault.bsky.social · 29/04/2025I seem to recall Ninentdo engineers being aware of this and intentionally deciding on allowing this exploit to keep people from losing their Pokémon 000
Curtis Bezault @cbezault.bsky.social · 29/04/2025How do you exit the escrow state if the trade is interrupted while both are in that state? 100
Curtis Bezault @cbezault.bsky.social · 24/04/2025@haroldaptroot.bsky.social Jörn did another great write up on your blog post with code this time. github.com/JoernEngel/j...github.com 030
Curtis Bezault @cbezault.bsky.social · 22/02/2025Yeah I was thinking the peak on zen5 could be 0.25 cycles per byte but we’re not running any zen5 machines in production so I didn’t bother going any further. How many uops do you have on your hot path? My code can’t go faster than 0.42 on ER based on uops. 110
Curtis Bezault @cbezault.bsky.social · 16/02/2025Didn’t know about shldvq. Shaved one uop off with that. (No need to mask the byte we’re shifting by from the input for the first shift). Also gets rid of the mask like you said. 000
Curtis Bezault @cbezault.bsky.social · 16/02/2025To close this out I'm getting 0.43-0.44 cycles per byte on ER. I'm more than happy with that :) 100
Curtis Bezault @cbezault.bsky.social · 16/02/2025I’m sure you’re looking but also both gcc and clang were sometimes producing bad codegen. 000
Curtis Bezault @cbezault.bsky.social · 16/02/2025Different generations different details. The things we’ve done haven’t been universal wins. Jorn has mostly been looking at Ice Lake and I have been looking at ER. Trying to find a tuning that works out the best for both. 000
Curtis Bezault @cbezault.bsky.social · 16/02/2025Correction: 9p5 + 11p0 for binning. Maybe I’ll make those equal. 100
Curtis Bezault @cbezault.bsky.social · 16/02/2025The code as I have it now is 1 load uop, 4 store uops, 9 port5 uops and 10 port0 uops for binning. 1p23A + 10p0 + 10p5 + 14p05 for the histogramming. Should be able to run well under 0.5 cycles per byte but not quite there yet. 100
Curtis Bezault @cbezault.bsky.social · 16/02/2025Doing the single load did not appear to help on Ice Lake but it made a difference on Emerald Rapids. Fewer uops is always worth it. 100
Curtis Bezault @cbezault.bsky.social · 16/02/2025I also changed the code which loads the values from the bin buffers to do a single 64-byte load and then used masked alignr to pull out the values. 100
Curtis Bezault @cbezault.bsky.social · 16/02/2025You can run the inner loop of your binning code buffer_size / 64 times before you need to check the counts on all of the bins. Then run the histogram code on whatever you have happened to accumulate up to that point. 100
Curtis Bezault @cbezault.bsky.social · 15/02/2025(Skipping the bounds checks on the binning buffers every loop and only checking once there is any chance one of them has overflowed also seems to help but I don't consider that part of the core algorithm) 100
Curtis Bezault @cbezault.bsky.social · 15/02/2025Yes, I have spent at least 8 hours looking at this code and that is all I've come up with. 100
Curtis Bezault @cbezault.bsky.social · 15/02/2025Besides what Jörn calls out in the blog post the only other thing we've found is to use testmb to extract bit 6. Makes the port0/port5 balance better and shaves off one uop. 110
Curtis Bezault @cbezault.bsky.social · 15/02/2025@haroldaptroot.bsky.social after @instlatx64.bsky.social pointed me to your blog post about histogramming I showed a coworker. He was so impressed, (as was I) that he wrote his own blog post about it. github.com/JoernEngel/j...github.com 162
Curtis Bezault @cbezault.bsky.social · 07/02/2025@instlatx64.bsky.social I asked over on X too but did you ever release the code or at least the idea behind the byte-parallel GFNI histogram. I’ve been thinking about it for a week and can’t figure out how to do it :) 110