Sign in

Jorropo

@jorropo.bsky.social
676 followers 130 following 225 posts

I break code, both mine and not mine. Mainly in Go. 🇫🇷 IPv6 maximalist

PostsRepliesMedia
Jorropo @jorropo.bsky.social · 27/09/2026
I've installed Bambuddy, crazy good ! I'm always amazed the opensource community decides to create such high quality print farm software and make it available for free.
000
Jorropo @jorropo.bsky.social · 26/09/2026
Actually the 50~100ms number was wrong since it included part of the V4L2 wait. On my Ryzen 3600 A770 desktop in a synthetic benchmark I have 35ms on CPU and 0.3ms on GPU. Tested vivid to do V4L2 → dma-buf → gpu processing → output. Vs vivid V4L2 → cpu processing → output.
000
Jorropo @jorropo.bsky.social · 26/09/2026
The 20X number I gave is just the load for the camera display in the UI, I didn't measured the latency for this.
000
Jorropo @jorropo.bsky.social · 26/09/2026
Yes, it depends on what is happening. For a classic bottom vision workload was 200ms on the CPU, while now it's in the 50~100ms. The camera runs at 5fps and is the biggest bottleneck currently.
200
Jorropo @jorropo.bsky.social · 26/09/2026
Now a lot of cleaning up to do to make that upstream-able. Claude needed a moderate amount of guidance, at first it picked OpenCV T-API since move OpenCV CPU to OpenCV GPU only seems logical. However one it found the 5X improvement, it kept getting stuck failing at improving the OpenCL usage.
000
Jorropo @jorropo.bsky.social · 26/09/2026
OpenPnP GPU processing results: - baseline, upstream's code, opencv on CPU + some java image kernels 100% CPU usage (1 core) - OpenCV T-API (OpenCL) 20% CPU usage with lots of GPU driver CPU overhead - Vulkan with command queue caching 5% CPU usage 20X performance improvement all done by claude.
210
Jorropo @jorropo.bsky.social · 24/09/2026
2026 satisfaction is when the LLM finishes with 99% of quota used.
Claude code Current session quota 99% used.
040
Jorropo @jorropo.bsky.social · 22/09/2026
I think your last try got reverted because fp → uint was too confusing ? 😄 So maybe a proposal is warranted ?
100
Jorropo @jorropo.bsky.social · 19/09/2026
Ah I see I pulled the trigger. well ...
000
Jorropo @jorropo.bsky.social · 19/09/2026
Sadly for perfect portability the only solution is fixed point math or software floats. Or maybe types like archfloat.x86, archfloat.riscv, .. that make sure to portably implement floats according to one exact ISA.
100
Jorropo @jorropo.bsky.social · 19/09/2026
I wonder if there should be a test mode that randomize let's say the lowest 3 bits after any float operation. FMA is the visible tip of the iceberg, there are edge cases where some CPUs just implement IEEE754 at partial precision on some costly operations.
100
Jorropo @jorropo.bsky.social · 17/09/2026
Wished so
010
Jorropo @jorropo.bsky.social · 08/09/2026
So about that
000
Jorropo @jorropo.bsky.social · 08/09/2026
Someone using valetudo ?
110
Jorropo @jorropo.bsky.social · 07/09/2026
What's the size of the diff ?
110
Jorropo @jorropo.bsky.social · 05/09/2026
pcrypto is opt-in unlike wg which has it's own default-on parallel crypto engine.
000
Jorropo @jorropo.bsky.social · 05/09/2026
I just bought a second hand GX 480 NIC to offload IPsec to the FPGA for 99$ ! For software idk, IPsec now supports the pcrypto module. So the architecture should be identical as wg, per tunnel serial → parallel crypto → serial output altho I wouldn't be surprised if wg's implementation is lighter.
200
Jorropo @jorropo.bsky.social · 05/09/2026
After handshake the daemon would give all cryptographic material to the kernel which would efficiently tunnel packets using the kernel's IPsec impl or even using your NIC IPsec offload capabilities which doesn't exists for wireguard.
100
Jorropo @jorropo.bsky.social · 05/09/2026
I'm thinking to make a spiritual V2 to wireguard. I love it but it lacks Post-Quantum encryption unless you use PSK and even then it still lack PQ-PFS. I'm thinking a userland IPsec daemon that use wireguard's config format and options but with TLS1.3 handshake and X25519MLKEM768 keys.
100
Jorropo @jorropo.bsky.social · 31/08/2026
What happen if you do: f := error.Error For concrete types it give you a function who's first argument is the receiver, and given the cursed newly taught knowledge you just gave me, I guess it also works for interfaces ?
110
Jorropo @jorropo.bsky.social · 30/08/2026
Next time I hear « say the word » I'll do something, idk what but it'll be ugly for some weights.
211
Jorropo @jorropo.bsky.social · 25/08/2026
min & max functions should have been named smallest and biggest tbh.
120
Jorropo @jorropo.bsky.social · 19/08/2026
It's 2026 and github added stacked PR support. Pasting diffs to generate a PR sounds like a 2036 github feature.
120
Jorropo @jorropo.bsky.social · 17/08/2026
Are you also upset because the computer did what you told it to do ?
110
Jorropo @jorropo.bsky.social · 14/08/2026
The biggest offender is the compiler's ssa uber package. It contains half a million of code alone (most of it generated peephole rules matchers). Michael Matloob is doing awesome work on splitting it up and on his branch it compiles in 6m 20s with a 16x multicore ratio.
110
Jorropo @jorropo.bsky.social · 14/08/2026
My 6 core Ryzen 3600 emulating RISC-V (qemu's userland jit) is 4 times faster at building Go than the 64 core beast. Altho Go's bootstrapping doesn't scale well to 64 cores. The parallelism factor is 13x. So assuming a 64x parallelism factor it could be 2minutes which wouldn't be bad at all.
100
Jorropo @jorropo.bsky.social · 14/08/2026
#golang compiler maintainer life: I dusted off my SG2042 64 core RISC-V server. Would be great to run Go's CI. It was broken (wouldn't boot), turns out 1/4 stick of ram works good enough to pass training, but hardlocks the boot process. The single core performance is abysmal, multicore is ok.
100
Jorropo @jorropo.bsky.social · 11/08/2026
Happy to have +2-ed you.
010
Jorropo @jorropo.bsky.social · 07/08/2026
a != null && b == null && a == b Thus null == !null hummmmmmmmmmmmmmmmmmmmm
000
Jorropo @jorropo.bsky.social · 06/08/2026
XCHG AX, ... is 1 byte shorter than MOV AX, ... And on Ryzen CPUs it execute as 2 register renames so there isn't a major performance cost to using it.
010
Jorropo @jorropo.bsky.social · 05/08/2026
We all hate it
010
Jorropo @jorropo.bsky.social · 05/08/2026
Does he knows that bee keepers will pay him to come and take hive rather than exterminating the bees ?
110
Jorropo @jorropo.bsky.social · 31/07/2026
5 5
010
Jorropo @jorropo.bsky.social · 10/07/2026
How was your evening ?
000
Jorropo @jorropo.bsky.social · 07/07/2026
Termux user spotted
020
Jorropo @jorropo.bsky.social · 04/07/2026
I've just finished a deep-dive on Intel APX github.com/golang/go/is... and the new CFCMOVcc and CCMPcc (as well as CTESTcc) instructions look very cool. You can remove branches from && and || even if they have (basic) memory side effects.
github.com
cmd/compile: intel APX meta issue · Issue #80257 · golang/go
I semi regularly get asked, what even is there to do in the toolchain ? Thus this issue track all the ideas that probably help performance of programs in the new Intel APX instruction set, so I can...
0100
Jorropo @jorropo.bsky.social · 02/07/2026
I don't have an intel CPU to try with, which is where uops.info is suspicious so take my results with a grain of salt.
uops.info
uops.info
000
Jorropo @jorropo.bsky.social · 02/07/2026
On intel arrow lake P it uses 3 uops somehow with a TP of 2. For E cores they use 2 uops and TP 0.4 which is more what is expected. I have no idea how intel uses more than 2 uops to implement this. Other backends (like LLVM) also love using leave so I would be surprised if it has negative impact.
110
Jorropo @jorropo.bsky.social · 02/07/2026
On AMD nothing. The CPU basically translate leave into a mov and a pop. I can measure some single % gains in constructed benchmarks where the denser code saves on I-CACHE / no longer cross a cache boundary. 1/2
110
Jorropo @jorropo.bsky.social · 02/07/2026
New surprisingly good go optimization, I've reduced the codesize by 2% with a simple change ! This changes the end of most functions from 2 instructions: add %rsp, $offset pop rbp To a single: leave Once encoded we go from 5 bytes to 1 byte. go-review.googlesource.com/c/go/+/548317
total 45787786 44851997 -935789 -2.044%
1140
Jorropo @jorropo.bsky.social · 02/07/2026
I am trying to move my browser to TLS1.3 exclusively and the only thing that broke today is bsky.app, please fix.
> Firefox error failing to complete a TLS handshake: PR_END_OF_FILE_ERROR
010
Jorropo @jorropo.bsky.social · 30/06/2026
The lack of location make it sounds like you want her to move out to an other country. 😄
100
Jorropo @jorropo.bsky.social · 29/06/2026
After doing a bit of research the whole NSA arc was a distraction to my understanding of the story. The options are not how Bernstein act like they are: ietf-tls-mlkem vs ietf-tls-ecdhe-mlkem They are infact: ietf-tls-mlkem + ietf-tls-ecdhe-mlkem vs ietf-tls-ecdhe-mlkem
020
Jorropo @jorropo.bsky.social · 28/06/2026
*I have no idea if it's true or not, just saying that what I think Bernstein wants to say.
100
Jorropo @jorropo.bsky.social · 28/06/2026
What ¿? > NSA is now overtly paying for standardization of "ietf-tls-mlkem", a weakening of the much more sensible "ietf-tls-ecdhe-mlkem". I'm not a native, please someone explain to me that his thesis isn't that the NSA is buying out peoples to push a weaker alternative to hybrid.*
100
Jorropo @jorropo.bsky.social · 25/06/2026
I'll read it
020
Jorropo @jorropo.bsky.social · 16/06/2026
What is the new benchmark result now that you've removed the elision ?
100
Jorropo @jorropo.bsky.social · 06/06/2026
But I LOVE monospace fonts.
110
Jorropo @jorropo.bsky.social · 27/05/2026
And I love it
110
Jorropo @jorropo.bsky.social · 27/05/2026
Yeah no idea, it just stopped working at some point. dd copies at 1GiB/s which is impossible due to the nature of USB2.0, I wonder if it's a bug in laptop's USB mass storage drive. I've worked around it by using the ttyACM driver and then doing a wget -O - ... | dd ... into the serial shell.
010