Sign in

Falvyu

@falvyu.bsky.social
31 followers 52 following 51 posts

PhD | French | Hardware-aware Algorithm design | Image Processing | HPC SIMD-friends: #SSE, #AVX512, #NEON, #RVV (Opinions are my own)

PostsRepliesMedia
Reposted by Falvyu
Tom Forsyth @tomforsyth.bsky.social · 21/11/2025
Recent discussion about the perils of doors in gamedev reminded me of a bug caused by a door in a game you may have heard of called "Half Life 2". I wrote it up over on Mastodon (I find it's better at long threads): mastodon.gamedev.place/@TomF/115589...
mastodon.gamedev.place
Tom Forsyth (@TomF@mastodon.gamedev.place)
Attached: 1 image Recent discussion about the perils of doors in gamedev reminded me of a bug caused by a door in a game you may heard of called "Half Life 2". Are you sitting comfortably? Then I sha...
18361115
Reposted by Falvyu
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
Tenstorrent decided to publish the first benchmark data for Ascalon's RVV implementation using the instruction throughput benchmark of my rvv-bench benchmark suite. <3 camel-cdr.github.io/rvv-bench-re... Overall, the results look really good so far:
133
Reposted by Falvyu
InstLatX64 @instlatx64.bsky.social · 17/05/2025
The most underutilized #AVX512 instruction, VPSHUFBITQMB, is actually very useful for compactly generating a mask with arbitrary criteria based on b[5:0], e.g. to select special characters for parsing: if the second operand is a broadcasted 64-bit vector, where the ones represent the searched values
3176
Reposted by Falvyu
Underfox @underfox3.bsky.social · 15/04/2025
Intel researchers have proposed Earth, an efficient architecture for RISC-V vector memory access patterns, enabling coalesced strided instruction memory access and buffer-free segment instruction processing. arxiv.org/pdf/2504.08334
2136
Reposted by Falvyu
HGPU group @hgpu.bsky.social · 30/03/2025
Analyzing Modern NVIDIA GPU cores #HardwareArchitecture hgpu.org?p=29837
hgpu.org
Analyzing Modern NVIDIA GPU cores
GPUs are the most popular platform for accelerating HPC workloads, such as artificial intelligence and science simulations. However, most microarchitectural research in academia relies on GPU core …
022
Reposted by Falvyu
Tom Forsyth @tomforsyth.bsky.social · 29/03/2025
Oh shit. I just realised no - it really is. I thought I was joking. This is going to require a long and complex backstory. Sorry absolutely not sorry it'll be funny maybe but also you'll learn something.
15013
Reposted by Falvyu
FCLC (being a silly goose) @fclc.bsky.social · 27/03/2025
Hey Folks! cool little FOSS community tool released for simulating the way @tenstorrent.bsky.social Hardware works! Namely the way the tensix instruction stream operates: github.com/tenstorrent/...
github.com
GitHub - tenstorrent/tensix-isa-simulator
Contribute to tenstorrent/tensix-isa-simulator development by creating an account on GitHub.
091
Reposted by Falvyu
Raph Levien @raphlinus.bsky.social · 21/03/2025
I've now published my blog post, "I want a good parallel computer." raphlinus.github.io/gpu/2025/03/... . Thanks much to all the feedback on the draft, I'd like to think I've clarified some things that might have been confusing.
raphlinus.github.io
I want a good parallel computer
The GPU in your computer is about 10 to 100 times more powerful than the CPU, depending on workload. For real-time graphics rendering and machine learning, you are enjoying that power, and doing those...
26218
Reposted by Falvyu
InstLatX64 @instlatx64.bsky.social · 19/03/2025
Another round of #Intel ISA defragmentation: Finally, #AVX10_2 is a common ISA for all CPU segments!
182
Reposted by Falvyu
Phoronix @phoronix-poster.bsky.social · 19/03/2025
Intel AVX10 Drops Optional 512-bit: No AVX10 256-bit Only E-Cores In The Future - www.phoronix.com/news/Intel-AVX10-D…
phoronix.com
Intel AVX10 Drops Optional 512-bit: No AVX10 256-bit Only E-Cores In The Future
Intel updated their AVX10 whitepaper and associated open-source compiler patches around this next Advanced Vector Extensions standard... While AVX10 had intended to allow either 256-bit or 612-bit modes depending upon processor capabilities, Intel has dropped the 256-bit-only approach and going for 512-bit everywhere. Thus it would seem to indicate that Intel E cores of the future will properly support AVX 512-bit operation!..
0103
Reposted by Falvyu
NSF–DOE Rubin Observatory @rubinobservatory.org · 12/03/2025
Lights, camera, action! The world's largest digital camera has been installed at NSF–DOE Rubin Observatory! 🤩 The LSST Camera was the final major component of the observatory. With it in place, Rubin officially enters its final phase of testing!🔭 🔗: rubinobservatory.org/news/lsst-camera-installed
521570
Reposted by Falvyu
FCLC (being a silly goose) @fclc.bsky.social · 24/02/2025
"All lessons in tech will be relearned every 25 years" -Me, looking at us relearning the same thing again
2102
Reposted by Falvyu
Björkus "No time_t to Die" Dorkus @thephd.dev · 07/02/2025
I gotta see so many posts saying "Rust is slower" and like I think these people just evaluate languages and designs based on vibes and not based on like the material output from the compilers.
61074
Reposted by Falvyu
Jonathan O'Callaghan @astrojonny.bsky.social · 04/02/2025
The chance of asteroid 2024 YR4 hitting out planet in 2032 is now 1.5%, or 1 in 67. The OVERWHELMING likelihood is that the asteroid will miss Earth. But for the first time ever, we might have to seriously consider a deflection mission soon. Let me explain. 🧵 (1/x)
media.tenor.com
a picture of the earth and an asteroid with the national geographic logo
Alt: a picture of the earth and an asteroid heading towards it
95806325
Falvyu @falvyu.bsky.social · 01/02/2025
I don't know anything about Raku, but these ads/posters are quite cool. #FOSDEM
000
Reposted by Falvyu
FCLC (being a silly goose) @fclc.bsky.social · 01/02/2025
Hey folks! Will be talking about the wonderful, wacky world of floating point Sunday afternoon in the HPC, Big data, & Data science dev room at 15:30! #fosdem
1279
Reposted by Falvyu
Kenneth Hoste (boegel) @boegel.bsky.social · 01/02/2025
“Welcome to the largest open source conference on earth!” #FOSDEM
411210
Falvyu @falvyu.bsky.social · 31/01/2025
Hello Brussels ! #FOSDEM
000
Reposted by Falvyu
Björkus "No time_t to Die" Dorkus @thephd.dev · 24/01/2025
Mmmmmnnnn, data.... thephd.dev/the-big-arra...
thephd.dev
Results! - The Big Array Size Survey for C
Happy New Year! It is time report the results of the Array Size Operator survey and answer some comments people have been asking for!
1475
Reposted by Falvyu
Phoronix @phoronix-poster.bsky.social · 21/01/2025
SDL 3 Official Release Now Available With New APIs - www.phoronix.com/news/SDL3-Official…
phoronix.com
SDL 3 Officially Released With New APIs, Better HiDPI & Improved Audio Handling
In addition to the Wine 10.0 stable release today, making the day very exciting as well for Linux gamers is the first official SDL 3.0 release!..
083
Reposted by Falvyu
Phoronix @phoronix-poster.bsky.social · 17/01/2025
Linux 6.14 To Add Support For SpacemiT's "Energy Efficient AI" RISC-V CPUs - www.phoronix.com/news/Linux-6.14-Sp…
phoronix.com
Linux 6.14 To Add Support For SpacemiT's "Energy Efficient AI" RISC-V CPUs
The upcoming Linux 6.14 kernel is poised to introduce initial support for SpacemiT platforms, the Chinese computing chip company developing "next-generation RISC-V high-performance CPUs." For this next Linux kernel release the SpaceMiT Key Stone K1 octa-core RISC-V AI CPU with SpacemiT X60 cores will see support...
181
Reposted by Falvyu
Gaia Mission @gaiamission.bsky.social · 15/01/2025
So long and thanks for all the fish! I just observed my last star this morning... #GaiaDPAC and @esa.int Gaia teams will do their best to make some great data releases from the data I gathered! #GaiaDR4 in 2026, #GaiaDR5 around the end of the decade. Keep posted for our news today at 10:00 CET.
giphy.com
rotating european space agency GIF - Find & Share on GIPHY
Discover & share this rotating european space agency GIF with everyone you know. GIPHY is how you search, share, discover, and create GIFs.
0114
Reposted by Falvyu
NSF–DOE Rubin Observatory @rubinobservatory.org · 14/01/2025
When you turn it on and it works spectacularly! The Rubin team has successfully completed the first on-sky engineering tests using the 144-Mpx engineering camera. Up next: installing the largest digital camera in the world, LSST Camera! 🔭🧪 rubinobservatory.org/news/rubin-completes-comcam-tests
115538
Reposted by Falvyu
八咫烏 @yatagarasu.bsky.social · 01/01/2025
today i learned x264 has a setting specifically for encoding touhou games
61173506
Reposted by Falvyu
efroach76.bsky.social @efroach76.bsky.social · 30/12/2024
As Harold mentions, Prefix-Or is just v | -v. Segment-scan-or can also be simplified, for example uint64_t s = (v | m) >> 1; return s ^ (s - v) (4 instructions instead of 7).
121
Falvyu @falvyu.bsky.social · 29/12/2024
Some neat tricks for computing bit-wise prefix-or and segmented-prefix-or within scalar registers.
341
Reposted by Falvyu
camel-cdr.bsky.social @camel-cdr.bsky.social · 27/12/2024
Removing single line comments in multiple lines in parallel with RVV. If anyone has a good AVX512 solution, please share. Inspired by this reddit post: www.reddit.com/r/simd/comme... #RVV #RISC-V #SIMD
/* Note: '\' is '\n' */
0123456789ABCDEFGHIJKLMNOPQRSTUVWX | vid
Line 1#  # # ## #\Line2 # #\Line 3 | v0 = input
0000001001010110110000001011000000 | m = vmor(vmseq(v0, '#'), vmseq(v0, '\n'))
000000111223345567777777889AAAAAAA | iota = vslide1down(viota(m)) /* vl-=1 */
0######\##\                        | v1 = vslide1up(vcompress(v0, m), 0)
000000###########\\\\\\\###\\\\\\\ | seg-cp-scan = vrgather(v1, iota)
0000001111111111100000001110000000 | seg-cp-scan == '#'
Line 1\Line2\Line 3                | vcompress(v0, seg-cp-scan)
272
Reposted by Falvyu
FCLC (being a silly goose) @fclc.bsky.social · 23/12/2024
Mini thread on developer trust for #AMD. Effectively: the default assumption for AMD is that the stack will be broken. The community tried to help, but as Dave Airlie said years ago “throwing source code over the fence doesn’t make you FOSS” #HPC #ROCm
2245
Reposted by Falvyu
halvarflake.bsky.social @halvarflake.bsky.social · 23/12/2024
Ok, reading semianalysis.com/2024/12/22/m... it seems that *most* of AMDs problems could indeed be remedied by having a team of 4-10 very strong engineers focus on always keeping default PyTorch fast on AMD hardware, with nightly regression tests.
semianalysis.com
MI300X vs H100 vs H200 Benchmark Part 1: Training – CUDA Moat Still Alive
Intro SemiAnalysis has been on a five-month long quest to settle the reality of MI300X. In theory, the MI300X should be at a huge advantage over Nvidia’s H100 and H200 in terms of specifications an…
4175
Reposted by Falvyu
FCLC (being a silly goose) @fclc.bsky.social · 17/12/2024
Curious what’s going on with the Nvidia CPU team. Will they try and bring the server cores down to desktop? Or launch an embedded core via a next gen Jetson before scaling that up? The new Jetson sure as heck looks like them realizing few hobbyist were interested in Orin nano at 500
161
Falvyu @falvyu.bsky.social · 17/12/2024
Updating the Orin Nano from a $500 & 15W 'low-power' computer vision systems to a $250 & 25W 'Generative AI Supercomputer' is an interesting decision. Looks like they are moving from low-power industrial applications (who might have preferred the NX and AGX) to advertising it to 'AI hobbyists'.
000
Reposted by Falvyu
ST010-14 @st01014.bsky.social · 11/12/2024
sbaziotis.com/compilers/co...
sbaziotis.com
Common Misconceptions about Compilers
A curated list of misconceptions about mainstream compilers.
182
Reposted by Falvyu
Luke Lau @lukel97.bsky.social · 11/12/2024
Trying to find the slowest possible RISC-V instruction. This single vlse8.v with a stride of 65536 bytes takes 66 million cycles on a Banana Pi F3. That's 0.04 seconds @1.6GHz #risc-v
A screenshot of a terminal:
luke@bananapif3:~/slowest-instr$ cat main.S
	.section .rodata
str:	.asciz "Cycles: %d\n"
foo:	.zero 256 * STRIDE
	.section .text
	.global main

main:
	addi	sp, sp, -8
	sd	ra, 0(sp)

	rdcycle	s1
	rdcycle s2
	sub	s3, s2, s1 	# rdcycle overhead

	la	a0, foo
	li	a1, STRIDE
	vsetvli t0, zero, e8, m8, tu, mu
	rdcycle s1
	vlse8.v	v8, (a0), a1
	rdcycle	s2

	sub	s1, s2, s1
	sub	s1, s1, s3
	la	a0, str
	mv	a1, s1
	call	printf

	ld	ra, 0(sp)
	addi	sp, sp, 8
	ret
luke@bananapif3:~/slowest-instr$ clang main.S -DSTRIDE=65536 -march=rv64gv 
luke@bananapif3:~/slowest-instr$ perf stat -e cycles:u ./a.out 
Cycles: 66640979

 Performance counter stats for './a.out':

        78,064,581      cycles:u                                                           

       0.049648957 seconds time elapsed

       0.000000000 seconds user
       0.049907000 seconds sys
4235
Reposted by Falvyu
Felipe O. Carvalho @felipe.rs · 06/12/2024
Yet another case of “hardware companies should be creating compilers and new programming models instead of just hardware.” NVIDIA is a hardware and compiler company.
Lisa Su (AMD): I think our strategy is that we have to increase performance continually. For gaming in particular, the gaming software developers have not necessarily used all the cores from time to time. There is no physical reason we couldn't go past 16 cores. The key is that we're going at a pace that the software guys can and do utilise it.
082
Falvyu @falvyu.bsky.social · 06/12/2024
There's been a lot of discourse on fixed-length SIMD (e.g. SSE, AVX512, NEON) vs variable-length/vector (SVE, RVV) ISAs. It's probably too early for a definite answer. But as I've designed SIMD image processing algorithms, I'll share a few results.
173
Reposted by Falvyu
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/12/2024
Here are some slightly tricky RVV mask patterns.
'RVV mask tricks'

# broadcast nth bit
vmand.mm v8, in, mNth
vcpop.m t0, v8
sub t0, x0, t0
vmv.v.x v8, t0

# prefix xor
viota.m v8, in
vand.vi v8, v8, 1
vmsne.vi v8, v8, 0
vmor.mm v0, v8, in # can often be omitted

# move nth bit to first
vmand.mm v8, in, mNth
vcpop.m t0, v8
vmv.v.x v8, t0
vmsof.m v0, v8

# move mask to GPR
vmv.x.s t0, v0
# move GPR to mask
vmv.s.x v0, t0
# assuming vl<=64, set SEW=64 before

# these two should really be dedicated instructions
# shift mask up by 1
vslide1up.vx v8, in, x0
vsrl.vi v8, v8, 7
vmadd.vx v0, 2, v8

# shift mask up by 1
vslide1down.vx  v8, in, x0
vadd.vv v0, in, in
vmacc.vx v0, 128, v8
173
Falvyu @falvyu.bsky.social · 03/12/2024
I defended my PhD last Friday (design of efficient data-dependent image processing algorithms). I'd like to thank everyone who has been involved in it (jury members, advisors, colleagues, and of course my family). This has been a long adventure, and I'm now looking forward to the next thing.
231
Reposted by Falvyu
FCLC (being a silly goose) @fclc.bsky.social · 29/11/2024
Not quite the first, but definitely one of them! Doing RVV1.0 in an OoO machine isn't for the faint of heart. As for ISA, I can confirm that it's fully RVV1.0 compliant. I don't know off hand if we support the custom opcodes we use on Tensix Vector, but I don't believe so!
221
Reposted by Falvyu
FCLC (being a silly goose) @fclc.bsky.social · 28/11/2024
2x256, full RVV1.0 as well as a fair few of the optional extras to RVV1.0! Phoronix article here: www.phoronix.com/news/LLVM-20... LLVM patches here: github.com/llvm/llvm-pr... One Pager: cdn.sanity.io/files/jpb4ed...
phoronix.com
LLVM Merges Support The For Tenstorrent TT-Ascalon-D8 RISC-V CPU
Adding to the interesting code building up for next spring's release of the LLVM 20 compiler stack is having the Tenstorrent TT-Ascalon D8 as the newest RISC-V processor target.
4121
Reposted by Falvyu
Underfox @underfox3.bsky.social · 25/11/2024
In this paper is presented OMP4Py, the first pure Python implementation of OpenMP, allowing developers to write parallel code with the same level of control and flexibility as in C, C++, or Fortran. #HPC arxiv.org/pdf/2411.14887
14413
Reposted by Falvyu
Jeff Geerling @jeffgeerling.com · 26/11/2024
👀 github.com/geerlingguy/...
github.com
Add Intel Arc A750 · Issue #510 · geerlingguy/raspberry-pi-pcie-devices
I am testing an Intel Arc A750 8GB graphics card: It requires external PCIe power, so needs to be used with a suitable power supply. Intel seems to only officially support Ubuntu 22.04, and by defa...
51042
Reposted by Falvyu
아리다 남편 @parkbot.bsky.social · 24/11/2024
Superintendent Chalmers: Good lord! What is happening in there?

Principal Skinner: we are building a universal processor

Chalmers: a chip that’s a paradigm shift in performance, replaces the CPU, GPU, DSP, and FPGA, is lower power, and will tape out in 2026?

Skinner: Yes.

Chalmers: May I see it?

Skinner: No.
4428
Reposted by Falvyu
François Tessier @hephtaicie.bsky.social · 25/11/2024
I'm launching the hashtag #HPCTales to share ridiculous articles about #HPC. I'm starting with news from a French defense (!!) website that headlines: "The 2nd most powerful supercomputer in the world (Frontier) is now able to create its own Matrix-like universe thanks to Exascale" 🤦
2106
Reposted by Falvyu
InstLatX64 @instlatx64.bsky.social · 19/11/2024
#Intel projects after 2024 v23
152
Reposted by Falvyu
HPC Guru @hpcguru.bsky.social · 18/11/2024
64th #Top500 is out! o #ElCapitan at LLNL is the new #1 o We now have 3 #Exascale systems officially o #HPC6 at Eni is new at #5 o Tuolumne, a smaller version of #ElCapitan, rounds up the #Top10 top500.org/news/el-capi... #HPC #AI #SC24
59336
Reposted by Falvyu
Karl Podesta @karlpodesta.bsky.social · 17/11/2024
Mandatory things to ask my #Microsoft #Azure colleagues at #SC24 : “I’m having a problem with <obscure-hardware-driver> on my Windows laptop - can you fix it?” 💻 “Got any free Xboxes?” 🎮 “When will you rename Copilot to Clippy 2.0?” 🤖 #we-also-do #HPC 🤦 ❤️🚀
71204
Reposted by Falvyu
FCLC (being a silly goose) @fclc.bsky.social · 16/11/2024
First cut of an #SIMD/ #ASM starter pack! I’ve definitely missed folks, please let me know if you want to be added to the list! go.bsky.app/7iKAaSv
123