Sign in

Gabriel

@dssgabriel.bsky.social
19 followers 88 following 19 posts

PhD candidate, HPC Software Engineering @cea.fr / DAM MSc HPC & Simulation from @univparissaclay.bsky.social Architecture, microbenchmarking & SIMD sorcery. Research on distributed computing, data structures & memory layouts at exascale. RTFM 👹

PostsRepliesMedia
Reposted by Gabriel
Claudia Blaas-Schenner @claudiablaas.bsky.social · 06/10/2026
We had a successful and really productive #MPI_Forum meeting the past two days at @tuwien.at #MPI #HPC
MPI Forum members at the October 2026 meeting.
043
Reposted by Gabriel
EuroMPI Conference @eurompiconf.bsky.social · 07/10/2026
Kicking of EuroMPI & IWOMP 2026 at TU Wien.
Michael Klemm (AMD, OpenMP ARB) and Claudia Blaas-Schenner (ASC / TU Wien) kicking of EuroMPI & IWOMP 2026 at TU Wien.
152
Reposted by Gabriel
High Performance Software Foundation (HPSF) @hpsf.bsky.social · 18h
The #HPSFcon 2027 CFP is open! Share your HPC software, community & AI dev work in Montreal, April 12-16. Submit your proposal by Sunday, January 17. Learn more + submit: bit.ly/3VAd5FI
012
Reposted by Gabriel
Torsten Hoefler 🇨🇭 @thoefler.bsky.social · 27/07/2026
Honored that our paper won the 2025 IEEE TPDS Best Paper Award! buff.ly/4AiBN0S The main idea is to use differential analysis on dependence graphs to causally determine root causes of performance issues. Credits to Yuyang for the hard work, Jidong for leadership, and to all of our co-authors!
053
Reposted by Gabriel
chipsandcheese.bsky.social @chipsandcheese.bsky.social · 14/07/2026
Hello you fine Internet folks, Today we are doing a look into the upcoming x86 ISA, ACE, which is x86's Outer Product Matrix Multiplication along with the differences between ACE, Intel's AMX, and Arm's SME. Hope y'all enjoy! chipsandcheese.com/p/is-x86-rea...
chipsandcheese.com
Is x86 ready to ACE it?
CPU designs must evolve to keep up with changing workloads.
062
Reposted by Gabriel
Ben Leonard @benleonard.bsky.social · 11/05/2026
Web 3d explorer for @oxide.computer. Built in three.js/r3f. Tonnes of little interactions and details in this. Poke around, and take a look at the guided tours. explorer.oxide.computer
515423
Reposted by Gabriel
CEA @cea.fr · 27/04/2026
#Fusion🔥| 50 000 000°C : c'est la température qui règne à l’intérieur du tokamak WEST en fonctionnement. ⛔Un endroit où (presque) personne ne peut entrer… Nous avons donné accès à Gauthier DM Science qui vous offre un reportage au ❤️ des coulisses de la fusion 👇 youtu.be/E39QEuU8R9s?...
youtu.be
La machine de fusion française qui explose tous les records scientifiques
YouTube video by Gauthier Dm - Science
0192
Reposted by Gabriel
Pavel Laptev @pavellaptev.bsky.social · 27/03/2026
Got some new prints in the @gitbutler.com office
122
Reposted by Gabriel
Phoronix @phoronix-poster.bsky.social · 09/03/2026
NVIDIA Adds Official Support For RHEL-Compatible Distributions Like AlmaLinux With CUDA 13.2 - www.phoronix.com/news/NVIDIA-Offici…
phoronix.com
NVIDIA Adds Official Support For RHEL-Compatible Distributions Like AlmaLinux With CUDA 13.2
With CUDA 13.2 that is now shipping, NVIDIA has provided official support for Red Hat Enterprise Linux compatible distributions/downstreams like AlmaLinux to CUDA. With this official NVIDIA CUDA support for these RHEL-compatible distributions, NVIDIA is also allowing the NVIDIA packages to be distributed directly from the OS package repositories...
0102
Reposted by Gabriel
GCC - GNU Toolchain @gnutools.bsky.social · 20/02/2026
Defer available in GCC and Clang gustedt.wordpress.com/2026/02/15/d...
gustedt.wordpress.com
Defer available in gcc and clang
About a year ago I posted about defer and that it would be available for everyone using gcc and/or clang soon. So it is probably time for an update. Two things have happened in the mean time: A tec…
001
Reposted by Gabriel
High Performance Software Foundation (HPSF) @hpsf.bsky.social · 28/01/2026
Join us for the HPSF Community Summit 2026 in Braunschweig, Germany, February 25-27! 💚 Learn what’s new with HPSF projects, give us feedback on your use of HPSF software, meet with project communities, and tell us how to grow and improve them. Details: hpsf.io/event/hpsf-c...
023
Reposted by Gabriel
GitButler ⧓ @gitbutler.com · 11/02/2026
Love the GitButler GUI but miss your CLI? Have we got the solution for you! youtu.be/Jg8L3SbgZ3o?...
youtu.be
Intro to the GitButler CLI
YouTube video by GitButler
0113
Reposted by Gabriel
Steve Klabnik @steveklabnik.com · 08/01/2026
#jj-vcs 0.37.0 came out yesterday! im intrigued by the new divergent change syntax, seems very neat github.com/jj-vcs/jj/re...
github.com
Release v0.37.0 · jj-vcs/jj
About jj is a Git-compatible version control system that is both simple and powerful. See the installation instructions to get started. Release highlights A new syntax for referring to hidden and...
4607
Reposted by Gabriel
Underfox @underfox3.bsky.social · 06/01/2026
Please note: Any claims of AI Exascale, AI Zettascale or beyond computing power are just baloney. Real computing power is measured in FP64. Period. AMD embraced utter stupidity by adopting this terminology by the leather jacket man. It's a really shame! #CES2026 #AMD
283
Reposted by Gabriel
Phoronix @phoronix-poster.bsky.social · 30/12/2025
LLVM 22 Lands NVIDIA Olympus CPU Scheduling Model - www.phoronix.com/news/NVIDIA-Olympu…
phoronix.com
LLVM 22 Lands NVIDIA Olympus CPU Scheduling Model
NVIDIA's Olympus are the ARM64 cores found within the upcoming Vera CPU that will be paired with Rubin. Olympus cores are claimed to be twice as fast as NVIDIA's current CPU cores found in Grace and based on Neoverse-V2. Earlier this year the open-source compilers landed initial support for Olympus while now a proper CPU scheduling model has been upstreamed into LLVM 22...
011
Reposted by Gabriel
Pete 🙆‍♂️ @petelevasseur.com · 28/12/2025
ever curious why people that work in safety-critical systems want to use Rust? here's the title slide for the talk i'll give at @rustnationuk.bsky.social about this
title slide of talk being given at Rust Nation UK:

[Title] Rust for Foundational SW or: Safety-Critical Software in Rust
2278
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 24/12/2025
Day 24 of #AoCO2025! A loop summing 0+1+2+...+n. GCC unrolls it. Clang does something jaw-dropping: the loop vanishes entirely, replaced by a direct calculation. How?! xania.org/202512/24-cu... youtu.be/V9dy34slaxA
xania.org
When compilers surprise you — Matt Godbolt’s blog
Sometimes compilers can surprise and delight even a jaded old engineer like me
3314
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 23/12/2025
Day 23 of #AoCO2025! Switch → jump table? Sometimes. Other times: arithmetic, bitmasks, or something cleverer. Compilers have more tricks than you think. xania.org/202512/23-sw... youtu.be/aSljdPafBAw
xania.org
0222
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 22/12/2025
Day 22: String comparison against "ABCDEFG" should call memcmp, but Clang inlines it with some clever memory tricks. How does it compare 7 bytes so efficiently? xania.org/202512/22-me... youtu.be/kXmqwJoaapg #AoCO2025
xania.org
Clever memory tricks — Matt Godbolt’s blog
We learn that compilers have tricks to access memory efficiently
0234
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 21/12/2025
Day 21: Summing integers? Compiler vectorises beautifully—8 at a time! Switch to floats? It refuses, doing each add individually. Same code, totally different output. Why? 🤔 xania.org/202512/21-ve... youtu.be/lUTvi_96-D8 #AoCO2025
xania.org
When SIMD Fails: Floating Point Associativity — Matt Godbolt’s blog
Why floating point maths doesn't vectorise like integers, and what to do about it
2223
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 20/12/2025
Day 20: Process 65,536 integers one at a time? Nah. The compiler vectorises it to handle 8 at once — same code, 8× faster! SIMD auto-vectorisation is compiler magic 🚀 xania.org/202512/20-si... youtu.be/d68x8TF7XJs #AoCO2025
xania.org
1264
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 19/12/2025
Day 19: Recursive functions calling themselves endlessly — stack growth? Nope! The compiler turns recursion into loops. Tail call optimisation is magic ✨ xania.org/202512/19-ta... youtu.be/J1vtP0QDLLU #AoCO2025
xania.org
Chasing your tail — Matt Godbolt’s blog
The art of not (directly) coming back: tail call optimisation
2243
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 18/12/2025
Day 18: Function with fast & slow paths. Inline = code bloat. Don't inline = slow fast path. Can't have both—or can you? The compiler finds a surprising way out of this dilemma. xania.org/202512/18-pa... youtu.be/STZb5K5sPDs #AoCO2025
xania.org
Partial inlining — Matt Godbolt’s blog
Inlining doesn't have to be all-or-nothing
0264
Reposted by Gabriel
InstLatX64 @instlatx64.bsky.social · 18/12/2025
Actually, this die configuration is not new information, it was already mentioned on this removed slide: (Although the CPU die's CBB name is seems still new.)
012
Reposted by Gabriel
Dead Code @deadcode.website · 17/12/2025
Listen: shows.acast.com/dead-code/e...
shows.acast.com
Deferred Conflict (with Steve Klabnik) | Dead Code
A podcast about how the software industry got this way
053
Reposted by Gabriel
Gergely Orosz @gergely.pragmaticengineer.com · 17/12/2025
How have servers and the cloud evolved in the last 30 years, and what might be next? @bcantrill.bsky.social has been at the thick of the industry since the Dotcom Boom, and shares fascinating stories. Bryan is one of my all-time favorite people to talk with - don't miss this one. (cont'd)
3587
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 17/12/2025
Day 17: Inlining — the ultimate optimisation ✨ A function gets inlined, half vanishes. The assembly is cleaner than hand-written. How does copy-paste make code disappear? xania.org/202512/17-in... youtu.be/JFHfFTvMPp0 #AoCO2025
xania.org
Inlining - the ultimate optimisation — Matt Godbolt’s blog
Copy paste can sometimes be a good thing, at least if the compiler does it for you
0193
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 16/12/2025
Day 16: Calling conventions matter! Pass 8 chars as separate args: stack spillage. Pack them in a struct: single register. Sometimes structs are MORE efficient than separate parameters! xania.org/202512/16-ca... youtu.be/Yaw8AMoP4sI #AoCO2025
xania.org
Calling all arguments — Matt Godbolt’s blog
Knowing how compilers call functions can help with design - and optimisation
2425
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 15/12/2025
Day 15: Two nearly identical loops—one writes to memory every iteration, the other stays in registers. Same code, wildly different performance. The culprit? Aliasing! xania.org/202512/15-al... youtu.be/PPJtJzT2U04 #AoCO2025
xania.org
Aliasing — Matt Godbolt’s blog
Knowing when the compiler can't optimise is important too
0274
Reposted by Gabriel
Glenn K. Lockwood @glennklockwood.com · 15/12/2025
Does this mean no more dirt-cheap NRE from Slurm? Or will Slurm development no longer be coin-operated? Would love to see serious engineering effort go into modernizing Slurm, but this could go in many directions.
201
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 14/12/2025
Day 14: Add ONE global counter to your loop and watch LICM vanish—strlen called every iteration! Why would incrementing an unrelated variable break the optimisation? 🤔 xania.org/202512/14-li... youtu.be/OwFNblEEAXo #AoCO2025
xania.org
When LICM fails us — Matt Godbolt’s blog
When aliasing can prevent loop-invariant code motion
1284
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 13/12/2025
Day 13 of Advent of Compiler Optimisations! 🔄 Loop calling a function whose result never changes? One compiler hoists it out automatically. The other… doesn't. Even with hints! xania.org/202512/13-li... youtu.be/dIwaqJG0WDo #AoCO2025
xania.org
Loop-Invariant Code Motion — Matt Godbolt’s blog
The compiler can move code outside of loops to speed things up
0163
Reposted by Gabriel
Shafik Yaghmour @shafik.bsky.social · 12/12/2025
Cursed code: void* f(void *p) { return p + 1; } Both gcc and clang support void* arithmetic as an extension in C: gcc.gnu.org/onlinedocs/g... -pedantic FTW! Godbolt: godbolt.org/z/rcrqWvMGW #Programming
gcc.gnu.org
Pointer Arith (Using the GNU Compiler Collection (GCC))
Pointer Arith (Using the GNU Compiler Collection (GCC))
142
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 12/12/2025
Day 12 of Advent of Compiler Optimisations! A loop that checks the same thing every time. The compiler's solution? Make the code bigger to make it faster. Wait, what? xania.org/202512/12-lo... youtu.be/-VCrYshE7iQ #AoCO2025
xania.org
Unswitching loops for fun and profit — Matt Godbolt’s blog
Duplicating loops around can yield some decent optimisations
0243
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 11/12/2025
Day 11: A clever bit-counting loop using the "clear bottom bit" trick. Change one compiler flag and... wait, what just happened to my loop?! Pattern recognition at its finest. xania.org/202512/11-po... youtu.be/Hu0vu1tpZnc #AoCO2025
xania.org
2295
Reposted by Gabriel
High Performance Software Foundation (HPSF) @hpsf.bsky.social · 11/12/2025
Kokkos 5.0 is officially out. ✨ Details: - Moves the project to C++20 - Retires older interfaces, reducing complexity for future work - Ideal time for teams to review workflows Read the full update here: hpsf.io/blog/2025/ko...
021
Gabriel @dssgabriel.bsky.social · 10/12/2025
Where did you guys get the info for the facility power draw and cooling limits? 👀 Was it publicly announced somewhere?
000
Reposted by Gabriel
chipsandcheese.bsky.social @chipsandcheese.bsky.social · 10/12/2025
Hello you fine Internet folks, At SC25, some of the specs of the upcoming Alice Recoque supercomputer were announce from which I have tried estimated the rough FP64 FLOPs of AMD's upcoming MI430X. Hope, y'all enjoy! chipsandcheese.com/p/sc25-estim... old.chipsandcheese.com/2025/12/10/s...
chipsandcheese.com
SC25: Estimating AMD’s Upcoming MI430X’s FP64 and the Discovery Supercomputer
Hello, you fine Internet folks,
111
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 10/12/2025
Day 10: Fixed loop count? Compiler transforms code surprisingly. At 50 iterations it switches strategy. What's the cutoff and why? xania.org/202512/10-lo... youtu.be/HvF3tF2efEA #AoCO2025
xania.org
Unrolling loops — Matt Godbolt’s blog
Learning when the compiler decides to unroll loops for performance
1245
Reposted by Gabriel
Steve Klabnik @steveklabnik.com · 09/12/2025
"why i think #jj-vcs is worth your time" schpet.com/note/why-i-t...
schpet.com
why i think jj-vcs is worth your time
14710
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 09/12/2025
Day 9: Why does the compiler keep an "expensive" multiply in the loop instead of using clever addition tricks? The slower operation enables something more valuable. What's the hidden benefit? xania.org/202512/09-in... youtu.be/vZk7Br6Vh1U #AoCO2025
xania.org
Induction variables and loops — Matt Godbolt’s blog
Compilers can rewrite loops to avoid expensive calculations
1316
Reposted by Gabriel
NextSilicon @nextsilicon.com · 08/12/2025
The Spectra supercomputer at Sandia National Laboratories is now live - powered by 128 #NextSilicon Maverick-2 accelerators. A first-of-its-kind system built on our Intelligent Compute Architecture for faster, more efficient HPC. 🔗 Full report here: tinyurl.com/yc7fy5x8
021
Reposted by Gabriel
Shafik Yaghmour @shafik.bsky.social · 08/12/2025
Cursed code of the day #programming
godbolt screen shot

Code:

#include <stdio.h>

volatile double d = -32.0;
  
int main(void) {
  printf("%f\n%lx\n%lx\n", d, 
    (unsigned long)d,
    (unsigned long)-32.0);
}


clang armv8 output:

-32.000000
0
ffffe8ed40f8

gcc ARM64 output:

-32.000000
0
0

Clang x86-64 output:

-32.000000
ffffffffffffffe0
7ffde9ce61f8

gcc x86-64 output:

-32.000000
ffffffffffffffe0
ffffffffffffffe0
042
Reposted by Gabriel
HGPU group @hgpu.bsky.social · 07/12/2025
Microbenchmarking NVIDIA’s Blackwell Architecture: An in-depth Architectural Analysis #PTX #CUDA #Benchmarking #Blackwell #HPC hgpu.org?p=30437
hgpu.org
Microbenchmarking NVIDIA’s Blackwell Architecture: An in-depth Architectural Analysis
As GPU architectures rapidly evolve to meet the overcoming demands of exascale computing and machine learning, the performance implications of architectural innovations remain poorly understood acr…
022
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 08/12/2025
Day 8 of Advent of Compiler Optimisations! 🔄 Index-based for vs pointer while vs range-for vs std::accumulate—which is fastest? Three produce identical assembly, but one doesn't! xania.org/202512/08-go... youtu.be/FB8Hgj3TpJM #AoCO2025
xania.org
Going loopy — Matt Godbolt’s blog
Exploring the ways optimisers deal with loop constructs
13011
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 07/12/2025
Day 7: Dividing by 10 compiles to... no division at all! How does the compiler pull this off? Magic constants, bit shifts, and a clever arithmetic trick. xania.org/202512/07-di... youtu.be/V9Pvv1tkocM #AoCO2025
xania.org
Multiplying our way out of division — Matt Godbolt’s blog
How compilers avoid expensive division with multiplication tricks
0347
Reposted by Gabriel
InstLatX64 @instlatx64.bsky.social · 07/12/2025
#Intel released the "Intel #Xeon6 vs #AMD #EPYC Competitive Infographic" pdf: cdrdv2-public.intel.com/859022/intel...
011
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 06/12/2025
Day 6 of Advent of Compiler Optimisations! Divide by 512—just a shift, right? But the compiler adds extra instructions. Why? A subtle difference between what you asked and what you meant! xania.org/202512/06-di... youtu.be/7Rtk0qOX9zs #AoCO2025
xania.org
Division — Matt Godbolt’s blog
Division doesn't have to be slow with some clever tricks
13511
Reposted by Gabriel
Matt Godbolt @matt.godbolt.org · 05/12/2025
Day 5 of Advent of Compiler Optimisations! x86 has LEA, but ARM has the barrel shifter—instructions can shift operands cheaply. The compiler uses this to multiply without multiplying! xania.org/202512/05-ba... youtu.be/TZubUyr2UEY #AoCO2025
xania.org
ARM's barrel shifter tricks — Matt Godbolt’s blog
The ARM architecture has a cool feature, and compilers know how to use it
2296