Sign in

Luke Lau

@lukel97.bsky.social
152 followers 72 following 13 posts

LLVM at Igalia

PostsRepliesMedia
Reposted by Luke Lau
LLVM Weekly @llvmweekly.org · 25/09/2026
Remembering Johannes Doerfert: "It is with great sadness that we share the news of the passing of Johannes Doerfert. [...] Johannes was one of the most prolific and respected contributors to the LLVM compiler project, and his loss will be deeply felt." blog.llvm.org/posts/2026-0...
blog.llvm.org
Remembering Johannes Doerfert
It is with great sadness that we share the news of the passing of Johannes Doerfert, on September 17, 2026, at the age of 36, after a battle with cancer.
026
Reposted by Luke Lau
Nicolò Ribaudo @nicr.dev · 21/09/2026
Today is Igalia's 25th birthday! 🎂 www.igalia.com/2026/09/21/T...
igalia.com
Twenty-Five Years Upstream | Igalia
Igalia is an open source consulting firm specialised in the development of innovative projects and solutions. Our engineers have expertise in a wide range of technological areas, including browsers an...
28212
Reposted by Luke Lau
Joyee Cheung @joyeecheung.bsky.social · 23/06/2026
Had a great Igalia week last week, @lukel97.bsky.social, Cathie and I fed other Igalians some Zongzi, told them the story of the Dragon Boat Festival and taught them how to write Chinese words related to the festival! (Forgot to take a picture of our table but I have a picture of preps ^^)
Preparing for the Chinese writing activity
1171
Reposted by Luke Lau
Ujjwal Sharma @ryzokuken.dev · 01/05/2026
I've been working for a long time alongside different folks to spread the word about web standards, how JavaScript is standardized and help improve how responsive we are to the needs of developers. For more @tc39.es lore, read my first post in "What even is Ecma?" www.ryzokuken.dev/blog/about-e...
ryzokuken.dev
What even is Ecma? (Part 1)
Ujjwal Sharma — Developer Advocate at Igalia, TC39 Co-chair, ECMA-402 Co-editor.
12512
Luke Lau @lukel97.bsky.social · 27/01/2026
One of the nice parts of #llvm is that often times you'll find yourself needing to do some sort of non-trivial analysis, but usually there's already a pass for it. Here's how you can reuse a block frequency analysis to make a chess engine 7% faster on #riscv: lukelau.me/2026/01/26/c...
lukelau.me
Closing the gap, part 2: Probability and profitability
Welcome back to the second post in this series looking at how we can improve the performance of RISC-V code from LLVM.
0111
Luke Lau @lukel97.bsky.social · 10/12/2025
Does LLVM produce slower RISC-V code than GCC? Currently, yes. Can we make LLVM produce faster code? Also, yes! lukelau.me/2025/12/10/c... #llvm #riscv
lukelau.me
Closing the LLVM RISC-V gap to GCC, part 1
At the time of writing, GCC beats Clang on several SPEC CPU 2017 benchmarks on RISC-V1: Compiled with -march=rva22u64_v -O3 -flto, running the train ↩
0157
Reposted by Luke Lau
camel-cdr.bsky.social @camel-cdr.bsky.social · 23/10/2025
"How NOT To Program an Out-of-order Vector Processor" slides are public. static.sched.com/hosted_files...
111
Reposted by Luke Lau
Igalia @igalia.com · 17/10/2025
We're looking forward to the RISC-V Summit North America next week where Mikhail Gadelha (one of our compiler engineers) will be presenting "Unlocking 15% More Performance: A Case Study in LLVM Optimization for RISC-V". Be sure to catch his talk next Thurs riscvsummit2025.sched.com/event/28OTp/...
A title card with a photo of Mikhail and the same information, but adding 11:50am (in Santa Clara)
0105
Reposted by Luke Lau
Hong Kong Free Press HKFP @hongkongfp.com · 04/06/2025
Police have deployed an armoured vehicle in Hong Kong's commercial heart, amidst an ongoing heavy security presence on the 36th anniversary of the Tiananmen Square crackdown. In full: buff.ly/f4hVB50
buff.ly
In Pictures: HK police deploy armoured vehicle on Tiananmen anniversary
Police have deployed an armoured vehicle in Hong Kong's commercial heart, amidst an ongoing heavy security presence on the 36th anniversary of the Tiananmen Square crackdown.
0149
Reposted by Luke Lau
Alex Bradbury @asbradbury.org · 14/05/2025
I'm delighted to see two of @igalia.com's projects for RISE highlighted at the RISC-V Summit Europe. Find out more about our work on both LLVM optimisation and testing/CI on the RISE blog (with more to come in the future!): riseproject.dev/2025/05/08/p... riseproject.dev/2024/10/15/w...
Picture of a presenter showing a slide that details outcomes of RISE funded RISC-V software ecosystem projects.
063
Reposted by Luke Lau
Alex Bradbury @asbradbury.org · 12/04/2025
We're looking forward to EuroLLVM next week in Berlin. Be sure to check out talks from my colleague @lukel97.bsky.social and myself on: * Work to further improve RISC-V vector codegen (extending the VL Optimizer), and * Work done with the support of RISE to improve RISC-V LLVM testing.
094
Reposted by Luke Lau
Paulo Matos @ocmatos.com · 06/03/2025
What if I told you 3DNow! square root recíprocals are defined for negative numbers?... Also the amazing FEX 2503 is out. Read about some of my work and the work of other FEX maintainers' in the release notes: fex-emu.com/FEX-2503/ #fex #igalia #gaming #linux #arm64
fex-emu.com
FEX 2503 Tagged
Here we are again, another month and some more cool changes with FEX. Let’s dive in and see what has changed!
142
Reposted by Luke Lau
Alex Bradbury @asbradbury.org · 27/02/2025
Some notes on ccache+LLVM. Summary: if you do a lot of builds across different checkouts/worktrees/builddirs, be sure to set the base_dir option and -DLLVM_USE_RELATIVE_PATHS_IN_DEBUG_INFO=ON muxup.com/2025q1/ccach...
muxup.com
ccache for LLVM builds across multiple directories
TL;DR: ccache base_dir saves the day
094
Reposted by Luke Lau
chipsandcheese.bsky.social @chipsandcheese.bsky.social · 26/01/2025
Hello you fine Internet folks, Today's article is on SiFive's P550 microarchitecture. The P550 core is one of the fastest RISC-V cores available currently and is claimed to be comparable to ARM's Cortex A75. Hope y'all enjoy! old.chipsandcheese.com/2025/01/26/i... open.substack.com/pub/chipsand...
old.chipsandcheese.com
Inside SiFive’s P550 Microarchitecture
RISC-V is a relatively young and open source instruction set. So far, it has gained traction in microcontrollers and academic applications. For example, Nvidia replaced the Falcon microcontrollers …
0125
Reposted by Luke Lau
Joyee Cheung @joyeecheung.bsky.social · 11/01/2025
New blog post covering the mysterious 10ms startup regression of Node.js on macOS, the journey of investigating the issue with various performance tools, and figuring out the fix (which also helped making the binary smaller). joyeecheung.github.io/blog/2025/01...
joyeecheung.github.io
Executable loading and startup performance on macOS
Recently, I fixed a startup performance regression in Node.js on macOS after an extensive investigation. Along the way, I learned a lot about tools on macOS and Node.js compilation workflows that don’
312618
Luke Lau @lukel97.bsky.social · 27/12/2024
A Simple ELF 4zm.org/2024/12/25/a...
4zm.org
A Simple ELF - The Ivory Tower
The Ivory Tower is a blog about software engineering and development philosophy by Anders Sundman.
000
Reposted by Luke Lau
Joyee Cheung @joyeecheung.bsky.social · 16/12/2024
After two months of chasing, finally found out what's happening behind this mysterious startup time regression on macOS from Node.js v20.x - it's missing -fvisibility=hidden 😅 (I guess that's what happens when the build configs become dusty enough) github.com/nodejs/node/...
github.com
build: build v8 with -fvisibility=hidden on macOS by joyeecheung · Pull Request #56275 · nodejs/node
V8 should be built with -fvisibility=hidden, otherwise the resulting binary would contain unnecessary symbols. In particular, on macOS, this leads to 5000+ weak symbols resolved at runtime, leading...
3598
Reposted by Luke Lau
Joe Cutler @alphaconvert.bsky.social · 12/12/2024
Recently I came across this treatise by Stephen Dolan github.com/ocaml/ocaml/...
github.com
Abnormally slow loop (25x) under OCaml 5 / macOS / arm64 · Issue #13262 · ocaml/ocaml
Hello, I am using macOS Ventura 13.6.7 with an Apple M2 Max processor. A loop that writes values into an integer array is about 20x slower with OCaml 5 than with OCaml 4. Using Array.set versus Arr...
2235
Luke Lau @lukel97.bsky.social · 11/12/2024
Trying to find the slowest possible RISC-V instruction. This single vlse8.v with a stride of 65536 bytes takes 66 million cycles on a Banana Pi F3. That's 0.04 seconds @1.6GHz #risc-v
A screenshot of a terminal:
luke@bananapif3:~/slowest-instr$ cat main.S
	.section .rodata
str:	.asciz "Cycles: %d\n"
foo:	.zero 256 * STRIDE
	.section .text
	.global main

main:
	addi	sp, sp, -8
	sd	ra, 0(sp)

	rdcycle	s1
	rdcycle s2
	sub	s3, s2, s1 	# rdcycle overhead

	la	a0, foo
	li	a1, STRIDE
	vsetvli t0, zero, e8, m8, tu, mu
	rdcycle s1
	vlse8.v	v8, (a0), a1
	rdcycle	s2

	sub	s1, s2, s1
	sub	s1, s1, s3
	la	a0, str
	mv	a1, s1
	call	printf

	ld	ra, 0(sp)
	addi	sp, sp, 8
	ret
luke@bananapif3:~/slowest-instr$ clang main.S -DSTRIDE=65536 -march=rv64gv 
luke@bananapif3:~/slowest-instr$ perf stat -e cycles:u ./a.out 
Cycles: 66640979

 Performance counter stats for './a.out':

        78,064,581      cycles:u                                                           

       0.049648957 seconds time elapsed

       0.000000000 seconds user
       0.049907000 seconds sys
4235
Reposted by Luke Lau
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/12/2024
Here are some slightly tricky RVV mask patterns.
'RVV mask tricks'

# broadcast nth bit
vmand.mm v8, in, mNth
vcpop.m t0, v8
sub t0, x0, t0
vmv.v.x v8, t0

# prefix xor
viota.m v8, in
vand.vi v8, v8, 1
vmsne.vi v8, v8, 0
vmor.mm v0, v8, in # can often be omitted

# move nth bit to first
vmand.mm v8, in, mNth
vcpop.m t0, v8
vmv.v.x v8, t0
vmsof.m v0, v8

# move mask to GPR
vmv.x.s t0, v0
# move GPR to mask
vmv.s.x v0, t0
# assuming vl<=64, set SEW=64 before

# these two should really be dedicated instructions
# shift mask up by 1
vslide1up.vx v8, in, x0
vsrl.vi v8, v8, 7
vmadd.vx v0, 2, v8

# shift mask up by 1
vslide1down.vx  v8, in, x0
vadd.vv v0, in, in
vmacc.vx v0, 128, v8
173