Sign in

Ibraheem Ahmed

@ibraheem.ca
195 followers 49 following 41 posts

Software developer interested in building fast, concurrent, and robust systems. @ibraheemdev on twitter.

PostsRepliesMedia
Ibraheem Ahmed @ibraheem.ca · 19/06/2025
It's been really exciting collaborating with the rust-analyzer folks on salsa (which we use internally in ty)! A lot of the improvements end up making their way into both projects, I think salsa has a bright future ahead.
0141
Ibraheem Ahmed @ibraheem.ca · 19/06/2025
I think teaching barriers first would reinforce the "nothing can be reordered past this barrier" idea, instead of the barrier being a way to *upgrade* a subsequent store to release the memory preceding the barrier. Reordering kind of misses the entire point of how atomic synchronization works.
100
Ibraheem Ahmed @ibraheem.ca · 19/06/2025
I really dislike the idea of teaching atomic orderings in terms of potential compiler/hardware reordering optimizations. That's not how the spec defines them, and that mental model can get problematic quickly. Atomic orderings prescribe visibility relationships that are actually pretty intuitive.
100
Reposted by Ibraheem Ahmed
Charlie Marsh @crmarsh.com · 13/05/2025
Today, we’re announcing the preview release of ty, an extremely fast type checker and language server for Python, written in Rust. In early testing, it's 10x, 50x, even 100x faster than existing type checkers. (We've seen >600x speed-ups over Mypy in some real-world projects.)
1433284
Ibraheem Ahmed @ibraheem.ca · 14/02/2025
I was recently made aware that every single additional *cycle* in the hash function used by rustc increases its runtime by a whopping 0.25%
040
Ibraheem Ahmed @ibraheem.ca · 18/01/2025
Very high quality code here too
the patch just removes random closing brackets throughout the codebase
040
Ibraheem Ahmed @ibraheem.ca · 18/01/2025
No I don't want AI generated PRs on my repository thank you very much
an AI generated patch for a bug on github
3132
Ibraheem Ahmed @ibraheem.ca · 08/01/2025
Unfortunately all of my short blog post ideas evolve into books
060
Ibraheem Ahmed @ibraheem.ca · 08/01/2025
Going to try and publish more than one blog post this year...
230
Ibraheem Ahmed @ibraheem.ca · 05/01/2025
000
Ibraheem Ahmed @ibraheem.ca · 25/12/2024
I wonder if the bitmask approach of github.com/cloudflare/t... has any potential here?
110
Ibraheem Ahmed @ibraheem.ca · 02/12/2024
There's a 10us state machine based solution on a local Discord leaderboard 😄.
100
Ibraheem Ahmed @ibraheem.ca · 02/12/2024
If you go in windows of three you can try cheating once by skipping the middle item, and every iteration after that you compare the last two items in the tuple instead of the first two. Though if you do it this way you need to handle skipping the first and second items as a separate case.
110
Ibraheem Ahmed @ibraheem.ca · 24/11/2024
Nice to see people experimenting with alternative scheduling algorithms! nurmohammed840.github.io/posts/announ...
nurmohammed840.github.io
Announcing Nio
040
Ibraheem Ahmed @ibraheem.ca · 21/11/2024
I see, does your implementation essentially translate '.' to '/'?
110
Ibraheem Ahmed @ibraheem.ca · 21/11/2024
Libraries like github.com/ibraheemdev/... are probably the best middle ground because it allows people who don't want async to still benefit from the ecosystem. It will be slower but unless someone is willing to write a professional grade synchronous HTTP 1/2/3 implementation, it's your best option.
github.com
GitHub - ibraheemdev/astra: Rust web servers without async/await.
Rust web servers without async/await. Contribute to ibraheemdev/astra development by creating an account on GitHub.
010
Ibraheem Ahmed @ibraheem.ca · 21/11/2024
Async might be a net negative for a lot of people, but unfortunately those people are also less likely to be investing time into maintaining quality infrastructure for synchronous code, so it's kind of of an impossible problem.
120
Ibraheem Ahmed @ibraheem.ca · 21/11/2024
I think all the talk about async being hard somewhat misses the point that high-throughput low-latency systems *are* hard, and the people building those systems are also the people developing and maintaining the async ecosystem.
130
Ibraheem Ahmed @ibraheem.ca · 21/11/2024
Would this "just work" if matchit supported parameters with static prefixes? e.g. `/{foo}.bar` I have a work in progress PR that implements this github.com/ibraheemdev/... but it needs some more work for conflict handling
110
Ibraheem Ahmed @ibraheem.ca · 21/11/2024
Do I spy matchit's new routing syntax?
110
Ibraheem Ahmed @ibraheem.ca · 17/11/2024
I wrote a bit about the design of papaya here: ibraheem.ca/posts/design.... This release will mostly be microoptimizations that contribute to a ~15% performance improvement.
110
Ibraheem Ahmed @ibraheem.ca · 16/11/2024
I think there's room for everyone to be more open about async not being the end all and be all of performance.
000
Ibraheem Ahmed @ibraheem.ca · 16/11/2024
There is a little bit of nuance around async being more expensive at lower concurrency and async synchronization primitives being more complex. async-sync interop is not free, and so there is a cost on users forced to use an async library in a synchronous context.
100
Ibraheem Ahmed @ibraheem.ca · 16/11/2024
You missed the "Windows CI is failing though"
010
Ibraheem Ahmed @ibraheem.ca · 16/11/2024
Maybe! Though papaya's higher baseline memory usage would be an important consideration.
010
Ibraheem Ahmed @ibraheem.ca · 16/11/2024
Hoping to release soon. github.com/ibraheemdev/...
github.com
GitHub - ibraheemdev/papaya: A fast and ergonomic concurrent hash-table for read-heavy workloads.
A fast and ergonomic concurrent hash-table for read-heavy workloads. - ibraheemdev/papaya
030
Ibraheem Ahmed @ibraheem.ca · 16/11/2024
The next version of papaya reaches over 2x the read throughput of dashmap!
382
Ibraheem Ahmed @ibraheem.ca · 16/11/2024
The other part of this is you can't really ignore the other 0.2us measurement, because that implies a huge savings on userspace synchronization (channels, mutexes, etc.)
000
Ibraheem Ahmed @ibraheem.ca · 16/11/2024
Exactly. epoll_wait is a user->kernel->user roundtrip just like a blocking recvfrom. If every IO op corresponds to a epoll_wait you're probably worse than blocking IO due to extra regostration syscalls. If 300 IO ops correspond to a given epoll_wait you're down to 1/300th of the expensive roundtrips
100
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
That's why saying that the cost of context switches caused by IO-readiness are equal between async and threads, as the benchmark repository did, is somewhat misleading.
100
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
The point about IO readiness is that as concurrency grows, a single call to epoll_wait starts returning more events, meaning the context switch is amortized, whereas blocking IO pays this for every operation.
100
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
I was saying that this can show up much earlier and more quantifiably than the actual context switching cost. Cooperative multitasking solves this inherently, and Tokio and other userspace schedulers are tuned to ensure fairness between processing I/O and executing tasks to avoid similar issues.
000
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
The latency issue is a somewhat unrelated (though more expensive context switching may still be a factor) answer to the thread-scaling limit. The problem is that threads can get unlucky and preempted *before* they are able to make their next I/O request, leading to uncontrollable tail latency.
100
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
Even then you will probably start seeing tail latency issues that you cannot control. There's also a lot of OS tooling that's going to start breaking down at that number of threads.
100
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
Tens or even a hundred thousand threads is definitely feasible, but there is a limit where it becomes impractical, and it's probably not memory usage.
100
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
Yeah, there are ways to get around the limits, even on a 16GB machine you should theoretically get to millions of threads. The level of scale I'm talking about is an AWS/Cloudflare load-balancer or HFT server where latency is critical.
100
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
By kernel memory running out you mean scheduling millions of threads?
100
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
On the flip side this means that async is typically more expensive at lower concurrency. I generally agree with the sentiment that the "perf advantage" is hard to quantify (which is why in ibraheem.ca/posts/too-ma... I introduced async as a more *powerful* programming model, rather than performant).
010
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
For most programs this is irrelevant, but it does matter at the extremes. Also note that async makes communication across tasks *significantly* cheaper (you save a syscall and kernel switch), so this can manifest in different ways. You also get more control over your scheduler, for example.
110
Ibraheem Ahmed @ibraheem.ca · 15/11/2024
This is a little misleading. At extreme scale the memory usage isn't really a big deal, it's actually the context switch time and scheduler latency that matters. The "I/O readiness switch" mentioned in the repo is amortized at high levels of concurrency.
210
Ibraheem Ahmed @ibraheem.ca · 13/11/2024
A little annoyed that everyone's username is slightly different
020
Ibraheem Ahmed @ibraheem.ca · 12/11/2024
I guess it's time to start cross posting here.
250