Sign in

Jonathon Belotti

@jonathonbelotti.bsky.social
64 followers 132 following 28 posts

Peeling back the layers @modal_labs. Previously data & ML platform @canva. dms open.

PostsRepliesMedia
Jonathon Belotti @jonathonbelotti.bsky.social · 19/03/2026
Miserable on Friday afternoon when everyone has left from GTC.
020
Jonathon Belotti @jonathonbelotti.bsky.social · 02/03/2026
Are there counter examples of companies that scaled their system this fast and had better reliability?
010
Jonathon Belotti @jonathonbelotti.bsky.social · 02/02/2026
10 months later, we're hiring for a reliability focused engineer :) jobs.ashbyhq.com/modal/84467a... If you know someone great that'd want to be a founding reliability eng for Modal's GPU-focused cloud platform, lmk!
jobs.ashbyhq.com
Member of Technical Staff - Reliability Engineering
Member of Technical Staff - Reliability Engineering
030
Jonathon Belotti @jonathonbelotti.bsky.social · 08/01/2026
Any book recommendations for stuff about aviation safety improvements? I’m currently reading Engineering a Safer World and I’m keen to read more about aviation and other engineering industry’s successes.
010
Jonathon Belotti @jonathonbelotti.bsky.social · 08/01/2026
Damn ok that’s surprising.
100
Jonathon Belotti @jonathonbelotti.bsky.social · 08/01/2026
How come @bcantrill.bsky.social didn’t like Drift into Failure by Dekker?
120
Jonathon Belotti @jonathonbelotti.bsky.social · 23/05/2025
Modal was originally built with CPU batch in mind, but the GenAI wave certainly swept us up :) We love working with computational bio/chemistry customers on their mostly CPU-intensive batch workloads. modal.com/use-cases/co... If you want a demo let me know :)
030
Jonathon Belotti @jonathonbelotti.bsky.social · 23/05/2025
When I was at Canva I had to get a batch job to run demographics analysis vision models on over 90 million images. It took a couple weeks, but felt like with the right infra it could be done in a couple hours. The infra now exists :)
modal.com
Introducing Modal Batch: Process 1 Million Jobs with 1 Line of Code
Modal Batch is a new interface backed by a new durable queue system built specifically to make job processing easy, scalable, and fault-tolerant.
020
Reposted by Jonathon Belotti
Steve Klabnik @steveklabnik.com · 20/05/2025
Hypervisor as a Library #rustlang seiya.me/blog/hypervi...
seiya.me
Hypervisor as a Library
1316
Jonathon Belotti @jonathonbelotti.bsky.social · 07/05/2025
We won't tell you where the deep and cheap GPU capacity is, but we will teach you how to fish.
modal.com
Linear Programming for Fun and Profit
How we use an eighty-year-old algorithm to find arbitrages in the cloud market.
020
Jonathon Belotti @jonathonbelotti.bsky.social · 05/04/2025
Is The Soul of a New Machine the best book about the computer industry because Kidder is.. - an outsider - a full-time writer - simply a first rate writer Other industries have had ‘homegrown’ Pulitzers—medicine, law, finance, bilogy—but not computing.
000
Jonathon Belotti @jonathonbelotti.bsky.social · 04/04/2025
LLMs have shifted my blog post drafting back to my static-site repository away from Notion. I can play with HTML and JS so easily during iteration with Cursor. Markdown+Blocks is hamstringing.
010
Jonathon Belotti @jonathonbelotti.bsky.social · 31/03/2025
Is there a SRE book but for startups? Google’s book is great but it doesn’t contend with the tradeoffs and constraints of a young growing startup.
151
Jonathon Belotti @jonathonbelotti.bsky.social · 24/03/2025
History-posting once again: thundergolfer.com/blog/the-fir...
thundergolfer.com
The First LLM
A tracing of the history of GPT-1 and its predecessors.
011
Jonathon Belotti @jonathonbelotti.bsky.social · 25/02/2025
My brilliant @modal-labs.bsky.social colleague Charles won't stop until we're all pushing our GPUs to their limits: modal.com/blog/gpu-uti...
modal.com
'I paid for the whole GPU, I am going to use the whole GPU': A high-level guide to GPU utilization
A guide to maximizing the utilization of GPUs, from cloud allocations to FLOP/s.
000
Jonathon Belotti @jonathonbelotti.bsky.social · 18/02/2025
Oh was “it’s an OOM larger” referring to the training cluster size?
110
Jonathon Belotti @jonathonbelotti.bsky.social · 18/02/2025
What’s the param count?
100
Jonathon Belotti @jonathonbelotti.bsky.social · 16/02/2025
The Tom Brady of Youtube educational content is at it again. One of those rare lectures which deserves to be called enlightening. I hope @karpathy.bsky.social is steering clear of helicopters, submersibles, and smoking. We need him.
youtube.com
Deep Dive into LLMs like ChatGPT
YouTube video by Andrej Karpathy
000
Jonathon Belotti @jonathonbelotti.bsky.social · 06/02/2025
What didn’t you like about Modal’s exp? We’re actively working on it, being dissatisfied about certain areas.
000
Jonathon Belotti @jonathonbelotti.bsky.social · 02/02/2025
At @modal-labs.bsky.social your customer questions sometimes get a whole blog post :) Why does an NVIDIA H100 SXM 80GB card offer 85.52 GB? thundergolfer.com/blog/nvidia-...
000
Jonathon Belotti @jonathonbelotti.bsky.social · 01/02/2025
Found out you get 220MiB more H100 HBM3e VRAM on Oracle Cloud compared to GCP. GCP's instances have more `reserved` memory and I can't figure out if this is configurable from the guest.
000
Jonathon Belotti @jonathonbelotti.bsky.social · 29/01/2025
To learn about container snapshotting I highly recommend Tristan Hume's telefork repostory. I forked it and started hacking on file descriptor restore: github.com/thundergolfe.... Once PID restore works there's a good chance a simple NVIDIA GPU checkpoint/restore would pass!
github.com
GitHub - thundergolfer/telefork: Like fork() but teleports the forked process to a different computer!
Like fork() but teleports the forked process to a different computer! - thundergolfer/telefork
010
Jonathon Belotti @jonathonbelotti.bsky.social · 29/01/2025
Wrote about how a warmed up container can be saved to disk and later restored for a 2.5x cold start performance boost. Saving live container processes to disk turns out to be pretty whacky and interesting! modal.com/blog/mem-sna...
modal.com
Memory Snapshots: Checkpoint/Restore for Sub-second Startup
Serializing container state to disk for aggressive cold start optimization.
2102
Jonathon Belotti @jonathonbelotti.bsky.social · 22/01/2025
This contains the shortest and clearest explanation of S3’s $0.02/gib/month cloud economics that I’ve seen, and lots more good stuff. tailscale.com/blog/living-...
tailscale.com
Living in the future, by the numbers
Instead of making the traditional New Year predictions, let’s talk instead about the beautiful technological future we live in: the one that exists right now but we don’t always notice.
021
Jonathon Belotti @jonathonbelotti.bsky.social · 20/01/2025
Ha all too true. I wouldn’t even get 100% and I just reviewed the post. Still, we aspire…
010
Jonathon Belotti @jonathonbelotti.bsky.social · 19/01/2025
post url: thundergolfer.com/latency-numb...
thundergolfer.com
Beyond ‘latency numbers every programmer should know’
Took 10 years, but there's finally a better list.
010
Jonathon Belotti @jonathonbelotti.bsky.social · 19/01/2025
Quick post about those 'latency numbers every programmer must know'. It became mostly an affirmation of @sirupsen.bsky.social 's napkin-math repo (bookmark it).
2153
Jonathon Belotti @jonathonbelotti.bsky.social · 07/01/2025
New goal in 2025 is to write one software essay with the zany energy of an All Souls College essay prompt: Does the moral character of an orgy change when the participants wear Nazi uniforms? thundergolfer.com/all-souls-so...
thundergolfer.com
An ‘All Souls Examination’ of the machine
Inspiration towards a better software essay.
000
Jonathon Belotti @jonathonbelotti.bsky.social · 12/12/2024
GPUs can be understood: modal.com/gpu-glossary...
modal.com
README | GPU Glossary
020