Sign in

AJ Stuyvenberg

@ajs.bsky.social
1.9K followers 238 following 325 posts

AWS Hero Staff Eng @ Datadog Streaming at: twitch.tv/aj_stuyvenberg Videos at: youtube.com/@astuyve I write about serverless minutia at aaronstuyvenberg.com/

PostsRepliesMedia
AJ Stuyvenberg @ajs.bsky.social · 31/07/2026
We have real numbers! The full blog is here: www.datadoghq.com/blog/enginee... But we went from roughly 450ms to 80ms
140
AJ Stuyvenberg @ajs.bsky.social · 12/02/2026
The clanker wars have begun:
040
AJ Stuyvenberg @ajs.bsky.social · 18/09/2025
Yeah of course
020
AJ Stuyvenberg @ajs.bsky.social · 18/09/2025
This includes full Datadog tracing/logs/metrics by the way. There's no compromising on observability, not even cold starts.
110
AJ Stuyvenberg @ajs.bsky.social · 18/09/2025
I use AWS a ton but Lambda still astounds me. Throw some code in a function, send 1m requests as fast as you can. It ate up all available file descriptors on my little t3 box and still ran 18k RPS with a p99 of 0.3479s. Not many services can go from 0 to 18k RPS instantaneously with this p99.
1113
AJ Stuyvenberg @ajs.bsky.social · 05/08/2025
The faster your cold starts are, the cheaper these will be!
010
AJ Stuyvenberg @ajs.bsky.social · 05/08/2025
Lambda now charges for init time, so it's useful to count sandboxes which are proactively initialized but never receive a request. Here's what happens after a 10k request burst. Hundreds of sandbox shutdowns, along with 22 sandboxes which were spun up but never received a request.
150
AJ Stuyvenberg @ajs.bsky.social · 01/08/2025
Happy Lambda Init Billing day to those who celebrate. Fix your cold starts!
150
AJ Stuyvenberg @ajs.bsky.social · 01/08/2025
Yeah!
010
AJ Stuyvenberg @ajs.bsky.social · 01/08/2025
aws.amazon.com/about-aws/wh...
aws.amazon.com
AWS Lambda response streaming now supports 200 MB response payloads - AWS
Discover more about what's new at AWS with AWS Lambda response streaming now supports 200 MB response payloads
000
AJ Stuyvenberg @ajs.bsky.social · 01/08/2025
NEW: Lambda can now send up to 200mb payloads using response streaming! I assume this is mostly directed at LLM inference workloads, where chatbots can stream large amounts of data over the wire as it becomes available.
220
AJ Stuyvenberg @ajs.bsky.social · 30/07/2025
www.datadoghq.com/blog/monitor...
datadoghq.com
Monitor Lambda-hosted web apps with the Lambda Web Adapter integration | Datadog
Learn how Datadog makes it easy to monitor legacy web apps running in AWS Lambda by automatically capturing logs, metrics, and traces through the Lambda Web Adapter.
000
AJ Stuyvenberg @ajs.bsky.social · 30/07/2025
That said, I'm excited to share that @Datadog's Serverless monitoring product now supports LWA! Thanks to Harold and AWS Labs for collaborating with us on the PRs, and huge thanks to Alex Gallotta for driving this work.
100
AJ Stuyvenberg @ajs.bsky.social · 30/07/2025
I've long been an advocate for the Lambda Web Adapter project which lets anyone pretty easily ship an app to Lambda without learning about the event model/API. Honestly AWS should simply support this natively.
121
AJ Stuyvenberg @ajs.bsky.social · 28/07/2025
I'm a big fan of continuous profiling/measuring your software against real world use cases. This is also how I often learn about in to new system changes in AWS early, heh. Great episode of Software Huddle w/ @alexbdebrie: www.youtube.com/watch?v=JAw9...
youtube.com
Operational Excellence Is the Moat with Sam Lambert
Today, Sam Lambert from Planetscale is back for a third time. Planetscale just announced Planetscale Postgres, so we had to get Sam back to tell us how and why they decided to add support for Postgres. It's always great to have Sam on -- he brings great stories about real customers and honest insight about the state of the database industry. In this episode, we talk about the road to Postgres and how operational excellence is the only true advantage in database providers. Sam walks us through the current Planetscale Postgres offering, along with details on Nova, a new sharded Postgres project that Planetscale is working on. Along the way, we get updates on Planetscale Metal, how demand has been for Planetscale Postgres, and future plans for Planetscale. *Timestamps* 01:16 Start 06:37 The Timeline 15:15 Not Much IP in the Database Market 21:48 PSBouncer 24:17 Zonal affinity 27:38 Query Insights 29:34 How to sign up 32:02 Convex 34:37 Other data stores? 56:18 Acquisitions
010
AJ Stuyvenberg @ajs.bsky.social · 28/07/2025
"We run benchmarks continually across all of our competitors, not just queries - even connections, ensuring we don't add any latency at all." @isamlambert Performance is such a competitive advantage which easily slips away if you're not constantly paying attention to it.
130
Reposted by AJ Stuyvenberg
ScyllaDB @scylladb.com · 23/07/2025
Datadog rewrote its AWS Lambda Extension from #Golang to #Rustlang with no prior Rust experience. @ajs.bsky.social will share how they achieved an 80% Lambda cold start improvement along with a 50% memory footprint reduction at our free and virtual #P99CONF. www.p99conf.io?latest_sfdc_... #ScyllaDB
063
AJ Stuyvenberg @ajs.bsky.social · 22/07/2025
In our case the secret is a Datadog API key which isn't required until we actually flush data, so deferring it to that point saves us over 50ms.
021
AJ Stuyvenberg @ajs.bsky.social · 22/07/2025
Here's another 33% cold start reduction, which comes from deferring expensive decryption calls made to AWS Secrets Manager until the secret is actually needed. Lazy loading is great!
131
AJ Stuyvenberg @ajs.bsky.social · 18/07/2025
50% lol was not reading the profile carefully
000
AJ Stuyvenberg @ajs.bsky.social · 18/07/2025
By switching to a memory arena, we preallocate a slab of memory and virtually eliminate the linear growth of malloc syscalls, which cuts down kernel mode switches, improving latency. Thank you profiling (and jemalloc)!
000
AJ Stuyvenberg @ajs.bsky.social · 18/07/2025
Here's how to visualize a 100% memory allocation improvement! A recent stress test revealed that malloc calls bottlenecked when sending > 100k spans through the API and aggregator pipelines in Lambda.
210
AJ Stuyvenberg @ajs.bsky.social · 17/07/2025
Now writing a job to a log and then using a subscription filter to run them async is deeply fucking cursed though omg
110
AJ Stuyvenberg @ajs.bsky.social · 17/07/2025
I think OP's intentions were pretty pure until they felt they were mistreated by AWS. So many people end up taking to social media in those instances so in my opinion it was mostly fine.
100
AJ Stuyvenberg @ajs.bsky.social · 17/07/2025
Thanks Corey!
000
AJ Stuyvenberg @ajs.bsky.social · 17/07/2025
cc @quinnypig.com, as I saw this post in your newsletter
110
AJ Stuyvenberg @ajs.bsky.social · 17/07/2025
aaronstuyvenberg.com/posts/does-l...
aaronstuyvenberg.com
Does AWS Lambda have a silent crash in the runtime?
Understanding what’s happening in the “AWS Lambda Silent Crash” blog post, what went wrong, and how to fix it
110
AJ Stuyvenberg @ajs.bsky.social · 17/07/2025
NEW: A recent blog post went viral in the AWS ecosystem, about how there's a silent crash in AWS Lambda's NodeJS runtime. Today I'll step you through the actual Lambda runtime code which causes this confusing issue, and walk you through how to safely perform async work in Lambda:
370
AJ Stuyvenberg @ajs.bsky.social · 14/07/2025
I assume this helps with capacity, especially as so many functions are triggered on schedules for the top of the hour or on a routine call schedule. I've long suspected that I could get faster cold starts/placements by scheduling a function at the 58th minute instead of the top of the hour.
010
AJ Stuyvenberg @ajs.bsky.social · 14/07/2025
Now after only ~4 invocation/shutdown cycles, Lambda shuts down the sandbox 1 minute after my request:
110
AJ Stuyvenberg @ajs.bsky.social · 14/07/2025
Lambda's fleet management shutdown algorithm is learning faster! I'm calling this function every 8 or so minutes. At first the gap from invocation to shutdown is about 5-6 minutes, which was the fastest I've observed during previous experiments.
130
AJ Stuyvenberg @ajs.bsky.social · 11/07/2025
NEW: AWS is rolling out a new free tier beginning July 15th!! New accounts get $100 in credits to start and can earn $100 exploring AWS resources. You can now explore AWS without worrying about incurring a huge bill, this is great! docs.aws.amazon.com/awsaccountbi...
062
AJ Stuyvenberg @ajs.bsky.social · 10/07/2025
You should care about your p99! By improving the function cold start time, the service on the left performed: 2x faster in RPS and thus, duration. p99 from 1.52s -> .949s The code and functionality is identical, but improving the cold start from 816ms to 301ms made all the difference.
010
AJ Stuyvenberg @ajs.bsky.social · 08/07/2025
Quick PSA to make sure you're using a DLQ and setting a max receive count for SQS, otherwise you may find yourself looking at a flamegraph like this. Hundreds of attempts, multiple messages in queue and not burning down and average age of message ticking up! Seems common knowledge, but...
030
AJ Stuyvenberg @ajs.bsky.social · 03/07/2025
of client devices, or used over webrtc. Check it out!
000
AJ Stuyvenberg @ajs.bsky.social · 03/07/2025
Real-time audio processing is an extremely interesting corner of performance engineering. You've gotta process data faster than it streams in, otherwise the work is wasted. The engineering team behind CoScreen just open sourced a new library, dtln-rs, suitable for removing noise from all kinds
datadoghq.com
How we built a real-time, client-side noise suppression library without server dependencies | Datadog
Learn how we implemented and open-sourced a noise filter for real-time audio chat without compromising performance. Better yet, try the demo and add it to your own project today.
131
AJ Stuyvenberg @ajs.bsky.social · 29/05/2025
Monitoring Lambda sandbox shutdowns reveals an interesting scaling behavior After a request spike, Lambda waits ~10m before reaping 2/3rds of sandboxes. 5m later it begins reaping the rest. Presumably this helps smooth latency during retry storms, or if traffic returns!
020
AJ Stuyvenberg @ajs.bsky.social · 21/05/2025
Which means you have to keep track of all inflight flush requests over the course of multiple invocations. If the sandbox detects that the request rate is slowing, you've gotta drive those requests synchronously and block /next until they're done (or risk losing data!)
000
AJ Stuyvenberg @ajs.bsky.social · 21/05/2025
So the strategy looks something like this:
100
AJ Stuyvenberg @ajs.bsky.social · 21/05/2025
The hard part is figuring out what to do when function invocation rates slow down and you need to adapt back to flushing data and blocking /next until it completes, because otherwise you'll drop data. Lambda functions can't coordinate so you have to figure this out another way
100
AJ Stuyvenberg @ajs.bsky.social · 21/05/2025
How? In short, I cheated. If the invocations are frequent enough (more frequent than the TCP Keep Alive/server timeout), you can drive one flush request asynchronously across multiple invocations and it costs you $0 in billed time. That's not the hard part though.
110
AJ Stuyvenberg @ajs.bsky.social · 21/05/2025
This is the result of a new flushing algorithm for our Next-Generation Lambda Extension. As you can see this strategy delivers a nearly ZERO overhead flushing solution to busy Lambda functions. Focusing on those p99 events can deliver huge wins, so keep an eye on what makes your service slow!
110
AJ Stuyvenberg @ajs.bsky.social · 21/05/2025
Here's another huge p99 cliff! One unreasonably effective way to lower the average latency of a service is to minimize the causes of p99 events. Here, we've managed to absolutely crush the Max Post Runtime Duration from ~80ms to 500µs!
110
AJ Stuyvenberg @ajs.bsky.social · 13/05/2025
We've actually already rewrote several high throughput (and high-alloc) services in Rust. My colleague Artem wrote about one service in our metrics index pipeline here: www.datadoghq.com/blog/enginee...
datadoghq.com
Timeseries indexing at scale | Datadog
Learn how we implemented a new timeseries indexing strategy when the amount of data we ingested increased significantly.
030
AJ Stuyvenberg @ajs.bsky.social · 13/05/2025
Nope, we didn't bother. The memory safety benefits and lack of a garbage collector made the choice pretty clear without needing a direct comparison.
120
AJ Stuyvenberg @ajs.bsky.social · 06/05/2025
Truly incredible to learn that Epic Games isn't self hosted on a VPS.
020
AJ Stuyvenberg @ajs.bsky.social · 01/05/2025
You also can ship logs to S3 or kinesis firehose instead of cloudwatch for $.25/GB (with discounts with volume): aws.amazon.com/blogs/comput...
aws.amazon.com
AWS Lambda introduces tiered pricing for Amazon CloudWatch logs and additional logging destinations | Amazon Web Services
Effective logging is an important part of an observability strategy when building serverless applications using AWS Lambda. Lambda automatically captures and sends logs to Amazon CloudWatch Logs. This...
000
AJ Stuyvenberg @ajs.bsky.social · 01/05/2025
I believe this is only for logs coming from Lambda. But yeah!
000
AJ Stuyvenberg @ajs.bsky.social · 01/05/2025
PRICE CUT: Lambda slashes the price of cloudwatch logs for high volume. From $.50/gb down to $0.05/gb after 50TB. You've gotta be spending a decent chunk of $$$ on cloudwatch logs for this to help, but still – a price cut is a price cut!
130
AJ Stuyvenberg @ajs.bsky.social · 01/05/2025
yeah in this case I had to compile the program with debug symbols enabled (unstripped)
010