Sign in

The Scale Factory

@scalefactory.com
215 followers 860 following 352 posts

Award-winning consultancy, empowering SaaS teams to deliver more on AWS. We're an AWS Advanced Consulting Partner, and a Kubernetes Certified Service Provider. Part of Ten10.

PostsRepliesMedia
The Scale Factory @scalefactory.com · 05/12/2025
And now, behind the scenes is where Werner is going (albeit 15 minutes late - it wouldn't be a Vogels keynote without running long!). And that's the last AWS re:Invent for this year. Thanks for following along with me.
020
The Scale Factory @scalefactory.com · 05/12/2025
Have pride in your work: most people will never see the majority of your effort, which goes on behind the scenes.
110
The Scale Factory @scalefactory.com · 05/12/2025
The fifth quality of the Renaissance developer: they must be a polymath - someone whose knowledge spans many subjects.
100
The Scale Factory @scalefactory.com · 05/12/2025
Werner encourages us to keep doing person to person code reviews. AI will change many things, but the craft of software is still learned person to person.
010
The Scale Factory @scalefactory.com · 05/12/2025
(On that topic, if you're using AI, you still need to own the outcomes. The work is yours, not the tool's)
200
The Scale Factory @scalefactory.com · 05/12/2025
The fourth quality of the Renaissance Developer: they are an owner. They own the quality of their software.
100
The Scale Factory @scalefactory.com · 05/12/2025
(Taking a detour via Kiro to highlight the importance of communication when talking to the AI about how it should build your software)
100
The Scale Factory @scalefactory.com · 05/12/2025
The third quality of a Renaissance Developer is that they communicate clearly.
100
The Scale Factory @scalefactory.com · 05/12/2025
The second quality The Renaissance Developer has is that they think in systems. Not computer systems per se, but other systems of interconnected things.
100
The Scale Factory @scalefactory.com · 05/12/2025
Werner shares some of the learnings he's come to, travelling the world, meeting customers, and seeing people change the world using their skills.
100
The Scale Factory @scalefactory.com · 04/12/2025
Werner's attempting to define "The Renaissance Developer". His first suggestion, a developer must by curious. I agree entirely, curiosity leads to learning.
100
The Scale Factory @scalefactory.com · 04/12/2025
Bezos believes we're at the convergence point of multiple golden ages, in AI, robotics, space travel, and development in each area accelerates the others. Werner reckons the other era this reminds him of is the renaissance. Ok, let's see where he's going with this.
100
The Scale Factory @scalefactory.com · 04/12/2025
We've had assembly, compilers, structured programming, object oriented programming, monolithic architectures, service-oriented architectures. Things have always changed, the new shift in focus to AI is no different.
100
The Scale Factory @scalefactory.com · 04/12/2025
But where AI might make some *jobs* obsolete, it won't make *you* obsolete, if you evolve.
100
The Scale Factory @scalefactory.com · 04/12/2025
The other elephant in the room: Will AI take my job? Werner's answer? Maybe. Some tasks will be automated, some skills made obsolete.
100
The Scale Factory @scalefactory.com · 04/12/2025
Werner starts by addressing the elephant in the room: Werner has given re:Invent keynotes since 2012, but this is his last one. He's not leaving Amazon, but he believes after so many years, the keynote audience deserves a fresh new set of AWS voices.
100
The Scale Factory @scalefactory.com · 04/12/2025
Lots of Amazon frugality on show in this particular video!
100
The Scale Factory @scalefactory.com · 04/12/2025
As is traditional, we're starting with some nerd wish fulfillment video content. In previous years, this has seen Werner starring in something like The Matrix, or Fear & Loathing. This time... Back To The Future.
100
The Scale Factory @scalefactory.com · 04/12/2025
Ok, @topper.me.uk here for the final AWS re:Invent keynote of the year. This one is Werner Vogels, CTO of Amazon.com, and unusually it's just an hour long - which is probably for the best, because it's pretty late here in the UK.
101
The Scale Factory @scalefactory.com · 04/12/2025
The closing keynote from Werner Vogels is just an hour long this year, and kicks off at 11.30pm London time. I'll see you back here for that.
000
The Scale Factory @scalefactory.com · 04/12/2025
And after a demo of "live visual intelligence", applying AI models to live content, we're done. This keynote always pushes my geek buttons, there are some very smart people doing some very cool things inside AWS.
110
The Scale Factory @scalefactory.com · 04/12/2025
In Q1, Trainium will be becoming PyTorch native, so that you don't have to learn a new platform in order to make use of Trainium hardware. Just target 'neuron' instead of 'cuda' in that code.
110
The Scale Factory @scalefactory.com · 04/12/2025
Compared with Trainium 2, the Trainium 3 chip can process 5 times the number of tokens per megawatt. This performance was recorded using the aforementioned profiling tools.
100
The Scale Factory @scalefactory.com · 04/12/2025
Trainium has on-chip capabilty for profiling workloads, without impacting performance. Neuron Explorer is a tool for reporting on those performance stats.
100
The Scale Factory @scalefactory.com · 04/12/2025
Now, the toolchain: Generally Available in Q1, the Neuron Kernel Interface (NKI). This provides direct access to AWS Trainium NeuronCore instruction set architecture, supporting Trainium 3.
000
The Scale Factory @scalefactory.com · 04/12/2025
Trainium chips have a number of micro-optimisations, built with real world AI workloads specifically in mind - not just benchmarks.
200
The Scale Factory @scalefactory.com · 04/12/2025
The Elastic Fabric Adapters in on these sleds provides high throughput networking across the whole UltraServer cluster.
100
The Scale Factory @scalefactory.com · 04/12/2025
The server sleds in these UltraServers use all three AWS hardware platforms: Nitro, Graviton, and Trainium. Everything on this sled is "top servicable", which means they can be assembled robotically, and that when they need maintenance, this can be performed quickly and efficiently.
100
The Scale Factory @scalefactory.com · 04/12/2025
Trainium chips provide efficient hardware for training models. A Trainium UltraServer provides 144 Trainium3 Chips across two racks, with dedicated "neuron switches" providing networking optimsed for model training workloads.
100
The Scale Factory @scalefactory.com · 04/12/2025
Customer TwelveLabs on stage now, talking about how they use AWS' s3 vector capabilites to build models that understand video in the way that humans do, providing plain text search for that video content.
100
The Scale Factory @scalefactory.com · 04/12/2025
The problem with S3 Vector indices is that these are stored on disc, which makes lookups slower than in-memory stores. At write time, S3 calculates a dataset of nearest neighbours. When you perform vector lookups, this nearest neighbour data is pulled into memory and queried there - quickly!
100
The Scale Factory @scalefactory.com · 04/12/2025
OpenSearch has vector search capabilites now. It's also now easy to add vector indices to S3 buckets, with the recently GA service, S3 Vectors, the first cloud object storage with native support to store and query vector data.
110
The Scale Factory @scalefactory.com · 04/12/2025
DeSantis is back on stage, giving an overview of how vectors and vector databases work in generative AI workloads. AWS have integrated vector capabilities across their range of services.
100
The Scale Factory @scalefactory.com · 04/12/2025
Finetuning capabilities are also logged on the journal. If inference requests spike, finetuning jobs can be paused, and resumed later when load drops off.
100
The Scale Factory @scalefactory.com · 04/12/2025
Requests in Mantle use a request journal, so that failed requests can be tried again, making Bedrock more reliable.
100
The Scale Factory @scalefactory.com · 04/12/2025
As well as latency tiering, Mantle provides fair scheduling, preventing other tenants under heavy load from impacting the performance of your requests.
100
The Scale Factory @scalefactory.com · 04/12/2025
We're onto AI Elasticity now. AWS built Bedrock using a traditional load balancer / webservice design. Those learnings led to a better understanding of how to build inference platforms. This led to Project Mantle, which underpins many of Bedrock's models today. This is how they offer latency tiers.
100
The Scale Factory @scalefactory.com · 04/12/2025
Lambda Managed Instances allow you to launch Lambdas on EC2 nodes provisioned to the spec of your choosing. You still don't manage the servers yourself. "Serverless is the absence of server management", Brown says, answering that question once and for all.
100
The Scale Factory @scalefactory.com · 04/12/2025
Brown is back, talking through the history of AWS Lambda, initially conceived of in 2013. A couple of years ago, the Lambda and EC2 teams were brought together. It's that coming together that led to the release this week of AWS Lambda Managed Instances.
100
The Scale Factory @scalefactory.com · 04/12/2025
A speaker from Apple tells us they've started using their language Swift to develop backend systems, and deploy these to Graviton. Amazon Linux now has a native toolchain for developing Swift, the first distribution to provide one.
100
The Scale Factory @scalefactory.com · 04/12/2025
Graviton 4 also provided a coherent link between cores, to enable 192 vCPU cores, but this introduced new latency paths. Graviton 5 improves that, doubling the number of cores, and increasing the L3 cache by 5.3x. These are now available in preview, as part of M9g EC2 instances.
100
The Scale Factory @scalefactory.com · 04/12/2025
Each Graviton generation is improved based on data about how the previous one performed on real workloads. L2 cache misses in Graviton 3 were having a real impact on performance, so Graviton 4 doubled the size of the L2 cache per core, added more cores, and increased the size of the L3 cache too.
100
The Scale Factory @scalefactory.com · 04/12/2025
As well as custom silicon, Graviton chips are cooled in a custom way too - the lid of the chip is removed, and the heat sink applied directly. This improves cooling efficiency and reduces power usage.
100
The Scale Factory @scalefactory.com · 04/12/2025
Dave Brown, VP, AWS Compute & ML Services takes the stage to talk about Graviton, Amazon's custom ARM-based chip which delivers excellent price/performance compute on AWS.
100
The Scale Factory @scalefactory.com · 04/12/2025
The latest edition of compsci classic "Computer Architecture: A Quantitative Approach" covers the architecture of Nitro and the AWS Graviton hardware. Neat.
100
The Scale Factory @scalefactory.com · 04/12/2025
DeSantis tells the story of how original EC2 performance was subject to jitter, as a result of the virtualisation mechanism. To solve this, they built AWS Nitro, the hardware virtualisation platform that underpins the whole AWS cloud.
100
The Scale Factory @scalefactory.com · 04/12/2025
Billions of dollars are being invested in AI right now. AWS are working on reducing the cost of running AI workloads.
000
The Scale Factory @scalefactory.com · 04/12/2025
Compute demand for AI is accelerating. DeSantis tells us that AWS wants to make it as easy to build and scale AI workloads as it is to use S3.
100
The Scale Factory @scalefactory.com · 04/12/2025
We start by asking "What does AI mean for the cloud?" so I guess we're getting another damn AI keynote? Come on guys.
100
The Scale Factory @scalefactory.com · 04/12/2025
This isn't always super easy to distill into "skeet" form, but as usual I'll have a go here.
100