Sign in

Glenn K. Lockwood

@glennklockwood.com
1.3K followers 222 following 1.4K posts

I am a supercomputing enthusiast, but I usually don't know what I'm talking about. I post about large-scale infrastructure for #HPC and #AI.

PostsRepliesMedia
Glenn K. Lockwood @glennklockwood.com · 23/09/2026
Spent a few months preparing the tech docs for DataEnclave, VAST's new confidential computing runtime. W/ help from VAST R&D, I learned how attestation etc worked well enough to describe and defend it. Short version: www.vastdata.com/blog/introdu... 15-pager: www.vastdata.com/resources/wh... #AI
vastdata.com
Introducing VAST DataEnclave Confidential AI
VAST DataEnclave runs AI workloads in hardware-isolated environments, so proprietary models and sensitive data can meet on untrusted infrastructure.
031
Glenn K. Lockwood @glennklockwood.com · 22/09/2026
By total coincidence, two of my colleagues are running a webinar on this topic on Wednesday. The lead speaker, Anat, is really sharp and worth hearing speak: www.vastdata.com/events/demys...
vastdata.com
Demystifying KV Cache Expert Advice on architectures, best practices, and sizing
Learn what KV cache is, why it matters for inference and agentic workflows, and the tools and benchmarks you need to build a resilient KV cache architecture.
021
Glenn K. Lockwood @glennklockwood.com · 21/09/2026
I wrote this doc to explain KV caches and KV cache offloading many months ago. I found salespeople couldn't confidently tell when their customers might have a problem, because all the content online is straight marketing or way too in the weeds. blog.glennklockwood.com/2026/09/what...
blog.glennklockwood.com
What are KV caches really?
KV caching lies sufficiently deep within the guts of how transformers work that really understanding how it can be used goes beyond any si...
161
Glenn K. Lockwood @glennklockwood.com · 17/09/2026
Malaysia is allegedly evaluating Huawei's Ascend 910C for their sovereign AI buildout. If Malaysia picks Huawei over NVIDIA, it'd be the first big buy of these chips outside of China. Egypt also name-dropped as a nation being pursued by Huawei. www.freemalaysiatoday.com/category/hig...
freemalaysiatoday.com
Malaysia weighs Huawei chips for sovereign AI push
The country is seriously considering using Huawei’s AI hardware as the backbone of a RM2 billion initiative aimed at giving Malaysia greater control over national data.
000
Glenn K. Lockwood @glennklockwood.com · 10/09/2026
The IO500 committee just stripped Sugon of its #1 position in the Production list because it appears that their ParaStor F9000 isn't a real product, or it could not be verified as one. Verify here: io500.org/list/isc26/p... My notes on ParaStor: glennklockwood.com/garden/paras... #HPC #notreally
io500.org
IO500 - ISC26 - Production List
062
Glenn K. Lockwood @glennklockwood.com · 03/09/2026
I recently had the opportunity talk to @nicolehemsoth.bsky.social about a few failures in #HPC and #AI infrastructure that taught me some important lessons about designing systems and managing risks. Have a look, have a listen. Hope others find it interesting! www.vastdata.com/blog/failure...
vastdata.com
Failure is a Finding: The Hard Lessons Behind Computing's Biggest Bets
VAST delivers the first AI Operating System, unifying storage, database, and compute to drive agentic computing and data intensive workloads. Learn more.
063
Glenn K. Lockwood @glennklockwood.com · 01/09/2026
MLPerf Storage results are out and StorageReview pointed out the interesting disconnect between MLPerf leaderboard winners and companies with the largest #storage footprints in the #AI clouds. Suggests this benchmark may not reflect real needs. www.storagereview.com/news/mlperf-...
storagereview.com
MLPerf Storage v3.0: 877 GiB/s Checkpoints, a Cloud First, and a Leaderboard Turned Over
MLPerf Storage v3.0 results: 877 GiB/s checkpoints from Everpure FlashBlade//EXA, new KV cache and vector DB tests, 143 submissions.
140
Glenn K. Lockwood @glennklockwood.com · 23/08/2026
Just reflects my lack of familiarity with Flux. Most of what I encounter in the wild is Slurm and Airflow, both of which suffer the same scalability limitations.
000
Glenn K. Lockwood @glennklockwood.com · 22/08/2026
I recently wrote up an argument about why batch schedulers like Slurm are architecturally insufficient for autonomous scientific laboratories and agentic workflows. Event-driven orchestration is the scalable alternative. www.vastdata.com/blog/autonom... #HPC
vastdata.com
Autonomous Science is Inheriting a Workflow Infrastructure Problem
VAST delivers the first AI Operating System, unifying storage, database, and compute to drive agentic computing and data intensive workloads. Learn more.
170
Glenn K. Lockwood @glennklockwood.com · 13/08/2026
Funny, there’s a reason NERSC took its datacenter out of this location. It’s not a great place to put one. www.datacenterdynamics.com/en/news/form...
datacenterdynamics.com
Former LBNL supercomputing facility in Oakland, California, targeted for 20MW data center
Behring looks to turn vacant site back into live data center
132
Glenn K. Lockwood @glennklockwood.com · 13/08/2026
I have a theory that every panel about “how AI will affect XYZ” will devolve into nonspecific discussion around “here is where the AI tools I use today fail” and “how do we trust AI?” regardless of what XYZ is.
020
Glenn K. Lockwood @glennklockwood.com · 13/08/2026
This is the diagram I’m most proud of. Stacking ~700 different LLM training jobs observed from the VAST fleet on a common timescale to figure out whether there is a “feeding the GPU” problem anywhere close to the write intensity of periodic checkpoints. There is not. Only read burst is at job start.
022
Glenn K. Lockwood @glennklockwood.com · 13/08/2026
I am not ashamed of my disdain
010
Glenn K. Lockwood @glennklockwood.com · 12/08/2026
The lack of shame some people have in copypasting prose straight out of Claude into Powerpoint is shocking. Were these people just recycling their students’ slides before, and now just tap a different source of material?
141
Glenn K. Lockwood @glennklockwood.com · 09/08/2026
Problem is, it's not my day job to do this kind of analysis. So here I am, frantically banging on signal analysis and statistical models on a Sunday, to have something compelling to say 😩 As with all longitudinal workload analysis, people are messy and hard to understand or predict.
120
Glenn K. Lockwood @glennklockwood.com · 09/08/2026
I am attending the ModSim conference for the first time next week, and I'm presenting analysis and design modeling of I/O performance requirements for frontier model training based on a year's worth of production data across the VAST hyperscale fleet (>900 training jobs, >140K GPU-years of compute)
221
Glenn K. Lockwood @glennklockwood.com · 07/08/2026
I recently upgraded to Claude Max and started using Opus 5 for exploratory analysis. If this is what the AI community things approaches "PhD-level capability," I don't know what to say. It's sloppy, inconsistent, and walks back conclusions at the lightest of questioning.
110
Glenn K. Lockwood @glennklockwood.com · 04/08/2026
Surprised to see a solution so heavily focused on tiering to HDD for capacity now tier to…more HDD for capacity. Curious what Wasabi does that PanFS doesn’t (and vice versa) to make this pairing complementary.
010
Glenn K. Lockwood @glennklockwood.com · 01/08/2026
yes you’re right; it reduces compute. it does also open another way to parallelize so you can effectively aggregate more GPUs’ HBM, but it doesn’t reduce the need for HBM capacity as I erroneously wrote.
000
Reposted by Glenn K. Lockwood
v @vsoch.bsky.social · 30/07/2026
I gave the invited talk for the Systems track at #PEARC26. If you were not able to attend, you're in luck! 🍀 I walked back to my room and recorded a second time so all would have access in the present and into the future. youtu.be/qiKfTgYoKkE?...
youtu.be
🌀 The Agentic HPC Center: Orchestrating the Future of Science
YouTube video by vsoch
152
Glenn K. Lockwood @glennklockwood.com · 29/07/2026
I’m glad someone liked it! Hard needle to thread in crafting a technical plenary that is interesting, accessible, and thematic to “resilient roots.” Also relieved that my SDSC colleagues didn’t mind me spilling the tea on one of Gordon’s failures. #HPC #alittleAI
151
Glenn K. Lockwood @glennklockwood.com · 29/07/2026
Nobody came 😩 Just kidding. Was mic check. Here’s the actual crowd.
030
Glenn K. Lockwood @glennklockwood.com · 29/07/2026
231
Glenn K. Lockwood @glennklockwood.com · 28/07/2026
It’s my first PEARC. I was at XSEDE’13 but one of the PEARC chairs told me that doesn’t count. No swag from me; I’m actually attending to give the Wednesday plenary (pearc.acm.org/pearc26/plen...) which will have nothing to do with VAST.
pearc.acm.org
Plenary Speakers - PEARC26 - Resilient Roots + Empowered Communities
Meet our Plenary Speakers Tuesday, July 28 – 9:00 am – 10:15 am Opening Ceremony Keynote: Amanda RandlesDuke University Amanda Randles is the Director of the Center for Computational and Digital Healt...
000
Glenn K. Lockwood @glennklockwood.com · 28/07/2026
I am ready! Almost. #PEARC26
382
Reposted by Glenn K. Lockwood
parasharmanish.bsky.social @parasharmanish.bsky.social · 22/07/2026
Another great addition to Utah’s amazing pro-human AI innovation ecosystem built on public-private partnerships.
011
Glenn K. Lockwood @glennklockwood.com · 21/07/2026
Surprising to see a software company go into the hardware business, selling an unusual 2U70 dual-controller flash server. It's hard to source flash these days, so there are strong market incentives to ship software on whatever hardware your customers can find.
000
Glenn K. Lockwood @glennklockwood.com · 15/07/2026
We are now in the age of AI-generated conferences. Today’s keynote by Gropp prominently features distinctly Claude-generated plots; yesterday had an entire presentation delivered by an AI voiceover of slides. And slides are full of Claude design tells.
060
Glenn K. Lockwood @glennklockwood.com · 15/07/2026
GPUs only, no storage. So no problem.
100
Glenn K. Lockwood @glennklockwood.com · 14/07/2026
First time I’ve seen someone with a PhD actually predict that data centers in space will be a real thing. Granted, it’s a CS PhD, not a physical sciences one. #HPC #notreallythough
2223
Glenn K. Lockwood @glennklockwood.com · 10/07/2026
Infinia hardware should be cheaper since it doesn’t require HA pairs. Software is software and can be however much or little as your sales rep likes you.
010
Glenn K. Lockwood @glennklockwood.com · 09/07/2026
With Infinia gaining support for files, is there a reason for anyone to buy EXAScaler anymore? Why would someone choose Lustre over Infinia if Infinia seems to support everything that Lustre does but with enterprisey glitter? Curious to see how DDN positions these two lines. #HPC
210
Glenn K. Lockwood @glennklockwood.com · 09/07/2026
Block off your calendar for Wednesday at 10am! I'll be there.
041
Reposted by Glenn K. Lockwood
OGAWA, Tadashi @ogawa-tadashi.bsky.social · 09/07/2026
=> Architecting for Frontier AI Labs: How Anthropic scales Claude on TPUs, Google & Anthropic, Google Cloud Next 2026, Apr 22 youtube.com/watch?v=LEPY... content-cdn.sessionboard.com/content/3e8V... Automated fault recovery & intelli scheduling to maintain continuous uptime across massive pod-scale
111
Glenn K. Lockwood @glennklockwood.com · 08/07/2026
Keen eye! Fixed that typo
110
Glenn K. Lockwood @glennklockwood.com · 08/07/2026
Yes I am giving the technical plenary at PEARC. More details should be publicized soon!
010
Glenn K. Lockwood @glennklockwood.com · 08/07/2026
I wrote up some notes after my week at #ISC26: blog.glennklockwood.com/2026/07/isc2... Interested to hear others' perspectives on highlights and things I may have missed. #HPC
blog.glennklockwood.com
ISC'26 recap
Last month was the 2026 ISC High Performance Conference in the beautiful (and sweltering) Hamburg, Germany. It was my fifth time att...
3102
Glenn K. Lockwood @glennklockwood.com · 02/07/2026
I somehow missed that China has developed its own implementation of InfiniBand, called ScaleFabric 400, to compete/supplant NVIDIA/Mellanox IB. Oddly enough, its manufacturer (Sugon) is not a member of IBTA though. en.eeworld.com.cn/mp/EEWorld/a... #HPC
en.eeworld.com.cn
The first domestically produced InfiniBand has been launched; real-world test data tells you just how powerful its performance is.-Electronics Headlines-EEWORLD
Recently, another technological high ground long monopolized by foreign countrie......view more!
031
Glenn K. Lockwood @glennklockwood.com · 26/06/2026
Interesting that C-DAC, whose indigenous Indian interconnect Trinetra is based around libfabric, opted to implement verbs support to make Lustre work instead of adopting Cray’s kfabric LND. #ISC26
010
Glenn K. Lockwood @glennklockwood.com · 26/06/2026
1. General LLMs beat coding-specific models at writing Fortran, likely bc code models are overfitted to Python 2. This conf begins the age of Claude-generated pptx. Clear tells 3. My phone doesn’t like these #ISC26 projectors
160
Reposted by Glenn K. Lockwood
Robert Henschel @roberthenschel.bsky.social · 26/06/2026
Now happening at #ISC26 #HPC
011
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
He has a chance to mention the Ozaki scheme but instead cites a 20-year-old algorithm that is the problematic core of HPL-MxP.
120
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
Don’t trust any of them. If the source is a tweet from a guy who has a history of having a loose relationship with the truth, then it shouldn’t be in a conference keynote.
020
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
This is the worst Jensen keynote ever
020
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
I wonder if ground has been broken on this datacenter yet. Because it takes two years to build such a facility without the federal land red tape part. How commercially meaningful will Blackwells be in two years?
130
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
Funny that “software adapt to whatever hardware shows up” is exactly how everyone’s favorite AI models are developed. Model architecture iterates much faster than chip design. This only works if your problems stand still.
200
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
How many wrong statements can you fit into a single sentence? Holy cow.
220
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
I can’t wait to hear what Jack Dongarra thinks about the future of #HPC at #ISC26’s closing keynote.
120
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
Dumb stuff that results in weird requirements in #HPC RFPs. Ultra dense GPU racks with all liquid requires unusual datacenter construction considerations. #ISC26
362
Glenn K. Lockwood @glennklockwood.com · 25/06/2026
Ogawa-san is how I first came to know about China’s new exascale system a while ago, and I assembled a quick spec page a few weeks back (glennklockwood.com/garden/syste...). Notably has a BIG liquid cooled storage system (67 racks, 10 TB/s).
glennklockwood.com
LineShine (NSCC-SZ)
LineShine is an all-CPU exascale supercomputer at the National Supercomputing Center in Shenzhen (NSCC-SZ), built entirely from domestically produced Chinese hardware with no reliance on foreign chips...
051