Sign in

Shivansh Vij

@shivanshvij.com
1.7K followers 193 following 72 posts

CEO @loopholelabs.io

PostsRepliesMedia
Shivansh Vij @shivanshvij.com · 13/06/2026
We have pci passthrough!
110
Shivansh Vij @shivanshvij.com · 13/06/2026
I should say, the coolest security perk of this approach unlike Edera is that the GPU can be on a completely different host! Or in its own isolated vm!
120
Shivansh Vij @shivanshvij.com · 13/06/2026
Oh, yes! We have a full CUDA API virtualization layer. You don’t even need to install the CUDA drivers / runtime in the VMs to use it!
120
Shivansh Vij @shivanshvij.com · 13/06/2026
What do you mean? Like what does the “start a vm” contract look like?
100
Shivansh Vij @shivanshvij.com · 12/06/2026
For now, yes - lots of our magic tech is deeply embedded in it (like 100ms constant time live migrations!)
110
Shivansh Vij @shivanshvij.com · 12/06/2026
FYI: both Firecracker and Cloud Hypervisor are also warm restores. Our cold restores (snapshot not already on the host) still beat FC’s warm restores. Because we can, and frankly because we should. Follow our live per-commit benchmarks here: benchmarks.substrate.so/?view=snapshot
benchmarks.substrate.so
Substrate — Benchmarks
Continuous performance benchmarks for the Substrate hypervisor, tracked on every commit to main.
100
Shivansh Vij @shivanshvij.com · 12/06/2026
I may be a bit late to the “time to interactive” agent sandbox party, but I wanted to make sure our hypervisor absolutely MOGGED everyone else. 10x faster than Firecracker and 425x faster than Cloud Hypervisor. agx.so is the best agent infra platform, period.
Benchmarking the Substrate Hypervisor vs Firecracker and Cloud Hypervisor. 

For warm restores, Substrate is 10x faster than Firecracker and 425x faster than Cloud Hypervisor
120
Shivansh Vij @shivanshvij.com · 22/05/2026
Once the sandbox providers figure out performance and scale, the next hills to climb will be: - agent/harness integration (ie. Shelley from @exe.dev) - agent to agent communication (ie. side channels between your sandboxes) - agent tracing integrations (think raindrop AI)
041
Shivansh Vij @shivanshvij.com · 22/04/2026
Yep, that's the name of the game for sure - and it's why CRIU needs to be overhauled. It just doesn't support enough stuff, it's painfully bloated, and it's increasingly difficult to debug or add features. We were even able to replicate a pretty good chunk of its feature set in under a week!
010
Shivansh Vij @shivanshvij.com · 22/04/2026
I think that workload portability will be the next big AI sandbox unblocker. With the ridiculous wave of sandbox compute being deployed right now, having the ability to move long-running workloads between nodes (spot instances!) or even clouds is *the* silver bullet for both availability and cost.
000
Shivansh Vij @shivanshvij.com · 22/04/2026
In zig
010
Shivansh Vij @shivanshvij.com · 22/04/2026
I think it’s time to rewrite CRIU
210
Shivansh Vij @shivanshvij.com · 22/04/2026
Nah, good old C with a kernel module
010
Shivansh Vij @shivanshvij.com · 22/04/2026
Nah you’re right, that’s Claude being an idiot. Can’t use it to write the code but it is good for parsing benchmarking logs (sometimes)
000
Shivansh Vij @shivanshvij.com · 21/04/2026
And if you have a NIC with TLS offload support, even the decryption happens fully in the NIC!
130
Shivansh Vij @shivanshvij.com · 21/04/2026
The end result is awesome, data flows from the NIC directly to the page that the kernel returns to the reader, no separate copy or userspace context switch
120
Shivansh Vij @shivanshvij.com · 21/04/2026
Better than NFS actually! The SIGv4 signing happens in userspace using AWS’s official libraries, and TLS is handled by kTLS; not a custom implementation!
120
Shivansh Vij @shivanshvij.com · 21/04/2026
In reference to a new in-kernel filesystem I've been working on that embeds an S3 client directly in the kernel (alongside a WAL, a caching mechanism, etc.) (Both modes are S3 backed, just with different fsync behaviour)
220
Shivansh Vij @shivanshvij.com · 21/04/2026
At Google Cloud Next this week if folks want to chat about virtualization, filesystems, live migration, and eBPF!
020
Shivansh Vij @shivanshvij.com · 21/04/2026
In-kernel filesystem are going to destroy the FUSE ecosystem. Bringing your custom logic into the VFS directly is absolutely incredible for performance and latency.
281
Shivansh Vij @shivanshvij.com · 30/01/2026
Perfect, let’s definitely hang out then! I’ll dm you!
010
Shivansh Vij @shivanshvij.com · 30/01/2026
Who’s going to FOSDEM in Brussels next week? DM me if you'd like to hang out, talk about high performance cloud infrastructure, live migration, eBPF, Spot Instances - or all of the above! Looking forward to meeting all the cool OSS folks!
110
Shivansh Vij @shivanshvij.com · 29/11/2025
Who’s going to re:invent in Las Vegas next week? DM me if you'd like to hang out, talk about high performance cloud infrastructure, live migration, eBPF, Spot Instances - or all of the above! Can’t wait to see everyone! #aws #reinvent #reinvent2025
091
Reposted by Shivansh Vij
Taras Glek @taras.glek.net · 16/11/2025
Nice that one can use network-bypass techniques without crazy vendor frameworks or kernel mods. Even cooler that one can recreate iptables functionality like this for faster inter-host networking
064
Shivansh Vij @shivanshvij.com · 15/11/2025
loopholelabs.io/blog/xdp-for...
loopholelabs.io
Using XDP for Egress Traffic
XDP only works for ingress. We found a loophole that lets it work for egress. Here's how we did the impossible.
040
Shivansh Vij @shivanshvij.com · 15/11/2025
People aren’t ready for the next blog post from @loopholelabs.io If you thought the XDP post was interesting (linked below) this next one is going to blow your mind. Here’s a hint: What do chess engines and NAT’s have in common?
172
Reposted by Shivansh Vij
Luiz Aoqui @luiz.aoqui.dev · 04/11/2025
Great post by @alex.sorlie.io
0104
Shivansh Vij @shivanshvij.com · 04/11/2025
We finally decided to write out how we use XDP in our network plane for live network migrations and more specifically to process outgoing packets! This is the first in a series of promised deep dives into how Loophole's live migration tech works!
0114
Reposted by Shivansh Vij
Damon Cortesi @dacort.velocipig.com · 07/10/2025
🙋‍♂️ kccncna2025.sched.com/event/27Fd6/...
kccncna2025.sched.com
KubeCon + CloudNativeCon North America 2025: Spark on Kubernetes, a Practical Guide -...
View more about this event at KubeCon + CloudNativeCon North America 2025
022
Shivansh Vij @shivanshvij.com · 07/10/2025
Who is coming to KubeCon NA 2025 in Atlanta?
132
Shivansh Vij @shivanshvij.com · 29/04/2025
Marked your account for priority!
020
Shivansh Vij @shivanshvij.com · 29/04/2025
We’re not just putting this together for CI, we’ll also be connecting it to a CLI so you can just throw your local environment into a cloud VM to run really quickly and be torn down!
140
Shivansh Vij @shivanshvij.com · 13/04/2025
RWX looks amazing! I wonder if there’s an opportunity for collaboration there…
020
Shivansh Vij @shivanshvij.com · 04/03/2025
I can’t wait to open source our filesystem this month, it’s ridiculous the kind of stuff we jammed into the Linux kernel
030
Reposted by Shivansh Vij
Chris @chris.blue · 03/03/2025
I'm getting more and more excited about drop-in, frictionless infrastructure. Lots going on in observability, containers, cloud compute, etc. eBPF is part of it, so is monkey patching. Some interesting ones: @loopholelabs.io, @polarsignals.com, junctionlabs.io ﹩, subtrace.dev ﹩
3252
Shivansh Vij @shivanshvij.com · 18/02/2025
- How to run VMs anywhere (even without KVM support) - How to do L3/4 Checksums in eBPF XDP hooks - How to write an in-kernel filesystem (and why) - Deduplicating memory between VMs - Reverse Engineering the CUDA API
030
Shivansh Vij @shivanshvij.com · 18/02/2025
Request: Topics for the @loopholelabs.io Blog I want to start posting some short-form blogs that go through the technical challenges of building @loopholelabs.io and how we’ve solved them. If folks have requests for content they’d like covered now’s the time to ask! Topics I have so far 👇
142
Reposted by Shivansh Vij
Trezy @trezy.codes · 19/01/2025
Just got to Malta for the company getaway! It’s really convenient that these EU outlets have a whole THREE USB-C outlets! So great.
3241
Shivansh Vij @shivanshvij.com · 14/01/2025
boinc.berkeley.edu You can volunteer your unused compute capacity towards research projects run by Berkeley For things like discovering pulsars, studying climate change, etc.
boinc.berkeley.edu
BOINCCompute for Science
BOINC is an open-source software platform for computing using volunteered resources
000
Shivansh Vij @shivanshvij.com · 10/01/2025
Donate to BOINC!
100
Shivansh Vij @shivanshvij.com · 01/01/2025
Yessir - the price to performance ratio alone makes them exciting but once you factor in power draw it becomes a no-brainer
010
Shivansh Vij @shivanshvij.com · 01/01/2025
If Asahi Linux ends up supporting M4 chips this year, I think adoption in datacenters will go through the roof
240
Reposted by Shivansh Vij
Shivansh Vij @shivanshvij.com · 27/12/2024
I think the most excited thing about this blog post is the one we get to publish next - a deep dive into how Architect works under the hood. There were so many “impossible” problems we had to solve to make Architect a reality - and now that we have I’m having a blast explaining how we did it.
393
Shivansh Vij @shivanshvij.com · 27/12/2024
Here’s a demo of us migrating between spot instances across cloud providers all of which have different CPUs (generations, clock speeds, etc.) loophole.sh/kc2024
110
Shivansh Vij @shivanshvij.com · 27/12/2024
Yep - and in those situations we can swap to a different instance type silently. We prefer that over a different region. One of the more “impossible” problems we had to solve was migrating between different cpu types/generations seamlessly - that’s going to be the highlight of the next post.
110
Shivansh Vij @shivanshvij.com · 27/12/2024
So the nice thing about architect is that we can always fall back to on-demand instances if there’s no spot available - or we can switch to a different region, or if we’re really in a pinch we can swap to a completely different cloud provider. Users can configure what level of fallback they want.
010
Shivansh Vij @shivanshvij.com · 27/12/2024
For sure - and we do recommend some level of overprovisioning to help against this, but because we run in your cloud account we generally only have to worry about the size of the individual customer’s workload on a per-region basis
100
Reposted by Shivansh Vij
James Bayer @jambay.bsky.social · 27/12/2024
@evanphx.dev this looks pretty cool loopholelabs.io/blog/rethink... built on OSS live migration tech that forked firecracker and related virtualization technology github.com/loopholelabs...
loopholelabs.io
Rethinking Spot Instances - How We Solved the Preemption Problem
Run any application on Spot Instances with zero downtime and 75%+ cost savings - no code changes required. Architect eliminates the reliability concerns that previously made Spot Instances impractical...
173
Shivansh Vij @shivanshvij.com · 27/12/2024
Your mental model is spot on (no pun intended)- though to make it work across spot instances and even cloud providers we had to go far beyond the capabilities of vmotion
120
Reposted by Shivansh Vij
James Bayer @jambay.bsky.social · 27/12/2024
This looks super cool. Get up to 75% cost-savings of using spot instances without data loss or interruption. My mental model of what this does is the equivalent of vSphere vMotion, but for spot instance workloads running in the Cloud.
141