Sign in

Jaz

@jaz.sh
64K followers 377 following 6.8K posts

Jasmine (Jaz) Gender Nomad IRC made me gay Reject drab, wear bright colors and play fast Former Platform Engineering Lead at Bsky 29. she/her 🏳️‍⚧️ BSky Stats- bsky.jazco.dev/stats github.com/jazware

PostsRepliesMedia
Jaz @jaz.sh · 7h
Upgrade should be smooth but you may lose a registered TOTP from a one-time migration, you should be able to sort it out in the operator dash if something goes wrong. Rolling back isn't safe at this point if you do go with the upgrade, there were some backward-incompatible user metadata changes.
000
Jaz @jaz.sh · 8h
What commit are you on now in your deployed version? i would potentially be open to issues but things are definitely still in flux as you can see in the repo. i would update to latest first and see if i’ve already addressed them. I can test an upgrade from your version if you want just in case
200
Jaz @jaz.sh · 06/10/2026
Oh yeah vlPDS now has full support for the atproto spaces alpha btw with tooling for users to manage their space data and operators to see the infrastructure-relevant info about spaces sync for their vlPDS instance. vlpds.jazco.dev/docs/spaces
vlpds.jazco.dev
Spaces · vlpds docs
Permissioned data on vlpds: every member keeps a private, signed space repo on their own PDS. It never reaches the firehose, and reading it takes a short-lived credential chain.
2676
Reposted by Jaz
Jaz @jaz.sh · 05/10/2026
I truly believe that Object Storage is the only real datastore and everything else should be cache and compute on top of it. Databases are fake, Object Storage is the last database. All your systems should be built on top of it and use NVME and RAM as caches and local compute.
1283
Reposted by Jaz
Jaz @jaz.sh · 05/10/2026
If you need low latency, simply support shard replicas and accept writes as durable when a quorum of replicas have written the write into memory and let them slowly flush to object storage in efficient batches that amortize operation costs. That's durable enough for pretty much anything.
1111
Reposted by Jaz
Jaz @jaz.sh · 05/10/2026
Of course for ^ to be durable you want your replicas to be on different physical nodes and maybe different racks but yeah. Horizontally scalable DBs use RAM as a cache and NVME as durable storage with writes durable at R=2 nodes reporting in-memory writes. Now we have RAM -> NVME -> Object Storage.
2111
Reposted by Jaz
Jaz @jaz.sh · 05/10/2026
This is nice because it allows single node operation and also allows recovery and seamless handoff by pulling from object storage instead of having to do heavy node-to-node reads to catch up or onboard a new node. Your durability is also a lot higher and you can run your local NVME hot.
091
Jaz @jaz.sh · 05/10/2026
This is nice because it allows single node operation and also allows recovery and seamless handoff by pulling from object storage instead of having to do heavy node-to-node reads to catch up or onboard a new node. Your durability is also a lot higher and you can run your local NVME hot.
091
Jaz @jaz.sh · 05/10/2026
Of course for ^ to be durable you want your replicas to be on different physical nodes and maybe different racks but yeah. Horizontally scalable DBs use RAM as a cache and NVME as durable storage with writes durable at R=2 nodes reporting in-memory writes. Now we have RAM -> NVME -> Object Storage.
2111
Jaz @jaz.sh · 05/10/2026
If you need low latency, simply support shard replicas and accept writes as durable when a quorum of replicas have written the write into memory and let them slowly flush to object storage in efficient batches that amortize operation costs. That's durable enough for pretty much anything.
1111
Jaz @jaz.sh · 05/10/2026
I truly believe that Object Storage is the only real datastore and everything else should be cache and compute on top of it. Databases are fake, Object Storage is the last database. All your systems should be built on top of it and use NVME and RAM as caches and local compute.
1283
Jaz @jaz.sh · 05/10/2026
Yes it's on the list!
040
Jaz @jaz.sh · 05/10/2026
💜
030
Jaz @jaz.sh · 05/10/2026
it is horizontally scalable, but each node increases the baseline load on object store operations so cost-wise it generally scales better vertically.
130
Reposted by Jaz
Erica Fischer @enf.bsky.social · 04/10/2026
There are numerous elevators
441309232
Jaz @jaz.sh · 04/10/2026
many are asking this
050
Jaz @jaz.sh · 04/10/2026
Another fun fact, a single node can handle 300k+ proxied requests per second to various AppViews with realistic payload/response sizes and latencies!
1210
Jaz @jaz.sh · 04/10/2026
Geo routing writes for a user should be trivial to do with GeoDNS or Anycast, though I suppose you'd want some mechanism for shards to rebalance themselves to nodes where their repos are being most heavily accessed and that doesn't exist yet.
060
Jaz @jaz.sh · 04/10/2026
On a crash recovery, the new owning node claims the shard, replays the relevant data from the crashed node's WAL, and catches up the SlateDB state before accepting writes to repos assigned to that shard again, all within 12 seconds.
2350
Jaz @jaz.sh · 04/10/2026
vlPDS uses a per-node in-bucket WAL complementing the per-shard SlateDB to better amortize Class-A operations against your bucket without sacrificing durability. Writes get a 200 only once they're durable in the WAL in object storage and flush from SlateDB's memtables within 10 seconds.
1360
Jaz @jaz.sh · 04/10/2026
The primary focus here is on reliability, scalability, and ease of operation. Things should be very set-it-and-forget and ideally in the worst case scenario of losing an entire node you'd have at most a 12 second blip of some of your shards write availability. It also supports manual shard splitting
1530
Jaz @jaz.sh · 04/10/2026
I probably wouldn't want to do that with SlateDB, storage is an LSM built on top of object storage, not an object per record or anything like that.
110
Jaz @jaz.sh · 04/10/2026
It is MIT licensed and FOSS :) github.com/jazware/vlpds
github.com
GitHub - jazware/vlpds: Very Large PDS, a world-scale, incredibly efficient, incredibly resilient, atproto personal data server
Very Large PDS, a world-scale, incredibly efficient, incredibly resilient, atproto personal data server - jazware/vlpds
080
Jaz @jaz.sh · 04/10/2026
I think it's already feasible to have geo distributed nodes if you use a global bucket. Writes are prefixed by node ID or by repo so a well balancing global bucket object store should be able to make writes local enough.
150
Jaz @jaz.sh · 04/10/2026
Right now it supports an entirely in-browser migration flow that lets you bring an account over to the new PDS. I don't currently have a mechanism for in-place replacing the reference implementation and all its data with vlPDS.
110
Jaz @jaz.sh · 04/10/2026
Oh yeah and that setup can support 10x the scale of today's production load comfortable (including the 3x traffic spike overhead)!!
4690
Jaz @jaz.sh · 04/10/2026
At Bluesky production scale, hosting the entire network with room for 3x peak commit spikes during the day, vlPDS can be hosted on 3 x 6vCPU machines with 32GB RAM and 400GB local NVME. Adding in the Object Storage costs and it comes to roughly $1.5k/mo (plus blob storage costs).
2730
Jaz @jaz.sh · 04/10/2026
I've tested vlPDS on an instance as small as a 2vCPU 4GB RAM, 40GB NVME OVH VPS at which scale it could host 10,000 of the most active accounts on the network comfortably with overhead. Operating costs at that scale are $4.50/mo for the instance and ~$75 on R2 for ops plus storage.
3780
Jaz @jaz.sh · 04/10/2026
Hi protocol friends! I've been cooking up a Very Large PDS (vlpds) implementation capable of scaling up to billions of accounts and down to just one. It uses Object Storage as its only state, supports HA operation, live scaling up and down, is feature-complete with the reference PDS, and is FOSS.
vlpds.jazco.dev
Overview · vlpds docs
An atproto PDS whose only durable storage is an object store. Any node serves any request, one node can run a personal server for free, and a few nodes can carry all of Bluesky's write load.
2349271
Jaz @jaz.sh · 04/10/2026
🫂💜
010
Jaz @jaz.sh · 04/10/2026
shard handoff
grafana panel of a shard handoff between two nodes
040
Jaz @jaz.sh · 04/10/2026
yep :) zero downtime live migration between nodes they share the same DNS name and serve the same data, backed by object storage
130
Jaz @jaz.sh · 04/10/2026
(im cooking on a cool piece of atproto infra coming soon)
1310
Jaz @jaz.sh · 04/10/2026
they don't know my repo just live migrated between vlPDS nodes, that was pretty cool
3520
Jaz @jaz.sh · 29/09/2026
i've been there a couple of times, it's fine it just isn't always open during stated business hours cause i guess not many people go there outside of lunch time M-F
110
Jaz @jaz.sh · 29/09/2026
incredible, i can hear it from here
011
Jaz @jaz.sh · 28/09/2026
conceptually related conversations basically, yeah conversations that are nearby in embedded vectorspace but not close enough to be clustered together get tied with lines like that
010
Jaz @jaz.sh · 28/09/2026
i'm glad you like it! it was about time the ol' atlas got a second pass with a lot more experience to draw from :)
030
Jaz @jaz.sh · 28/09/2026
you can enable adult topics by scrolling to the bottom of the list, they just get summarized as "adult" until you tick the box
010
Reposted by Jaz
Jaz @jaz.sh · 27/09/2026
Okay so I've built a new Atlas, this time it lets you explore the conversations on Bluesky over the past 7 days graphically. It should update every ~6 hours and help you navigate the information environment that exists on here. I've learned a LOT about what people do here today... atlas.jazco.dev
atlas.jazco.dev
Atlas
A living map of what Bluesky is talking about: the past week's conversations, grouped into topics and regions and rebuilt every six hours.
1161425423
Jaz @jaz.sh · 28/09/2026
there will likely be one big shift coming soon cause im swapping embedding models to support more than just english language posts but after that the regions should remain stable across updates
010
Jaz @jaz.sh · 28/09/2026
Try here atlas.jazco.dev?sel=topic-15...
atlas.jazco.dev
Atlas
A living map of what Bluesky is talking about: the past week's conversations, grouped into topics and regions and rebuilt every six hours.
030
Jaz @jaz.sh · 28/09/2026
the counts are for the full range of data for now (7 days)
010
Jaz @jaz.sh · 28/09/2026
until a conversation falls out of the 7d window ideally
020
Jaz @jaz.sh · 28/09/2026
thank you! i am no longer active bluesky staff but i still have a lot of love for the platform and the protocol 💜
050
Jaz @jaz.sh · 28/09/2026
Thank you Liz!! I'm glad you like it!!!! 💜
1100
Jaz @jaz.sh · 28/09/2026
Projects like this are fun because I always end up with more questions than I had when I started, such as: - What is "Feral Flannel Friday"? - What is "Drag'n Wash" and why is there backlash about it? (the summary helped answer this one) - Who is making Hazbin Hotel Omegaverse art? - Furry pilot?
4425
Jaz @jaz.sh · 28/09/2026
Yeah I've got a WIP explainer for the methodology sitting in my drafts right now. This was about a day and a half of effort (a nice sized weekend project) and drew heavily on the fact that i've already got a bsky archive in clickhouse for exporting and i have some background in ml pipelines
2120
Jaz @jaz.sh · 28/09/2026
yeah! you can get an idea of how it's working based on the web requests in your network tab
000
Jaz @jaz.sh · 28/09/2026
it's webGL accelerated
050