Sign in

Sebastian Galkin

@functionth.bsky.social
22 followers 34 following 6 posts
PostsRepliesMedia
Sebastian Galkin @functionth.bsky.social · 10/07/2025
So happy with this milestone. Lots of work went into this one!
010
Reposted by Sebastian Galkin
Earthmover @earthmover.io · 20/05/2025
Our latest fundamentals blog post provides an overview of @zarr.dev and its open-source ecosystem. Read more: earthmover.io/blog/what-is...
earthmover.io
Fundamentals: What Is Zarr? A Cloud-Native Format for Tensor Data - Earthmover
What Zarr is, and how it enables fast, scalable access to multidimensional array data in the cloud.
0113
Reposted by Sebastian Galkin
Earthmover @earthmover.io · 14/05/2025
𝐻𝑜𝑤 𝑑𝑜𝑒𝑠 𝐼𝑐𝑒𝑐ℎ𝑢𝑛𝑘 𝑎𝑣𝑜𝑖𝑑 𝑟𝑒𝑑𝑢𝑛𝑑𝑎𝑛𝑡 𝑠𝑡𝑜𝑟𝑎𝑔𝑒 𝑏𝑒𝑡𝑤𝑒𝑒𝑛 𝑑𝑎𝑡𝑎 𝑣𝑒𝑟𝑠𝑖𝑜𝑛𝑠? Icechunk stores only new or changed chunks for each version —no redundant copies or rewrites. You get instant time travel, branching, and efficient updates, all with negligible storage overhead. More: bit.ly/3F1XFST
earthmover.io
Icechunk: Efficient storage of versioned array data - Earthmover
We recently got an interesting question in Icechunk’s community Slack channel (thank you Iury Simoes-Sousa for motivating this post): I’m new to Icechunk. How is the storage managed for redundant info...
024
Reposted by Sebastian Galkin
Earthmover @earthmover.io · 12/05/2025
Our latest blog post dives into the chaos of the status quo - where every tweak means regenerating the 𝑤ℎ𝑜𝑙𝑒 𝑑𝑎𝑡𝑎𝑠𝑒𝑡 and collaboration and experimentation is often stifled by silos and secret knowledge. Check out the full post: earthmover.io/blog/tensoro...
earthmover.io
TensorOps: Scientific Data Doesn't Have to Hurt - Earthmover
Curious how your team scores on the "Data Pain Survey"? Wondering why your teams are building Rube Goldberg machines just to put some data on a map? Or just want to see our plan to bring order to your...
033
Sebastian Galkin @functionth.bsky.social · 05/05/2025
After months of Rust, I wrote some Python this weekend. I immediately got burned by global mutable state
070
Sebastian Galkin @functionth.bsky.social · 29/04/2025
Last week @deepakcherian.bsky.social gave a fascinating talk at NCAR on data sharing and open-data. The historic perspective, the achievements and failures past and present, how to learn and move forward to fulfill the promises. Remarkable and illuminating www.youtube.com/watch?v=JZT3...
youtube.com
CISL Seminar: Deepak Cherian (Earthmover)
YouTube video by NCAR Computational and Information Systems Laboratory (CISL)
010
Sebastian Galkin @functionth.bsky.social · 23/04/2025
Had the idea of using Icechunk (an multi-dimensional array database) for something I would never use Icechunk for
000
Reposted by Sebastian Galkin
Earthmover @earthmover.io · 17/04/2025
1/ 💡 Our latest blog post in the fundamentals series, written by @tegnicholas.bsky.social, demystifies cloud-optimized scientific data formats! Read more: earthmover.io/blog/fundame...
earthmover.io
Fundamentals: What is Cloud-Optimized Scientific Data?
What cloud-optimized data really means, and how Zarr and Icechunk enable fast access to massive scientific datasets in cloud object storage.
2159
Reposted by Sebastian Galkin
TEGNicholas.bsky.social @tegnicholas.bsky.social · 10/04/2025
You could also do this for arbitrarily large scientific array datasets using Xarray + Icechunk + R2/Tigris juhache.substack.com/p/0-data-dis...
juhache.substack.com
0$ Data Distribution
Ju Data Engineering Weekly - Ep 78
001
Reposted by Sebastian Galkin
Earthmover @earthmover.io · 09/04/2025
📣 Blog post alert! 𝐄𝐱𝐩𝐥𝐨𝐫𝐢𝐧𝐠 𝐈𝐜𝐞𝐜𝐡𝐮𝐧𝐤 𝐬𝐜𝐚𝐥𝐚𝐛𝐢𝐥𝐢𝐭𝐲: 𝐮𝐧𝐭𝐚𝐧𝐠𝐥𝐢𝐧𝐠 𝐒𝟑'𝐬 𝐩𝐫𝐞𝐟𝐢𝐱 𝐬𝐭𝐨𝐫𝐲. This technical post by @functionth.bsky.social dives deep into the internals of how S3 shards data, showing that distributed Icechunk can easily perform 230,000 object reads/sec and beyond. earthmover.io/blog/explori...
earthmover.io
Exploring Icechunk scalability: untangling S3's prefix story | Earthmover
We show Icechunk can scale to extremely high concurrency levels, and explain how it achieves this in modern object stores.
254
Reposted by Sebastian Galkin
Joe Hamman @jhamman.bsky.social · 03/04/2025
We often see folks try to convince tabular data tools to perform well with multi-dimensional array data. This post by @rabernat.bsky.social explains, from first principles, why this rarely works. Its a good one! 👇👇👇
131
Sebastian Galkin @functionth.bsky.social · 28/03/2025
I've worked on Icechunk almost exclusively for the last six months. I'm very proud of the result; you should check it out.
030
Reposted by Sebastian Galkin
Earthmover @earthmover.io · 20/02/2025
1/ Check out our latest blog post earthmover.io/blog/xarray-... to learn about the dramatic improvement and performance of Xarray’s Zarr backend. We achieved improved the “time to first byte” metric, building on Zarr-Python’s new asyncio internals.
earthmover.io
Accelerating Xarray with Zarr-Python 3 | Earthmover
We have recently dramatically improved the performance of Xarray’s Zarr backend. This post explores how we’ve improved the “time to first byte” metric, building on Zarr-Python’s new asyncio internals.
143