Sign in

Blosc Development Team

@blosc.org
30 followers 4 following 77 posts

Announcements about Blosc2 developments blosc.org

PostsRepliesMedia
Blosc Development Team @blosc.org · 8h
📢 Python-Blosc2 4.14 is out! A common interface to stream, query & slice remote Parquet, #PyTables, #HDF5, #Zarr & #Blosc2 data over HTTP/S3: ⚡️ Zero full downloads ⚡️ Cloud index pushdown ⚡️ DuckDB/Polars streaming! Read the tour: blosc.org/posts/remote... Compress Better, Stream Smarter 🚀
001
Blosc Development Team @blosc.org · 22/09/2026
We are making a lot of progress on accessing remote PyTables tables via Blosc2 and fsspec. Stay tuned!
011
Blosc Development Team @blosc.org · 15/09/2026
📢 Python-Blosc2 4.13.1 is out! Continuing our "not an island" mood: stream remote datasets (HDF5, Zarr, B2Z) over HTTP/S3 lazily, and compute across formats in one expression: b2z + zarr * h5! Plus tiered caching & Win ARM64 wheels. github.com/Blosc/python... Compress Better, Compute Bigger 🚀
021
Blosc Development Team @blosc.org · 11/09/2026
📢 Python-Blosc2 4.13.0 is out! One lazy API for unite all remote data: browse B2Z/Zarr/HDF5 hierarchies over S3/GCS/HTTP, slice only what you need, cache in RAM or on disk, and export portable snapshots. b2view demo below👇 github.com/Blosc/python... Compress Better, Compute Bigger 🚀
022
Blosc Development Team @blosc.org · 01/09/2026
📢 Python-Blosc2 4.12.0 is out! Remote arrays work over S3/GCS/HTTP via fsspec. lazy=True fetches only the blocks needed and reuses a validated cache — 5–17x faster in S3 benchmarks. Also: concurrent Caterva2 writers + UTF-8 indexes. github.com/Blosc/python... Compress Better, Compute Bigger 🚀
000
Blosc Development Team @blosc.org · 25/08/2026
💡 Blosc2 tip: repeated text? use dictionary(): one int32 code per row, one copy of each value. On 1 Mrow with 100 distinct titles: group_by 22x faster, isin() 5.7x, 5.6x smaller on disk. blosc.org/python-blosc2/guides/optimization_tips.html#repeated-text-store-it-as-dictionary Enjoy data 🚀
blosc.org
Optimization tips - Python-Blosc2 documentation
000
Blosc Development Team @blosc.org · 13/08/2026
📣 Python-Blosc2 4.11.0 continues aligning with Arrow conventions. New CTable nullability on a validity mask instead of a sentinel value ✨️ Also, wheels are now abi3 compliant w/ CPython 3.11+, including 3.15 (tested) 🚅 More info: github.com/Blosc/python... Compress Better, Compute Bigger 🚀
github.com
Releases · Blosc/python-blosc2
A high-performance library for compressed ND arrays and columnar tables, with compute and indexing engines - Blosc/python-blosc2
001
Blosc Development Team @blosc.org · 11/08/2026
💡 Blosc2 tip: @blosc2.jit can compile complex control flow with loops and conditionals. A Mandelbrot escape loop written plainly: ~155x faster than plain Python, ~2.3x than hand-vectorized NumPy. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀 #Python #Blosc2
000
Blosc Development Team @blosc.org · 07/08/2026
💡 Blosc2 tip: for unbounded text, use utf8(). blosc2.field(blosc2.utf8()) string(200) costs 800 B/row on every read; utf8() rows cost what they weigh: ~4× less memory. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀 #Python #DataScience #Blosc2
001
Blosc Development Team @blosc.org · 04/08/2026
💡 Blosc2 tip: build big on-disk arrays from small ones. (rows + cols).compute(urlpath="big.b2nd", mode="w") Chunk by chunk: similar speed, ~26× less peak memory. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀 #Python #Blosc2
001
Reposted by Blosc Development Team
francescalted.bsky.social @francescalted.bsky.social · 31/07/2026
Most compression libraries ask you to move in. We think that's backwards. Blosc2 4.9.1: DuckDB, Polars and pandas 3 read a CTable directly via Arrow's PyCapsule protocol — no conversion step. Cooperation, not completeness. blosc.org/posts/not-an... Compress Better, Compute Bigger 🚀 #DataScience
blosc.org
Not an island: bringing compression to the tabular ecosystem
A compression library can go one of two ways: try to be everything (its own dataframe, its own query engine, its own format that nothing else reads), or be a fast, compact layer that slots underneath
022
Blosc Development Team @blosc.org · 29/07/2026
💡 Blosc2 tip: mmap read-only opens — skip the syscalls. >>> arr = blosc2.open(path, mmap_mode="r") Maps once, reads pages directly. ~15% faster for scattered reads, same memory. Free win, and scales with concurrent readers 🚀 blosc.org/python-blosc... #Python #DataScience #Blosc2
001
Blosc Development Team @blosc.org · 28/07/2026
💡 Blosc2 tip: `where=` for filtered reductions. t["temperature"].sum(where=t.region == 3) ~1.9× faster, ~7× less memory than mask-then-index. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀 #Python #DataScience #Blosc2
001
Blosc Development Team @blosc.org · 27/07/2026
💡 Python-Blosc2 tip #7: don't slice a column to reduce it. t["val"].sum() reduces the compressed column chunk-wise, never materialising it: 1.7× faster, ~12× less peak memory on 50M rows. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀
001
Blosc Development Team @blosc.org · 21/07/2026
💡 Python-Blosc2 tip #6: generate arrays with DSL kernels, not NumPy A @blosc2.dsl_kernel handed to lazyudf() compiles to native code, filling an NDArray chunk by chunk, in parallel. ~3.6x faster, ~2x less peak memory on 200M elements. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀
012
Blosc Development Team @blosc.org · 17/07/2026
💡 Python-Blosc2 tip #5: sorted top-k? Stream it from a FULL index CTable.sort_by(view=True)[:k] and NDArray.iter_sorted(start=-k) read just that slice from the index sidecar instead of sorting everything: up to ~74× less time, ~193× less peak memory. blosc.org/python-blosc... Enjoy data! #Python
020
Blosc Development Team @blosc.org · 16/07/2026
💡 Python-Blosc2 tip #4: let SUMMARY indexes answer min()/max() directly CTable auto-builds per-block min/max indexes for its scalar columns, so Column.min()/max() never decompress the column: ~4× faster, essentially no extra memory. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀
012
Blosc Development Team @blosc.org · 15/07/2026
💡 Python-Blosc2 tip #3: align your reads with the double partition A chunk-aligned read decompresses 1 chunk instead of 2 (~2.2× faster), and chunk-aligned slice() copies chunks as-is with no decompression at all (~4.9× faster). Same principle at block level. blosc.org/python-blosc... Enjoy data!
011
Blosc Development Team @blosc.org · 14/07/2026
💡 Python-Blosc2 tip #2: many readers, one file? Open with mmap_mode="r" in every reader — all share one set of mapped pages instead of a syscall + copy per read. 8 concurrent readers: ~4.5x faster, half the CPU, still ~one copy of the file in RAM. blosc.org/python-blosc... Enjoy data! 🚀
blosc.org
Optimization tips - Python-Blosc2 documentation
011
Blosc Development Team @blosc.org · 13/07/2026
💡 Python-Blosc2 tip #1: skip the NumPy detour. blosc2.linspace(0, 1, N) fills a compressed array chunk by chunk — at 200M float64, ~25x less peak memory than asarray(np.linspace(...)), comparable speed. Same for arange() and fromiter(). blosc.org/python-blosc... Compress Better, Compute Bigger 🚀
001
Blosc Development Team @blosc.org · 10/07/2026
📢 Python-Blosc2 4.8.0: share on-disk compressed containers across processes! SWMR readers follow a writer's appends in near real-time, and opt-in locking makes updates atomic — no torn reads, no retries. 🎬 One writer appending, three readers chasing it live: www.blosc.org/python-blosc... #Python
000
Blosc Development Team @blosc.org · 30/06/2026
Python-Blosc2 now has a 🔥 DSL-to-JavaScript transpiler! Newton-Raphson runs 2.7× faster on the browser! 💨 With #Pyodide, it’s not just Python in the browser: it’s the PyData stack 📊 too, #Blosc2 tacking compression 📦 + computing 🧮 duties. Demo 👇 blosc.org/demos/newton...
blosc.org
Newton fractal — blosc2's JIT, live in your browser
011
Blosc Development Team @blosc.org · 17/06/2026
✨ Tabular data visualization, fast filtering, smart navigation, remote plotting, and more... with just a text terminal! This is the new b2view Text User Interface that comes with latest Python-Blosc2 release. github.com/Blosc/python... Enjoy!
121
Blosc Development Team @blosc.org · 11/06/2026
🚀 New Blosc post: one selective query, 24.3M taxi trips, five tools. ⚡ Cold: CTable/.b2z ~1.9x faster than DuckDB, 9.4x vs pandas 🎯 Block indexes skip ~90% of the data 🤝 Warm: dead heat with DuckDB 💾 Only +2% disk vs Parquet blosc.org/posts/ctable... #Python #DuckDB
032
Blosc Development Team @blosc.org · 09/06/2026
📣 Python-Blosc2 4.4.x is out! 🖥️ b2view: interactive TUI browser 🎯 SUMMARY indexes → fast WHERE 🧮 DSL kernels as CTable columns ⚡ Faster chunk-by-chunk writes 🔀 where() via miniexpr Compressed, NumPy-native arrays & tables. 📝 github.com/Blosc/python... Enjoy data! #Python #NumPy
050
Blosc Development Team @blosc.org · 25/05/2026
🚀 C-Blosc2 3.1.0 is here! New in this release: sparse coordinate getters for faster random access, header-only b2nd metalayer access for plugins (no libblosc2 link needed), and new codec IDs for J2K/HTJ2K support. 🧩⚡ Release notes: github.com/Blosc/c-blos... Compress better, compute bigger! 💥
github.com
010
Blosc Development Team @blosc.org · 15/05/2026
Reaching 5 milions monthly downloads in PyPI is something to celebrate! 🥳 🥳 Thanks to all our users! This makes us happy for continuing working tirelessly for providing high quality compression for storage, computation for arrays, and now, for tabular data too! 🙏 #DataCompression #Computation
021
Blosc Development Team @blosc.org · 12/05/2026
Did you know that in recent Python-Blosc2 4.2.0 we released extremely efficient indexing engines that can store data in (guess what) ...compressed state? As a result, much larger tables can be indexed. Look at how this fares against other good indexing engines in plots below. Enjoy! #TabularData
021
Blosc Development Team @blosc.org · 08/05/2026
Python-Blosc2 4.2.0 is out 🎉 Meet `CTable`: a new compressed columnar container with null support, fast queries, and easy Arrow/Parquet interoperability. Compress better, analyze faster. Release notes: github.com/Blosc/python...
000
Blosc Development Team @blosc.org · 28/04/2026
🚀 C-Blosc2 3.0.0 is here! With variable length support, better compression, streamlined threading and more safety protections, C-Blosc2 3.0.0 makes compressed, persistent binary data storage more scalable, flexible, and production-ready. Release notes: github.com/Blosc/c-blos... Enjoy!
github.com
Releases · Blosc/c-blosc2
A fast, compressed, persistent binary data store library for C. - Blosc/c-blosc2
020
Blosc Development Team @blosc.org · 22/04/2026
📣 C-Blosc2 3.0.0-rc2 is out 🚀 Big highlight: a new shared managed thread pool for parallel execution. ✅ With this and previous work, C-Blosc2 will become specially useful for storing variable length data, using extremely low resources. github.com/Blosc/c-blos... #CBlosc2 #Compression #OpenSource
032
Blosc Development Team @blosc.org · 27/03/2026
🚀 C-Blosc2 3.0.0 RC1 is out! What's new? 📦 VL-blocks for irregular data (strings, JSON) 🗜️ LZ4/LZ4HC dictionary compression 🛡️ Important safety fixes Check out the full release notes here: github.com/Blosc/c-blos... Give it a spin and provide feedback! #Blosc2 #OpenSource #C
github.com
032
Reposted by Blosc Development Team
francescalted.bsky.social @francescalted.bsky.social · 25/03/2026
Look at the bump in performance that we will see with the next Python-Blosc2 release. Matrix multiplication has been speeded up by using blocks and Blosc2 prefilters and its own efficient and multithreaded engine. Expect between 5x and 6x better speed for matrices with no padding.
042
Blosc Development Team @blosc.org · 10/03/2026
We designed DSL kernels in Blosc2 so that they can be used in the main platforms; nowadays this necessarily includes Web Assembly (WASM) in the browser. We also implemented a new JIT compiler for WASM specially meant for Blosc2 DSL kernels. Play with it at: cat2.cloud/demo/roots/@... Have fun!
021
Reposted by Blosc Development Team
francescalted.bsky.social @francescalted.bsky.social · 06/03/2026
New DSL kernels in Python-Blosc2, being parallelizable and JIT-compiling capable, can accelerate code quite a bit 🚀 For example, linspace() has been rewritten to use DSL kernels, and performance boost is well beyond expectations.
001
Reposted by Blosc Development Team
luke.shaw@ironarray.io @lshaw-ironarray.bsky.social · 04/03/2026
Blosc2 4.1 Release! We've packed this minor release: optimised compression and funcs for unicode arrays; cumulative reductions; memory-map support for store containers like `DictStore` ; and a DSL kernel functionality for faster, compiled, user-defined funcs!👇️ Notebook - github.com/Blosc/python...)
022
Blosc Development Team @blosc.org · 20/02/2026
Super-efficient DSL kernels are coming with forthcoming Python-Blosc 4.1.0. A DSL kernel is a function that takes ndarray objects as input and returns a ndarray object as output. Performance with classic Mandelbrot is kind of mind-blowing. See github.com/Blosc/python... for full notebook.
001
Blosc Development Team @blosc.org · 05/02/2026
The new engine in Python-Blosc2 4.0 delivers quite awesome performance results. Processing a 400 MB array can be up to 9.6x faster than NumPy 🤯 The best part is that it fully supports NumPy's array and ufunc interfaces. High performance with zero friction! 🏎️💨 ironarray.io/blog/miniexp... Enjoy!
064
Blosc Development Team @blosc.org · 02/02/2026
Python-Blosc2 just became a compute powerhouse. ⚡️ v4.0 integrates `miniexpr`: computing on cache-friendly blocks instead of RAM-heavy chunks. 🚀 Result? Massive speedups for memory-bound tasks (up to 4.5x in Cat2Cloud). See benchmarks: ironarray.io/blog/miniexp... #Python #HPC #DataEngineering
021
Reposted by Blosc Development Team
luke.shaw@ironarray.io @lshaw-ironarray.bsky.social · 30/01/2026
📢 Blosc2 🤝 OpenZL 🔁 Use the OpenZL compression 🗜️ library from Blosc2: github.com/Blosc/blosc2...! 𝚙̲𝚒̲𝚙̲ ̲𝚒̲𝚗̲𝚜̲𝚝̲𝚊̲𝚕̲𝚕̲ ̲𝚋̲𝚕̲𝚘̲𝚜̲𝚌̲𝟸̲–̲𝚘̲𝚙̲𝚎̲𝚗̲𝚣̲𝚕̲ and just like that, OpenZL compression + Blosc2 compute engine! Check out the OpenZL project openzl.org. Inside story on the plugin ⏩️ blosc.org/posts/openzl...
022
Reposted by Blosc Development Team
ironArray SLU @ironarray.io · 21/01/2026
🚀 Our #PyDataGlobal 2025 tutorial on modern #Blosc2 & #Caterva2 features is out! Learn how compression boosts array performance & enables cloud computing without downloads. ☁️ Watch here: 👇 www.youtube.com/watch?v=tUvS... #Python #HPC #DataScience #BigData
youtube.com
Francesc Alted+Luke Shaw - Hands-on with Blosc2: Accelerating Your Python Data Workflows-PyData 2025
YouTube video by PyData
002
Reposted by Blosc Development Team
francescalted.bsky.social @francescalted.bsky.social · 27/11/2025
🤔 Can a library that computes on compressed data actually outperform performance heavyweights like NumPy and NumExpr? Find some surprising answers 👉 www.blosc.org/posts/roofli... During the walk, I'm also introducing a funny anecdote back when NumPy was a 👶 Enjoy! #Blosc2 #HPC #NumPy #Numexpr
173
Reposted by Blosc Development Team
luke.shaw@ironarray.io @lshaw-ironarray.bsky.social · 25/11/2025
To round off our presentation of the shiny new Cat2Cloud product, we have a final 💊 ironPill 💊on data handling from the web client 🌐 . See how to move ➡️ , copy ©️ and perform other file management operations easily from the prompt box! Stop downloading, start doing! 🔨
031
Reposted by Blosc Development Team
luke.shaw@ironarray.io @lshaw-ironarray.bsky.social · 19/11/2025
ironArray is on a roll 🚅! With Cat2Cloud released this week we continue our 💊ironPill series with another video, showing you how to browse your cloud-hosted data ☁️ in the bespoke web client! Visualise and manage all kinds of data formats - from arrays to jpgs 🖼️, markdown to PDFs 📄 !
032
Blosc Development Team @blosc.org · 20/11/2025
Excited to be part of this data story! ❤️
000
Reposted by Blosc Development Team
ironArray SLU @ironarray.io · 18/11/2025
Tired of downloading large datasets? Cat2Cloud is now live! 🚀 Bring computation directly to your data. Host, run, and share workflows with incredible speed. ⚡ Read the announcement: ironarray.io/blog/cat2clo... #Cat2Cloud #BigData #DataScience #Python #HPC
031
Reposted by Blosc Development Team
ironArray SLU @ironarray.io · 12/11/2025
Stop downloading. Start computing. Everything changes on November 18, 2025. #Cat2Cloud is coming.
032
Blosc Development Team @blosc.org · 07/11/2025
Our Blosc2 tutorial for @euroscipy.bsky.social 2025 is up! Learn how compression boosts tensor data handling, sharing & computing. www.youtube.com/watch?v=BdpT... #DataScience #BigData #HPC #OpenSource
youtube.com
18.08.2025 Compress, Compute, and Conquer: Python-Blosc2 for Efficient Data Analysis
YouTube video by EuroSciPy
042
Reposted by Blosc Development Team
luke.shaw@ironarray.io @lshaw-ironarray.bsky.social · 03/11/2025
📢🚨Blosc2 3.11.0 Released! 🚨 📢 Blosc2 is part of the NumPy array library ecosystem: numpy.org In v3.11, Blosc2 has become even more flexible; all functions can accept basically any array object: 𝚋𝚕𝚘𝚜𝚌𝟸.𝚖𝚊𝚝𝚖𝚞𝚕(𝙰, 𝙱) for 𝙰, 𝙱 PyTorch, Jax, Tensorflow, Zarr... arrays! #Compute #Tensor #PyTorch
022
Reposted by Blosc Development Team
luke.shaw@ironarray.io @lshaw-ironarray.bsky.social · 28/10/2025
💊IronPill 2💊 See how Blosc2 powers heavy-duty linear algebra (100GB!) workflows ⚡1.5-2x faster than PyTorch + h5py! 🧱 optimised chunking for your cache hierarchy 🐍 one line syntax 𝚋𝚕𝚘𝚜𝚌𝟸.𝚖𝚊𝚝𝚖𝚞𝚕(𝙰, 𝙱, 𝚞𝚛𝚕𝚙𝚊𝚝𝚑='𝚘𝚞𝚝.𝚋𝟸𝚗𝚍') See blog here: ironarray.io/blog/la-blosc #Blosc2 #Data #LinearAlgebra
041