Sign in

Blosc Development Team

@blosc.org
30 followers 4 following 77 posts

Announcements about Blosc2 developments blosc.org

PostsRepliesMedia
Blosc Development Team @blosc.org · 30/09/2026
📢 Python-Blosc2 4.14 is out! A common interface to stream, query & slice remote Parquet, #PyTables, #HDF5, #Zarr & #Blosc2 data over HTTP/S3: ⚡️ Zero full downloads ⚡️ Cloud index pushdown ⚡️ DuckDB/Polars streaming! Read the tour: blosc.org/posts/remote... Compress Better, Stream Smarter 🚀
012
Blosc Development Team @blosc.org · 22/09/2026
We are making a lot of progress on accessing remote PyTables tables via Blosc2 and fsspec. Stay tuned!
011
Blosc Development Team @blosc.org · 15/09/2026
📢 Python-Blosc2 4.13.1 is out! Continuing our "not an island" mood: stream remote datasets (HDF5, Zarr, B2Z) over HTTP/S3 lazily, and compute across formats in one expression: b2z + zarr * h5! Plus tiered caching & Win ARM64 wheels. github.com/Blosc/python... Compress Better, Compute Bigger 🚀
021
Blosc Development Team @blosc.org · 11/09/2026
📢 Python-Blosc2 4.13.0 is out! One lazy API for unite all remote data: browse B2Z/Zarr/HDF5 hierarchies over S3/GCS/HTTP, slice only what you need, cache in RAM or on disk, and export portable snapshots. b2view demo below👇 github.com/Blosc/python... Compress Better, Compute Bigger 🚀
022
Blosc Development Team @blosc.org · 01/09/2026
📢 Python-Blosc2 4.12.0 is out! Remote arrays work over S3/GCS/HTTP via fsspec. lazy=True fetches only the blocks needed and reuses a validated cache — 5–17x faster in S3 benchmarks. Also: concurrent Caterva2 writers + UTF-8 indexes. github.com/Blosc/python... Compress Better, Compute Bigger 🚀
000
Blosc Development Team @blosc.org · 11/08/2026
💡 Blosc2 tip: @blosc2.jit can compile complex control flow with loops and conditionals. A Mandelbrot escape loop written plainly: ~155x faster than plain Python, ~2.3x than hand-vectorized NumPy. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀 #Python #Blosc2
000
Blosc Development Team @blosc.org · 07/08/2026
💡 Blosc2 tip: for unbounded text, use utf8(). blosc2.field(blosc2.utf8()) string(200) costs 800 B/row on every read; utf8() rows cost what they weigh: ~4× less memory. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀 #Python #DataScience #Blosc2
001
Blosc Development Team @blosc.org · 04/08/2026
💡 Blosc2 tip: build big on-disk arrays from small ones. (rows + cols).compute(urlpath="big.b2nd", mode="w") Chunk by chunk: similar speed, ~26× less peak memory. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀 #Python #Blosc2
001
Blosc Development Team @blosc.org · 29/07/2026
💡 Blosc2 tip: mmap read-only opens — skip the syscalls. >>> arr = blosc2.open(path, mmap_mode="r") Maps once, reads pages directly. ~15% faster for scattered reads, same memory. Free win, and scales with concurrent readers 🚀 blosc.org/python-blosc... #Python #DataScience #Blosc2
001
Blosc Development Team @blosc.org · 28/07/2026
💡 Blosc2 tip: `where=` for filtered reductions. t["temperature"].sum(where=t.region == 3) ~1.9× faster, ~7× less memory than mask-then-index. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀 #Python #DataScience #Blosc2
001
Blosc Development Team @blosc.org · 27/07/2026
💡 Python-Blosc2 tip #7: don't slice a column to reduce it. t["val"].sum() reduces the compressed column chunk-wise, never materialising it: 1.7× faster, ~12× less peak memory on 50M rows. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀
001
Blosc Development Team @blosc.org · 21/07/2026
💡 Python-Blosc2 tip #6: generate arrays with DSL kernels, not NumPy A @blosc2.dsl_kernel handed to lazyudf() compiles to native code, filling an NDArray chunk by chunk, in parallel. ~3.6x faster, ~2x less peak memory on 200M elements. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀
012
Blosc Development Team @blosc.org · 17/07/2026
💡 Python-Blosc2 tip #5: sorted top-k? Stream it from a FULL index CTable.sort_by(view=True)[:k] and NDArray.iter_sorted(start=-k) read just that slice from the index sidecar instead of sorting everything: up to ~74× less time, ~193× less peak memory. blosc.org/python-blosc... Enjoy data! #Python
020
Blosc Development Team @blosc.org · 16/07/2026
💡 Python-Blosc2 tip #4: let SUMMARY indexes answer min()/max() directly CTable auto-builds per-block min/max indexes for its scalar columns, so Column.min()/max() never decompress the column: ~4× faster, essentially no extra memory. blosc.org/python-blosc... Compress Better, Compute Bigger 🚀
012
Blosc Development Team @blosc.org · 15/07/2026
💡 Python-Blosc2 tip #3: align your reads with the double partition A chunk-aligned read decompresses 1 chunk instead of 2 (~2.2× faster), and chunk-aligned slice() copies chunks as-is with no decompression at all (~4.9× faster). Same principle at block level. blosc.org/python-blosc... Enjoy data!
011
Blosc Development Team @blosc.org · 13/07/2026
💡 Python-Blosc2 tip #1: skip the NumPy detour. blosc2.linspace(0, 1, N) fills a compressed array chunk by chunk — at 200M float64, ~25x less peak memory than asarray(np.linspace(...)), comparable speed. Same for arange() and fromiter(). blosc.org/python-blosc... Compress Better, Compute Bigger 🚀
001
Blosc Development Team @blosc.org · 10/07/2026
📢 Python-Blosc2 4.8.0: share on-disk compressed containers across processes! SWMR readers follow a writer's appends in near real-time, and opt-in locking makes updates atomic — no torn reads, no retries. 🎬 One writer appending, three readers chasing it live: www.blosc.org/python-blosc... #Python
000
Blosc Development Team @blosc.org · 18/06/2026
Shorter and annotated video on newer b2view capabilities👇 You are welcome!
010
Blosc Development Team @blosc.org · 17/06/2026
✨ Tabular data visualization, fast filtering, smart navigation, remote plotting, and more... with just a text terminal! This is the new b2view Text User Interface that comes with latest Python-Blosc2 release. github.com/Blosc/python... Enjoy!
121
Blosc Development Team @blosc.org · 11/06/2026
🚀 New Blosc post: one selective query, 24.3M taxi trips, five tools. ⚡ Cold: CTable/.b2z ~1.9x faster than DuckDB, 9.4x vs pandas 🎯 Block indexes skip ~90% of the data 🤝 Warm: dead heat with DuckDB 💾 Only +2% disk vs Parquet blosc.org/posts/ctable... #Python #DuckDB
032
Blosc Development Team @blosc.org · 09/06/2026
📣 Python-Blosc2 4.4.x is out! 🖥️ b2view: interactive TUI browser 🎯 SUMMARY indexes → fast WHERE 🧮 DSL kernels as CTable columns ⚡ Faster chunk-by-chunk writes 🔀 where() via miniexpr Compressed, NumPy-native arrays & tables. 📝 github.com/Blosc/python... Enjoy data! #Python #NumPy
050
Blosc Development Team @blosc.org · 15/05/2026
Reaching 5 milions monthly downloads in PyPI is something to celebrate! 🥳 🥳 Thanks to all our users! This makes us happy for continuing working tirelessly for providing high quality compression for storage, computation for arrays, and now, for tabular data too! 🙏 #DataCompression #Computation
021
Blosc Development Team @blosc.org · 12/05/2026
Did you know that in recent Python-Blosc2 4.2.0 we released extremely efficient indexing engines that can store data in (guess what) ...compressed state? As a result, much larger tables can be indexed. Look at how this fares against other good indexing engines in plots below. Enjoy! #TabularData
021
Blosc Development Team @blosc.org · 08/05/2026
Python-Blosc2 4.2.0 is out 🎉 Meet `CTable`: a new compressed columnar container with null support, fast queries, and easy Arrow/Parquet interoperability. Compress better, analyze faster. Release notes: github.com/Blosc/python...
000
Blosc Development Team @blosc.org · 22/04/2026
📣 C-Blosc2 3.0.0-rc2 is out 🚀 Big highlight: a new shared managed thread pool for parallel execution. ✅ With this and previous work, C-Blosc2 will become specially useful for storing variable length data, using extremely low resources. github.com/Blosc/c-blos... #CBlosc2 #Compression #OpenSource
032
Blosc Development Team @blosc.org · 10/03/2026
We designed DSL kernels in Blosc2 so that they can be used in the main platforms; nowadays this necessarily includes Web Assembly (WASM) in the browser. We also implemented a new JIT compiler for WASM specially meant for Blosc2 DSL kernels. Play with it at: cat2.cloud/demo/roots/@... Have fun!
021
Blosc Development Team @blosc.org · 20/02/2026
Super-efficient DSL kernels are coming with forthcoming Python-Blosc 4.1.0. A DSL kernel is a function that takes ndarray objects as input and returns a ndarray object as output. Performance with classic Mandelbrot is kind of mind-blowing. See github.com/Blosc/python... for full notebook.
001
Blosc Development Team @blosc.org · 05/02/2026
The new engine in Python-Blosc2 4.0 delivers quite awesome performance results. Processing a 400 MB array can be up to 9.6x faster than NumPy 🤯 The best part is that it fully supports NumPy's array and ufunc interfaces. High performance with zero friction! 🏎️💨 ironarray.io/blog/miniexp... Enjoy!
064
Blosc Development Team @blosc.org · 02/02/2026
Python-Blosc2 just became a compute powerhouse. ⚡️ v4.0 integrates `miniexpr`: computing on cache-friendly blocks instead of RAM-heavy chunks. 🚀 Result? Massive speedups for memory-bound tasks (up to 4.5x in Cat2Cloud). See benchmarks: ironarray.io/blog/miniexp... #Python #HPC #DataEngineering
021
Blosc Development Team @blosc.org · 22/09/2025
🚀 Boost your Python computations on NumPy arrays with @blosc.jit! Let Blosc2's compute engine accelerate your code without changing containers. ⚡ For ultimate speed, native Blosc2 containers are still king. 🏎️ Code: gist.github.com/FrancescAlte... #Python #DataScience #NumPy #Performance #Blosc
110
Blosc Development Team @blosc.org · 20/08/2025
🚀 We're thrilled to announce *TreeStore*, a new class in Python-Blosc2! Endow your datasets with a hierarchical structure! ⚡️ 📝 We've blogged about it: www.blosc.org/posts/new-tr... It's in beta, and available in Python-Blosc2 v3.7.2. Enjoy! #Python #Blosc2 #TreeStore #DataScience #OpenSource
022
Blosc Development Team @blosc.org · 29/07/2025
Python-Blosc2 comes with an extensively tested partition sizes algorithm for automatically setting the best chunk and block sizes for you. But you can always bypass this mechanism for more appropriate fine-tuning in different scenarios 🚀 www.blosc.org/python-blosc... #Compression #Performance
010
Blosc Development Team @blosc.org · 18/07/2025
🗣️ Announcing Python-Blosc2 3.6.1 ¡Unlock new levels of data manipulation with #Blosc2! 🚀 We've tamed the complexity of fancy indexing to make it intuitive, efficient, and consistent with NumPy's behavior. 👉 www.blosc.org/posts/blosc2... Enjoy #Python #DataScience #BigData #NumPy #Performance
New Blosc2 fancy indexing performance against other competitors, including NumPy. Blosc2 is quite efficient.
010
Blosc Development Team @blosc.org · 14/07/2025
🚀 Blosc2 supports memory-mapped files for super-efficient data access! 🚀 ✨ Why memory-mapping? 1️⃣ No system call overhead for each read/write 2️⃣ Data goes straight from page cache to user space—much faster than traditional I/O! 👉 github.com/Blosc/python... #DataScience #Performance #BigData 🚀💾
020
Blosc Development Team @blosc.org · 08/07/2025
#Blosc2 now runs directly in your browser! Leveraging the power of #WASM, #Pyodide, and #JupyterLite, you can harness efficient, adaptable compression through the web's universal interface. Compress Better, Compute Bigger, Share Faster #WebAssembly #DataCompression #WebDevelopment #DataScience
031
Blosc Development Team @blosc.org · 02/07/2025
🗣️ Announcing Python-Blosc2 3.5.1 🚀 This introduces significant performance and memory optimizations, enhancing the experience of computing with large, compressed datasets using lazy expressions. Compress Better, Compute Bigger! #Performance #MemoryEfficiency #DataScience
Evaluate slices of lazy expressions (inner dims)
011
Blosc Development Team @blosc.org · 01/07/2025
Struggling with slow I/O in #HDF5? Try #Blosc2 as a filter or I/O data handler — faster data, less pain! 👉 www.blosc.org/posts/pytabl... #Performance #DataScience
010
Blosc Development Team @blosc.org · 24/06/2025
📣Python-Blosc2 3.5.0 is out! 🚀 This release introduces the powerful stack() function! 🥞 You can now stack multiple Blosc2 NDarrays along a new axis, similar to np.stack(), while getting great performance. We have blogged about it here: www.blosc.org/posts/blosc2... Compress Better, Compute Bigger
001
Blosc Development Team @blosc.org · 23/06/2025
📢 We are pleased to announce the integration of a new stack feature in #Blosc2 🚀, which allows for stacking large arrays along a new axis. We've updated our recent blog post: Check it out! 👇 www.blosc.org/posts/blosc2... Compress Better, Compute Bigger #Python #DataScience #Performance #OpenSource
031
Blosc Development Team @blosc.org · 05/06/2025
#Python-Blosc2 is hitting 1 million weekly downloads on PyPI! 🎉 Users are rapidly adopting #Blosc2, which now accounts for over 95% of downloads compared to Blosc1. 📈 Thanks to our amazing users. 🙏 🚀 Our motto: Compress Better, Compute Bigger! 💪 #Milestone #CommunitySupport #DataCompression
010
Blosc Development Team @blosc.org · 20/05/2025
💡 Did you know you can supercharge your #HDF5 datasets with #Blosc2? 🚀 Leverage hdf5plugin (hdf5plugin.readthedocs.io) to integrate Blosc2 as a filter within HDF5. Create and read data using popular Python wrappers like h5py or PyTables, while achieving excellent performance! 💨 Compress Better!
100
Blosc Development Team @blosc.org · 30/04/2025
🗣️New blog entry of Ricardo Sales Piquer, our intern working in linear algebra problems: Make NDArrays Transposition Fast (and Compressed!) in #Blosc2 🚀 www.blosc.org/posts/optimi... Great work Ricardo! 🎉 #Compression #LinearAlgebra #Optimization #DataScience Compress Better, Compute Faster 😀
Performance of Blosc2 transposition depending on the chunk size
023
Blosc Development Team @blosc.org · 09/04/2025
📣 Announcing the release of Python-Blosc2 3.3.0: 🔄 New blosc2.transpose() ⚡️ New fast path for NDArray.slice() - up to 40x speedup 🔧 NDArray.slice() now preserves compression params 📝 Improved documentation Thanks to all contributors! github.com/Blosc/python... Compress Better, Compute Bigger! 😀
Transposing a compressed array with Blosc2
011
Blosc Development Team @blosc.org · 28/03/2025
📢 🔥 Article on Blosc2: Compute with TB-sized datasets on your own hardware! 🚀 Outperforms NumPy by 10x ~ 100x for large computations 🐍 Integrates seamlessly with the Python data science ecosystem 💻 Works both in-memory and on-disk with minimal performance differences ironarray.io/blog/compute...
Up to 75x speedups of Blosc2 vs NumPy and Numba when computing with generic NumPy functions
022
Blosc Development Team @blosc.org · 26/03/2025
🔥 Blosc2 3.2.1: NumPy's best friend! 🔥 blosc2.jit + NumPy array interface = blazing fast, direct compressed data crunching! See how: github.com/Blosc/python-blosc2/blob/main/examples/ndarray/jit-numpy-funcs.py #DataScience #Python #Compression Compress Better, Compute Bigger!
The plot shows the performance of the new Blosc2 compute capabilities. It can be more than 10x faster than NumPy and numba.
022
Blosc Development Team @blosc.org · 07/03/2025
Did you know that you can use the @blosc2.jit decorator to accelerate your computations with NumPy by 10x and more by just adding a single line to your function? No need to add loops manually, just add @blosc2.jit and it will do its magic! Compress Better, Compute Bigger!
blosc2.jit can accelerate computations a lot, and very easily.
022
Blosc Development Team @blosc.org · 25/02/2025
Blosc compression helped fighting memory bottlenecks for 15 years now! 🎂 See how Dask + Zarr benefits from it. But vertical integration between compression and the compute engine in newest Python-Blosc2 makes a big difference in terms of speed ⚡ and scalability 🤟 #Compress better, #Compute bigger!
Performance of Blosc2 vs Dask+Zarr+Blosc for a complex math calculation.
011
Blosc Development Team @blosc.org · 19/02/2025
Did you know that with python-blosc2 you can select an arbitrary precision of floating point data? This allows for a much better compression ratio: from 2x up to 100x. And still performing like a champ ⚡⚡ More info: www.blosc.org/python-blosc2/ Compress better, compute bigger! #compute #compress
Compute performance vs arbitrary precision
010
Blosc Development Team @blosc.org · 14/02/2025
📢 For Python-Blosc2 3.1.1 we have optimized even more the shiny new compute engine. See how the new version can reach great performance even for the case that the operands largely exceed (by up to 20x) the RAM of your computer ⚡️ www.blosc.org/python-blosc2/ Compress better, compute bigger!
Compute performance scales well beyond the available RAM
000
Blosc Development Team @blosc.org · 12/02/2025
🗣️ C-Blosc2 2.16.0 is out. We have added new features, like: * _fseeki64/_ftelli64/_stat64 on Windows for large file support. * 12-byte unshuffle for avx2 and sse2. * Better description of the Blosc2 format as a whole. Compress better, Compute faster!
Double partitioning diagram
001