Profile
Back to NewsBack
GitHub Trending 4 min
Reader Mode
Blosc/python-blosc2: A high-performance library for compressed ND arrays and columnar tables, with compute and indexing engines

Blosc/python-blosc2: A high-performance library for compressed ND arrays and columnar tables, with compute and indexing engines

15 hours ago

============= Python-Blosc2 =============

A fast & compressed ndarray library with a flexible compute engine ==================================================================

:Author: The Blosc development team :Contact: [email protected] :Github: https://github.com/Blosc/python-blosc2 :Actions: |actions| :PyPi: |version| :NumFOCUS: |numfocus| :Code of Conduct: |Contributor Covenant|

.. |version| image:: https://img.shields.io/pypi/v/blosc2.svg :target: https://pypi.python.org/pypi/blosc2 .. |Contributor Covenant| image:: https://img.shields.io/badge/Contributor%20Covenant-v2.0%20adopted-ff69b4.svg :target: https://github.com/Blosc/community/blob/master/code_of_conduct.md .. |numfocus| image:: https://img.shields.io/badge/powered%20by-NumFOCUS-orange.svg?style=flat&colorA=E1523D&colorB=007D8A :target: https://numfocus.org .. |actions| image:: https://github.com/Blosc/python-blosc2/actions/workflows/build.yml/badge.svg :target: https://github.com/Blosc/python-blosc2/actions/workflows/build.yml

What is Python-Blosc2? =======================

Python-Blosc2 is a high-performance compressor, compute engine, and format for binary data containers that are portable and open-source. It comes with a lazy expression engine allowing for complex calculations on compressed data, whether stored in memory, on disk, or over the network (e.g., via Caterva2 _). It is especially optimized for storing and retrieving data from N-dimensional arrays (NDArray) and columnar tables (CTable), complemented by a query/indexing layer. The main use case is fast, compressed, out-of-core numerical data — especially when data is too large to fit comfortably in RAM.

C-Blosc2 _ is used under the hood as its compression backend. Written in C, and building on its predecessor C-Blosc _, C-Blosc2 aims to be an extremely fast meta-compressor for binary data, supporting a diverse set of strategies, and with an extensible plugin architecture for a wide range of codecs and filters.

More info: https://www.blosc.org/python-blosc2/getting_started/overview.html

Installing ==========

Binary packages are available for major OSes (Win, Mac, Linux) and platforms. Install from PyPI using `pip:

.. code-block:: console

pip install blosc2 --upgrade

Conda users can install from conda-forge:

.. code-block:: console

conda install -c conda-forge python-blosc2

Command line tools ==================

Two CLI tools are installed along with the package:

  • b2view: an interactive terminal browser (TUI) for TreeStore bundles
(.b2d directories or .b2z files), with paged views of NDArray and CTable data of any size (walkthrough _; requires pip install "blosc2[tui]").
  • parquet-to-blosc2: converts Parquet files to Blosc2 columnar table
stores, and back (walkthrough _; requires pip install "blosc2[parquet]").

Documentation =============

The documentation is available here:

https://blosc.org/python-blosc2/python-blosc2.html

You can find examples at:

https://github.com/Blosc/python-blosc2/tree/main/examples

A tutorial from PyData Global 2025 is available at:

https://github.com/Blosc/PyData-Global-2025-Tutorial

(Click here _ to watch the video recording of the tutorial)

It contains Jupyter notebooks explaining the main features of Python-Blosc2.

License =======

This software is licensed under a 3-Clause BSD license. A copy of the python-blosc2 license can be found in LICENSE.txt _.

Discussion forum ================

Discussion about this package is welcome at:

https://github.com/Blosc/python-blosc2/discussions

Social feeds ------------

Stay informed about the latest developments by following us in Mastodon _, Bluesky _ or LinkedIn _.

Thanks ======

Blosc2 is supported by the NumFOCUS foundation _, the LEAPS-INNOV project _ and ironArray SLU _, among many other donors. This allowed the following people to have contributed in an important way to the core development of the Blosc2 library:

  • Francesc Alted
  • Marta Iborra
  • Luke Shaw
  • Aleix Alcacer
  • Oscar Guiñón
  • Juan David Ibáñez
  • Ivan Vilata i Balaguer
  • Oumaima Ech.Chdig
  • Ricardo Sales Piquer
In addition, other people have participated in the project in different aspects:
  • Jan Sellner, contributed the mmap support for NDArray/SChunk objects.
  • Dimitri Papadopoulos, contributed a large bunch of improvements to
many aspects of the project. His attention to detail is remarkable.
  • And many others that have contributed with bug reports, suggestions and
improvements.

Developed using JetBrains IDEs.

.. image:: https://resources.jetbrains.com/storage/products/company/brand/logos/jetbrains.svg :target: https://jb.gg/OpenSource :alt: JetBrains logo.

Citing Blosc ============

You can cite our work on the various libraries under the Blosc umbrella as follows:

.. code-block:: console

@ONLINE{blosc, author = {{Blosc Development Team}}, title = "{A fast, compressed and persistent data store library}", year = {2009-2026}, note = {https://blosc.org} }

Support Blosc for a Sustainable Future ======================================

If you find Blosc useful and want to support its development, please consider making a donation or contract to the Blosc Development Team `_. Thank you!

Compress Better, Compute Bigger

Chat with me