Hi,

For several years there has been an ongoing effort in order to break
down the Arrow C++ library into smaller modules to break monolithic
installations. [1][2][3]
Currently we have the following libraries (Arrow Core, Arrow Compute,
Arrow Acero, Arrow Dataset, Arrow Cuda, Arrow Gandiva, Arrow
Substrait, Arrow Flight, Arrow Flight SQL, Arrow Flight SQL ODBC,
Arrow S3, Parquet).

There's still work to do in order to move GCS and Azure out of the
core libarrow, similar to the work we've done to move S3 and the AWS
SDK.

There was a lot of work also on the conda recipes in order to be able
to install those separately and provide PyArrow installations where
those can also be installed separately. [4]

PyArrow wheel size on PyPI has been a topic of discussion both as a
maintainer's burden and as a user problem. [5][6][7]

For context, a release today is ~2.1 GB on wheels size. 49 wheels 7
Python ABIs * 7 platforms.

There have been several investigations in order to break down the
installation into several packages. [8][9]

I have started working on a PoC to have a multi-wheel installation
[10]. This is green on CI, built wheels and tested for Manylinux,
Musl, macOS, Windows both free-threaded and non-free-threaded.

My approach has been to follow on what we did for conda, create
optional side-wheels that only contain the shared library for the
specific library (libarrow_s3 on the test case).

The approach is, build PyArrow with everything turned on so all Cython
extensions get built then bundle everything but the specific shared
library that goes into a different wheel (libarrow_s3). That one is
then bundled on a specific pyarrow-s3 wheel.

The new wheel pins the exact same version of PyArrow and PyArrow wheel
declares the new optional dependency with the exact same version.

So the following:
- pip install pyarrow
Would install only core (no libarrow_s3.so/dll/dylib).
- pip install pyarrow[s3]
Would install also pyarrow-s3.
- pip install pyarrow-s3
Would pull PyArrow core wheel too.

We probably want `pip install pyarrow` to also install `pyarrow-s3`
initially when we do the split and make it optional once the wheel
split is battle tested. The [s3] extra should be declared from day one
even if it's required so downstream projects can write pyarrow[s3]
where necessary.
I've adapted the import mechanism to either import pyarrow-s3 which
loads libarrow_s3 globally or fails with a message in the following
lines:

```
>>> import pyarrow.fs
>>> pyarrow.fs.S3FileSystem
ImportError: The pyarrow installation is not built with support for
'S3FileSystem'
 (libarrow_s3.so.2600: cannot open shared object file: No such file or
directory.
  If pyarrow was installed from PyPI, S3 support is provided by the
separate 'pyarrow-s3' package)
```

We do not name mangle libarrow or any of the currently built
libraries, this is why an effort like consolidatewheels [9] is not
required for them.

Splitting wheels as a wrapper for a shared library is something that I
took inspiration from Openblas [11][12] and has been proven on runtime
loading of NVIDIA libraries[13][14] on pytorch [15][16].

This PoC reduces the Linux wheels size by ~5.5MB/~6MB (the
libarrow_s3.so size). macOS and Windows reductions are smaller due to
AWS SDK footprint being smaller on those ~1.5MB/~2MB.

Some of those side wheels could be cheap because they don't rely on
Python ABI (at least for filesystems) so 7 wheels instead of 49. This
alone doesn't solve the PyPI quota problem (~9% wheel reduction) but
it compounds with the abi3 efforts [17].

The idea is to further reduce it by moving the rest of filesystems,
Azure, GCS to their own wheels but the next candidate I'd try is
flight as it would grant the maximum benefit as the largest optional
library.

Some of the downsides I see from this:
- Managing the extra complexity of library loading from multiple wheels.
- Versions must match, this might make dependency management on some
projects with several dependencies pulling Arrow much more complex.
- Would require managing multiple PyPI projects with the release
management costs associated (not a big issue but worth mentioning).
- pip and conda mixing has new failure modes. Installing pyarrow-core
from conda and pip installing pyarrow-s3 reports pyarrow requirement
as already satisfied from conda and conda's pyarrow already contains
s3.

This email is to bring attention to this effort and help with design
considerations. Please, come to the PoC PR [10] and help shape this.

Regards,
Raúl

[1] https://github.com/apache/arrow/issues/49399
[2] https://github.com/apache/arrow/issues/25025
[3] https://github.com/apache/arrow/issues/15280
[4] https://github.com/conda-forge/arrow-cpp-feedstock/issues/1035
[5] https://github.com/pypi/support/issues/4409
[6] https://github.com/pypi/support/issues/10331
[7] https://github.com/pandas-dev/pandas/issues/54466#issuecomment-1674229577
[8] https://github.com/apache/arrow/issues/24688
[9] https://github.com/amol-/consolidatewheels
[10] https://github.com/apache/arrow/pull/51459
[11] https://pypi.org/project/scipy-openblas64/
[12] 
https://github.com/MacPython/openblas-libs/blob/2b6df42890977bb66d203ab1f2c85c7879a36f7a/local/scipy_openblas64/__init__.py
[13] https://pypi.org/project/nvidia-cublas-cu12/
[14] https://pypi.org/project/nvidia-cudnn-cu12/
[15] 
https://github.com/pytorch/pytorch/blob/48b7fc30407ee1818b59718732e60968e2d0ca96/torch/__init__.py#L343-L358
[16] 
https://github.com/pytorch/pytorch/blob/48b7fc30407ee1818b59718732e60968e2d0ca96/torch/__init__.py#L402
[17] https://github.com/apache/arrow/issues/50398

Reply via email to