Hi all, Thank you Raúl for working on this topic and providing us with such a detailed and informative email!
I am very much looking forward to this improvement as I have also witnessed on multiple occasions that pyarrow is being rejected as a dependency due to the size of the package. I think this work will be well received in the Python community! I do have some questions about the downsides. In the first item you mention extra complexity of the loading. Is that meant from the user perspective as with this change downstream libraries would need to manage multiple downloads instead of one? Or is there any extra complexity that will be new for maintainers on our side (or is this covered by item 3?). If it is the former I think that is a tradeoff between having everything in one wheel and having the option to reduce size. And as already mentioned, the community seems willing to take on the extra complexity in order to save on wheel sizes. If it is the latter, I am hoping the extra burden on the release managers will not be big. The conda and pip mix seems ok if conda is already including a pyarrow package one would like to install with pip. What if one would like to install a package with pip that is not yet included in the conda environment? Does that work? Will have a look at the PR and try to help in the discussion as much as I can. Best, Alenka V V pet., 25. sep. 2026 ob 14:16 je oseba Raúl Cumplido <[email protected]> napisala: > Hi, > > For several years there has been an ongoing effort in order to break > down the Arrow C++ library into smaller modules to break monolithic > installations. [1][2][3] > Currently we have the following libraries (Arrow Core, Arrow Compute, > Arrow Acero, Arrow Dataset, Arrow Cuda, Arrow Gandiva, Arrow > Substrait, Arrow Flight, Arrow Flight SQL, Arrow Flight SQL ODBC, > Arrow S3, Parquet). > > There's still work to do in order to move GCS and Azure out of the > core libarrow, similar to the work we've done to move S3 and the AWS > SDK. > > There was a lot of work also on the conda recipes in order to be able > to install those separately and provide PyArrow installations where > those can also be installed separately. [4] > > PyArrow wheel size on PyPI has been a topic of discussion both as a > maintainer's burden and as a user problem. [5][6][7] > > For context, a release today is ~2.1 GB on wheels size. 49 wheels 7 > Python ABIs * 7 platforms. > > There have been several investigations in order to break down the > installation into several packages. [8][9] > > I have started working on a PoC to have a multi-wheel installation > [10]. This is green on CI, built wheels and tested for Manylinux, > Musl, macOS, Windows both free-threaded and non-free-threaded. > > My approach has been to follow on what we did for conda, create > optional side-wheels that only contain the shared library for the > specific library (libarrow_s3 on the test case). > > The approach is, build PyArrow with everything turned on so all Cython > extensions get built then bundle everything but the specific shared > library that goes into a different wheel (libarrow_s3). That one is > then bundled on a specific pyarrow-s3 wheel. > > The new wheel pins the exact same version of PyArrow and PyArrow wheel > declares the new optional dependency with the exact same version. > > So the following: > - pip install pyarrow > Would install only core (no libarrow_s3.so/dll/dylib). > - pip install pyarrow[s3] > Would install also pyarrow-s3. > - pip install pyarrow-s3 > Would pull PyArrow core wheel too. > > We probably want `pip install pyarrow` to also install `pyarrow-s3` > initially when we do the split and make it optional once the wheel > split is battle tested. The [s3] extra should be declared from day one > even if it's required so downstream projects can write pyarrow[s3] > where necessary. > I've adapted the import mechanism to either import pyarrow-s3 which > loads libarrow_s3 globally or fails with a message in the following > lines: > > ``` > >>> import pyarrow.fs > >>> pyarrow.fs.S3FileSystem > ImportError: The pyarrow installation is not built with support for > 'S3FileSystem' > (libarrow_s3.so.2600: cannot open shared object file: No such file or > directory. > If pyarrow was installed from PyPI, S3 support is provided by the > separate 'pyarrow-s3' package) > ``` > > We do not name mangle libarrow or any of the currently built > libraries, this is why an effort like consolidatewheels [9] is not > required for them. > > Splitting wheels as a wrapper for a shared library is something that I > took inspiration from Openblas [11][12] and has been proven on runtime > loading of NVIDIA libraries[13][14] on pytorch [15][16]. > > This PoC reduces the Linux wheels size by ~5.5MB/~6MB (the > libarrow_s3.so size). macOS and Windows reductions are smaller due to > AWS SDK footprint being smaller on those ~1.5MB/~2MB. > > Some of those side wheels could be cheap because they don't rely on > Python ABI (at least for filesystems) so 7 wheels instead of 49. This > alone doesn't solve the PyPI quota problem (~9% wheel reduction) but > it compounds with the abi3 efforts [17]. > > The idea is to further reduce it by moving the rest of filesystems, > Azure, GCS to their own wheels but the next candidate I'd try is > flight as it would grant the maximum benefit as the largest optional > library. > > Some of the downsides I see from this: > - Managing the extra complexity of library loading from multiple wheels. > - Versions must match, this might make dependency management on some > projects with several dependencies pulling Arrow much more complex. > - Would require managing multiple PyPI projects with the release > management costs associated (not a big issue but worth mentioning). > - pip and conda mixing has new failure modes. Installing pyarrow-core > from conda and pip installing pyarrow-s3 reports pyarrow requirement > as already satisfied from conda and conda's pyarrow already contains > s3. > > This email is to bring attention to this effort and help with design > considerations. Please, come to the PoC PR [10] and help shape this. > > Regards, > Raúl > > [1] https://github.com/apache/arrow/issues/49399 > [2] https://github.com/apache/arrow/issues/25025 > [3] https://github.com/apache/arrow/issues/15280 > [4] https://github.com/conda-forge/arrow-cpp-feedstock/issues/1035 > [5] https://github.com/pypi/support/issues/4409 > [6] https://github.com/pypi/support/issues/10331 > [7] > https://github.com/pandas-dev/pandas/issues/54466#issuecomment-1674229577 > [8] https://github.com/apache/arrow/issues/24688 > [9] https://github.com/amol-/consolidatewheels > [10] https://github.com/apache/arrow/pull/51459 > [11] https://pypi.org/project/scipy-openblas64/ > [12] > https://github.com/MacPython/openblas-libs/blob/2b6df42890977bb66d203ab1f2c85c7879a36f7a/local/scipy_openblas64/__init__.py > [13] https://pypi.org/project/nvidia-cublas-cu12/ > [14] https://pypi.org/project/nvidia-cudnn-cu12/ > [15] > https://github.com/pytorch/pytorch/blob/48b7fc30407ee1818b59718732e60968e2d0ca96/torch/__init__.py#L343-L358 > [16] > https://github.com/pytorch/pytorch/blob/48b7fc30407ee1818b59718732e60968e2d0ca96/torch/__init__.py#L402 > [17] https://github.com/apache/arrow/issues/50398 >
