FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

ci: Use uv torch backend for CPU installs by franciscojavierarceo · Pull Request #6588 · feast-dev/feast · GitHub

Repository navigation

ci: Use uv torch backend for CPU installs - #6588

Merged
franciscojavierarceo merged 1 commit into
feast-dev:masterfrom
franciscojavierarceo:codex/fix-ci-torch-index
Jul 8, 2026
Merged

franciscojavierarceo merged 1 commit into
feast-dev:masterfrom
franciscojavierarceo:codex/fix-ci-torch-index

Conversation

Copy link
Copy Markdown
Member

Summary

  • replace the Linux CI PyTorch CPU extra index install with uv's targeted --torch-backend cpu option
  • keep dependency resolution on PyPI for unrelated packages like databricks-sdk

Root cause

The failing unit-test-python (3.10, ubuntu-latest) leg on master died before pytest during make install-python-dependencies-ci. The old command used --extra-index-url https://download.pytorch.org/whl/cpu --index-strategy unsafe-best-match, which caused uv to query the PyTorch CPU index for unrelated packages. In the failing run, that made CI fetch https://download.pytorch.org/whl/cpu/databricks-sdk/, which returned HTTP 503.

Validation

  • git diff --check
  • PYTHON_VERSION=3.10 uv pip sync --dry-run --torch-backend cpu sdk/python/requirements/py3.10-ci-requirements.txt

Signed-off-by: Francisco Javier Arceo <farceo@redhat.com>
franciscojavierarceo marked this pull request as ready for review July 8, 2026 03:53
franciscojavierarceo merged commit 66547dc into feast-dev:master Jul 8, 2026
35 of 36 checks passed
patelchaitany added a commit to patelchaitany/feast that referenced this pull request Sep 10, 2026
The CI requirement files are compiled without a torch backend, so their
hashes describe PyPI wheels, but Makefile:110 syncs them on Linux with
--torch-backend cpu, which serves torch==2.13.0+cpu and
torchvision==0.28.0+cpu from download.pytorch.org. Those are different
builds with different hashes:

    Failed to download `torchvision==0.28.0+cpu`
    Hash mismatch for `torchvision==0.28.0+cpu`

The divergence has existed since feast-dev#6588 added the flag. Both the old and
the new uv redirect to the CPU index; what changed is verification. uv
0.12.11 began to "apply hashes from public-version pins to matching
local versions when no exact local-version hash is provided", so the
PyPI hashes recorded for torchvision==0.28.0 are now checked against the
+cpu wheel. Earlier releases found no hash for the local version and
skipped the check. uv is unpinned in the workflows, so CI moved from
0.12.7 to 0.12.11 and every job began failing in the install step.

Set torch-backend in [tool.uv] rather than passing it on the command
line. The setting is read by every uv pip command, so compile and sync
can no longer disagree about where torch comes from. An explicit index
with [tool.uv.sources] would not serve, because the pip interface
applies sources only at compile time and sync would then fail to find
the +cpu builds at all.

Compile the CI locks universally so that one file can serve both the
Linux and macOS runners, pinning each variant with its own hashes:

    torch==2.14.0+cpu ; sys_platform != 'darwin'
    torch==2.14.0 ; sys_platform == 'darwin'

The install-time override is then redundant and comes out, and pixi
raises its uv floor to a release that understands the setting.

Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
patelchaitany added a commit to patelchaitany/feast that referenced this pull request Sep 12, 2026
Every CI job fails in the install step:

    x Failed to download `torchvision==0.28.0+cpu`
    `-> Hash mismatch for `torchvision==0.28.0+cpu`

The CI locks are compiled without a torch backend, so their hashes
describe PyPI wheels, but `install-python-dependencies-ci` synced them on
Linux with `--torch-backend cpu`, which serves the `+cpu` builds from
download.pytorch.org. Different artifacts, different hashes.

The divergence has existed since feast-dev#6588 added the flag to sync only; what
changed is verification. uv 0.12.11 began applying hashes from public
version pins to matching local versions, so the PyPI hashes recorded for
`torchvision==0.28.0` are now checked against the `+cpu` wheel. uv is
unpinned in the workflows, so CI moved from 0.12.7 to 0.12.11 and every
job started failing.

Compile and sync now both pass `--torch-backend cpu`, so they can no
longer disagree about where torch comes from. The CI locks are compiled
`--universal` so one file serves both the Linux and macOS runners,
pinning each variant with its own hashes:

    torch==2.13.0           ; sys_platform == 'darwin'
    torch==2.13.0+cpu       ; sys_platform != 'darwin'
    torchvision==0.28.0     ; sys_platform == 'darwin'
    torchvision==0.28.0+cpu ; sys_platform != 'darwin'

Without `--universal` the lock pins an unmarked `torch==2.13.0` carrying
PyPI hashes. Linux still resolves that to `2.13.0+cpu`, because
`==2.13.0` matches the local version under PEP 440, and the download then
fails verification -- the original breakage in a different guise.

No `nvidia-*` package remains in any CI requirements file, which is what
feast-dev#6588 set out to achieve.

Recompiling the locks also exposed five dependencies with no upper bound,
which float into regressions on every recompile:

  * mcp -- 1.30 drops the `mcp-session-id` response header, so the
    feature server handshake in mcp-feature-server-runtime fails.
    98e5bca already pinned 1.29.0 in the lock files; the bound makes
    that durable instead of resettable on the next recompile.
  * ray -- 2.56+ returns its own extension dtypes from ray.data, failing
    17 type assertions in test_universal_types.py. feast declares ray
    directly only on 3.10, so the bound is repeated for 3.11+ where ray
    arrives transitively through codeflare-sdk.
  * pyarrow, mlflow -- capped at the versions 3.10 and 3.11 already lock
    to, so all three interpreters agree.

The caps hold every package at a version already present on master, and
bring py3.10 (ray 2.56.1) and py3.12 (pyarrow 24.0.0, mlflow 3.14.0) back
in line with the other interpreters. pixi.lock moves pyarrow 24.0.0 ->
21.0.0 and nothing else.

pixi raises its uv floor to 0.6.9, the first release with
`--torch-backend`. The uv already locked in infra/scripts/pixi/pixi.lock
satisfies it, so that lock is untouched.

Verified on a macOS host, so the full Linux install was not exercised
end to end:

  * `uv pip sync --dry-run --python-platform x86_64-unknown-linux-gnu
    --torch-backend cpu` against py3.11-ci-requirements.txt resolves all
    445 packages and selects torch==2.13.0+cpu and torchvision==0.28.0+cpu.
  * A real hash-checked install of the CPU torchvision wheel for the
    Linux platform, under uv 0.12.13, downloads from download.pytorch.org
    and verifies clean -- the specific artifact that was failing.
  * The same install against master's lock entry reproduces the failure,
    confirming the universal marker split is load-bearing.
  * `pixi lock --check` is clean on master and reports pyarrow as the
    only movement under the new bounds.

CI on the PR is the actual proof.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
patelchaitany added a commit to patelchaitany/feast that referenced this pull request Sep 14, 2026
Every CI job fails in the install step:

    x Failed to download `torchvision==0.28.0+cpu`
    `-> Hash mismatch for `torchvision==0.28.0+cpu`

The CI locks are compiled without a torch backend, so their hashes
describe PyPI wheels, but `install-python-dependencies-ci` synced them on
Linux with `--torch-backend cpu`, which serves the `+cpu` builds from
download.pytorch.org. Different artifacts, different hashes.

The divergence has existed since feast-dev#6588 added the flag to sync only; what
changed is verification. uv 0.12.11 began applying hashes from public
version pins to matching local versions, so the PyPI hashes recorded for
`torchvision==0.28.0` are now checked against the `+cpu` wheel. uv is
unpinned in the workflows, so CI moved from 0.12.7 to 0.12.11 and every
job started failing.

Compile and sync now both pass `--torch-backend cpu`, so they can no
longer disagree about where torch comes from. The CI locks are compiled
`--universal` so one file serves both the Linux and macOS runners,
pinning each variant with its own hashes:

    torch==2.13.0           ; sys_platform == 'darwin'
    torch==2.13.0+cpu       ; sys_platform != 'darwin'
    torchvision==0.28.0     ; sys_platform == 'darwin'
    torchvision==0.28.0+cpu ; sys_platform != 'darwin'

Without `--universal` the lock pins an unmarked `torch==2.13.0` carrying
PyPI hashes. Linux still resolves that to `2.13.0+cpu`, because
`==2.13.0` matches the local version under PEP 440, and the download then
fails verification -- the original breakage in a different guise.

No `nvidia-*` package remains in any CI requirements file, which is what
feast-dev#6588 set out to achieve.

Recompiling the locks also exposed five dependencies with no upper bound,
which float into regressions on every recompile:

  * mcp -- 1.30 drops the `mcp-session-id` response header, so the
    feature server handshake in mcp-feature-server-runtime fails.
    98e5bca already pinned 1.29.0 in the lock files; the bound makes
    that durable instead of resettable on the next recompile.
  * ray -- 2.56+ returns its own extension dtypes from ray.data, failing
    17 type assertions in test_universal_types.py. feast declares ray
    directly only on 3.10, so the bound is repeated for 3.11+ where ray
    arrives transitively through codeflare-sdk.
  * pyarrow, mlflow -- capped at the versions 3.10 and 3.11 already lock
    to, so all three interpreters agree.

The caps hold every package at a version already present on master, and
bring py3.10 (ray 2.56.1) and py3.12 (pyarrow 24.0.0, mlflow 3.14.0) back
in line with the other interpreters. pixi.lock moves pyarrow 24.0.0 ->
21.0.0 and nothing else.

pixi raises its uv floor to 0.6.9, the first release with
`--torch-backend`. The uv already locked in infra/scripts/pixi/pixi.lock
satisfies it, so that lock is untouched.

Verified on a macOS host, so the full Linux install was not exercised
end to end:

  * `uv pip sync --dry-run --python-platform x86_64-unknown-linux-gnu
    --torch-backend cpu` against py3.11-ci-requirements.txt resolves all
    445 packages and selects torch==2.13.0+cpu and torchvision==0.28.0+cpu.
  * A real hash-checked install of the CPU torchvision wheel for the
    Linux platform, under uv 0.12.13, downloads from download.pytorch.org
    and verifies clean -- the specific artifact that was failing.
  * The same install against master's lock entry reproduces the failure,
    confirming the universal marker split is load-bearing.
  * `pixi lock --check` is clean on master and reports pyarrow as the
    only movement under the new bounds.

CI on the PR is the actual proof.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
ntkathole pushed a commit that referenced this pull request Sep 15, 2026
* ci: Pin CPU-only torch and bound floating dependencies

Every CI job fails in the install step:

    x Failed to download `torchvision==0.28.0+cpu`
    `-> Hash mismatch for `torchvision==0.28.0+cpu`

The CI locks are compiled without a torch backend, so their hashes
describe PyPI wheels, but `install-python-dependencies-ci` synced them on
Linux with `--torch-backend cpu`, which serves the `+cpu` builds from
download.pytorch.org. Different artifacts, different hashes.

The divergence has existed since #6588 added the flag to sync only; what
changed is verification. uv 0.12.11 began applying hashes from public
version pins to matching local versions, so the PyPI hashes recorded for
`torchvision==0.28.0` are now checked against the `+cpu` wheel. uv is
unpinned in the workflows, so CI moved from 0.12.7 to 0.12.11 and every
job started failing.

Compile and sync now both pass `--torch-backend cpu`, so they can no
longer disagree about where torch comes from. The CI locks are compiled
`--universal` so one file serves both the Linux and macOS runners,
pinning each variant with its own hashes:

    torch==2.13.0           ; sys_platform == 'darwin'
    torch==2.13.0+cpu       ; sys_platform != 'darwin'
    torchvision==0.28.0     ; sys_platform == 'darwin'
    torchvision==0.28.0+cpu ; sys_platform != 'darwin'

Without `--universal` the lock pins an unmarked `torch==2.13.0` carrying
PyPI hashes. Linux still resolves that to `2.13.0+cpu`, because
`==2.13.0` matches the local version under PEP 440, and the download then
fails verification -- the original breakage in a different guise.

No `nvidia-*` package remains in any CI requirements file, which is what
#6588 set out to achieve.

Recompiling the locks also exposed five dependencies with no upper bound,
which float into regressions on every recompile:

  * mcp -- 1.30 drops the `mcp-session-id` response header, so the
    feature server handshake in mcp-feature-server-runtime fails.
    98e5bca already pinned 1.29.0 in the lock files; the bound makes
    that durable instead of resettable on the next recompile.
  * ray -- 2.56+ returns its own extension dtypes from ray.data, failing
    17 type assertions in test_universal_types.py. feast declares ray
    directly only on 3.10, so the bound is repeated for 3.11+ where ray
    arrives transitively through codeflare-sdk.
  * pyarrow, mlflow -- capped at the versions 3.10 and 3.11 already lock
    to, so all three interpreters agree.

The caps hold every package at a version already present on master, and
bring py3.10 (ray 2.56.1) and py3.12 (pyarrow 24.0.0, mlflow 3.14.0) back
in line with the other interpreters. pixi.lock moves pyarrow 24.0.0 ->
21.0.0 and nothing else.

pixi raises its uv floor to 0.6.9, the first release with
`--torch-backend`. The uv already locked in infra/scripts/pixi/pixi.lock
satisfies it, so that lock is untouched.

Verified on a macOS host, so the full Linux install was not exercised
end to end:

  * `uv pip sync --dry-run --python-platform x86_64-unknown-linux-gnu
    --torch-backend cpu` against py3.11-ci-requirements.txt resolves all
    445 packages and selects torch==2.13.0+cpu and torchvision==0.28.0+cpu.
  * A real hash-checked install of the CPU torchvision wheel for the
    Linux platform, under uv 0.12.13, downloads from download.pytorch.org
    and verifies clean -- the specific artifact that was failing.
  * The same install against master's lock entry reproduces the failure,
    confirming the universal marker split is load-bearing.
  * `pixi lock --check` is clean on master and reports pyarrow as the
    only movement under the new bounds.

CI on the PR is the actual proof.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>

* fix: Drop the redundant pyarrow upper bound

mlflow 3.2.0 already requires `pyarrow<22,>=4.0.0`, so the `mlflow<3.3`
bound implies the pyarrow cap wherever mlflow takes part in the
resolution. The explicit `pyarrow>=16.1.0,<22` added nothing, and it
broke test_flink_extra_constrains_shared_pyarrow_dependency, which
compares the core pyarrow spec by exact string.

Dropping it leaves all three CI requirement files byte-identical -- the
resolution is unchanged, including py3.12 coming down to pyarrow 21.0.0,
which mlflow drives rather than the removed bound.

pixi.lock shrinks to a metadata-only update: the recorded requires_dist
for the editable feast package now lists the four remaining bounds, and
no package version moves. It keeps pyarrow 24.0.0 in the pixi
environments, as master does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>

* fix: Keep numpy dtypes coming out of the Ray offline store

Ray 2.56 changed BlockAccessor.to_pandas to map Arrow types onto
pandas Arrow-backed dtypes, so RayRetrievalJob.to_df() began returning
pd.ArrowDtype columns where every other offline store returns
numpy/object. That fails 17 assertions in test_universal_types, and
more importantly changes the dtypes users get back from
get_historical_features(...).to_df() on the Ray store alone.

Ray 2.56.1 added DataContext.enable_arrow_backed_pandas_conversion as
the opt-out. Set it at each of the three initialization sites, next to
the enable_tensor_extension_casting already pinned there for the same
reason. The hasattr guard doubles as a version check: the attribute
exists only from 2.56.1, and below that the behaviour it disables does
not exist either.

Fixing the store is what lets the dependency caps come off, rather than
loosening the universal type assertions to accommodate one store:

- ray<2.56 and codeflare-sdk<0.39 are dropped; codeflare-sdk pins ray
  exactly, so the two moved in lockstep anyway.
- mlflow relaxes from <3.3 to <4, matching the major-version bound every
  other extra in the file carries. The <3.3 cap held mlflow 14 minor
  releases back on no demonstrated breakage.

The CI locks were recompiled in place with --upgrade-package limited to
those three, so all three interpreters now agree on ray 2.58.0,
codeflare-sdk 0.39.0 and mlflow 3.16.0. The torch and torchvision pins
and their hashes are untouched, and no nvidia-* package reappears.

Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>

* chore: Drop the mlflow upper bound

Review feedback on #6825: the `<4` cap adds maintenance for no benefit.

mlflow's latest release is 3.16.0, so the bound never constrained
resolution. Recompiling all three CI locks with it removed leaves them
byte-identical, and the pyarrow cap that originally motivated aligning
mlflow across interpreters is already gone.

pixi.lock only mirrors the requires-dist metadata; `pixi lock --check`
(pixi 0.75.0, the version CI pins) is clean.

Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>

---------

Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. Learn more about bidirectional Unicode characters
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant


Back | FazBrowse Home | New Git URL