FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Test issue 2271 by katosh · Pull Request #2277 · scverse/anndata · GitHub

Test issue 2271 - #2277

Closed
katosh wants to merge 18 commits into
scverse:mainfrom
settylab:test-issue-2271
Closed

katosh wants to merge 18 commits into
scverse:mainfrom
settylab:test-issue-2271

Conversation

katosh commented Jan 5, 2026

Copy link
Copy Markdown
Contributor

Add regression test for issue #2271 (expected to fail)

This PR adds a regression test that demonstrates the bug reported in #2271.

Purpose

This PR is intentionally expected to fail CI to confirm that issue #2271 exists on main and not just on my local machine. The fix is in PR #2272.

The Bug

read_lazy() returns obs_names/var_names as byte-string representations ("b'cell_A'") instead of properly decoded strings ("cell_A").

Test Output on main

FAILED tests/lazy/test_read.py::test_nullable_string_index_decoding

    assert obs_names == expected_obs
E   assert ["b'cell_A'", "b'cell_B'", ...] == ['cell_A', 'cell_B', ...]
E     At index 0 diff: "b'cell_A'" != 'cell_A'

Related

katosh commented Jan 5, 2026 •
edited
Loading

Copy link
Copy Markdown
Contributor Author

Note on version requirements

The regression test initially passed in CI due to version dependencies. This bug only manifests with pandas 3.0+ AND numpy < 2.2.3.

See failed test in https://github.com/scverse/anndata/actions/runs/20719835864/job/59480330959

Why these versions matter

  • pandas 3.0: anndata has been actively working to support pandas 3.0 (see feat: Support pandas 3.0 upcoming changes #2133: "Support pandas 3.0 upcoming changes").
  • numpy < 2.2.3: NumPy 2.2.3 fixed incorrect bytes-to-string coercion (numpy#28282), which accidentally masks this bug. Before 2.2.3, bytes like b'cell_A' were incorrectly converted to "b'cell_A'" instead of "cell_A".

Version matrix

pandas numpy Bug manifests?
2.2.x any No
3.0.x < 2.2.3 Yes
3.0.x >= 2.2.3 No (masked by numpy fix)

Why the fix in #2272 is still needed

Even though numpy 2.2.3+ masks this bug, the fix in #2272 may still be the correct solution because:

  1. Users may still be on numpy < 2.2.3
  2. The anndata code should properly decode bytes rather than relying on numpy's coercion behavior
  3. Using .asstr() explicitly handles HDF5 byte strings correctly

Reproducing the bug

To demonstrate the bug exists on main, the test must run with:

  • pandas >= 3.0
  • numpy < 2.2.3

katosh closed this Jan 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. Learn more about bidirectional Unicode characters
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant


Back | FazBrowse Home | New Git URL