FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Redis online store: HGETALL for wide reads, and an opt-in GIL-free client (valkey-glide) for get_online_features · Issue #6856 · feast-dev/feast · GitHub

Repository navigation

Redis online store: HGETALL for wide reads, and an opt-in GIL-free client (valkey-glide) for get_online_features #6856

Description

Is your feature request related to a problem? Please describe.

The Python get_online_features path against the Redis online store is slow for wide reads, and the cost is on the client, not the server. Our shape is 4 FeatureViews, ~90 fields, 350 to 500 entities per call, served from a FastAPI process that does other work at the same time.

Two separate problems show up:

  1. HMGET on listpack-encoded hashes scans once per requested field. Redis and Valkey keep a hash as a listpack up to hash-max-listpack-entries (default 128 on Redis 7 and Valkey 8). Our hashes have ~120 fields, so every HMGET of ~90 fields does ~90 linear scans. Measured on Valkey 8 with INFO commandstats, same 500 keys:

    command server time per command server time per 500-entity read
    HMGET, 94 named fields 25.8 us 12.9 ms
    HGETALL (123 fields) 5.6 us 2.8 ms

    HGETALL returns about a third more bytes and is still 4.6x cheaper for the server. This is the optimization HGETALL optimization for Redis retrieval in Python #3337 asked for in 2022 (the Java server switches to HGETALL above 50 features); that issue was closed by the stale bot without a change.

  2. The Python client work holds the GIL in thousands of short bursts per read. redis-py packs every command, hiredis parses every reply, and each socket receive releases and reacquires the GIL. Then the SDK turns every ValueProto blob into a Python object. In a process with other threads running Python, every one of those reacquires waits out the switch interval, so the read time depends on how busy the rest of the process is.

    Same machine, same data, Feast 0.66 (which already has the single-pipeline read from feat: Addresses performance issues in the Redis online store #6337), 500 entities, ~90 fields across 4 views, p50 of 40 runs:

    read path alone while another thread runs pure-Python work
    get_online_features (redis-py + hiredis) 140 ms 946 ms
    same field plan, one HGETALL per entity through valkey-glide (Rust core), decode in Python 43 ms 160 ms

    The 6x gap under contention is the part that matters in production. It does not show up in a single-threaded benchmark.

Describe the solution you'd like

Two changes to RedisOnlineStore, independent of each other:

  1. In _read_features_per_fv / the batched get_online_features, issue HGETALL instead of HMGET when the number of requested hash fields crosses a threshold (the Java server uses 50), and pick the requested fields out of the reply. Everything else stays the same: same _redis_key, same _mmh3 field names, same _ts:<view> presence check.

  2. An opt-in client for the Redis online store that runs the pipeline off the GIL. valkey-glide (valkey-glide-sync on PyPI) is an official Valkey client with a Rust core; a non-atomic Batch is one FFI call, so the whole fetch runs with the GIL released and returns once. It speaks RESP to Redis and Valkey and supports TLS and cluster mode. A connection_string-compatible client: glide option on RedisOnlineStoreConfig would keep feature_store.yaml as the single place the store is configured.

Describe alternatives you've considered

  • The Go feature server. It needs the file registry (we use the SQL registry) and a separate Python transformation service for on-demand views, and the Python wheel does not ship the Go library.
  • get_online_features_async. Same Python parsing on the same GIL; measured within a few percent of the sync path.
  • Lowering hash-max-listpack-entries on the server so hashes become hashtables. That does remove most of the HMGET cost (7.4 ms to 1.1 ms server time in our test) but costs ~40% more memory and is an operator-side change; HGETALL gets most of the same win from the client.
  • Reading OnlineResponse.proto directly instead of to_dict(). Worth about 20%, but the redis-py fetch is still the part that stalls under contention.

We have both changes running in a fork of the read path and are happy to contribute either as a PR if there is interest in the direction.

Additional context

Related: #3337 (HGETALL request, closed stale), #4711 and #6337 (single pipeline across feature views, merged; the numbers above are with that in place), #3649 (registry from_proto overhead, a separate cost).

Activity

  1. patelchaitany commented on Sep 23, 2026

    Contributor

    Hey @chandlerok, thanks for the writeup. I reproduced the GIL half closely (181 ms idle, 749 ms with one background thread, 8438 ms with eight), so that one clearly holds up.

    I couldn't reproduce the HGETALL half I measured 1.3x–2.0x server-side rather than 4.6x, and end-to-end it's usually a net loss, since Feast splits reads per feature view so each HMGET asks for ~22 fields rather than 90. Could you share how you measured the 5.6 vs 25.8 µs — fields per HMGET, hash width, and whether it was raw Redis or through get_online_features?

    Please do take the valkey-glide PR if you'd like it you have the production context for it, and I'd be glad to help however is most useful. On HGETALL, would it be reasonable to hold off until we've worked out where our measurements diverge? Entirely your call, and happy to go whichever way you prefer.

  2. chandlerok commented on Sep 25, 2026

    ContributorAuthor

    hey @patelchaitany! Thanks for taking a look!

    I'm guessing you probably measured a narrower shape than I did. My number is for one 123-field listpack hash:

    • HMGET asks for ~94 fields
    • HGETALL reads the whole hash once

    If you benchmark a single ~20-field HMGET, or swap each per-view HMGET for its own HGETALL, the win mostly disappears because you either aren’t measuring the wide scan or you’re returning the same full hash multiple times.

    To reproduce: seed 500 listpack hashes with ~123 fields, reset INFO commandstats, run 500 HMGETs for ~94 fields, reset again, run 500 HGETALLs, then compare cmdstat_hmget.usec_per_call vs cmdstat_hgetall.usec_per_call.

  3. chandlerok commented on Sep 25, 2026

    ContributorAuthor

    Here's the exact setup behind those numbers, plus the two things that decide whether you see a large ratio or a small one.

    Your three questions: raw Redis (not through get_online_features), 94 named fields per HMGET, 123-field hashes.

    The setup

    500 hashes x 123 fields, then 500 HMGETs of 94 fields, then 500 HGETALLs, reading cmdstat_*.usec_per_call after each phase. This script is runnable as-is.

    It seeds into database 9 and calls FLUSHDB on that database, so point DB at a scratch database before running it:

    import random
    
    import redis
    
    HOST, DB = "127.0.0.1", 9
    N_KEYS, HASH_FIELDS, HMGET_FIELDS, ROUNDS = 500, 123, 94, 500
    
    client = redis.Redis(host=HOST, port=6379, db=DB)
    client.flushdb()
    random.seed(0)
    
    # Values must stay under hash-max-listpack-value (64 bytes by default) or the
    # hash converts to a hashtable and the effect disappears.
    value = "v" * 4
    fields = [f"f{i}" for i in range(HASH_FIELDS)]
    pipe = client.pipeline(transaction=False)
    for k in range(N_KEYS):
        pipe.hset(f"repro:{k}", mapping={f: value for f in fields})
    pipe.execute()
    
    print("encoding:", client.object("encoding", "repro:0"))  # must print listpack
    
    # Real Feast field names are mmh3 hashes, so their position in the listpack is
    # arbitrary. Requesting the first N is the cheapest case and understates this.
    requested = random.sample(fields, HMGET_FIELDS)
    
    
    def phase(queue, stat):
        client.execute_command("CONFIG", "RESETSTAT")
        pipe = client.pipeline(transaction=False)
        for k in range(N_KEYS):
            queue(pipe, k)
        pipe.execute()
        stats = client.info("commandstats")[f"cmdstat_{stat}"]
        assert stats["calls"] == ROUNDS, "another client is issuing this command"
        return stats["usec_per_call"]
    
    
    hmget_us = phase(lambda pipe, k: pipe.hmget(f"repro:{k}", requested), "hmget")
    hgetall_us = phase(lambda pipe, k: pipe.hgetall(f"repro:{k}"), "hgetall")
    
    print(f"HMGET    94 of 123: {hmget_us:6.1f} us/call")
    print(f"HGETALL       123: {hgetall_us:6.1f} us/call")
    print(f"ratio: {hmget_us / hgetall_us:.2f}x")

    The assert stats["calls"] == ROUNDS matters: if anything else on the instance issues the same command, INFO commandstats is server-wide and the number is polluted. Also worth checking the limits on your instance, since they gate everything below:

    CONFIG GET hash-max-listpack-entries hash-max-listpack-value
    

    On the Valkey 9 box I measured on that returned entries=512, value=64, so a 123-field hash stays listpack either way.

    What actually decides the ratio

    HMGET over a listpack hash does one linear scan from the front per requested field, so 94 fields is ~94 scans. HGETALL is a single scan. Two conditions control how large that gap is.

    1. The hash has to still be listpack-encoded. If any value exceeds hash-max-listpack-value (64 bytes by default), that entity's hash converts to a hashtable, HMGET becomes O(1) per field, and HGETALL still reads everything, so the advantage inverts. Median of 3 runs, call counts verified at exactly 500:

    value bytes encoding fields requested HMGET HGETALL ratio
    4 listpack first 94 of 123 45.6 us 10.1 us 4.36x
    4 listpack random 94 of 123 52.1 us 10.1 us 5.17x
    4 listpack last 94 of 123 63.4 us 10.0 us 6.36x
    60 listpack random 94 of 123 57.0 us 10.4 us 5.41x
    100 hashtable random 94 of 123 10.1 us 11.8 us 0.85x

    The absolute microseconds are hardware-dependent. The ratio is the portable part, and it is the part that reproduces.

    2. How deep in the listpack the requested fields sit. The scan starts at the front, so the first 94 fields is the cheapest case (4.4x) and the last 94 is the most expensive (6.4x). Feast's field names are mmh3 hashes, so their positions are effectively arbitrary, which is the middle row (~5.2x).

    The shape that reproduces ~4.6x is therefore: listpack hash, ~123 fields, ~94 fields requested per HMGET scattered through the hash, measured raw via INFO commandstats.

    Why you likely measured 1.3-2.0x

    Three things, all of which pull the number toward 1x:

    • Mixed encodings within one dataset. Encoding is per hash, so if some entities carry a ValueProto blob that pushes a value past 64 bytes, those entities become hashtables while the rest stay listpacks. A real dataset blends both and lands between the 0.85x and 5.4x rows above. This is my best guess at what you hit, and it is easy to check: OBJECT ENCODING on a sample of entity keys will show listpack vs hashtable directly.
    • Where the requested fields sit in the listpack. Fields near the front scan cheaply.
    • Measuring through get_online_features. Feast splits reads per feature view, so each HMGET asks for ~22 fields rather than ~94. That is a genuinely cheaper shape, and it is why the end-to-end number can be a net loss even when the per-command server time improves.

    One caveat on my own numbers: the table above is from a different box than the figures in my original comment, so the absolute microseconds differ. I would compare the ratio.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL