| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
FetchedPart decoded a hollow part's single parquet row group into one arrow batch and kept it alive until the last row was emitted. With `persist_part_decode_batch_rows` set, readers that go through FetchedPart (persist_source via FetchedBlob::parse, Listen, snapshot_and_stream) keep the encoded bytes and decode, validate, and normalize at most that many rows at a time, so one decoded batch per part is resident. Consolidation via the peek stash carries across batch boundaries. Inline parts and EncodedPart consumers (compaction, consolidating iteration, inspect) still decode whole parts. The flag defaults to 0, which keeps whole-part decoding, and CI runs with 16384. CPU-296 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Decode whole parts through `BlobTraceBatchPartReader`, so the format dispatch exists once. Pick whole or batched decoding of a hollow part in `PartSource::from_hollow_blob`, keep a `FetchedPart`'s current batch and its cursor in one `DecodedRows`, and read all `BatchFetcherConfig` values at fetch time. Add a test for batched decoding of a part with a rewritten description. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| Back | FazBrowse Home | New Git URL |
FetchedPart decodes a hollow part's single parquet row group into one arrow batch and keeps it alive until the last row is emitted. During hydration, each worker therefore holds a whole decoded part (persist targets 128 MiB encoded parts), and more than one when parts queue in persist_source.
With the new dyncfg persist_part_decode_batch_rows set to a positive value, readers that go through FetchedPart keep the encoded bytes and decode, validate, and normalize at most that many rows at a time. These readers are persist_source (via FetchedBlob::parse), Listen, and snapshot_and_stream. Consolidation through the peek stash carries across batch boundaries. Inline parts and EncodedPart consumers (compaction, consolidating iteration, inspect) still decode whole parts. The flag defaults to 0, which keeps today's behavior, and CI runs with 16384 (plus 0 and 7 as variants).
Whole-part decoding now goes through the same BlobTraceBatchPartReader with a batch size of the part's row count, so both paths share one format dispatch. As a result, a part with no rows decodes to empty updates instead of failing.
Two new tests, fetch::tests::batched_part_decode and fetch::tests::batched_part_decode_ts_rewrite, read one 10-row hollow part with batch sizes 0, 1, 3, 4, 10, and 16384 through both read paths. They cover duplicates that merge across a batch boundary and a pair that cancels across one. The second test reads a part whose description was rewritten, so every batch is truncated and has its timestamps rewritten.
Measured on a TPC-H SF100 lineitem index hydrating on an 8 GiB, 8-worker replica (r8gd.8xlarge, file-backed buffer pool, lgalloc off, fetch semaphore at 0.1): whole-part decode was OOM-killed, and 16384-row batches hydrated in 303 s. In heap profiles, the roughly 1.1 GB of parquet decode output no longer appears. Details are in CPU-296.
CPU-296
🤖 Generated with Claude Code