PR #28 established protobuf-over-JNI as the transport for SessionContext configuration. This PR applies the same pattern to the CSV and Parquet read paths.
Before this change, registering or reading a CSV file passed 14–16 raw JNI arguments — booleans, byte values, nullable-encoded as xxx_set / xxx_value pairs, -1L sentinels for "unset" longs, and FileCompressionType shipped as its name() string. Parquet had the same shape with 7–9 args.
After: each call takes a single serialized CsvReadOptionsProto / ParquetReadOptionsProto byte array plus an optional Arrow-IPC schema byte array. Nullability, enums, and field evolution are now native to the wire format. The contributor guide documents the proto-over-JNI convention so future structured JNI calls follow the same pattern.
What changes are included in this PR?
New proto/csv_read_options.proto and proto/parquet_read_options.proto, mirroring the structure of session_options.proto. FileCompressionType is a proto3 enum with prefixed values and a _UNSPECIFIED = 0 sentinel.
CsvReadOptions.toBytes() and ParquetReadOptions.toBytes() serialize the Java options through the generated builders.
with_csv_options and with_parquet_options on the Rust side decode the proto via prost and fold the fields into DataFusion's option structs. The Unspecified compression arm returns an error rather than silently defaulting.
Four JNI methods collapse to 4 or 5 arguments each: (handle, [name,] path, byte[] optionsBytes, byte[] schemaIpcBytesOrNull).
New native/src/schema.rs::decode_optional_schema replaces two copies of identical Arrow-IPC schema-decode logic.
Renamed Rust module session_options → proto_gen since the single generated file now contains the types for all three protos (they share package datafusion_java;).
New contributor-guide section Passing structured options across the JNI boundary documents the convention, including proto3 enum-prefix and _UNSPECIFIED = 0 requirements.
The public Java API is unchanged: every public setter on CsvReadOptions / ParquetReadOptions and every register* / read* method on SessionContext keeps the same signature.
Are these changes tested?
Yes:
CsvReadOptionsTest (4 tests) and ParquetReadOptionsTest (3 tests) round-trip through toBytes() / Proto.parseFrom(...), verifying every field, default presence/absence, and all five FileCompressionType values.
The existing SessionContextCsvTest and SessionContextParquetOptionsTest continue to exercise the public API end-to-end through JNI without modification — strong evidence that the new wire format reaches the Rust side correctly.
Full ./mvnw test passes (49 run, 0 failed, 12 skipped — skips are pre-existing tpch-data integration tests).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
No tracking issue — follow-up to #28.
Rationale for this change
PR #28 established protobuf-over-JNI as the transport for SessionContext configuration. This PR applies the same pattern to the CSV and Parquet read paths.
Before this change, registering or reading a CSV file passed 14–16 raw JNI arguments — booleans, byte values, nullable-encoded as xxx_set / xxx_value pairs, -1L sentinels for "unset" longs, and FileCompressionType shipped as its name() string. Parquet had the same shape with 7–9 args.
After: each call takes a single serialized CsvReadOptionsProto / ParquetReadOptionsProto byte array plus an optional Arrow-IPC schema byte array. Nullability, enums, and field evolution are now native to the wire format. The contributor guide documents the proto-over-JNI convention so future structured JNI calls follow the same pattern.
What changes are included in this PR?
The public Java API is unchanged: every public setter on CsvReadOptions / ParquetReadOptions and every register* / read* method on SessionContext keeps the same signature.
Are these changes tested?
Yes:
Are there any user-facing changes?
No. Public Java API is unchanged.