| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
Closes apache#37 Mirrors the existing parquet/csv/json reader pattern for Arrow IPC files. Adds: - proto/arrow_read_options.proto with the ArrowReadOptionsProto message (file_extension only; explicit Arrow schema rides on the existing schema-IPC byte channel rather than this proto, matching the other formats) - ArrowReadOptions Java builder with fileExtension default ".arrow" and schema(Schema) - SessionContext.registerArrow(name, path[, options]) and readArrow(path[, options]) overloads with null-argument validation per Andy's apache#47 review feedback - native/src/arrow.rs JNI module that decodes the proto and dispatches to upstream SessionContext::register_arrow / read_arrow Note on ArrowReadOptions construction: upstream's ArrowReadOptions exposes file_extension as a public field (not a builder setter), unlike the other format options. The native side uses struct-update syntax to set it without tripping clippy's field_reassign_with_default lint. Tests cover proto round-trip, schema-by-reference, register/read on a fixture written by arrow-vector's ArrowFileWriter (the canonical Arrow IPC file format DataFusion's source supports), custom file extension, explicit Arrow schema, and null-argument rejection on both register and read. Out of scope: tablePartitionCols (no parquet/csv/json analog on the Java side yet). Arrow IPC carries body compression inside the file format itself, so unlike CSV and NDJSON there is no FileCompressionType on this options class.
There was a problem hiding this comment.
LGTM. Thanks @LantaoJin
Sorry, something went wrong.
# Conflicts: # core/src/main/java/org/apache/datafusion/SessionContext.java # native/build.rs
| Back | FazBrowse Home | New Git URL |
Which issue does this PR close?
Rationale for this change
DataFusion 53.x supports Arrow IPC files via SessionContext::register_arrow / read_arrow, but the Java bindings only expose Parquet, CSV, and (in #47) NDJSON. Since JVM results already come back as Arrow batches via the C Data Interface, an Arrow IPC reader on the Java side closes the natural round-trip: Java callers can write Arrow IPC to disk with arrow-vector's ArrowFileWriter, then read it back through DataFusion without going through Parquet or any other intermediate format. Today they have to fall back to CREATE EXTERNAL TABLE … STORED AS ARROW via SQL, which works but bypasses the typed builder pattern.
This PR is the Java surface for the existing upstream functionality. Issue #37 tracks it; the implementation follows the same proto-over-JNI pattern as #47 (NDJSON), #29 (the CSV/Parquet refactor), and the merged CSV/Parquet readers.
What changes are included in this PR?
Out of scope (for follow-ups):
Are these changes tested?
Yes, 9 new tests across ArrowReadOptionsTest and SessionContextArrowTest.
Are there any user-facing changes?
Yes, purely additive. New public API:
The new org.apache.datafusion.protobuf.ArrowReadOptionsProto generated class is also exposed via the protobuf-Java output, consistent with how CsvReadOptionsProto, NdJsonReadOptionsProto, and ParquetReadOptionsProto are exposed. No API removals, no deprecations, no behavior change for existing callers.