| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
Mirror the parquet pattern (apache#18, apache#19): add CsvReadOptions builder, four SessionContext methods (registerCsv x2, readCsv x2), and JNI plumbing in native/src/csv.rs. Builder exposes the subset of DataFusion's Rust CsvReadOptions that has a parquet analog already on the Java side -- header, delimiter, quote, escape, terminator, comment, newlinesInValues, schemaInferMaxRecords, fileExtension, fileCompressionType, and explicit Arrow schema. Tests cover: option-builder fluent API, header-inferred schema with SQL round-trip, explicit-schema header-less file with custom delimiter, and a custom file extension with a tab delimiter. table_partition_cols, file_sort_order, null_regex and truncated_rows are intentionally deferred -- they have no parquet-side counterpart yet on the Java side.
|
@andygrove glad to see we have a official java binding repo, I'd like to do some contributions. |
Sorry, something went wrong.
There was a problem hiding this comment.
LGTM. Thanks @LantaoJin
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Which issue does this PR close?
Summary
Mirror the parquet pattern (#18, #19): add CsvReadOptions builder, four SessionContext methods (registerCsv x2, readCsv x2), and JNI plumbing in native/src/csv.rs. Builder exposes the subset of DataFusion's Rust CsvReadOptions that has a parquet analog already on the Java side -- header, delimiter, quote, escape, terminator, comment, newlinesInValues, schemaInferMaxRecords, fileExtension, fileCompressionType, and explicit Arrow schema.
Tests cover: option-builder fluent API, header-inferred schema with SQL round-trip, explicit-schema header-less file with custom delimiter, and a custom file extension with a tab delimiter.
table_partition_cols, file_sort_order, null_regex and truncated_rows are intentionally deferred -- they have no parquet-side counterpart yet on the Java side.
Changes
datafusion::prelude::CsvReadOptions that has a Parquet analog already
exposed in this repo: header, delimiter, quote, escape, terminator,
comment, newlinesInValues, schemaInferMaxRecords, fileExtension,
fileCompressionType, and explicit Arrow schema.
matching the shape of registerParquet / readParquet.
schema-IPC-serialization helper already used by the Parquet path.
tablePartitionCols, fileSortOrder, nullRegex, and truncatedRows are intentionally deferred to a follow-up; they have no Parquet-side counterpart yet on the Java side and were called out in the issue as out-of-scope.
Test plan