| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
Seed the project with a minimal end-to-end JNI binding from the JVM to Apache DataFusion, plus the build, format, and license-check tooling needed for ongoing contribution. Java surface (org.apache.datafusion): - SessionContext: AutoCloseable session, sql(String) returning a lazy DataFrame, registerParquet(String, String) for registering local Parquet files as SQL tables. - DataFrame: AutoCloseable, collect(BufferAllocator) executes the plan and returns result batches as an Arrow ArrowReader via the Arrow C Data Interface. collect() consumes the DataFrame; close() releases the native plan if never collected. Native side (native/, crate datafusion-jni): - JNI entry points for SessionContext create/close/registerParquet/ createDataFrame and DataFrame collect/close. - Results are exported as FFI_ArrowArrayStream so the JVM reads batches without per-row JNI crossings or row-by-row copies. Build and contributor tooling: - pom.xml with Maven wrapper, JUnit 5, Arrow 19, JDK 17 toolchain. - apache-rat-plugin (license-header check) and spotless-maven-plugin (google-java-format) both bound to the verify phase. - Makefile targets for native build, JVM build, test, clean, and TPC-H SF1 test data generation via tpchgen-cli. - GitHub Actions workflow running spotless:check and cargo fmt --check on push and pull_request to main.
Co-authored-by: Oleks V <comphead@users.noreply.github.com>
|
|
||
| public static synchronized void loadLibrary() { | ||
| if (!loaded) { | ||
| System.loadLibrary("datafusion_jni"); |
There was a problem hiding this comment.
lets make the string as constant?
Sorry, something went wrong.
There was a problem hiding this comment.
updated
Sorry, something went wrong.
|
Note that no GitHub workflows will run until after this PR is merged. Also we cannot create any issues until this PR is merged because it contains the .asf.yaml change to enable GitHub issues. |
Sorry, something went wrong.
Re-enable datafusion's default features (parquet, sql) and add arrow dependency with the ffi feature so FFI_ArrowArrayStream, ctx.sql, and register_parquet compile again.
|
|
||
| private static native long createSessionContext(); | ||
|
|
||
| private static native long createDataFrame(long handle, String sql); |
There was a problem hiding this comment.
maybe we can specify what kind of handle expected in methods? like sessionHandle?
Sorry, something went wrong.
There was a problem hiding this comment.
Thanks @andygrove, you actually dont need an approve to merge the first PR :)
Sorry, something went wrong.
|
🎉 |
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Summary
Seed the project with a minimal end-to-end JNI binding from the JVM to Apache DataFusion, plus the build, format, and license-check tooling needed for ongoing contribution.
What is in this PR
Java surface (org.apache.datafusion)
Native side (native/, crate datafusion-jni)
Build and contributor tooling
Project status
This is the first code drop into a brand-new repository. The README labels the project as early development: the API is small and will change without notice, and there is no published release.
A Roadmap section in the README outlines near-term priorities: session configuration, full SessionContext/DataFrame API parity with the Rust side, JVM-side plan construction via DataFusion's Protobuf representation, and Java-defined vectorized expressions over Arrow.
Verification
Locally, on this branch:
The optional TPC-H integration test runs after make tpch-data (requires tpchgen-cli); it reads lineitem.parquet via registerParquet and asserts SELECT COUNT(*) returns 6,001,215.