| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
|
Here is the summary of changes. You are about to add 2 region tags.
This comment is generated by snippet-bot.
|
Sorry, something went wrong.
There was a problem hiding this comment.
This pull request introduces two new code snippets and their corresponding tests to demonstrate reading BigQuery query results in Apache Arrow format. Specifically, query_and_wait_arrow.py uses the query_and_wait method to fetch Arrow RecordBatches directly, while read_rows_query_job.py initiates a query job and streams the results using the BigQuery Storage Read API. A critical issue was identified in read_rows_query_job.py where the query job is started asynchronously, but the code immediately attempts to read from the stream without waiting for the job to complete. It is recommended to call job.result() to ensure the query finishes before reading.
Sorry, something went wrong.
|
From the Kokoro System Tests presubmit, =================================== FAILURES ===================================
__________________________ test_query_and_wait_arrow ___________________________
project_id = 'precise-truck-742'
def test_query_and_wait_arrow(project_id: str):
> batches = query_and_wait_arrow.query_and_wait_arrow(project_id=project_id)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
samples/snippets/query_and_wait_arrow_test.py:21:
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _
project_id = 'precise-truck-742'
def query_and_wait_arrow(
project_id: Optional[str] = None,
) -> Iterable[pyarrow.RecordBatch]:
"""Queries BigQuery and returns results as an iterable of Apache Arrow RecordBatches.
Args:
project_id (Optional[str]): The Google Cloud project ID to bill for the query.
If not specified, the project is inferred from the environment.
Returns:
Iterable[pyarrow.RecordBatch]: An iterable of Apache Arrow RecordBatch objects.
"""
# Initialize a BigQuery client.
client = bigquery.Client(project=project_id) if project_id else bigquery.Client()
query = """
SELECT name, number, state
FROM `bigquery-public-data.usa_names.usa_1910_current`
LIMIT 100000
"""
# Run the query and wait for results returned directly in Arrow format
# compressed with LZ4_FRAME.
results = client.query_and_wait(
query,
> query_results_format=enums.QueryResultsFormat.ARROW,
^^^^^^^^^^^^^^^^^^^^^^^^
compression_codec=enums.QueryResultsCompressionCodec.LZ4_FRAME,
)
E AttributeError: module 'google.cloud.bigquery.enums' has no attribute 'QueryResultsFormat'
samples/snippets/query_and_wait_arrow.py:48: AttributeError
- generated xml file: /tmpfs/src/github/google-cloud-python/packages/google-cloud-bigquery-storage/system_3.12_sponge_log.xml -
=========================== short test summary info ============================
FAILED samples/snippets/query_and_wait_arrow_test.py::test_query_and_wait_arrow
========================= 1 failed, 5 passed in 30.39s =========================
|
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Description
Adds documentation code snippets and system tests demonstrating high-performance query result retrieval in Apache Arrow format with LZ4 compression using the BigQuery Storage API:
query_and_wait() with Arrow format & LZ4 frame compression (query_and_wait_arrow.py):
Direct read_rows on query job default stream (read_rows_query_job.py):
Tests & Dependencies:
Follow-up to #18027
Related to #18047
Checklist