This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| LocalComputeEngine | Runs on Arrow + Pandas/Polars/Dask etc., designed for light weight transformation. | ✅ | |
| SparkComputeEngine | Runs on Apache Spark, designed for large-scale distributed feature generation. | ✅ | |
| TrinoComputeEngine | Runs on Trino, designed for scalable feature generation using Trino SQL. | ✅ | [docs](../../reference/compute-engine/trino.md) |
| SnowflakeComputeEngine | Runs on Snowflake, designed for scalable feature generation using Snowflake SQL. | ✅ | |
| LambdaComputeEngine | Runs on AWS Lambda, designed for serverless feature generation. | ✅ | |
| FlinkComputeEngine | Runs on Apache Flink, designed for distributed feature generation through PyFlink Table API. | ✅ | |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Trino Compute Engine provides a distributed execution engine for batch materialization operations (`materialize` and `materialize-incremental`) and historical retrieval operations (`get_historical_features`).
It is designed to handle large-scale data processing directly on Trino clusters without moving raw data to the client machine.
### Design
The Trino Compute engine is implemented as a subclass of `feast.infra.compute_engines.base.ComputeEngine`.
The engine supports the following features:
- **Pushdown SQL execution**: Compiles feature pipeline operations (filtering, deduplication, time-windowed aggregations, point-in-time joins, and transformations) into Trino SQL Common Table Expressions (CTEs) executed directly on the Trino cluster.
- **Streaming online materialization**: Streams query results in chunks as PyArrow record batches directly to the online store via a thread pool, avoiding loading entire datasets into client memory.
- **Atomic offline table writes**: Materializes offline tables using a temporary staging table and rename swap (`CREATE TABLE ...__staging AS ...; DROP TABLE ...; ALTER TABLE ...__staging RENAME TO ...`).
- **Configuration inheritance**: Inherits connection details from the Trino offline store when both are configured.
---
## Example
```yaml
project: feast_trino_project
registry: data/registry.db
provider: local
offline_store:
type: trino.offline
host: localhost
port: 8080
catalog: iceberg
dataset: feast_offline
user: feast_user
batch_engine:
type: trino.engine
batch_size: 10000
write_concurrency: 4
online_store:
type: redis
connection_string: localhost:6379
```
---
## Example in Python
```python
from datetime import timedelta
from feast import (
BatchFeatureView,
Entity,
Field,
)
from feast.aggregation import Aggregation
from feast.infra.offline_stores.contrib.trino_offline_store.trino_source import TrinoSource
from feast.transformation.mode import TransformationMode
from feast.transformation.trino_transformation import TrinoTransformation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
feat: Add Trino compute engine for batch retrieval and materialization #6898
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Are you sure you want to change the base?
feat: Add Trino compute engine for batch retrieval and materialization #6898
Filter by extension
Viewed files
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
There are no files selected for viewing
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.