Is your feature request related to a problem? Please describe.
Scalable data stores—Data warehouse and modern Data Lakes—are common today for offline store implementations or as a source of data ingestion. These stores are optimized and support ACID functionality where clean data can reside, and new data constantly can be updated and merged, accommodating data changes and observing schema evolutions. This ability to ingest data from or provide an offline store as a feast.data_source.FileSource from these stores extends
Feast's ecosystem to modern data lakes such as Delta Lake, Apache Hudi, or Apache Iceberg.
A similar feature request for HudiTableSource has been filed by @blvp
Describe the solution you'd like
Extend feast.data_source.FileSource(...) to take table names and locations to read from, for both local or remote sources.
Describe alternatives you've considered
I would have to save my Delta Lake tables as a single parquet file and use that as FileSource, which may defeat the purpose of being able to ingest point-in-time data from these modern data lake sources.
Reactions are currently unavailable
Is your feature request related to a problem? Please describe.
Scalable data stores—Data warehouse and modern Data Lakes—are common today for offline store implementations or as a source of data ingestion. These stores are optimized and support ACID functionality where clean data can reside, and new data constantly can be updated and merged, accommodating data changes and observing schema evolutions. This ability to ingest data from or provide an offline store as a feast.data_source.FileSource from these stores extends
Feast's ecosystem to modern data lakes such as Delta Lake, Apache Hudi, or Apache Iceberg.
A similar feature request for HudiTableSource has been filed by @blvp
Describe the solution you'd like
Extend feast.data_source.FileSource(...) to take table names and locations to read from, for both local or remote sources.
Describe alternatives you've considered
I would have to save my Delta Lake tables as a single parquet file and use that as FileSource, which may defeat the purpose of being able to ingest point-in-time data from these modern data lake sources.