| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
docs: Add Feast Chronon integration blog post
Signed-off-by: Francisco Javier Arceo <arceofrancisco@gmail.com>
fix: Read Kafka and Kinesis sources without a batch source from proto
batch_source is optional on KafkaSource and KinesisSource, but from_proto
checked it with a truthiness test. An unset proto sub-message is still
truthy, so the empty message was parsed as a data source and raised
"Could not identify the source type being added.", which made these
sources impossible to read back from the registry.
Use HasField("batch_source"), as PushSource.from_proto already does.
Signed-off-by: LuisFigueroaG <luis.h.figueroa.g@gmail.com>fix: Do not log /metrics requests from the pre-fork metrics server th…
…read
The metrics HTTP server runs as a thread in the Gunicorn master, which
forks the workers. Its default WSGIRequestHandler writes an access-log
line to stderr for every scrape. A fork while that thread holds the
stderr buffer lock leaves the new worker with the lock held forever: it
blocks on its first log line ("Booting worker") and never serves.
Use a request handler with a no-op log_message for both the IPv4 and
the dual-stack server built by _make_metrics_httpd, like
prometheus_client's own _SilentHandler.
Fixes #6928
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alexandr Borgatin <a.borgatin@yandex.ru>fix: Reject unknown resource types and actions in the REST permission…
… API apply_permission silently dropped any type or action name that was not in the name-to-enum maps. A typo or wrong casing in every type left the spec with an empty types list, which Permission treats as ALL_RESOURCE_TYPES, so a permission meant for one resource type ended up covering all of them. Partially invalid action lists were also narrowed without notice. Return a 400 listing the invalid names and the valid options instead. Signed-off-by: LuisFigueroaG <luis.h.figueroa.g@gmail.com>
feat: Add a non-JVM read path for Lance data sources (#6943)
* feat: Add a non-JVM read path for Lance data sources Adds LanceSource and teaches the DuckDB offline store to read it, completing the read half of #6899. LanceFormat landed in #6925 as a format descriptor; nothing read Lance until now. Lance already worked through SparkSource, which drives its reader generically from table_format.format_type.value and table_format.properties. What was missing is a path that needs no JVM, which is also the real test of whether the DataSource abstraction is engine-agnostic rather than Spark-agnostic in name only. Placement: a new source read by the existing DuckDB store, rather than a Lance offline store or an extension of FileSource. FileSource is the wrong host. Its format axis is already taken by file_format, so adding table_format would give one source two overlapping format axes. It is also read by two stores with incompatible contracts: duckdb._read_data_source dispatches on type, while dask._read_datasource has no dispatch seam and reads file_options.uri unconditionally as Parquet, and asserts isinstance(..., FileSource) in three places. Decisively, Lance's catalog addressing has no path to put in FileSource.path, so the catalog-based layer would not fit the class even if the path-based one did. A Lance offline store would be the wrong 120 lines. duckdb.py is a binding that injects reader and writer callbacks into the engine in ibis.py, so a Lance store would be a near-copy of it plus a repo_config entry, and would force a choice between Lance and Parquet instead of mixing them in one feature service. Reading a source as ibis.memtable(arrow_table) in the DuckDB store already has two precedents, IcebergSource and MlflowDatasetSource. Following them leaves ibis.py untouched, so the point-in-time join, TTL handling, field mapping and ODFVs work unchanged, and no edit to repo_config.py or data_source.py is needed because CUSTOM_SOURCE plus data_source_class_type is self-describing. Both addressing modes work: a uri, and catalog/namespace/table through namespace_client and table_id. Pin semantics follow what was argued on #5782 and #6925: a pin selects data, never shape. get_table_column_names_and_types reads the pinned schema so feast apply infers what reads will actually see; a pre-flight check fails with a message naming the pin when a pinned version cannot satisfy the declared schema; and vector widths go through _validate_vector_field_lengths from #6909 rather than a second validator. Tests use the dir namespace implementation, which exercises the same namespace_client and table_id code path as a remote catalog with no server required. 43 tests, including a demonstration that a tag pin returns earlier data after the dataset has been overwritten for the same entity and timestamp. Read-only for now: _write_data_source is untouched, so a LanceSource is not yet a persist target and there is no SavedDatasetLanceStorage. Signed-off-by: hao-xu5 <hxu44@apple.com> * fix: Address Lance read review feedback Signed-off-by: HaoXuAI <sduxuhao@gmail.com> --------- Signed-off-by: hao-xu5 <hxu44@apple.com> Signed-off-by: HaoXuAI <sduxuhao@gmail.com>
fix: Refresh the SQL registry cache after deleting a permission
delete_permission ran its own DELETE and returned, so unlike the other delete_* methods it never bumped the project's last-updated metadata or refreshed the cache in sync mode. list_permissions(allow_cache=True) kept returning the deleted permission. Route it through _delete_object like the other object types. That alone wasn't enough: _delete_object refreshed the cache inside the open write transaction, so the refresh read a snapshot that still contained the deleted row. Refresh after the transaction commits instead, matching _apply_object. This also fixes stale cached reads after delete_entity, delete_data_source and the other deletes that go through _delete_object. Signed-off-by: LuisFigueroaG <luis.h.figueroa.g@gmail.com>
fix: Map SQL Server bigint to INT64 and float to DOUBLE
mssql_to_feast_value_type mapped bigint to FLOAT, so 64-bit integer columns were inferred as Float32 and lost precision above 2^24. It also mapped float to FLOAT, but SQL Server float defaults to float(53), an 8-byte double. real (float(24)) stays FLOAT. Signed-off-by: LuisFigueroaG <luis.h.figueroa.g@gmail.com>
feat: Support Milvus partition keys in the Milvus online store
Multi-tenant feature views, such as a catalogue shared by many brands, benefit from Milvus partition keys: searches filtered on the key only scan the matching partitions. A field can be marked as the partition key with the feature view tag milvus.partition_key, or for every feature view containing the field with the new partition_key store config. The tag takes precedence. The field must be stored as VARCHAR or INT64. Partition keys only apply when a collection is created. Feast logs a warning when an existing collection lacks the configured key. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Simon Hearne <simon.hearne@gmail.com>
feat: Add standalone MCP server as a first-class component
* feat: Add standalone feast mcp server Adds `feast mcp`, a Model Context Protocol server that runs in its own process and proxies to a running Feast deployment over HTTP. It holds no registry or online store of its own; two sub-servers are mounted behind a single MCP endpoint, each only when its upstream URL is configured: - `features` -> the Python feature server (online features, vector search, push, materialization) - `registry` -> the REST registry server (feature views, entities, data sources, lineage) This is distinct from `feast serve` with `mcp_enabled: true`, which mounts an OpenAPI-derived MCP endpoint inside the feature server itself. Settings resolve CLI > environment > feast_mcp.yaml > defaults. The server enforces no RBAC of its own: it forwards the caller's bearer token upstream and lets Feast apply its permission model. `--auth-mode oidc` additionally fronts an OIDC provider so IDE clients can complete a browser login. The `feast mcp` subcommand resolves lazily so that FastMCP is not imported by unrelated commands such as `feast apply`, and degrades to a clear error when the optional `mcp-server` extra is absent. Two CI pins had to move to make the dependency set solvable. fastmcp requires httpx>=0.28.1, which the ci extra pinned down to 0.27.2 because python-keycloak <4.7.3 passed the `proxies` argument httpx removed in 0.28 (python-keycloak#623); raising the python-keycloak floor to >=4.7.3 lifts that constraint. virtualenv is unpinned from 20.23.0 to >=20.26,<21. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com> * feat: Add the MCP server as a served component in feast-operator Setting `spec.services.mcpServer` adds a dedicated `feast mcp` container to the FeatureStore deployment, exposed on its own Service on port 8100. The container runs in the same pod as the online and registry servers, so it reaches both over localhost. The operator owns only `--host` and `--port`, so they always match the generated Service. Everything else (transport, upstream URLs, auth, observability) is read from a feast_mcp.yaml supplied in a ConfigMap and mounted read-only at /etc/feast/mcp. A CEL rule enforces that at least one upstream is actually available: either onlineStore is present and not disabled, or registry.local.server.restAPI is true. Readiness is reported on the McpServer status condition, and the Service hostname on status.serviceHostnames.mcpServer. Operator-managed TLS is not yet supported for this container; the `tls` field is ignored. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com> * docs: Add standalone MCP server reference and example Adds the reference page for `feast mcp`, covering the CLI options, the feast_mcp.yaml schema and its environment equivalents, the available tools, both auth modes, and deployment with the Feast Operator. Links it from SUMMARY.md and the feature-servers index, and cross-references it from the existing MCP feature server page so the two are not confused. Adds examples/feast_mcp_server, an end-to-end walkthrough that starts a local feature store, runs the MCP server against it, and calls its tools from a client, plus the Kubernetes equivalent deployed by the operator. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com> * test: Add unit tests for the standalone feast mcp server Covers the parts that are easy to get wrong and invisible until someone else's deployment breaks: - config resolution, asserting CLI > env > yaml > default one rung at a time, because a stray CLI default silently outranks feast_mcp.yaml - token forwarding in both auth modes, including that a passthrough caller's bearer token still reaches Feast - that upstream denials surface as MCP tool errors rather than being handed to the model as feature data - a request matrix over every tool, asserting the upstream request each one builds and that optional arguments are omitted rather than sent as null Adds the operator-side controller test for the mcpserver container, Service, ConfigMap mount and status conditions. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com> * build: Relock Python requirements Regenerates the locked requirement sets for all supported Python versions so that the new `mcp-server` extra, and the dependency floors it forces, are reflected in the lock files. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com> * build: Add the feast-mcp container image Packages `feast mcp` as a container whose entrypoint is the server itself. It is a thin wrapper over the feature-server image, which already installs feast[minimal] and therefore everything the MCP server needs, so this only sets the entrypoint and inherits the UBI base and the arbitrary-uid permission setup. BASE_IMAGE and BASE_TAG select the feature-server image to build on, so the same Dockerfile serves a local build, a build against a published tag, and a downstream rebuild from another repository. TMPDIR is set to /dev/shm because Python needs a writable temp dir and the usual candidates are all read-only under readOnlyRootFilesystem. Also adds test-python-unit-mcp for iterating on the MCP server's tests alone; they are already covered by test-python-unit. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com> * fix: Address review feedback on the standalone MCP server - Treat an omitted onlineStore as available in the mcpServer CEL rule, matching the operator's default online feature server, and add admission tests for the omitted, disabled and REST-registry-only cases. - Decode JSON text in the demo client's CallToolResult fallback so registry discovery works on older fastmcp versions. - Describe the supported kubernetes auth mode in the example README and align the MCP docs, sample configs and API reference with the code. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com> * fix: Address review comments on the standalone MCP server - Reject unknown auth modes instead of falling back to passthrough - Percent-encode tool arguments used as URL path segments - Log the socket peer IP instead of trusting X-Forwarded-For/X-Real-IP - Build a fresh FastMCP instance per run and drop unused config helpers - Default the operator-managed MCP transport to http when no config is given - Reject an empty mcpServer.config.configMapRef.name at admission - Require BASE_TAG for the feast-mcp image instead of defaulting to latest - Keep the MCP API reference rows on one line and refresh the secrets baseline - Only recurse into real packages in the doctest walker so feast.mcp does not collide with the mcp SDK Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com> --------- Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
ci: Publish the feast-mcp image to quay.io
The MCP server image had a Dockerfile but nothing to build or push it, so the quay.io/feastdev/feast-mcp tags the docs reference were never produced. Add build-feast-mcp-docker and push-feast-mcp-docker, and a release job that runs them. The image wraps an already-published feature-server image, so it cannot go in the existing matrix -- its base has to be pushed first. BASE_TAG defaults to VERSION, and CI pins it to the version it just built. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
fix: Use a valid ShuffleStrategy for the Ray DatasetContext
Feast set `DatasetContext.shuffle_strategy = "sort"`, which is not a member of `ray.data.context.ShuffleStrategy`. Up to Ray 2.58 the setter stored the raw string and Ray fell through to the pull-based sort shuffle. Ray 2.59 coerces the value with `ShuffleStrategy(value)`, so `ensure_ray_initialized()` and `RayResourceManager.configure_ray_context()` now raise `ValueError: 'sort' is not a valid ShuffleStrategy`. Use `ShuffleStrategy.SORT_SHUFFLE_PULL_BASED`, which keeps the behavior on older Ray versions and is accepted by Ray 2.59. Signed-off-by: Yihang Chen <yhc0720@berkeley.edu>
fix: Read naive datetimes as UTC in UnixTimestamp values
Converting a naive datetime.datetime to a UnixTimestamp value called datetime.timestamp(), which reads a naive value in the machine's local timezone. A naive pd.Timestamp or np.datetime64 is read as UTC, as is a naive ZonedTimestamp, and Feast treats naive datetimes as UTC elsewhere (make_tzaware). So the same wall-clock value was stored at a different instant depending on its Python type and on the host timezone, e.g. a naive request timestamp passed to get_online_features came back shifted by the server's UTC offset. Attach UTC to naive datetimes before taking the timestamp. This covers UnixTimestamp scalars, lists and sets. Signed-off-by: Mohammad Hijjawi <mohammad.hijjawi1997@gmail.com>
fix: Treat datetimes with a None utcoffset as naive
A datetime is naive when utcoffset() returns None, which includes a tzinfo whose utcoffset() is None, not only tzinfo=None. Check utcoffset() so those values are also read as UTC instead of raising. Make the regression test stricter: skip it where time.tzset is not available, use the fixed POSIX offset UTC+8 so no timezone database is needed, assert the local offset really changed, and cover a custom tzinfo whose utcoffset() is None. Signed-off-by: Mohammad Hijjawi <mohammad.hijjawi1997@gmail.com>
ci: Publish the feast-mcp image on merge to master
The release workflow builds feast-mcp, but master merges push feature-server, feature-transformation-server and feast-operator to quay.io/feastdev-ci without it, so there is no :develop tag to test with. Add a job that wraps the feature-server image that job just pushed. It needs that image present, so it runs after the matrix rather than inside it. Unlike the release path this base comes from Dockerfile.dev, which installs from the lockfiles and local source, so it carries mcp-server before the first release that publishes it to PyPI. Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
fix: Initialize auth managers in standalone lineage server (#6873)
Signed-off-by: ntkathole <nikhilkathole2683@gmail.com>
fix: Map Spark smallint and tinyint columns to INT32
Signed-off-by: Raashish Aggarwal <94279692+raashish1601@users.noreply.github.com>
Unfortunately it looks like we can’t render this comparison for you right now. It might be too big, or there might be something weird with your repository.
You can try running this command locally to see the comparison on your machine:
git diff master@{1day}...master
| Back | FazBrowse Home | New Git URL |