Follow-up to #22702; #22703 proposes the same for Python.
MCP servers expose tools to LLM agents. Every argument of a tool, resource or prompt callback is chosen by the client, that is by a model that prompt injection can steer, or by anyone who can reach the server over HTTP. This PR models that input for the TypeScript SDKs as sources of the remote threat model, as data extensions in javascript/ql/lib/ext.
First commit, the MCP SDKs:
@modelcontextprotocol/sdk 1.x and @modelcontextprotocol/server 2.x, McpServer: the arguments of the callbacks of tool, registerTool, prompt and registerPrompt, and the URI and template variables of resource and registerResource.
The low-level Server: request.params in handlers of setRequestHandler, also when the server is reached through McpServer.server, and the parameters of custom methods in 2.x.
The request headers and the bearer token that the SDK hands to a callback: extra.requestInfo.headers and extra.authInfo.token in 1.x, ctx.http.req.headers.get() and ctx.http.authInfo.token in 2.x.
typeModel rows for the import paths with and without .js and for a server that arrives as a typed parameter.
Second commit, flow summaries for zod (independent of the first, can be split off): handlers of the low-level server get the raw arguments and validate them with a zod schema before use (Schema.parse(request.params.arguments)). There is no model of zod, so the flow ends at that call. Four rows let data flow from the argument of parse, safeParse, parseAsync and safeParseAsync to the result. I used kind value: js/path-injection does not follow a taint summary there (it is a data-flow configuration with flow states) and does follow a value one; I checked both on a small server.
Tests: library-tests/threat-models/sources gets mcp.ts, fastmcp.ts and zod.ts with inline expectations (threat-source, hasFlow); callbacks that are not registered with a server and a static resource have none. codeql test run passes locally with CLI 2.27.1.
Evidence: as a model pack, these rows take CodeQL 2.27.1 from 6 to 25 of 38 real, publicly disclosed vulnerabilities in TypeScript and JavaScript MCP servers in mcp-vulnbench. On the half of the cases that the models were not run on while they were written, plain CodeQL finds 3 of 21, the rows of the first commit 11, and both commits together 14 (the zod rows were measured with kind taint), at 2.62 instead of 0.08 alarms per KLOC. Known limits are documented there: servers kept in a field of a project class, handler tables built across modules, and sinks such as fetch from undici or Playwright's page.goto.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #22702; #22703 proposes the same for Python.
MCP servers expose tools to LLM agents. Every argument of a tool, resource or prompt callback is chosen by the client, that is by a model that prompt injection can steer, or by anyone who can reach the server over HTTP. This PR models that input for the TypeScript SDKs as sources of the remote threat model, as data extensions in javascript/ql/lib/ext.
First commit, the MCP SDKs:
Second commit, flow summaries for zod (independent of the first, can be split off): handlers of the low-level server get the raw arguments and validate them with a zod schema before use (Schema.parse(request.params.arguments)). There is no model of zod, so the flow ends at that call. Four rows let data flow from the argument of parse, safeParse, parseAsync and safeParseAsync to the result. I used kind value: js/path-injection does not follow a taint summary there (it is a data-flow configuration with flow states) and does follow a value one; I checked both on a small server.
Tests: library-tests/threat-models/sources gets mcp.ts, fastmcp.ts and zod.ts with inline expectations (threat-source, hasFlow); callbacks that are not registered with a server and a static resource have none. codeql test run passes locally with CLI 2.27.1.
Evidence: as a model pack, these rows take CodeQL 2.27.1 from 6 to 25 of 38 real, publicly disclosed vulnerabilities in TypeScript and JavaScript MCP servers in mcp-vulnbench. On the half of the cases that the models were not run on while they were written, plain CodeQL finds 3 of 21, the rows of the first commit 11, and both commits together 14 (the zod rows were measured with kind taint), at 2.62 instead of 0.08 alarms per KLOC. Known limits are documented there: servers kept in a field of a project class, handler tables built across modules, and sinks such as fetch from undici or Playwright's page.goto.