[ Web Proxy ]
URL:
Viewing: https://developers.cloudflare.com/workers-ai/platform/limits/ [Back]  [Original]

Limits Cloudflare Workers AI docsSkip to content
SearchCtrlKLog in
  1. Home
  2. /Workers AI
  3. /Platform
  4. /Limits

Limits

Last updated Aug 7, 2026Copy as MarkdownView as MarkdownAgent setup
OverviewRate limits by task typeAutomatic Speech RecognitionImage ClassificationImage-to-TextObject DetectionSummarizationText ClassificationText EmbeddingsText GenerationText-to-ImageTranslation

Workers AI is now Generally Available. We've updated our rate limits to reflect this.

Note that model inferences in local mode using Wrangler will also count towards these limits. Beta models may have lower rate limits while we work on performance and scale.

Custom requirements

If you have custom requirements like private custom models or higher limits, complete the Custom Requirements Form . Cloudflare will contact you with next steps.

Rate limits are default per task type, with some per-model limits defined as follows:

Rate limits by task type

  • 720 requests per minute
  • 3000 requests per minute
  • 720 requests per minute
  • 3000 requests per minute
  • 1500 requests per minute
  • 2000 requests per minute

Frontier models

The following limits apply per account, per model:

Model Standard Workers AI billing Prepaid AI Gateway credits
@cf/moonshotai/kimi-k2.6 20 requests per minute 50 requests per minute
@cf/moonshotai/kimi-k2.7-code 20 requests per minute 50 requests per minute
@cf/zai-org/glm-5.2 20 requests per minute 50 requests per minute

To receive the elevated limit, load prepaid AI Gateway credits and set the gateway's Workers AI billing setting to Unified billing. These limits are designed for typical agentic and coding workloads, where requests to frontier models can take longer to complete.

  • 720 requests per minute
PreviousData usageNextGlossary

Was this helpful?

YesNo
Edit pageReport issue
[]

Web Proxy Viewer  |  New URL  |  Original Page