| [ Web Proxy ] |
| Viewing: https://developers.cloudflare.com/workers-ai/changelog/ | [Back] [Original] |
@cf/zai-org/glm-5.2 is now available on Workers AI. Z.ai's flagship agentic coding model with a 262,144 token context window, function calling, and reasoning support. Read changelog to get started.@cf/moonshotai/kimi-k2.7-code now available on Workers AI. A frontier-scale 1T parameter MoE model optimized for coding, with a 262.1k context window, vision, multi-turn tool calling, and reasoning. Read changelog to get started.@cf/zai-org/glm-4.7-flash for fast tool-calling, @cf/google/gemma-4-26b-a4b-it for an efficient open model, or @cf/moonshotai/kimi-k2.6 for a capable tool-calling and vision model.@cf/moonshotai/kimi-k2.5 will be automatically aliased to @cf/moonshotai/kimi-k2.6, which has a higher price. The deprecation date was extended from May 10, 2026. Please review the K2.6 pricing and model capabilities prior to May 30, 2026.@cf/moonshotai/kimi-k2.5 --> @cf/moonshotai/kimi-k2.6@hf/meta-llama/meta-llama-3-8b-instruct@cf/meta/llama-3-8b-instruct@cf/meta/llama-3-8b-instruct-awq@cf/meta/llama-3.1-8b-instruct@cf/meta/llama-3.1-8b-instruct-awq@cf/meta/llama-3.1-70b-instruct@cf/meta/llama-2-7b-chat-int8@cf/meta/llama-2-7b-chat-fp16@cf/mistral/mistral-7b-instruct-v0.1@hf/mistral/mistral-7b-instruct-v0.2@hf/google/gemma-7b-it@cf/google/gemma-3-12b-it@hf/nousresearch/hermes-2-pro-mistral-7b@cf/microsoft/phi-2@cf/defog/sqlcoder-7b-2@cf/unum/uform-gen2-qwen-500m@cf/facebook/bart-large-cnn-fast and -lora variants of models will remain active. LoRA models may be deprecated in the future, and we will communicate when new LoRA models come online to give users time to train new LoRAs before we deprecate old ones.@cf/moonshotai/kimi-k2.6 now available on Workers AI. The latest frontier-scale model from Moonshot AI with improved reasoning, coding, and agentic capabilities. Read changelog to get started.chat_template_kwargs.thinking parameter to control reasoning (instead of chat_template_kwargs.enable_thinking) and returns reasoning content in the reasoning field (instead of reasoning_content).@cf/google/gemma-4-26b-a4b-it now available on Workers AI. A Mixture-of-Experts model with 26B total parameters and 4B active, featuring a 256K context window, vision, built-in thinking mode, and function calling. Read changelog to get started.@cf/moonshotai/kimi-k2.5 now available on Workers AI. A frontier-scale open-source model with a 256k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads. Read changelog to get started.x-session-affinity header to route requests to the same model instance and maximize prefix cache hit rates across multi-turn conversations.@cf/nvidia/nemotron-3-120b-a12b now available on Workers AI! A hybrid MoE model with 120B total parameters and 12B active, optimized for multi-agent and agentic AI workloads. Read changelog to get started.@cf/deepgram/nova-3 now supports 10 languages with regional variants for real-time transcription. Supported languages include English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, and Dutch with regional variants like en-GB, fr-CA, and pt-BR.@cf/openai/gpt-oss-120b and @cf/openai/gpt-oss-20b now support Chat Completions API format. Use /v1/chat/completions with a messages array, or use /ai/run which dynamically detects your input format and accepts Chat Completions (messages), legacy Completions (prompt), or Responses API (input).content field in message objects only accepted string values. The field now properly accepts both string content and array content (structured content parts for multi-modal inputs). This fix applies to all affected chat models including GPT-OSS models, Llama 3.x, Mistral, Qwen, and others.tool_call_id values that it generated itself, fixing issues with multi-turn tool calling conversations.content: null and tool_calls are now accepted in both the Workers AI binding and REST API (/v1/chat/completions), fixing tool call round-trip failures.finish_reason only on the usage chunk, matching OpenAI's streaming behavior and preventing duplicate finish events./v1/chat/completions now preserves original tool call IDs from models instead of regenerating them. Previously, the endpoint was generating new IDs which broke multi-turn tool calling because AI SDK clients could not match tool results to their original calls./v1/chat/completions now correctly reports finish_reason: "tool_calls" in the final usage chunk when tools are used. Previously, it was hardcoding finish_reason: "stop" which caused AI SDK clients to think the conversation was complete instead of executing tool calls.@cf/zai-org/glm-4.7-flash is now available on Workers AI! A fast and efficient multilingual text generation model optimized for multi-turn tool calling across 100+ languages. Read changelog to get started.@cloudflare/tanstack-ai package for using Workers AI and AI Gateway with TanStack AI.workers-ai-provider v3.1.1 adds transcription, text-to-speech, and reranking capabilities.@cf/black-forest-labs/flux-2-klein-9b now available on Workers AI! Read changelog to get started@cf/black-forest-labs/flux-2-klein-4b now available on Workers AI! Read changelog to get started@cf/deepgram/flux model page or pricing page@cf/black-forest-labs/flux-2-dev now available on Workers AI! Read changelog to get started@cf/qwen/qwen3-30b-a3b-fp8 and @cf/qwen/qwen3-embedding-0.6b now available on Workers AI@cf/deepgram/aura-2-en and @cf/deepgram/aura-2-es on how to use the new models.@cf/ibm-granite/granite-4.0-h-micro@cf/deepgram/flux and check out the changelog for in-depth examples.@cf/pfnet/plamo-embedding-1b creates embeddings from Japanese text.@cf/aisingapore/gemma-sea-lion-v4-27b-it is a fine-tuned model that supports multiple South East Asian languages, including Burmese, English, Indonesian, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai, and Vietnamese.@cf/ai4bharat/indictrans2-en-indic-1B is a translation model that can translate between 22 indic languages, including Bengali, Gujarati, Hindi, Tamil, Sanskrit and even traditionally low-resourced languages like Kashmiri, Manipuri and Sindhi..docx and .odt files.npm i -D wrangler@latest to update your packages.@cf/google/embeddinggemma-300m for more details. Now available to use for embedding in AI Search too.@cf/deepgram/aura-1 is a text-to-speech model that allows you to input text and have it come to life in a customizable voice@cf/deepgram/nova-3 is speech-to-text model that transcribes multilingual audio at a blazingly fast speed@cf/pipecat-ai/smart-turn-v2 helps you detect when someone is done speaking@cf/leonardo/lucid-origin is a text-to-image model that generates images with sharp graphic design, stunning full-HD renders, or highly specific creative direction@cf/leonardo/phoenix-1.0 is a text-to-image model with exceptional prompt adherence and coherent text@cf/deepgram/aura-1, @cf/deepgram/nova-3, @cf/pipecat-ai/smart-turn-v2gpt-oss-120b and gpt-oss-20b model pages for more information about schemas, pricing, and context windowsuser and assistant to function correctlymax_tokens defaults were not properly being respected - max_tokens now correctly defaults to 256 as displayed on the model pages. Users relying on the previous behaviour may observe this as a breaking change. If you want to generate more tokens, please set the max_tokens parameter to what you need.max_tokens accordingly depending on your prompt length and the context window length to ensure a successful response.@cf/black-forest-labs/flux-1-schnell is now available on Workers AIWorkers AI now suppoorts Meta Llama 3.1.
@cloudflare/ai-utils npm packageai-utils on GithubWorkers AI now natively supports AI Gateway.
We will be deprecating @cf/meta/llama-2-7b-chat-int8 on 2024-06-30.
Replace the model ID in your code with a new model of your choice:
@cf/meta/llama-3-8b-instruct is the newest model in the Llama family (and is currently free for a limited time on Workers AI).@cf/meta/llama-3-8b-instruct-awq is the new Llama 3 in a similar precision to your currently selected model. This model is also currently free for a limited time.If you do not switch to a different model by June 30th, we will automatically start returning inference from @cf/meta/llama-3-8b-instruct-awq.
-lora versionAdded OpenAI compatible API endpoints for /v1/chat/completions and /v1/embeddings. For more details, refer to Configurations.
const resp = await env.AI.run(modelName, inputs)@cloudflare/ai npm package. While existing solutions using the @cloudflare/ai package will continue to work, no new Workers AI features will be supported.
Moving to native AI bindings is highly recommended| Web Proxy Viewer | New URL | Original Page |