Skip to content
important New models verified update

Kimi K3 from Moonshot AI is newly available on the Amazon Bedrock platform

independently confirmed updated September 18, 2026

Kimi K3 (2.8 trillion parameters) from Moonshot AI has also been available on the Amazon Bedrock platform since 18 September 2026, with support for prompt caching and AWS data protection. The model was released in July 2026 as the first open model in the 3 trillion parameter class.

Kimi K3 from Moonshot AI is newly available on the Amazon Bedrock platform as an additional distribution channel alongside the API from Moonshot AI. According to AWS, the deployment on Bedrock is optimized for long programming and knowledge tasks, and it is the first open model on this platform to support explicit prompt caching, which reduces latency and input token costs when reusing the same context. According to AWS, data processed through Bedrock stays within the AWS data boundary, is not shared with the model provider, and is not used for training.

Moonshot AI first introduced Kimi K3 on 27 July 2026 as its most capable model to date, with 2.8 trillion parameters, which according to Moonshot AI makes Kimi K3 the first open model in the 3 trillion parameter class. It is a Mixture of Experts model with 896 experts, of which 16 are activated per token (approximately 104 billion active parameters), with a context window of 1 million tokens and native support for both text and images. Open weights are available on Hugging Face under the name moonshotai/Kimi-K3 in MXFP4 format; for self-hosting, AWS recommends a p6-b300 instance with 8 NVIDIA B300 GPUs.

Through the API from Moonshot AI, Kimi K3 costs 3 dollars per million input tokens and 15 dollars per million output tokens, a significant increase over the older Kimi K2.6 model (0.95 and 4 dollars per million tokens), putting Kimi K3 at the same price level as models in the Claude Sonnet series. According to benchmarks from Moonshot AI, the model outperforms Claude Opus 4.8 Max and GPT-5.5 High; according to an independent comparison from Artificial Analysis, cited by commentator Simon Willison, however, it trails Claude Fable 5 and GPT-5.6 Sol. On the Arena.ai benchmark for frontend code, Kimi K3 currently ranks first, including ahead of Claude Fable 5.

What changed

Why it matters

Companies looking to deploy open models for programming with long context can now use Kimi K3 through Amazon Bedrock with data security guarantees from AWS (without sharing data with the model provider), without changing their security model. They also need to account for a significantly higher API price compared with previous Kimi models and demanding infrastructure requirements (8 NVIDIA B300 GPUs) for self-hosting.

Release card

Kimi K3

Moonshot AI

open weights
Context
1 million tokens
Inputs
native text and image (multimodal)
Price
3 $/million input tokens and 15 $/million output tokens (API from Moonshot AI)
Availability
API from Moonshot AI, Amazon Bedrock platform, open weights on Hugging Face under the name moonshotai/Kimi-K3 in MXFP4 format, also available through OpenRouter
Documented measurements
  • Arena.ai – Frontend Code 1. místo, předstihuje Claude Fable 5 The model currently leads the Arena.ai leaderboard for frontend code, where it has also surpassed Claude Fable 5.
  • Efektivita škálování vůči modelu Kimi K2 přibližně 2,5násobné zlepšení According to AWS, the model achieves about a 2.5-fold improvement in scaling efficiency over its predecessor Kimi K2.
According to the sources, it is suitable for
  • long-running programming tasks across large repositories
  • working with knowledge and documents thanks to a context window of 1 million tokens
  • agentic tasks with tool calling and long conversations
Documented limits
  • only one level of reasoning (thinking effort), with no option to adjust it as with some competing models
  • self-hosting requires expensive infrastructure (p6-b300 instances with 8 NVIDIA B300 GPUs)
  • the API price (3 $/15 $ per million tokens) is significantly higher than for the older Kimi K2.6 model and matches the more expensive models in the Claude Sonnet series

According to benchmarks from Moonshot AI, the model outperforms Claude Opus 4.8 Max and GPT-5.5 High; according to an independent comparison from Artificial Analysis, however, it trails Claude Fable 5 and GPT-5.6 Sol.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. You can find a dated profile with supporting materials at model page. Model selection and other announcements →

What was added since the original report

Verified updates

  1. New verified information

    Kimi K3 is now available on Amazon Bedrock as a new distribution channel; The model supports prompt caching on Bedrock; Data in AWS Bedrock is protected by the AWS data boundary without being shared with the provider; The Bedrock deployment optimizes the model for long coding and knowledge workflows

    • Kimi K3 is now available on Amazon Bedrock as a new distribution channel
    • The model supports prompt caching on Bedrock
    • Data in AWS Bedrock is protected by the AWS data boundary without being shared with the provider
    • The Bedrock deployment optimizes the model for long coding and knowledge workflows
  2. New verified information

    API availability at $3/million input tokens and $15/million output tokens; Leading position in the Arena.ai Frontend Code benchmark; Beats Claude Opus 4.8 max and GPT-5.5 high in comparisons; Significantly higher prices than the previous Kimi K2.6 ($0.95/$4); The first open model in the 3 trillion parameter class

    • API availability at $3/million input tokens and $15/million output tokens
    • Leading position in the Arena.ai Frontend Code benchmark
    • Beats Claude Opus 4.8 max and GPT-5.5 high in comparisons
    • Significantly higher prices than the previous Kimi K2.6 ($0.95/$4)
    • The first open model in the 3 trillion parameter class

Two audiences, two different impacts

What this means

01

For individuals

Developers can try Kimi K3 through the API or through tools such as OpenCode for agentic programming tasks with long context, but must account for a higher price per token than with previous Kimi models.

What to do Try Kimi K3 through the API or a tool such as OpenCode for agentic programming tasks.
More practical updates →
02

For a business

Companies can deploy Kimi K3 through Amazon Bedrock with data protection from AWS without needing their own GPU infrastructure, or host the model themselves on instances with 8 NVIDIA B300 GPUs – but both options involve higher costs than earlier open Kimi models.

Development
What to decide Consider deploying Kimi K3 through Amazon Bedrock for long coding and knowledge tasks, taking into account the new prices and infrastructure requirements for self-hosting.
More business impacts →
Amazon Bedrock API AWS benchmarky Kimi K3 LLM MoE Moonshot AI open-weight open weights open weights Prompt caching vision vLLM

Check the original

Event sources

independently confirmed · 2 publishers, 1 independent. We count feeds from the same owner only once.

3
Simon Willison — AI tag (leading independent LLM commentator) community signal · first detected Kimi K3, and what we can still learn from the pelican benchmark AWS Machine Learning Blog primary source Deploying Kimi K3 on AWS AWS Machine Learning Blog primary source Introducing Kimi K3 on Amazon Bedrock