Kimi K3 from Moonshot AI is newly available on the Amazon Bedrock platform
Kimi K3 (2.8 trillion parameters) from Moonshot AI has also been available on the Amazon Bedrock platform since 18 September 2026, with support for prompt caching and AWS data protection. The model was released in July 2026 as the first open model in the 3 trillion parameter class.
Kimi K3 from Moonshot AI is newly available on the Amazon Bedrock platform as an additional distribution channel alongside the API from Moonshot AI. According to AWS, the deployment on Bedrock is optimized for long programming and knowledge tasks, and it is the first open model on this platform to support explicit prompt caching, which reduces latency and input token costs when reusing the same context. According to AWS, data processed through Bedrock stays within the AWS data boundary, is not shared with the model provider, and is not used for training.
Moonshot AI first introduced Kimi K3 on 27 July 2026 as its most capable model to date, with 2.8 trillion parameters, which according to Moonshot AI makes Kimi K3 the first open model in the 3 trillion parameter class. It is a Mixture of Experts model with 896 experts, of which 16 are activated per token (approximately 104 billion active parameters), with a context window of 1 million tokens and native support for both text and images. Open weights are available on Hugging Face under the name moonshotai/Kimi-K3 in MXFP4 format; for self-hosting, AWS recommends a p6-b300 instance with 8 NVIDIA B300 GPUs.
Through the API from Moonshot AI, Kimi K3 costs 3 dollars per million input tokens and 15 dollars per million output tokens, a significant increase over the older Kimi K2.6 model (0.95 and 4 dollars per million tokens), putting Kimi K3 at the same price level as models in the Claude Sonnet series. According to benchmarks from Moonshot AI, the model outperforms Claude Opus 4.8 Max and GPT-5.5 High; according to an independent comparison from Artificial Analysis, cited by commentator Simon Willison, however, it trails Claude Fable 5 and GPT-5.6 Sol. On the Arena.ai benchmark for frontend code, Kimi K3 currently ranks first, including ahead of Claude Fable 5.
Why it matters
Companies looking to deploy open models for programming with long context can now use Kimi K3 through Amazon Bedrock with data security guarantees from AWS (without sharing data with the model provider), without changing their security model. They also need to account for a significantly higher API price compared with previous Kimi models and demanding infrastructure requirements (8 NVIDIA B300 GPUs) for self-hosting.
Release card
Kimi K3
Moonshot AI
- Context
- 1 million tokens
- Inputs
- native text and image (multimodal)
- Price
- 3 $/million input tokens and 15 $/million output tokens (API from Moonshot AI)
- Availability
- API from Moonshot AI, Amazon Bedrock platform, open weights on Hugging Face under the name moonshotai/Kimi-K3 in MXFP4 format, also available through OpenRouter
- Arena.ai – Frontend Code 1. místo, předstihuje Claude Fable 5 The model currently leads the Arena.ai leaderboard for frontend code, where it has also surpassed Claude Fable 5.
- Efektivita škálování vůči modelu Kimi K2 přibližně 2,5násobné zlepšení According to AWS, the model achieves about a 2.5-fold improvement in scaling efficiency over its predecessor Kimi K2.
- long-running programming tasks across large repositories
- working with knowledge and documents thanks to a context window of 1 million tokens
- agentic tasks with tool calling and long conversations
- only one level of reasoning (thinking effort), with no option to adjust it as with some competing models
- self-hosting requires expensive infrastructure (p6-b300 instances with 8 NVIDIA B300 GPUs)
- the API price (3 $/15 $ per million tokens) is significantly higher than for the older Kimi K2.6 model and matches the more expensive models in the Claude Sonnet series
According to benchmarks from Moonshot AI, the model outperforms Claude Opus 4.8 Max and GPT-5.5 High; according to an independent comparison from Artificial Analysis, however, it trails Claude Fable 5 and GPT-5.6 Sol.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. You can find a dated profile with supporting materials at model page. Model selection and other announcements →
What was added since the original report
Verified updates
-
Kimi K3 is now available on Amazon Bedrock as a new distribution channel; The model supports prompt caching on Bedrock; Data in AWS Bedrock is protected by the AWS data boundary without being shared with the provider; The Bedrock deployment optimizes the model for long coding and knowledge workflows
- Kimi K3 is now available on Amazon Bedrock as a new distribution channel
- The model supports prompt caching on Bedrock
- Data in AWS Bedrock is protected by the AWS data boundary without being shared with the provider
- The Bedrock deployment optimizes the model for long coding and knowledge workflows
-
API availability at $3/million input tokens and $15/million output tokens; Leading position in the Arena.ai Frontend Code benchmark; Beats Claude Opus 4.8 max and GPT-5.5 high in comparisons; Significantly higher prices than the previous Kimi K2.6 ($0.95/$4); The first open model in the 3 trillion parameter class
- API availability at $3/million input tokens and $15/million output tokens
- Leading position in the Arena.ai Frontend Code benchmark
- Beats Claude Opus 4.8 max and GPT-5.5 high in comparisons
- Significantly higher prices than the previous Kimi K2.6 ($0.95/$4)
- The first open model in the 3 trillion parameter class
Two audiences, two different impacts
What this means
For individuals
Developers can try Kimi K3 through the API or through tools such as OpenCode for agentic programming tasks with long context, but must account for a higher price per token than with previous Kimi models.
For a business
Companies can deploy Kimi K3 through Amazon Bedrock with data protection from AWS without needing their own GPU infrastructure, or host the model themselves on instances with 8 NVIDIA B300 GPUs – but both options involve higher costs than earlier open Kimi models.
DevelopmentCheck the original
Event sources
independently confirmed · 2 publishers, 1 independent. We count feeds from the same owner only once.