Skip to content
worth noting Tools and apps

AWS has made GLM 5.3 available in Amazon Bedrock

only one source so far

GLM 5.3 is available to eligible corporate customers in Amazon Bedrock through managed APIs. The integration supports prompt caching and cross-region request processing, without requiring customers to run their own inference infrastructure.

AWS has made GLM 5.3 from Z.ai available to eligible corporate customers through Amazon Bedrock. The model has 753 billion parameters and a mixture-of-experts architecture. AWS describes it as a model focused on coding and long-running tasks involving AI agents. Access through managed APIs does not require running your own inference infrastructure.

The integration supports Responses and Chat Completions APIs compatible with the OpenAI API, as well as the Invoke and Converse APIs in Amazon Bedrock. For new applications, AWS recommends the compatible APIs because of their broader feature support. The model can also be tried through the chat interface in AWS Management Console. Cross-region request processing and different service tiers are available.

Automatic prompt caching, which reuses the stored beginning of a prompt, works by default. Developers can also explicitly mark portions for reuse; each marked prefix must contain at least 1 024 tokens. According to AWS, this feature can reduce latency and input token costs for repeated requests that share the same beginning of a prompt.

What changed

Why it matters

Eligible corporate customers gain another model for coding and AI agent tasks through a managed service. Compatible APIs allow them to integrate it through a supported interface. For long workflows that repeatedly send the same instructions or files, prompt caching can, according to AWS, reduce input token costs and the wait for a response.

Two audiences, two different impacts

What this means

01

For individuals

A developer at an eligible corporate customer can first try GLM 5.3 in the console’s chat interface without writing code or installing developer tools.

What to do If you have authorized access, try GLM 5.3 in the Test > Playground section of Amazon Bedrock.
More practical updates →
02

For a business

A company can use GLM 5.3 without provisioning its own inference infrastructure. For processes with repeated context, it also gains the ability to influence costs and latency through prompt caching.

Development
What to decide For workflows with repeated context, check whether explicit prompt caching is available and whether the marked prefix meets the minimum of 1 024 tokens.
More business impacts →
Amazon Bedrock GLM-5.3 Prompt caching Z.ai

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
AWS Machine Learning Blog primary source · first detected Introducing GLM 5.3 on Amazon Bedrock