Skip to content
Apixo
Blog
news· 3 min read· via AWS Machine Learning

Amazon Bedrock Adds Z.ai's GLM 5.3 for Coding and Agentic Workloads

Amazon Bedrock now hosts GLM 5.3, a 753B-parameter MoE model from Z.ai optimized for coding, prompt caching, and cybersecurity agent workflows.

Amazon Bedrock Adds Z.ai's GLM 5.3 for Coding and Agentic Workloads

Amazon Web Services has expanded its managed model catalog by bringing GLM 5.3, developed by Z.ai (Zhipu AI), to Amazon Bedrock. The open-weight mixture-of-experts (MoE) system features 753 billion parameters and is engineered specifically to handle software engineering tasks and long-horizon agentic workflows. By making GLM 5.3 available on Amazon Bedrock, AWS allows enterprise customers to run the model through fully managed APIs without provisioning, hosting, or maintaining custom inference infrastructure.

Upgraded benchmarks and cybersecurity features

GLM 5.3 introduces several performance gains compared to previous generations in the model line. According to Z.ai, the model delivers competitive results across multiple standard software benchmarks, including DeepSWE, Terminal Bench 3.0, and FrontierSWE. Internal testing by Z.ai demonstrates a 50 percent improvement over GLM 5.2 on proprietary coding assessments. Direct comparative figures against GLM 5 were omitted due to updates made to the benchmark suites following the release of GLM 5.1.

Beyond general code refactoring and multi-step reasoning, GLM 5.3 highlights specialized capabilities in defensive cybersecurity. Upon release, Z.ai reported a top score of 84.5 on the CyberGym benchmark. These capabilities make the model well-suited for automated vulnerability research. For example, the open-source penetration testing tool Strix currently uses GLM 5.3 as its default reasoning model. Developers can deploy Strix alongside local targets—such as the OWASP Juice Shop application—to automatically coordinate sub-agents that identify security flaws, validate vulnerabilities using proof-of-concept tests, and output remediation reports within an organization's existing AWS control boundaries.

API integration, prompt caching, and deployment tiers

Amazon Bedrock provides flexible developer access to GLM 5.3. Applications can interface with the model using OpenAI-compatible Responses and Chat Completions APIs, or via Amazon Bedrock's native Invoke and Converse endpoints. AWS recommends the OpenAI-compatible structure for new applications due to broader feature alignment. Authentication supports temporary AWS IAM credentials generated using tools like the aws-bedrock-token-generator library rather than relying on long-lived API keys.

To keep costs manageable during extended agent interactions, GLM 5.3 includes support for prompt caching. Implicit prompt caching is active by default, reducing input token fees and response latency for repeated prompt prefixes. Developers can also enable explicit prompt caching by specifying prompt_cache_options and inserting prompt_cache_breakpoint markers on static content blocks containing at least 1,024 tokens, such as large repository context files or system prompts.

Deployment is handled through US (us.zai.glm-5.3) and Global (global.zai.glm-5.3) cross-Region inference profiles, allowing AWS to route traffic securely across regions. To manage price and performance trade-offs, customers can choose between three service tiers: Flex for non-urgent tasks, Standard for standard operation, and Priority for latency-critical applications.

What it means for developers

For engineering teams, the launch of GLM 5.3 on Amazon Bedrock offers a scalable pathway to run complex, long-running AI agents without maintaining backend server clusters. Workflows that require analyzing entire codebases across hundreds of files or executing multi-hour diagnostic scripts frequently strain standard infrastructure. Managed prompt caching and granular service tiers directly address these cost bottlenecks by reusing context across conversation turns.

Furthermore, compatibility with standard OpenAI SDK formats allows teams to connect GLM 5.3 into existing developer workflows, terminal agents, and coding assistants like OpenCode. Enterprise organizations evaluating open-weight models can conduct security assessments and refactoring tasks inside their established AWS IAM security framework rather than transferring data to third-party providers.

Developers looking for affordable access to top AI models can also explore options through a single API key at https://apixoai.online, which offers a streamlined route to experiment with leading foundation models alongside cloud provider deployments.

Overall, the availability of GLM 5.3 on Amazon Bedrock reflects a continuing transition toward specialized, agentic models capable of executing complex system tasks, code generation, and automated security evaluations.


Source: Introducing GLM 5.3 on Amazon Bedrock | Amazon Web Services — AWS Machine Learning. Written by the Apixo team from that report.

#ai-news#aws#amazon-bedrock#glm-5-3#ai-agents#llm
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading