Introducing GLM 5.3 on Amazon Bedrock
4d
- Publicado
- Coletado

GLM 5.3 from Z.ai is now available on Amazon Bedrock: a 753B-parameter mixture-of-experts model built for coding and long-horizon agentic tasks. Learn how to invoke it with the OpenAI-compatible APIs, cut cost and latency with prompt caching, and run an authorized security test with the open-source Strix agent.
Coding and agentic workloads are asking more of AI models than ever: refactor a repository spanning hundreds of files, sustain a multi-hour agentic workflow without losing context, and reason through complex systems problems with tool use at every step. Meeting those demands with open-weight models has historically meant provisioning and operating your own inference infrastructure. GLM 5.3 from Z.ai (Zhipu AI) is now available on Amazon Bedrock . GLM 5.3, as published on Hugging Face Hub , is a…
O overfeed.news indexa e aponta. Publicamos um trecho curto — o artigo completo fica em AWS Machine Learning Blog.
Mais de AWS Machine Learning Blog
Entre para seguir esta fonteICYMI: What landed for AI builders in September 2026
A monthly recap of the latest Amazon Bedrock, Amazon Bedrock AgentCore, and Strands updates from September 2026: broader model choice, faster serverless agents with built-in evaluation, and automated knowledge base syncing with native enterprise connectors.
How Postman runs Agent Mode for 40 million developers on Amazon Bedrock
Building an AI agent that works in a demo is a different problem from running one for 40 million developers. Postman and AWS share the architectural patterns behind Agent Mode: controlling tool sprawl, exposing schema-based reads, and treating context as the real bottleneck, plus how it runs on Amazon Bedrock at scale.
Pay-per-inference for AI agents: How BlockRun and Incarna use Amazon Bedrock AgentCore payments
Amazon Bedrock AgentCore payments gives AI agents a managed way to pay for services on demand, with spending limits enforced by the infrastructure. See how Incarna's agents pay BlockRun for model inference one request at a time over x402, cutting the work of adding x402 payment support from months to days.
Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod
A reference architecture for securely sharing one Amazon SageMaker HyperPod EKS cluster across multiple teams, using AWS IAM Identity Center for authentication, per-team SageMaker Domains and Kubernetes namespaces for isolation, HyperPod Task Governance for fairness, and namespace-level cost allocation for chargeback.