Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2
1mês
- Publicado
- Coletado

Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts GPU infrastructure by 75% while holding sub-second latency at 92.1 requests per second per GPU.
This post is a collaboration between AWS, NVIDIA and Heidi . Reducing automatic speech recognition (ASR) inference costs on Amazon Elastic Compute Cloud (Amazon EC2) becomes critical when GPU utilization per request is low but latency requirements are strict. A single ASR inference request typically uses only 15–20 percent of a GPU’s compute capacity, yet the default time-slicing behavior in NVIDIA CUDA® forces sequential access, leaving 80 percent of the hardware idle. Heidi Health is an AI…
O overfeed.news indexa e aponta. Publicamos um trecho curto — o artigo completo fica em AWS Machine Learning Blog.
Mais de AWS Machine Learning Blog
Entre para seguir esta fonteICYMI: What landed for AI builders in September 2026
A monthly recap of the latest Amazon Bedrock, Amazon Bedrock AgentCore, and Strands updates from September 2026: broader model choice, faster serverless agents with built-in evaluation, and automated knowledge base syncing with native enterprise connectors.
How Postman runs Agent Mode for 40 million developers on Amazon Bedrock
Building an AI agent that works in a demo is a different problem from running one for 40 million developers. Postman and AWS share the architectural patterns behind Agent Mode: controlling tool sprawl, exposing schema-based reads, and treating context as the real bottleneck, plus how it runs on Amazon Bedrock at scale.
Pay-per-inference for AI agents: How BlockRun and Incarna use Amazon Bedrock AgentCore payments
Amazon Bedrock AgentCore payments gives AI agents a managed way to pay for services on demand, with spending limits enforced by the infrastructure. See how Incarna's agents pay BlockRun for model inference one request at a time over x402, cutting the work of adding x402 payment support from months to days.
Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod
A reference architecture for securely sharing one Amazon SageMaker HyperPod EKS cluster across multiple teams, using AWS IAM Identity Center for authentication, per-team SageMaker Domains and Kubernetes namespaces for isolation, HyperPod Task Governance for fairness, and namespace-level cost allocation for chargeback.