overfeed.news

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

1mês

Idade

Publicado
Coletado
Imagem: AWS Machine Learning Blog

Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts GPU infrastructure by 75% while holding sub-second latency at 92.1 requests per second per GPU.

Trecho da fonte

This post is a collaboration between AWS, NVIDIA and Heidi . Reducing automatic speech recognition (ASR) inference costs on Amazon Elastic Compute Cloud (Amazon EC2) becomes critical when GPU utilization per request is low but latency requirements are strict. A single ASR inference request typically uses only 15–20 percent of a GPU’s compute capacity, yet the default time-slicing behavior in NVIDIA CUDA® forces sequential access, leaving 80 percent of the hardware idle. Heidi Health is an AI…

Leia o artigo completo em aws.amazon.com

O overfeed.news indexa e aponta. Publicamos um trecho curto — o artigo completo fica em AWS Machine Learning Blog.

Entre para seguir esta fonte
Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2…