← 返回事件
持续讨论AI

Introducing Amazon SageMaker HyperPod Inference Gateway

图:AWS Machine Learning

发生了什么

Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing add-on for Amazon EKS. It uses real-time GPU signals to send each inference request to the best-suited pod, cutting first-token latency by up to 82% with no changes to your model servers or client applications.

摘要按规则整理自下方来源原文

为什么在扩散

来源

一手来源