← 返回事件
基本结束AI

In-House LLM Serving at Netflix

发生了什么

By AI Platform’s Model Runtime team and Inference team Introduction Most organizations consume LLMs through hosted APIs. Netflix went further — we run the full stack ourselves, from model deployment through inference, inside our existing production environment rather than a separate ML silo. Some of those decisions weren’t obvious, and a few revealed their trade-offs only under production load. This post focuses on…

摘要按规则整理自下方来源原文

为什么在扩散

来源

一手来源