In-House LLM Serving at Netflix
- Netflix: 13 events in the last 90 days
- Previous: 3 days earlier · Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned
What happened
By AI Platform’s Model Runtime team and Inference team Introduction Most organizations consume LLMs through hosted APIs. Netflix went further — we run the full stack ourselves, from model deployment through inference, inside our existing production environment rather than a separate ML silo. Some of those decisions weren’t obvious, and a few revealed their trade-offs only under production load. This post focuses on…
Summary assembled by rule from the sources below