← Back to events
CoolingAI

An 8.6 GB model that serves only 7 requests a second

What happened

Understanding ML inference by breaking it apart

Summary assembled by rule from the sources below

Why it's spreading

Sources