← Back to events
ActiveAI

The efficient frontier of LLM inference

Photo: Hacker News

What happened

Inference techniques either move a deployment along the latency–throughput frontier or push the entire frontier out, creating more efficiency to allocate.

Summary assembled by rule from the sources below

Why it's spreading

Sources

Community