LLM latency: the number your users actually feel
· 5 min read
TTFT, tokens per second, total completion, end-to-end — four numbers that diverge, and your interface decides which one matters. Why the tail is fat, why output length is the real lever, and why a latency gap between arms is usually an SRM.