Latency

Time elapsed from sending a request to receiving a complete response.

Why it matters

User-facing apps and batch jobs have different latency budgets.

Example

Measuring time-to-first-token vs total completion time across models.