Latency
Time elapsed from sending a request to receiving a complete response.
Why it matters
User-facing apps and batch jobs have different latency budgets.
Example
Measuring time-to-first-token vs total completion time across models.
Time elapsed from sending a request to receiving a complete response.
User-facing apps and batch jobs have different latency budgets.
Measuring time-to-first-token vs total completion time across models.