Reading Latency Percentiles Without Fooling Yourself
By Priya Raman · May 7, 2026 · Operations
Arithmetic averages hide the extreme tail latencies that percentiles make evident. Even if an endpoint posts an average duration of fifty milliseconds, one out of twenty calls might experience a grueling two-second delay; customers subjected to cold paths and overloaded database shards are the ones who submit urgent issue reports.
Interpreting percentile curves carries distinct analytical hazards. Tracking p95 shifts across deployment boundaries only holds meaning if query distributions remain static; sudden traffic shifts distort the curve through pure demographic variance. Furthermore, observing p99 fluctuations across tiny sixty-second windows reflects short-term jitter rather than genuine system health.
Maintain compact logarithmic histograms at proxy edges for downstream aggregation. Recording histogram distributions allows operators to calculate accurate percentiles across arbitrary timeframes, while arithmetic means discard variance permanently.