Menu
The New Stack·September 27, 2026

Advanced Performance Engineering: Beyond P99s with Adrian Cockcroft

This article features Adrian Cockcroft's insights on advanced performance engineering, moving beyond traditional P99 metrics. It emphasizes the importance of understanding full response time distributions and using AI-powered custom tooling to identify nuanced performance issues that single-point metrics often miss. The discussion highlights the architectural implications of accurately monitoring and diagnosing system performance in complex distributed environments.

Read original on The New Stack

The Limitations of P99 Percentiles in Modern Systems

Adrian Cockcroft, a veteran in architecting high-performance systems, argues that relying solely on P99 (99th percentile) or other single-point percentiles can be misleading when diagnosing latency and performance issues in modern web services. These metrics often fail to reveal the underlying complexity of response time distributions, which typically exhibit multiple peaks due to factors like cache hits/misses, different service paths, or varied upstream dependencies. An average or a single percentile value can obscure critical operational details, making it harder to identify and resolve root causes.

⚠️

Beware of Misleading Percentiles

While P99s are useful for a quick overview, they can mask bimodal or multimodal distributions. Understanding the full histogram of response times is crucial for uncovering nuanced performance bottlenecks that affect a significant portion of users, even if they don't hit the 99th percentile.

Deeper Analysis: Examining Response Time Peaks

Instead of just percentiles, Cockcroft advocates for analyzing the distribution of response time peaks in a histogram. For instance, a system might show a fast peak for cache hits and a slower peak for cache misses. As the cache hit rate changes, the positions of these peaks remain constant (representing inherent latency modes), but their heights fluctuate. Traditional metrics like averages or P99s would change significantly, but only reflect a shift in hit rate, not a change in the fundamental latency characteristics of the fast or slow paths. Architectural decisions around caching strategies, for example, directly impact these distributions.

  • Understanding Bimodal Distributions: Identify distinct performance modes (e.g., fast path vs. slow path, synchronous vs. asynchronous processing).
  • Tracking Peak Fluctuations: Observe how the heights of these peaks change over time, indicating shifts in operational patterns (e.g., cache effectiveness, database query distribution).
  • Investigating Underlying Causes: Connect observed distribution changes back to specific architectural components or system behaviors to pinpoint issues.

AI and Custom Tooling for Performance Engineering

Cockcroft's approach involves building custom tools to dive deeper into performance data, a process significantly accelerated by modern AI capabilities. He utilized ChatGPT to develop a tool that identifies an arbitrary number of peaks in a response time distribution and tracks their fluctuations. This allows engineers to move beyond aggregated metrics and gain a fine-grained understanding of system behavior, enabling more precise performance optimization. Such custom tooling complements existing end-to-end tracing solutions by focusing on uncovering anomalies that are 'hidden in plain sight'.

performance engineeringlatencyobservabilitymetricsmonitoringAIdistributed systemssystem design

Comments

Loading comments...