This article details Netflix's approach to building an end-to-end (E2E) observability system, tackling the challenges of correlating user experience with system performance at massive scale (38M events/sec). They discuss transitioning from reactive monitoring to an AI-driven operational ontology, leveraging knowledge graphs and agentic workflows for automated triaging, root-cause analysis, and self-healing systems. The core challenge is framed as a data engineering problem to unify disparate telemetry data into an actionable knowledge graph.
Read original on InfoQ ArchitectureNetflix operates at an immense scale, handling 65 million concurrent streams, over 2 billion requests daily, and processing more than 38 million real-time logging events per second across thousands of microservices and diverse client platforms. Traditional reactive monitoring and alerting systems prove insufficient at this scale, leading to prolonged incident resolution times and significant engineering overhead. An average incident can involve nine teams and over 30 engineers, taking hours to resolve, highlighting the need for a more efficient and proactive approach to observability.
Netflix's vision shifts observability from reactive monitoring to a proactive insights engine. The goal is a system capable of automatically detecting, prioritizing, triaging, and root-causing issues across the entire stack within minutes. Ideally, it should predict issues before they impact users and even suggest or take corrective actions. This requires seamless visibility from user devices through networks to backend dependencies, all in real-time. The core principle to achieve this is connectedness.
The Power of Connected Data
The fundamental architectural shift involves unifying disparate telemetry data (MELT: Metrics, Events, Logs, Traces) into a single, integrated layer. This unified data layer forms the foundation for a consistent language across dashboards, alerts, debugging workflows, and incident response tools, significantly reducing duplicated effort and improving diagnostic accuracy.
To achieve connectedness, Netflix is building an operational ontology and an E2E knowledge graph. This involves standardizing the structure and relationships between all observability data points. By representing services, users, devices, applications, and infrastructure as nodes and their interactions as edges in a graph database, the system can perform complex queries to understand causality and impact dynamically. This knowledge graph, combined with AI/ML, enables advanced AIOps capabilities.