Menu
InfoQ Architecture·October 9, 2026

Netflix's Ontology-Driven Observability: Building an E2E Knowledge Graph for AIOps

This article details Netflix's approach to building an end-to-end (E2E) observability system, tackling the challenges of correlating user experience with system performance at massive scale (38M events/sec). They discuss transitioning from reactive monitoring to an AI-driven operational ontology, leveraging knowledge graphs and agentic workflows for automated triaging, root-cause analysis, and self-healing systems. The core challenge is framed as a data engineering problem to unify disparate telemetry data into an actionable knowledge graph.

Read original on InfoQ Architecture

The Observability Challenge at Netflix Scale

Netflix operates at an immense scale, handling 65 million concurrent streams, over 2 billion requests daily, and processing more than 38 million real-time logging events per second across thousands of microservices and diverse client platforms. Traditional reactive monitoring and alerting systems prove insufficient at this scale, leading to prolonged incident resolution times and significant engineering overhead. An average incident can involve nine teams and over 30 engineers, taking hours to resolve, highlighting the need for a more efficient and proactive approach to observability.

Limitations of Traditional Observability

  • Siloed Data Sources: Metrics, events, logs, and traces are stored disparately, lacking standardization and interoperability.
  • Disconnected Alerting: Alerts are often localized and non-contextual, preventing a holistic view of system health.
  • Complex Troubleshooting: Triage and debugging become arduous tasks due to the lack of integrated insights.
  • Undetermined User Impact: Downstream services often lack visibility into how their performance affects the end-user experience.

Vision for End-to-End Observability: Connectedness

Netflix's vision shifts observability from reactive monitoring to a proactive insights engine. The goal is a system capable of automatically detecting, prioritizing, triaging, and root-causing issues across the entire stack within minutes. Ideally, it should predict issues before they impact users and even suggest or take corrective actions. This requires seamless visibility from user devices through networks to backend dependencies, all in real-time. The core principle to achieve this is connectedness.

ℹ️

The Power of Connected Data

The fundamental architectural shift involves unifying disparate telemetry data (MELT: Metrics, Events, Logs, Traces) into a single, integrated layer. This unified data layer forms the foundation for a consistent language across dashboards, alerts, debugging workflows, and incident response tools, significantly reducing duplicated effort and improving diagnostic accuracy.

Ontology-Driven Knowledge Graph Architecture

To achieve connectedness, Netflix is building an operational ontology and an E2E knowledge graph. This involves standardizing the structure and relationships between all observability data points. By representing services, users, devices, applications, and infrastructure as nodes and their interactions as edges in a graph database, the system can perform complex queries to understand causality and impact dynamically. This knowledge graph, combined with AI/ML, enables advanced AIOps capabilities.

  • Data Ingestion & Unification: Centralizing and standardizing MELT data from diverse sources.
  • Knowledge Graph Construction: Building a graph database that models relationships between all system components and their telemetry.
  • Operational Ontology: Defining a semantic layer that provides context and meaning to the raw data, allowing for intelligent querying and analysis.
  • AI-Driven Workflows: Using machine learning to process the knowledge graph for automated anomaly detection, root cause identification, and predictive insights.
  • Agentic Systems: Developing AI agents (e.g., using LLMs like Claude) to interact with the knowledge graph, automate triage, and suggest self-healing actions.
observabilityAIOpsknowledge graphNetflixtelemetrymonitoringdistributed tracingincident management

Comments

Loading comments...