Menu
InfoQ Architecture·October 8, 2026

Architecting Multi-Agent AI Systems at Scale: Spotify's Advertising Platform

This article discusses Spotify's approach to building a production-grade multi-agent AI platform for advertising, emphasizing architectural patterns for scalability and reliability. It covers key considerations such as agent boundary definition, tool design for LLMs, deterministic guardrails, and tracing-based evaluation strategies. The insights offer valuable lessons for designing and operating complex AI-powered distributed systems.

Read original on InfoQ Architecture

Introduction to Spotify's Ads AI Platform

Spotify's Ads AI platform utilizes a multi-agent architecture to generate ad scripts, recommend audiences, and ensure policy compliance. This system has seen significant adoption, with over 70% of ads using AI tools and tens of thousands of creatives generated for thousands of advertisers. The core goal is to enable advertisers to describe their desired ad in natural language, abstracting away the complexities of copywriting and music generation. The system leverages Google's ADK (Agent Development Kit) for orchestration and integrates with Vertex AI for LLM capabilities.

Core Principles of the Multi-Agent Architecture

  • One Agent, One Package, One Owner: Each agent is an independent package with clear ownership, responsibility, monitoring, and prompt definition. This enforces modularity and prevents monolithic agent designs.
  • Deterministic Guardrails: Not everything should be controlled by agents. Critical business logic that can be deterministically coded should be, ensuring reliability and predictable behavior. This includes policy enforcement and data grounding via APIs.
  • Separation of Concerns: Safety and orchestration are managed at a separate layer from individual agents, providing consistent guardrails and common orchestration logic across all agents.

Architectural Components and Flow

The platform architecture consists of several layers:

  • Client Applications: Initiate gRPC service calls and manage sessions with the agent system.
  • gRPC Services: Handle session management and route requests to the AI agents layer.
  • AI Agents Layer: Contains individual, isolated agents (e.g., Ad Script Generation Agent, Ad Guardrail Agent, Audience Recommendation Agent). These agents run on a shared context managed by Google ADK.
  • Model and Tool Platform: Provides the LLM runtime (Vertex AI, internal GCP), plugins (for accessing Spotify's internal APIs for data grounding), and observability configurations for all agents.
ℹ️

Importance of Tooling and Data Grounding

A key architectural insight is the explicit grounding of LLM responses using external tools and APIs. Instead of allowing LLMs to "guess" data like geo-interest, specific API endpoints provide factual, deterministic data. This prevents hallucination and ensures accuracy, especially crucial in advertising where precision matters for targeting and budgeting.

Key Design Considerations for Multi-Agent Systems

  • Agent Boundary Definition: A critical and often challenging aspect is determining what responsibilities an agent should handle versus what should be deterministic code or external tool calls. The goal is to control blast radius and align ownership.
  • Tool Design: Tools define how agents interact with external systems. Their schemas are not just for developers but also guide the LLM's understanding and invocation of these tools. This is akin to prompt engineering for structured interactions.
  • Reliability via Evaluation and Tracing: Implementing robust tracing is fundamental for understanding agent behavior and for enabling effective evaluation strategies. This helps in debugging, performance monitoring, and ensuring the agents deliver expected value.

Spotify enforces modularity through compile-time checks (Bazel visibility) to prevent unauthorized agent imports, ensuring clear dependencies and preventing tight coupling between agent teams. Shared platform components provide essential cross-cutting concerns like metrics, traceability, and a moderation layer for policy enforcement.

multi-agent systemsLLM orchestrationAI architecturemicroagentsSpotifyGoogle ADKproduction AIadvertising platform

Comments

Loading comments...