Menu
Pinterest Engineering·October 9, 2026

Pinterest's User Journey Inference with LLMs: A Generative Approach to Personalization

This article details how Pinterest evolved its user journey inference system from a multi-stage clustering pipeline to a single, fine-tuned 4-billion-parameter LLM. The new generative approach reads chronological activity logs to directly output a ranked list of user journeys, significantly improving personalization quality and addressing limitations of the legacy system. The architecture relies on distilling large frontier models into smaller, production-ready student models and optimizing the serving stack for high throughput at notification scale.

Read original on Pinterest Engineering

Evolution of User Journey Inference

Pinterest's core challenge in personalization is understanding long-horizon user intents, or "journeys," which span multiple sessions and activities, rather than just optimizing for the next click. The initial system was a multi-stage pipeline involving keyword extraction, embedding, hierarchical clustering, and heuristic-based classification. While effective for cold-start and yielding significant improvements in engagement metrics, this legacy system suffered from fragmentation (splitting a single project into multiple journeys), generic naming, and sensitivity to incidental noise.

From Multi-Stage to Generative Model

To overcome the limitations of the clustering pipeline, Pinterest reimagined journey inference as a single generative act: reading user activity and directly writing intent. This involved moving to a fine-tuned 4-billion-parameter Large Language Model (LLM) that takes a structured prompt containing user profile and chronological activity log, then outputs a strict JSON object with ranked journey names. This shift eliminates multiple intermediate steps, simplifying the architecture and improving coherence.

📌

Generative Journey Inference Prompt Structure

The LLM prompt is crucial, acting as the ranking and safety policy. It includes: - User Profile: Gender, language, country, age. - User Owned Boards: Context from explicit user actions. - User Activities: Chronological log of actions (search, save, like, click, dislike) grouped by date/action, with specific delimiters for token economy. Key prompt instructions include: - Output Format: Strict JSON, fixed output schema. - Output Constraints: Up to 12 distinct 1-5 word journeys, in user's locale, ranked by engagement strength, repeated support, recency, coherence, and specificity. - Safety Policies: Avoid merging unrelated topics, drop one-off outliers, treat auto-generated tags as noisy, lightly diversify, allow empty output but forbid inventing unsupported journeys. This consolidates logic that was previously spread across multiple models and heuristics into the prompt itself, demanding rigorous review for prompt changes.

Model Distillation and Serving Architecture

Running large frontier models for hundreds of millions of users is prohibitively expensive. Pinterest addressed this through model distillation, using larger, more capable frontier models as 'teachers' to generate high-quality synthetic training examples. These examples are then used to fine-tune smaller, open-weight 'student' models (Qwen3 in this case). The 4-billion-parameter model was chosen for its optimal balance of output quality and serving throughput.

  • Supervised Fine-Tuning (SFT): Learning from teacher-generated examples.
  • Low-Rank Adaptation (LoRA): Efficiently updating small adapters instead of all model parameters.
  • Data Diversity: Prioritizing diverse and difficult consolidation/multilingual examples over simply more examples.

For serving at notification scale, Pinterest leverages an optimized batch inference system. While initial thought might lean towards offline batch, online inference won on throughput due to advanced decode optimizations: - PagedAttention: Efficient use of key-value cache by avoiding padding waste. - Continuous Batching: Immediately reassigning GPU slots from finished sequences. - Prefix Caching & Chunked Prefill: Reusing common system prompt prefixes across requests. The serving stack utilizes NVIDIA Dynamo with vLLM decode workers on L40S GPUs, with independent scaling for frontend and decode workers and KV cache-aware routing to maximize efficiency. This architecture sustains approximately 775 requests per second end-to-end with a cluster of around 100 GPUs, highlighting critical considerations for deploying large generative models in high-throughput, low-latency environments.

Impact and Future Implications

The generative approach significantly improved journey quality, leading to better consolidation of fragmented intents, more specific journey names, and multilingual support not possible with the legacy system. The domain-agnostic nature of this 'activity-to-intent' formulation suggests it can be adopted by other platforms with sequential user activity and structured intent targets, potentially collapsing classical multi-stage recommendation pipelines into a single generative model.

LLMmachine learningpersonalizationmodel distillationgenerative AIserving architecturehigh-throughputPinterest

Comments

Loading comments...