Netflix's GenRec is a novel recommendation system that leverages Large Language Models (LLMs) to enhance content personalization. It addresses the complexity and cost of traditional feature engineering by using LLMs to interpret user behavior and content in natural language. The architecture involves a two-phase training process to adapt a foundation LLM to Netflix's specific domain and then fine-tune it for ranking recommendations, optimizing for user satisfaction and operational efficiency.
Read original on ByteByteGoTraditional recommendation systems, like Netflix's original engine, relied heavily on thousands of hand-crafted features across users, items, and interactions. While powerful, this approach leads to significant architectural complexity, high onboarding costs for new use cases, and challenges in maintaining and evolving the system as content and user behavior change. The overhead of feature engineering — identifying relevant observations, calculating their values, and making them available for training and prediction — becomes substantial.
Large Language Models offer a promising alternative due to their broad world knowledge and language understanding capabilities. LLMs can inherently grasp relationships between diverse content and interpret user action sequences expressed in natural language (verbalization). This allows for a more flexible and less rigid approach to understanding user preferences and content themes, potentially reducing the need for extensive manual feature engineering.
LLM Challenges in Recommendation Systems
While powerful, general-purpose LLMs are not suitable off-the-shelf for recommendations. They may favor globally popular items, suggest out-of-catalog content, or lack personalization. Specific training and adaptation are crucial to align LLMs with business objectives and catalog constraints.
Netflix's GenRec overcomes general LLM limitations through a two-phase training approach: * Phase 1: Netflix-aware Foundation Development: An open-source foundation LLM is adapted using proprietary Netflix data. This phase makes the model familiar with Netflix content, user behavior, and domain-specific terminology. This foundational model is stable and updated less frequently. * Phase 2: Ranking Model Training: The foundational LLM is further trained on specific examples and objectives to properly rank content based on reward signals (e.g., viewing time, satisfaction) and cost targets. This phase is updated more frequently to account for new content releases and evolving viewer preferences.
The ultimate goal of GenRec is not just to predict clicks, but to recommend content that truly satisfies users and encourages continued engagement, thereby impacting key business metrics like retention.