This article details Airbnb's "Proximity Features" system, an architectural solution to the cold-start problem in personalization for new or logged-out users. By leveraging coarse geographic location derived from IP addresses, the system groups users into "proximity buckets" and aggregates collective behavior signals, enabling personalized recommendations without relying on individual user identities or persistent identifiers. The design addresses privacy constraints and delivers real-time, location-aware personalization.
Read original on Airbnb EngineeringPersonalization models typically rely on extensive user history. However, for a significant portion of users (e.g., first-time visitors, logged-out users, or those from marketing campaigns), this data is unavailable, leading to a "cold-start" problem. Traditional approaches fall back to generic content, resulting in a poor user experience. Additionally, evolving privacy regulations (like GDPR) and browser restrictions further limit the use of persistent user identifiers, compounding the challenge for personalized experiences.
Airbnb's solution, "Proximity Features," addresses the cold-start problem and privacy concerns by using aggregate activity patterns from geographically proximate users. The core insight is that users in the same local area often share similar travel interests. The system derives a coarse geographic location from a user's IP address, then assigns them to an "adaptive proximity bucket" containing approximately 1,000 nearby users. Features for personalization are then computed and served based on the collective behavior within these buckets, without ever needing a persistent individual user identifier at inference time.
A proximity key is a compact group key representing a local geographic cluster. It encodes a quantized latitude/longitude tile and, for dense areas, an IP hash bucket index. This key acts as an aggregation identifier, similar to a `user_id`, allowing ML models to condition on collective local behavior. Features computed daily for these buckets include short-term engagement (e.g., top destinations, median prices from recent searches), long-term booking patterns, and aggregate bucket metadata (e.g., size, geographic density). This allows a new user to effectively "borrow" signals from their local crowd.
System Design Insight
The use of a proximity key as an aggregation key enables a clean separation: ML models can consume feature vectors tied to a group identifier, abstracting away the complexities of individual user identity management and cold-start scenarios. This modularity is crucial for maintaining scalability and privacy compliance.
A key challenge is creating stable geographic buckets of consistent size (~1,000 users) given the vast differences in user density globally. Airbnb employs a two-phase adaptive clustering algorithm: for dense areas, it subdivides geographic tiles further using IP hash buckets; for sparse areas, it applies multi-pass coarsening, progressively widening the geographic tile until enough users accumulate. This creates a zoom-adaptive map, fine-grained in cities and coarser in rural regions. The clustering is designed for stability and runs efficiently at a global scale on geo-IP coordinate data.
At serving time, a user's IP is resolved in real-time to a proximity key, and features are fetched from a distributed key-value store. This lookup is implemented as a soft dependency, meaning personalization proceeds without it if the lookup times out, ensuring the core request path is never blocked.