← All writing
articleJun 01, 202519 min read

Travel Recommendation Engine: Why the Obvious Approach Fails and What Actually Works

Candidate generation, two-stage ranking, exploration, and the feedback loops that turn a travel recommender into a filter bubble.

Machine LearningSearchPersonalizationArchitecture
Travel Recommendation Engine: Why the Obvious Approach Fails and What Actually Works cover illustration

The first version of every travel recommender is the same three lines: index the hotels, run collaborative filtering, return the top ten. It works in a demo with 5,000 hotels and 50,000 users. It does not work in production, and the reasons it fails are structural, not tuning problems.

The scale, and what makes it hard

  a mid-sized travel marketplace
    2 million listings
      - 400,000 hotels
      - 1.2M vacation rentals
      - 380,000 restaurants
      - 25,000 experiences
    8 million users
    40 million booking events
    impressions: billions

  the interaction matrix at 8M x 2M
    density if every user viewed 30 items:
      240M filled cells / 16,000,000,000,000,000
      = 0.0000015%

  candidate generation must go from
    2,000,000 -> ~500 in under 100ms

Two properties of this data shape every design decision.

Extreme sparsity. With 0.0000015% density, collaborative filtering has almost no signal per user. A user with 30 views across 30 different cities is not a preference signal — it is noise. This is why content and context carry most of the weight, and why the best systems use collaborative filtering as one feature among many rather than as the model.

A long-tail catalogue. A small fraction of listings get most of the bookings. The head is well covered by collaborative filtering. The tail — which is most of the catalogue, and much of what makes the marketplace worth browsing — has no interaction history and must be recommended on attributes alone. This is a cold-start problem on the item side, and it is a permanent feature of the business, not a bootstrap phase you pass through.

Two systems, not one

  +----------------------------------------------------------+
  |  STAGE 2: ranking                                        |
  |  order ~500 candidates precisely for this user            |
  |  a learned model, ~20-50ms                              |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  STAGE 1: candidate generation                           |
  |  2,000,000 -> ~500                                       |
  |  must be fast, cheap, and have high recall               |
  |  ~20-80ms, runs on every request                         |
  +----------------------------------------------------------+

  the rule
    recall@500 of 1.0 matters far more than
    ranking accuracy

    if the perfect answer is not in the
    candidate set, no ranker can find it

Stage 1 is where most engineering effort goes, and it is where most projects underinvest. The metric that matters is recall@k: of the items the user would have engaged with, what fraction is in the candidate set. Recall of 0.9 with a mediocre ranker usually beats recall of 0.7 with a great ranker, because the ranker can only reorder what it receives.

Practical generation sources, run in parallel and merged:

Source Why it contributes
Recently viewed / booked city Strongest single signal in travel
Similar listings (“you may also like”) Item-to-item, works on the tail
Collaborative neighbours User-to-user, strong on the head
Content and attribute match The only source for brand-new listings
Popular in destination and dates The cold-start fallback that always works
Price and availability band match Users cluster tightly on price sensitivity

The merge is a deduplicated union, usually with each source contributing a variable number of slots — a fixed 50 from each source is wrong because source quality varies by user and by context.

What the ranker actually sees

  features, roughly grouped

  user
    - home city, travel history summary
    - average price band actually booked
    - average trip length
    - party size, children, pets
    - loyalty tier, if any

  item
    - category, brand, star rating
    - price, price per night
    - review score, review count
    - amenities
    - distance from city centre
    - property type

  context (the most valuable group)
    - destination
    - dates, and how far away
    - day of week
    - trip occasion: leisure vs business
    - weather at the destination
    - is this a first visit to the city
    - how late in the session it is

  interaction
    - impression count
    - click count
    - click-through rate, impression-weighted
    - last interacted timestamp
    - was this saved / shortlisted

The context group is where a travel recommender earns its keep, and it is consistently underweighted in first attempts. A user who books a 3-star city hotel in a warm destination in March and a 5-star beach resort in July in the same year is not contradictory. They are two trip types. A model without occasion and weather features will average these into mush.

Filtering is a constraint, not a feature

This is the single most common correctness bug in travel recommenders.

  WRONG
    features: price = 180, maxPrice = 150
    model scores candidates
    -> the 145-rated hotel scores highest
    -> the 180 one gets shown
    -> the user notices, loses trust permanently

  RIGHT
    stage 1 filter: price <= 150, rating >= 4.0,
                    2 bedrooms, airport <= 20km
    stage 2 ranks whatever survives
    -> the constraint is not negotiable
    -> the model optimises within it

A user who sets a hard budget and then sees a result above it does not treat it as a ranking error. They treat it as the system not listening, and the session is over. Hard filters go in candidate generation, before any scoring, and the filter set comes directly from the explicit UI controls.

There is a softer middle ground worth knowing: some constraints are preferences with weights (a user who has never booked above $200 but has “budget mode” off) and some are absolute. Learn the distinction from how the user engages — if they repeatedly widen a filter rather than scrolling past results, the filter may be a preference. If they filter and then never widen, it is a hard constraint.

Cold start, and why the first session matters

  no history: 0 views
    1. the user's stated intent
       (destination, dates, party, budget)
    2. destination-level popularity
       for those exact dates
    3. high-review-count listings
       in the user's price band
    4. editorial or algorithmic "best of"
       for the destination

  rule
    never show a brand-new listing with
    2 reviews to a user with no history

    a low-review-count item cannot be
    recommended on quality, only on
    attributes, and the user has no
    reason to trust the attributes yet

The first-session strategy that works: use the destination, not the user. “Best hotels in Lisbon for March under $200 with rating above 4.4” is a good first result set. It is a wide, safe, and immediately useful answer, and it gives the user something to click, which generates the signal that every later recommendation needs.

The listing cold start runs on a different clock. A new property has no reviews, no clicks, and no bookings. Give it a fixed exposure budget — display it in browse and search results at controlled rates, with elevated position when it matches strong contextual criteria, and cap the daily impressions. Track the conversion carefully: early impressions land on an audience that did not ask for this property, so a low early conversion rate is not evidence the property is bad.

Position bias and the click problem

  position 1: 12% CTR
  position 5: 5% CTR
  position 10: 2% CTR

  if you train on raw clicks, the model learns
    "position 1 is good"
  not
    "this item is good"

  and then it ranks position 1 higher
  -> which makes position 1 get more clicks
  -> a self-reinforcing loop with no
     information about quality in it at all

The standard fix is to make clicks position-aware. Train with an inverse propensity weight, where each click is down-weighted by how likely it was to happen at that position regardless of quality. Where you lack propensity data, inject randomised exploration: a small percentage of impressions deliberately randomised across positions, which both generates the propensities and produces an unbiased sample.

The signals worth more than the click, in order:

  1. Booking. The only signal that reliably reflects satisfaction, and it is sparse and delayed.
  2. Save to list / shortlist. An explicit statement of intent, cheap for the user, and much more predictive than a click.
  3. Detail view with dwell time. A click that stays for 40 seconds is a different signal from a click that bounces in 800ms.
  4. Click. Weak, biased, and used in quantity.

Two-stage ranking beats one big model

  stage A: 500 candidates
    a GBDT (LightGBM / XGBoost) over
    the full feature set
    ~20ms, handles thousands of features
    and non-linear interactions well

  stage B: top 50 from A
    a small neural model (two-tower or
    a shallow DNN) with embeddings
    - user embedding
    - item embedding
    - context embedding
    ~20-30ms
    captures taste that explicit features
    cannot express

  output: final ordering

Why split. The GBDT is strong on structured features, which is most of what stage A has, and it is fast and debuggable — you can print feature importances and see that a misconfigured date feature is dominating. The neural model is strong on learned representations, which is what stage B needs for the candidate set where the decision is genuinely hard. Running the neural model over 500 candidates is 10x the latency for no benefit, because the top 50 it would reorder are all similarly scored by stage A.

The value of the split shows up in operations. When quality regresses, stage A or stage B identifies which half changed. With a single model, you have an unexplained accuracy drop.

Exploration and the filter-bubble problem

  exploitation
    show the user what they clicked before
    -> they click it again
    -> the model confirms the model
    -> no information gained

  a travel recommender that only recommends
  the Lisbon boutique hotel they already
  saved will never introduce them to the
  thing they would have loved

A travel catalogue has enormous long-tail value, and an exploitation-only recommender cannot see any of it. The mechanism that works is constrained exploration: deliberately inject items that are different but defensible — a different neighbourhood, a different property type, a different price band, at the same quality level. The constraint is what makes it safe. Random exploration produces bad results and trains the user to ignore the surface.

The measurement is a genuine tension worth naming: the recommendation system’s short-term click metric will fall when you explore, and the long-term metric that pays for exploration is hard to attribute. The workable practice is to hold out a small exploration cohort permanently, so you have a permanently unbiased read on what the catalogue actually contains and not just what the previous model showed people.

Serving architecture

  OFFLINE, every few hours
    - user profile: recent interactions,
      aggregate preferences, trip history
    - item profile: attributes, quality
      scores, availability calendar summary
    - item embeddings, trained
    - user embeddings, trained
        |
        v
  +----------------------------------------------------------+
  |  CANDIDATE CACHE                                        |
  |  precompute neighbours, "similar to this",             |
  |  "people who stayed here also stayed",                  |
  |  for the top ~50k items users actually browse          |
  |  NOT for 2M items x 8M users                           |
  +----------------------------+-----------------------------+
                               |
  REQUEST
    filter (hard constraints from UI)
        |
        v
  +----------------------------------------------------------+
  |  generation: online sources + cache hits                |
  |  ~500 candidates in <80ms                               |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  stage A GBDT (20ms) -> stage B NN (25ms)               |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  post-processing                                          |
  |  - dedupe near-identical listings                       |
  |  - cap results per host / per neighbourhood              |
  |  - enforce inventory and availability                    |
  |  - the deterministic exploration slots                  |
  +----------------------------+-----------------------------+
                               |
                               v
  response + impression logging with full context

  watch: recall@500, CTR by position, booking conversion,
        exploration slot performance, cold-start survival rate

The candidate cache is the single biggest latency lever and it is cheap to build: precompute item-to-item neighbours and user-to-user neighbours for the small set of entities that carry real traffic, since the tail gets recommended through content and popularity instead.

Failure stories worth testing

Set max price to 150 and check the results

No result may exceed 150. This is the filter-integrity test and it is where hand-rolled ranking code usually fails.

Train on raw clicks with no position weighting

Measure CTR by true relevance rank. The model will look better on its own metric and worse in reality, which is the entire point of the test.

Remove the destination from the feature set

Performance collapses. Confirms context dominance and shows you how much of the “personalisation” is actually just “where are they going”.

Show a 2-review listing to a brand-new user as position 1

They bounce. Confirms that low-review-count items need an exposure ramp.

Fill the slots entirely from “recently viewed” similarities

Click-through will not drop. Booking rate will. This is the filter-bubble failure and the metric that detects it is the conversion metric, not the engagement metric.

Compare stage A ranking against the final stage A+B output

Any improvement here is free latency reduction. If the neural model reorders almost nothing, it is not earning its 25ms.

Serve the same destination to a user 40 times in a session

Engagement drops after roughly the fourth rotation. Confirms the need for diversified slots and exploration, not just top-N.

Turn off all exploration for two weeks

The permanently-explored cohort’s booking rate should be higher. That gap is the argument for keeping exploration on.

Duplicate a listing with 90% identical content into the results

Check the dedupe and host-diversity caps. A marketplace that shows eight rooms from one host reads as an ad, and users treat it as one.

Make availability 3 days stale

Confirm sold-out listings are filtered before ranking, not after. Filtering after ranking is how users end up clicking a dead listing.

Give a returning user a completely different first screen than a new user with identical filters

Measure whether the difference helps. It usually does, which means the user’s history is being underused.

Serve results with all candidates from one neighbourhood

Compare against a forced-neighbourhood-diverse set. This is the test for whether “diversity” means anything in your implementation.

A production-ready checklist

  1. Split generation from ranking and measure recall@k on generation. Optimise the ranker only after recall is high.
  2. Run multiple generation sources in parallel and merge, with source-specific slot budgets rather than a fixed quota per source.
  3. Apply hard filters in candidate generation, before any scoring. Never let a model outrank an explicit user constraint.
  4. Weight the context features heavily: destination, dates, occasion, weather, party composition. That is where travel-specific signal lives.
  5. Use bookings and saves as primary targets. Treat clicks as a biased, secondary signal and correct for position.
  6. Inject randomised exploration to obtain unbiased propensities, and keep a permanent exploration cohort.
  7. Rank the tail on content and popularity; do not attempt collaborative filtering on items with no interaction history.
  8. Give new listings a capped exposure budget and evaluate them on contextual conversion, not raw early CTR.
  9. Use a GBDT then a neural model rather than one large model, so failures are attributable and latency is controlled.
  10. Precompute neighbours for the entities that get real traffic, and serve the long tail from content and popularity.
  11. Enforce diversity in the post-processing stage: dedupe, cap per host, cap per neighbourhood.
  12. Log the full context with every impression, or the training data will be unrecoverable later.
  13. Track conversion as the primary metric alongside CTR, because exploration and diversity always look like a CTR loss.
  14. Keep a holdout group to measure what the catalogue truly contains, independent of what past models recommended.
  15. Version the features. A regression traced to a feature change is minutes of work; an unversioned model is days.

Common mistakes

Mistake What actually happens Better decision
Collaborative filtering as the model Nothing to learn from at 0.0000015% density Content, context, and popularity carry it
One stage over 2M items Unaffordable latency, no recall measurement Generate then rank, and measure recall
Filters applied as features The model violates them and trust is gone Filter in generation, before scoring
Training on raw clicks Learns position, not quality Propensity weights, plus exploration
Ignoring context Averages a business trip with a family holiday Weight occasion, weather, dates heavily
Recommending low-review listings to new users The user has no basis to trust them Exposure ramp, high-review items for cold users
Top-N from one source Filter bubble, no long-tail discovery Multiple sources, diversity caps, exploration
Exploiting only Never surfaces the long tail Constrained exploration, defensible alternatives
One large neural model Latency cost, unattributable regressions GBDT then neural, both cheap to debug
No dedupe or host caps Reads as an advertisement Post-processing diversity
Availability filtered after ranking Users click dead listings Filter availability before scoring
Judging by CTR alone Exploration looks like a loss always Conversion and long-term cohorts
No impression context logged Training data is unrecoverable Log the full context per impression
Same first screen for everyone Wasteful, and worse for new users Rank per stage of user maturity
Optimising only for the head Most of the catalogue is unreachable Content path for the tail
Forgetting booking lag Attribution lands on the wrong trip Delayed-conversion-aware training
Unversioned features A quality drop with no explanation Version features with the model

The complete story in one minute

A travel recommender is two systems. Candidate generation turns two million listings into five hundred, and its recall is the number that matters — a perfect ranker over a set that excludes the right answer is still wrong. Generation has to be fast, cheap, and cover the long tail, and the honest assessment is that most of the value is not collaborative filtering. At 0.0000015% interaction density, content and context do the work, and collaborative filtering is one feature among many. The context features are the ones people skip: occasion, weather, trip length, party composition. A user booking a 3-star city hotel in March and a 5-star beach resort in July is not contradictory, they are two trip types.

Hard filters go in generation, never as features. A user who sets a $150 ceiling and sees a $180 result has learned the system is not listening, and that is unrecoverable. The signals, ranked: booking, then save-to-list, then detail view with dwell, then click — because clicks are position-biased, so train with propensity weights and inject randomised exploration to get unbiased data and surface the long tail at the same time. A pure exploitation recommender will never show anyone the listing they would have loved, and that catalogue depth is what the marketplace is selling.

Two-stage ranking — a GBDT over the 500, then a small neural model over the top 50 — because the GBDT handles thousands of structured features fast and debuggable, and the neural model is needed only where the decision is hard. The operational benefit is attributability: when quality regresses, you know which half changed. Cold start is solved differently on each side: new users get destination-level answers, which are wide and safe and generate the first signal, while new listings get a capped exposure budget evaluated on contextual conversion, because early impressions land on an audience that did not ask for that property.

Serve it with a neighbour cache for the entities that get real traffic, content and popularity for the tail, post-processing dedupe and host caps so the page does not read as an ad, and a permanently held-out exploration cohort. Measure conversion, not click-through, because exploration and diversity will always look like a CTR loss and the only way to defend them is a number that shows what they are worth.

generate wide, measure recall, filter before ranking
context features over collaborative filtering
propensity-weighted clicks plus real exploration
GBDT then neural, cold users by destination, cold items by exposure ramp
conversion as the metric, diversity enforced in post-processing

The hard part was never scoring hotels. It was knowing which two hundred hotels to score, and being honest about the fact that a click at position one is mostly a statement about position.

What this team still owns

Log the candidate set, eligibility filters, feature/model versions, position, exploration policy, and outcome so ranking changes can be evaluated without pretending displayed items were random. Hard constraints—dates, occupancy, legal availability, accessibility requested by the user—run before scoring. Exploration has a quality floor and an exposure budget, and its long-term effects are measured by booking, cancellation, repeat use, supplier concentration, and destination coverage rather than CTR alone.

Technical references

Keep reading
Browse everything