← All writing
articleJul 17, 202519 min read

Flight Search System: Caching by Degradation, Not by Answer

Flight search at scale — GDS and aggregator fan-out, result caches that age gracefully, and the itinerary-combinatorics problem that makes searching hard.

SearchCachingTravelArchitecture
Flight Search System: Caching by Degradation, Not by Answer cover illustration

Flight search is presented as a search problem and is actually a distributed-fan-out-and-recombination problem with a caching requirement that fights the nature of the data. The price is a claim about a future that only the airline knows, and every design decision follows from taking that seriously.

The scale

  a flight search service
    500 queries per second average
    4,000 per second peak (a fare sale
    at 9am on a Monday)

    each query fans out to
      - 3 GDSs (Amadeus, Sabre, Travelport)
      - 15+ airline websites (scraped or
        NDC-connected)
      - 4 metasearch aggregators
      = 22 providers
      - p50 latency: 400ms
      - p95 latency: 4s
      - p99: 11s (or timeout)
      - failure rate: 2-8% per provider

    and the answer is combinatorial
      2 outbound options x 4 return
        options x 3 fare families
        = 24 bookable itineraries
      from 2,000+ raw offers

The combinatorial point is the one people miss. Providers return offers — a leg with a fare, a cabin, a fare family. The user wants itineraries — combinations of legs across providers that are actually connectable and actually bookable together. Assembling those is a real problem with real constraints, and it is the part that determines whether the results are usable or a list of legs that cannot be booked as a trip.

The fan-out

  request: SFO -> LIS, 2026-06-10 to 2026-06-17

  +----------+     +----------+     +----------+
  | GDS 1    |     | GDS 2    |     | airline  |
  | 380ms    |     | 1200ms   |     | website  |
  +----------+     +----------+     | 6.5s     |
       |                 |            +----------+
       v                 v                 |
  +-----------------------------------------+
  |  normaliser: provider format ->         |
  |  canonical offer                        |
  +---------------------+-------------------+
                        |
                        v
  +-----------------------------------------+
  |  combiner: offers -> connectable        |
  |  itineraries                           |
  +-----------------------------------------+

  the critical rule
    NEVER let a slow provider hold the
    whole response

The NEVER is the design rule the whole thing rests on. Three approaches, in increasing order of correctness:

Sequential with short timeout. A 1-second timeout per provider. Loses results on every provider slower than 1s, which is most of the direct-airline sites. Rejected.

Parallel with a global deadline. Fire everything at once, wait up to a total budget (say 2.5s), return what arrived. Simple, and it works — but the user experience depends on which providers happened to be fast, and that is a lottery.

Parallel with tiered response. Fire everything. Return the first tier (GDSs) as soon as the slowest GDS arrives — usually under 1s — and stream or silently merge the airline-direct results as they arrive. This is the right answer and it is what good aggregators do: the initial paint has the reliable providers, and the rest fill in.

The “silently merge” matters. Streaming results into the page causes the list to reorder under the user’s cursor, which is worse than a slightly slower complete list. The pattern that works is: show the GDS results immediately and completely, then merge in the rest and re-render only if the user has not interacted, or show a small “checking airline sites…” indicator with a count.

The combiner, which is the actual hard part

  offers from providers
    (legs, not itineraries)

  a valid itinerary requires
    1. legs chain geographically
       (LIS -> MAD -> LIS is fine,
        LIS -> MAD -> BCN is not)
    2. minimum connection time at each
       transfer
       - intra-EU: 45 min (often 1.5h+
         in practice)
       - long-haul: 2h minimum
       - the MCT depends on airport,
         terminal, airline, and whether
         the passenger has checked bags
    3. the total fare is bookable
       AS ONE TICKET
       - a GDS can combine its own
         offerings
       - mixing GDS offers with airline
         direct offers is often not
         bookable as a single PNR
    4. fare rules are consistent
       (baggage, change fees, whether
        the segments can be split)
    5. ticket validity and pricing are
       still live at booking time

  combinatorics
    2000 offers
    -> naive pairwise combination:
       millions of candidate pairs
    -> most are invalid
    -> filter by geography first,
       then by MCT, then by fare rule

The MCT (minimum connection time) is where naive implementations produce embarrassing results: a 25-minute connection at a large airport with a security queue and immigration, on a flight that arrives at 23:50 and the connection departs at 00:15. Real MCTs come from the provider, are airport-specific, and are longer for international arrivals than domestic ones. Using a global constant is the most common cause of “your itinerary is impossible” at booking.

The bookability constraint is the second big one. A price on a flight search page is a quote, and the quote is only meaningful if the itinerary can be purchased as a single ticket. Providers that do not support cross-provider combination must be kept in their own lane, which is why good flight search results often show a “booked by [provider]” label and why prices for the same itinerary differ between them. The system has to know, per provider, what it is capable of combining.

Caching, and the honesty problem

  the fundamental tension
    a flight search answer is only
    valid for a few seconds
    but a search UI gets hammered,
    and origin-destination-date is
    a small key space
    -> so cache, obviously

  the risk
    a user refreshes in 10 seconds
    and sees a DIFFERENT price
    -> they now distrust the site
    -> and they are right to

The resolution is to make the cache’s degradation visible and its behaviour principled:

Cache by provider offer, not by search result. A cached GDS offer for SFO-LIS on a date is far more reusable than a cached “search result” for the full query, because the same offer participates in many itineraries. This gives a much higher hit rate on stable data (fare rules, schedules) and forces only the volatile part (availability and price) to be live.

Two-tier cache. A short-lived availability cache (5-15s) for the volatile price/availability, and a long-lived static cache (hours to days) for schedule, route, and fare-rule data. Most of the hit rate comes from the static layer, and the static layer is the layer that is safe to cache hard.

Stamp every result with its age. The UI shows “prices as of 14:32” and, if a result is more than a couple of minutes old, says so. This is the cheapest possible honesty and it converts a trust problem into a non-issue.

Never serve a stale price as if it were live. If the availability cache misses and the provider is down, say “price unavailable” for that provider. Serving a 10-minute-old price as current is the failure that gets a flight site sued.

The cache key has to include everything that changes the answer: origin, destination, dates (as local dates at each airport, not UTC — a critical and frequently wrong detail), cabin, passenger counts and types (adult, child, infant — infant Laplace produces different results and different fares), and the currency and fare-filter settings. Under-key the cache and you will serve a business-class result to an economy search, which is a spectacular bug.

The date and timezone problem

  a flight departing SFO at 23:50
  on 2026-06-10 arrives in Tokyo at
  2026-06-12 (local)

  so
    - the search is for departure DATE
      2026-06-10
    - the arrival date in the destination
      is a different calendar day
    - "one way to Tokyo" has no
      destination-side date constraint
    - a "return on 2026-06-17" is a
      departure date from the destination,
      interpreted in the DESTINATION's
      timezone

  which means
    a search made in Tokyo and a search
    made in London for the same
    "SFO-LIS, Jun 10" are the same
    query
    but the *display* of the result
    depends on the viewer's timezone

Anyone who has seen flight search return a flight on the wrong day has usually hit a UTC-vs-local mismatch. The rule that works: the search key uses the origin’s local calendar date; the display converts using the current location of each airport; and the cache key uses the same local-date convention, not a UTC timestamp. This is boring and it is a whole class of bugs.

The “search is not a query” part

  a real user journey
    1. search SFO -> LIS, 2 adults
    2. see 200 results, refine by
       "direct only" (client-side filter
       is fine, it's the same result set)
    3. sort by price
    4. click a result, see fare rules
    5. go BACK, change dates by one day
    6. the whole search re-runs
    7. adjust the fare filter
    8. see 200 results again

  the engineering consequences
    - the sort and filter are client-side
      against a result set of a few
      hundred, not a re-query
    - the date change is a re-query
    - the whole flow needs to be fast
      on the FIRST query (skeleton
      state matters more than anything)
    - the user will press enter
      repeatedly; debounce and cancel
      in-flight requests, and never
      render a stale response

Two things get missed here. Client-side sort and filter is not an optimisation, it is the correct design — re-querying for a sort is slow and makes the list jump. And request cancellation is mandatory: users type dates and hit enter, and if responses arrive out of order the list flickers between searches. Every request needs a sequence number and the client must discard any response that is not the latest.

Failure stories worth testing

Make one provider return in 9 seconds

The initial render must not wait for it. This is the tiered-response test and the one that separates a usable system from a slow one.

Make one provider return an error

The error must not appear as “no flights found”. Confirm the error is scoped to that provider and the others still populate.

Make two providers disagree on the price of the same itinerary by 4%

Both should display, with the provider identified. Confirm the system does not pick one silently.

Serve a 20-minute-old availability cache

The UI must say so, and the price must be re-validated at booking. This is the trust test.

Search a 1-adult-and-1-infant query

Infant results differ (lap infant, no seat, different fare). Confirm the query differentiates and the key includes the infant.

Search “SFO to Tokyo, departing 23:50” and check the arrival date

The displayed arrival date must be in the destination’s timezone. This is the classic timezone bug.

Drop a cache key element (remove cabin, keep the rest)

Search business and confirm economy results appear. This is the under-keyed cache test.

Have a user press enter 5 times in 2 seconds

Only the last response may render. Any flicker means cancellation is not implemented.

Set a provider’s schedule data to be stale for a month

The static cache will happily serve a flight that no longer exists. Confirm schedules have a TTL and a validation path.

Force a 5,000 QPS spike

Measure whether the cache absorbs it or whether providers get hammered. The provider-facing rate limiter is the thing that keeps a traffic spike from becoming a partner outage.

Add a route with a 30-minute connection at a hub

Confirm the MCT filter rejects it, and confirm the message tells the user why.

Book an itinerary the search showed, after a 3-minute cache window

The booking flow must re-validate price and availability and fail gracefully. A price that changed at booking is not a bug, but crashing is.

A production-ready architecture

   user query
     (origin, dest, local dates,
      cabin, pax types, filters, currency)
        |
        v
  +----------------------------------------------------------+
  |  CACHE LAYER                                             |
  |  - static: schedule, routes, fare rules (hours-days)      |
  |  - volatile: availability + price (5-15s)                 |
  |  - key includes local dates, cabin, pax types, currency   |
  |  - every result carries an age                            |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  FAN-OUT (parallel, global deadline, per-provider        |
  |  isolation)                                               |
  |  - tier 1: GDSs -> return when slowest GDS lands         |
  |  - tier 2: airline direct -> merge silently on arrival    |
  |  - circuit breaker per provider                           |
  |  - rate limit per provider, shared budget                |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  NORMALISER                                                |
  |  provider offer -> canonical leg (times in local tz,     |
  |  carrier, flight numbers, equipment, fare family)        |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  COMBINER                                                 |
  |  - chain by geography and MCT (provider-supplied)        |
  |  - enforce single-PNR bookability per provider lane       |
  |  - build bookable itineraries with total fare            |
  +----------------------------+-----------------------------+
                               |
                               v
  +----------------------------------------------------------+
  |  RANK / PRESENT                                           |
  |  - sort, filter client-side                               |
  |  - "as of" timestamp visible                             |
  |  - provider attribution on price                         |
  |  - bookable vs "requires separate tickets" labelled       |
  +----------------------------+-----------------------------+

  watch: fan-out latency per provider, provider error rate,
        cache hit rate by tier, result age distribution,
        price change rate between search and booking

Delivery checklist:

  1. Fan out in parallel with a global deadline and per-provider isolation. A slow or erroring provider must never block or empty the response.
  2. Return the GDS tier first and merge the airline-direct tier silently; do not reorder the list under the user’s cursor.
  3. Cache by provider offer, not by search result, and split static data (schedules, fare rules) from volatile data (price, availability).
  4. Stamp every result with its age and show it. Never serve a stale price as current.
  5. Use provider-supplied minimum connection times, airport-specific and international-aware. A global constant produces impossible itineraries.
  6. Keep providers in lanes by bookability, and label which provider can actually sell the itinerary as one ticket.
  7. Key the cache on the origin’s local calendar date, display in each airport’s local timezone, and include cabin, passenger types, and currency in the key.
  8. Handle infants as a distinct query — lap infant pricing and results differ, and under-keying the cache here is a spectacular bug.
  9. Do sorting and filtering client-side against the fetched set. Re-querying for a sort makes the list jump.
  10. Cancel or sequence client requests so only the latest search renders.
  11. Have a circuit breaker and a shared rate-limit budget per provider, so a traffic spike does not become a partner outage.
  12. Re-validate price and availability at booking with a graceful failure, since a changed price is expected, not exceptional.

Common mistakes

Mistake What actually happens Better decision
Sequential provider calls Total latency is the sum, most timeouts Parallel fan-out, global deadline
No provider isolation One error looks like “no flights” Scope failures per provider
Cache the whole search result Low hit rate, and stales the whole set Cache offers; split static vs volatile
Serve a stale price silently The user sees a different price on refresh Age-stamp everything, re-validate at booking
Global minimum connection time Impossible 25-minute international connections Provider-supplied, airport-specific MCT
Cross-provider combination assumed Itineraries shown that cannot be booked Lane by bookability, label it
UTC dates in the cache key Same query, different result by an hour Origin-local calendar dates
Cabin/pax types missing from the key Economy results in a business search Full query in the key
Re-query on sort List jumps, feels broken Client-side sort and filter
No request cancellation Out-of-order responses, flicker Sequence numbers, discard stale
No circuit breaker A slow provider takes down search Per-provider breakers and budgets
Mean-ish provider timeouts Loses all the airline-direct results Tiered response, GDS first
Recomputing availability at booking without handling change Crash on a price change Graceful re-quote
Ignoring schedule staleness Static cache serves a flight that is gone TTL plus a validation path
No provider attribution on price Users cannot tell why prices differ Attribute the price to the seller

The complete story in one minute

A flight search is a fan-out to twenty providers with wildly different latencies, and the answer is a combination problem rather than a query problem. Providers return legs, not itineraries, and assembling bookable trips means chaining by geography, enforcing provider-supplied minimum connection times (a global constant is how you produce a 25-minute international connection), and respecting single-PNR bookability — which is why good results identify who can actually sell the trip. Return the GDS tier as soon as the slowest GDS lands, merge the airline-direct results silently as they arrive, and never let a slow or erroring provider block the response or appear as “no flights”.

Caching is where the honesty problem lives, and the answer is to cache by provider offer rather than by whole search result, split static data (schedules, routes, fare rules — hours to days) from volatile data (price and availability — five to fifteen seconds), and stamp every result with its age so the UI can say “as of 14:32”. A user who sees a different price on refresh is right to distrust you, and a ten-minute-old price presented as current is the failure that gets a flight site into trouble. Get the cache key right: origin-local calendar dates, not UTC; cabin; passenger types including infants, because lap-infant pricing differs and under-keying that is a spectacular bug.

The rest is the user journey. Sort and filter client-side against the fetched set — re-querying for a sort makes the list jump and feels broken — and cancel or sequence client requests so only the latest search renders, because users hit enter repeatedly. And re-validate price at booking: a changed price is expected, not exceptional, so it needs a graceful re-quote rather than a crash.

parallel fan-out, GDS tier first, per-provider isolation and breakers
cache offers; static vs volatile tiers; age-stamp and show
origin-local dates, provider MCTs, single-PNR lanes
client-side sort and filter, cancelled and sequenced requests

The hard part was never the query. It was that the price is a claim about a future only the airline knows, and every design decision in the system is really about how honest you are about that.

What this team still owns

Search owns a timestamped, explainable observation; the airline offer owns whether the itinerary can still be sold. Cache schedules, content, and indicative offers under separate freshness rules. Before payment, re-shop or price the exact passenger, segment, fare, ancillary, currency, and tax context and return a typed change when it differs. Do not mutate the displayed result into a booking behind the user’s back: preserve the original quote beside the confirmed order.

Technical references

Keep reading
Browse everything