Flight Search System: Caching by Degradation, Not by Answer
Flight search at scale — GDS and aggregator fan-out, result caches that age gracefully, and the itinerary-combinatorics problem that makes searching hard.

Flight search is presented as a search problem and is actually a distributed-fan-out-and-recombination problem with a caching requirement that fights the nature of the data. The price is a claim about a future that only the airline knows, and every design decision follows from taking that seriously.
The scale
a flight search service
500 queries per second average
4,000 per second peak (a fare sale
at 9am on a Monday)
each query fans out to
- 3 GDSs (Amadeus, Sabre, Travelport)
- 15+ airline websites (scraped or
NDC-connected)
- 4 metasearch aggregators
= 22 providers
- p50 latency: 400ms
- p95 latency: 4s
- p99: 11s (or timeout)
- failure rate: 2-8% per provider
and the answer is combinatorial
2 outbound options x 4 return
options x 3 fare families
= 24 bookable itineraries
from 2,000+ raw offers
The combinatorial point is the one people miss. Providers return offers — a leg with a fare, a cabin, a fare family. The user wants itineraries — combinations of legs across providers that are actually connectable and actually bookable together. Assembling those is a real problem with real constraints, and it is the part that determines whether the results are usable or a list of legs that cannot be booked as a trip.
The fan-out
request: SFO -> LIS, 2026-06-10 to 2026-06-17
+----------+ +----------+ +----------+
| GDS 1 | | GDS 2 | | airline |
| 380ms | | 1200ms | | website |
+----------+ +----------+ | 6.5s |
| | +----------+
v v |
+-----------------------------------------+
| normaliser: provider format -> |
| canonical offer |
+---------------------+-------------------+
|
v
+-----------------------------------------+
| combiner: offers -> connectable |
| itineraries |
+-----------------------------------------+
the critical rule
NEVER let a slow provider hold the
whole response
The NEVER is the design rule the whole thing rests on. Three approaches, in increasing order of correctness:
Sequential with short timeout. A 1-second timeout per provider. Loses results on every provider slower than 1s, which is most of the direct-airline sites. Rejected.
Parallel with a global deadline. Fire everything at once, wait up to a total budget (say 2.5s), return what arrived. Simple, and it works — but the user experience depends on which providers happened to be fast, and that is a lottery.
Parallel with tiered response. Fire everything. Return the first tier (GDSs) as soon as the slowest GDS arrives — usually under 1s — and stream or silently merge the airline-direct results as they arrive. This is the right answer and it is what good aggregators do: the initial paint has the reliable providers, and the rest fill in.
The “silently merge” matters. Streaming results into the page causes the list to reorder under the user’s cursor, which is worse than a slightly slower complete list. The pattern that works is: show the GDS results immediately and completely, then merge in the rest and re-render only if the user has not interacted, or show a small “checking airline sites…” indicator with a count.
The combiner, which is the actual hard part
offers from providers
(legs, not itineraries)
a valid itinerary requires
1. legs chain geographically
(LIS -> MAD -> LIS is fine,
LIS -> MAD -> BCN is not)
2. minimum connection time at each
transfer
- intra-EU: 45 min (often 1.5h+
in practice)
- long-haul: 2h minimum
- the MCT depends on airport,
terminal, airline, and whether
the passenger has checked bags
3. the total fare is bookable
AS ONE TICKET
- a GDS can combine its own
offerings
- mixing GDS offers with airline
direct offers is often not
bookable as a single PNR
4. fare rules are consistent
(baggage, change fees, whether
the segments can be split)
5. ticket validity and pricing are
still live at booking time
combinatorics
2000 offers
-> naive pairwise combination:
millions of candidate pairs
-> most are invalid
-> filter by geography first,
then by MCT, then by fare rule
The MCT (minimum connection time) is where naive implementations produce embarrassing results: a 25-minute connection at a large airport with a security queue and immigration, on a flight that arrives at 23:50 and the connection departs at 00:15. Real MCTs come from the provider, are airport-specific, and are longer for international arrivals than domestic ones. Using a global constant is the most common cause of “your itinerary is impossible” at booking.
The bookability constraint is the second big one. A price on a flight search page is a quote, and the quote is only meaningful if the itinerary can be purchased as a single ticket. Providers that do not support cross-provider combination must be kept in their own lane, which is why good flight search results often show a “booked by [provider]” label and why prices for the same itinerary differ between them. The system has to know, per provider, what it is capable of combining.
Caching, and the honesty problem
the fundamental tension
a flight search answer is only
valid for a few seconds
but a search UI gets hammered,
and origin-destination-date is
a small key space
-> so cache, obviously
the risk
a user refreshes in 10 seconds
and sees a DIFFERENT price
-> they now distrust the site
-> and they are right to
The resolution is to make the cache’s degradation visible and its behaviour principled:
Cache by provider offer, not by search result. A cached GDS offer for SFO-LIS on a date is far more reusable than a cached “search result” for the full query, because the same offer participates in many itineraries. This gives a much higher hit rate on stable data (fare rules, schedules) and forces only the volatile part (availability and price) to be live.
Two-tier cache. A short-lived availability cache (5-15s) for the volatile price/availability, and a long-lived static cache (hours to days) for schedule, route, and fare-rule data. Most of the hit rate comes from the static layer, and the static layer is the layer that is safe to cache hard.
Stamp every result with its age. The UI shows “prices as of 14:32” and, if a result is more than a couple of minutes old, says so. This is the cheapest possible honesty and it converts a trust problem into a non-issue.
Never serve a stale price as if it were live. If the availability cache misses and the provider is down, say “price unavailable” for that provider. Serving a 10-minute-old price as current is the failure that gets a flight site sued.
The cache key has to include everything that changes the answer: origin, destination, dates (as local dates at each airport, not UTC — a critical and frequently wrong detail), cabin, passenger counts and types (adult, child, infant — infant Laplace produces different results and different fares), and the currency and fare-filter settings. Under-key the cache and you will serve a business-class result to an economy search, which is a spectacular bug.
The date and timezone problem
a flight departing SFO at 23:50
on 2026-06-10 arrives in Tokyo at
2026-06-12 (local)
so
- the search is for departure DATE
2026-06-10
- the arrival date in the destination
is a different calendar day
- "one way to Tokyo" has no
destination-side date constraint
- a "return on 2026-06-17" is a
departure date from the destination,
interpreted in the DESTINATION's
timezone
which means
a search made in Tokyo and a search
made in London for the same
"SFO-LIS, Jun 10" are the same
query
but the *display* of the result
depends on the viewer's timezone
Anyone who has seen flight search return a flight on the wrong day has usually hit a UTC-vs-local mismatch. The rule that works: the search key uses the origin’s local calendar date; the display converts using the current location of each airport; and the cache key uses the same local-date convention, not a UTC timestamp. This is boring and it is a whole class of bugs.
The “search is not a query” part
a real user journey
1. search SFO -> LIS, 2 adults
2. see 200 results, refine by
"direct only" (client-side filter
is fine, it's the same result set)
3. sort by price
4. click a result, see fare rules
5. go BACK, change dates by one day
6. the whole search re-runs
7. adjust the fare filter
8. see 200 results again
the engineering consequences
- the sort and filter are client-side
against a result set of a few
hundred, not a re-query
- the date change is a re-query
- the whole flow needs to be fast
on the FIRST query (skeleton
state matters more than anything)
- the user will press enter
repeatedly; debounce and cancel
in-flight requests, and never
render a stale response
Two things get missed here. Client-side sort and filter is not an optimisation, it is the correct design — re-querying for a sort is slow and makes the list jump. And request cancellation is mandatory: users type dates and hit enter, and if responses arrive out of order the list flickers between searches. Every request needs a sequence number and the client must discard any response that is not the latest.
Failure stories worth testing
Make one provider return in 9 seconds
The initial render must not wait for it. This is the tiered-response test and the one that separates a usable system from a slow one.
Make one provider return an error
The error must not appear as “no flights found”. Confirm the error is scoped to that provider and the others still populate.
Make two providers disagree on the price of the same itinerary by 4%
Both should display, with the provider identified. Confirm the system does not pick one silently.
Serve a 20-minute-old availability cache
The UI must say so, and the price must be re-validated at booking. This is the trust test.
Search a 1-adult-and-1-infant query
Infant results differ (lap infant, no seat, different fare). Confirm the query differentiates and the key includes the infant.
Search “SFO to Tokyo, departing 23:50” and check the arrival date
The displayed arrival date must be in the destination’s timezone. This is the classic timezone bug.
Drop a cache key element (remove cabin, keep the rest)
Search business and confirm economy results appear. This is the under-keyed cache test.
Have a user press enter 5 times in 2 seconds
Only the last response may render. Any flicker means cancellation is not implemented.
Set a provider’s schedule data to be stale for a month
The static cache will happily serve a flight that no longer exists. Confirm schedules have a TTL and a validation path.
Force a 5,000 QPS spike
Measure whether the cache absorbs it or whether providers get hammered. The provider-facing rate limiter is the thing that keeps a traffic spike from becoming a partner outage.
Add a route with a 30-minute connection at a hub
Confirm the MCT filter rejects it, and confirm the message tells the user why.
Book an itinerary the search showed, after a 3-minute cache window
The booking flow must re-validate price and availability and fail gracefully. A price that changed at booking is not a bug, but crashing is.
A production-ready architecture
user query
(origin, dest, local dates,
cabin, pax types, filters, currency)
|
v
+----------------------------------------------------------+
| CACHE LAYER |
| - static: schedule, routes, fare rules (hours-days) |
| - volatile: availability + price (5-15s) |
| - key includes local dates, cabin, pax types, currency |
| - every result carries an age |
+----------------------------+-----------------------------+
|
v
+----------------------------------------------------------+
| FAN-OUT (parallel, global deadline, per-provider |
| isolation) |
| - tier 1: GDSs -> return when slowest GDS lands |
| - tier 2: airline direct -> merge silently on arrival |
| - circuit breaker per provider |
| - rate limit per provider, shared budget |
+----------------------------+-----------------------------+
|
v
+----------------------------------------------------------+
| NORMALISER |
| provider offer -> canonical leg (times in local tz, |
| carrier, flight numbers, equipment, fare family) |
+----------------------------+-----------------------------+
|
v
+----------------------------------------------------------+
| COMBINER |
| - chain by geography and MCT (provider-supplied) |
| - enforce single-PNR bookability per provider lane |
| - build bookable itineraries with total fare |
+----------------------------+-----------------------------+
|
v
+----------------------------------------------------------+
| RANK / PRESENT |
| - sort, filter client-side |
| - "as of" timestamp visible |
| - provider attribution on price |
| - bookable vs "requires separate tickets" labelled |
+----------------------------+-----------------------------+
watch: fan-out latency per provider, provider error rate,
cache hit rate by tier, result age distribution,
price change rate between search and booking
Delivery checklist:
- Fan out in parallel with a global deadline and per-provider isolation. A slow or erroring provider must never block or empty the response.
- Return the GDS tier first and merge the airline-direct tier silently; do not reorder the list under the user’s cursor.
- Cache by provider offer, not by search result, and split static data (schedules, fare rules) from volatile data (price, availability).
- Stamp every result with its age and show it. Never serve a stale price as current.
- Use provider-supplied minimum connection times, airport-specific and international-aware. A global constant produces impossible itineraries.
- Keep providers in lanes by bookability, and label which provider can actually sell the itinerary as one ticket.
- Key the cache on the origin’s local calendar date, display in each airport’s local timezone, and include cabin, passenger types, and currency in the key.
- Handle infants as a distinct query — lap infant pricing and results differ, and under-keying the cache here is a spectacular bug.
- Do sorting and filtering client-side against the fetched set. Re-querying for a sort makes the list jump.
- Cancel or sequence client requests so only the latest search renders.
- Have a circuit breaker and a shared rate-limit budget per provider, so a traffic spike does not become a partner outage.
- Re-validate price and availability at booking with a graceful failure, since a changed price is expected, not exceptional.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
| Sequential provider calls | Total latency is the sum, most timeouts | Parallel fan-out, global deadline |
| No provider isolation | One error looks like “no flights” | Scope failures per provider |
| Cache the whole search result | Low hit rate, and stales the whole set | Cache offers; split static vs volatile |
| Serve a stale price silently | The user sees a different price on refresh | Age-stamp everything, re-validate at booking |
| Global minimum connection time | Impossible 25-minute international connections | Provider-supplied, airport-specific MCT |
| Cross-provider combination assumed | Itineraries shown that cannot be booked | Lane by bookability, label it |
| UTC dates in the cache key | Same query, different result by an hour | Origin-local calendar dates |
| Cabin/pax types missing from the key | Economy results in a business search | Full query in the key |
| Re-query on sort | List jumps, feels broken | Client-side sort and filter |
| No request cancellation | Out-of-order responses, flicker | Sequence numbers, discard stale |
| No circuit breaker | A slow provider takes down search | Per-provider breakers and budgets |
| Mean-ish provider timeouts | Loses all the airline-direct results | Tiered response, GDS first |
| Recomputing availability at booking without handling change | Crash on a price change | Graceful re-quote |
| Ignoring schedule staleness | Static cache serves a flight that is gone | TTL plus a validation path |
| No provider attribution on price | Users cannot tell why prices differ | Attribute the price to the seller |
The complete story in one minute
A flight search is a fan-out to twenty providers with wildly different latencies, and the answer is a combination problem rather than a query problem. Providers return legs, not itineraries, and assembling bookable trips means chaining by geography, enforcing provider-supplied minimum connection times (a global constant is how you produce a 25-minute international connection), and respecting single-PNR bookability — which is why good results identify who can actually sell the trip. Return the GDS tier as soon as the slowest GDS lands, merge the airline-direct results silently as they arrive, and never let a slow or erroring provider block the response or appear as “no flights”.
Caching is where the honesty problem lives, and the answer is to cache by provider offer rather than by whole search result, split static data (schedules, routes, fare rules — hours to days) from volatile data (price and availability — five to fifteen seconds), and stamp every result with its age so the UI can say “as of 14:32”. A user who sees a different price on refresh is right to distrust you, and a ten-minute-old price presented as current is the failure that gets a flight site into trouble. Get the cache key right: origin-local calendar dates, not UTC; cabin; passenger types including infants, because lap-infant pricing differs and under-keying that is a spectacular bug.
The rest is the user journey. Sort and filter client-side against the fetched set — re-querying for a sort makes the list jump and feels broken — and cancel or sequence client requests so only the latest search renders, because users hit enter repeatedly. And re-validate price at booking: a changed price is expected, not exceptional, so it needs a graceful re-quote rather than a crash.
parallel fan-out, GDS tier first, per-provider isolation and breakers
cache offers; static vs volatile tiers; age-stamp and show
origin-local dates, provider MCTs, single-PNR lanes
client-side sort and filter, cancelled and sequenced requests
The hard part was never the query. It was that the price is a claim about a future only the airline knows, and every design decision in the system is really about how honest you are about that.
What this team still owns
Search owns a timestamped, explainable observation; the airline offer owns whether the itinerary can still be sold. Cache schedules, content, and indicative offers under separate freshness rules. Before payment, re-shop or price the exact passenger, segment, fare, ancillary, currency, and tax context and return a typed change when it differs. Do not mutate the displayed result into a booking behind the user’s back: preserve the original quote beside the confirmed order.


