Multi-Tenant SaaS Architecture: Isolation Tiers and the One You Should Start On
Shared schema, database-per-tenant, schema-per-tenant, and the row-level security middle ground — with the noisy-neighbour maths that picks the tier.

Multi-tenancy gets discussed as if it were one decision. It is a spectrum of four, each with different costs in code, in migrations, in support, and in the ceiling you hit when a customer asks for their own database.
The useful way to think about it is that you are choosing where the boundary is drawn, and every tier puts it somewhere different: in a WHERE clause, in a schema, in a database, or on a machine. The mistake is picking a tier for the average tenant when the distribution is what matters.
The distribution is the design input
A platform with a long tail, which is the normal shape of any self-serve SaaS business.
tier 1 ~10,000 tenants 1% of compute, 1 user each
tier 2 ~500 tenants 5% of compute, 5-50 users
tier 3 ~40 tenants 30% of compute, 50-5,000 users
tier 4 ~2 tenants 64% of compute, regulated or enormous
90th percentile tenant < 0.1% of total load
99.9th percentile tenant > 20% of total load
Three consequences follow immediately.
Optimise for the tail, not the mean. A capacity plan based on the average tenant is wrong by two orders of magnitude at the top end. A single tier-3 tenant running 30% of your compute means “one customer’s slow query” is a load-bearing incident for everyone else.
The tiers will not be uniform forever. Tenants grow into tiers and a hard boundary in the model is a migration project later. Whatever you build needs a path from tier 1 to tier 4 that does not involve downtime.
Isolation needs are correlated with size and with compliance, not with age. The two tenants that will ask for a dedicated database are the two largest and the two most regulated. Build that path early even if nobody uses it, because the first time it is needed is an urgent customer escalation.
Tier 1: shared schema, tenant_id in every table
Everything in one database, one set of tables, a tenant_id column, and it is read by the application on every query.
SELECT id, amount, status
FROM orders
WHERE tenant_id = $1
AND created_at > $2;
This is where almost everyone should start, and it is the right answer for a long time.
Why it wins: one database to operate, one migration to run, one connection pool, one backup, and cross-tenant queries are trivial. Adding a table is a normal migration. Restoring a backup restores every tenant. Cost scales linearly and simply.
The real risk is not performance, it is a forgotten predicate. One query without tenant_id and you have leaked one customer’s data to another. The failure is silent, it is discovered by a customer rather than by you, and it is the kind of incident that ends up in a regulator’s hands.
The defence is not discipline. Discipline is exactly what fails at 2am during an incident when someone is adding an admin query. Two structural mitigations:
Composite keys force the tenant into the identity of the row:
PRIMARY KEY (tenant_id, id)
Now SELECT * FROM orders WHERE id = $1 is ambiguous across tenants and a developer has to think about it.
Row-level security makes the database the enforcement point:
ALTER TABLE orders ENABLE ROW LEVEL SECURITY;
CREATE POLICY tenant_isolation ON orders
USING (tenant_id = current_setting('app.tenant_id')::uuid);
With PostgreSQL RLS on, the application sets tenant context for the transaction and the policy restricts visible rows. A forgotten application predicate can then return only policy-visible rows instead of every tenant. That protection has conditions: table owners normally bypass RLS unless FORCE ROW LEVEL SECURITY is enabled, superusers and roles with BYPASSRLS bypass it, and pooled session state must be reset. Use a minimally privileged application role, transaction-scoped context, forced RLS where appropriate, and tests that execute as the production role.
The costs of RLS are real: policies become part of query planning and execution, and connection pooling needs care because tenant context can outlive a request when stored at session scope. Prefer a transaction-local setting, reset on every transaction, and fail closed when it is absent. Then test cross-tenant reads, writes, foreign keys, background jobs, migrations, and table-owner behaviour with the exact roles used in production.
Tier 2: database-per-tenant
One database per tenant, same schema in all of them.
+-------------------------------------------+
| control plane (shared) |
| tenants, plans, provisioning, routing |
+------------------+------------------------+
| lookup
+------------+------------+
| | |
v v v
+---------+ +---------+ +---------+
| tenant | | tenant | | tenant |
| db | | db | | db |
| 1 | | 37 | | 9002 |
+---------+ +---------+ +---------+
Why upgrade: isolation is now structural rather than conventional. A query cannot leak across tenants because there is nothing to leak across. Backups, restores, and exports are per tenant. A noisy tenant can be moved, throttled, or given its own hardware. And a customer asking for their own database is a billing change rather than an engineering project.
The costs hit in specific places:
Connection pools. 10,000 tenants times even a 2-connection floor is 20,000 connections. Postgres’s default max_connections is 100. You need a pooler, and a pooler that knows tenant identity, and the operational reality that a pooler with a per-tenant session is a pooler that can serve one tenant’s session to another tenant.
Migrations. A schema change is now 10,000 migrations. They cannot run in one transaction, they must be resumable, and they must tolerate tenants being at different versions while you deploy. This is the single largest engineering cost of this tier and it deserves a proper migration tool, not a shell script.
Query plans. A shared table has one set of statistics. A per-tenant table has statistics from that tenant only, so a plan that is fine across millions of rows can be catastrophic on one tenant’s skewed data. This is the most-reported operational surprise in this model.
Cross-tenant queries. Analytics across all tenants now require either a warehouse path or fanning out over thousands of databases. Build that second path on purpose; do not discover it when the board asks for usage numbers.
Tier 3: schema-per-tenant
One database, one schema per tenant.
+------------------------------------------------+
| one database |
| +------------+ +------------+ +------------+ |
| | tenant_1 | | tenant_37 | | tenant_9002| |
| | tables | | tables | | tables | |
| +------------+ +------------+ +------------+ |
| shared: tenants, plans, provisioning |
+------------------------------------------------+
This is the compromise, and it is genuinely useful in a specific situation: you need real isolation of the data but not of the database, and you have more tenants than you can run migrations against comfortably.
A schema is a namespace, so a query cannot reference another tenant’s table by accident. Connection pools work normally, because there is one database. Migrations are still per-tenant, which is the main cost.
The drawbacks are PostgreSQL-specific knowledge wearing an architecture hat: a few hundred schemas is fine, a few thousand starts to show up in planning time and in pg_dump handling, and cross-schema queries need the schema name to be quoted and constructed safely, which is a string-concatenation bug waiting to happen. Many teams reach for this tier and then regret the quoting.
Tier 4: dedicated everything
Separate database, often separate instance, sometimes separate cluster, occasionally separate account in the cloud provider.
This is what you deploy when a customer contractually requires it, when a regulation requires data residency in a specific region, or when a single tenant’s load is large enough to distort everyone else’s. It is the correct answer for the two largest tenants in the distribution above.
The cost is operational multiplicity: more databases to patch, more backups to verify, more incident response surfaces, and every new feature that must work across four deployment shapes. The mitigation is that this tier should be a thin adapter — the same application, a different connection target, and nothing else — which means keeping the data access path behind a narrow interface so “dedicated” is a deployment concern rather than a code fork.
Quotas, because noisy neighbour is a capacity problem
None of the tiers above stop one tenant from consuming all the capacity. That is a capacity problem, and it is the one customers actually complain about.
per-tenant resource controls
concurrency
a semaphore per tenant on the request path
50 requests in flight, then queue or 429
rate limit
1,000 requests/sec, token bucket per tenant
+ 100/sec sustained refill
resource class
tier 3 tenants get a bigger pool share
so a big tenant does not need its own database
just to get predictable performance
query cost
statement_timeout set per connection from the
tenant's plan: 2s on starter, 30s on enterprise
queueing
per-tenant fair queue, so one tenant's burst
does not sit in front of everyone else's requests
Two of these are worth highlighting because they are the ones teams skip.
Per-tenant statement_timeout. A single unindexed query from one tenant can saturate a connection pool and affect every tenant. Setting the timeout per connection from the tenant’s plan turns a multi-hour incident into a 2-second error message, and it is a couple of lines of code.
Fair queueing rather than FIFO. A shared request queue means a tenant with high throughput delays tenants with low throughput. Per-tenant queues with a fair scheduler between them is what makes a small tenant on a big plan feel fine.
Provisioning is a feature, not a script
Tenant lifecycle is a state machine and it has all the same reliability problems as any other distributed workflow.
requested
|
v
creating ----> failed ---> (retry or manual repair)
|
v
provisioning schema / database
|
v
seeding default settings, initial data
|
v
active
|
v
suspending soft state, data retained
|
v
deleting ---> purge ---> gone
The case that is always missed is partial provisioning. The schema was created, the seed failed, and now there is a tenant that is half-real: it has a row in the tenants table, it has a database, it has no data, and it is in nobody’s mental model. Then a retry runs and conflicts with the existing schema.
The fix is idempotency plus an explicit state. Every provisioning step records what it did, so a retry resumes rather than restarts, and the tenant is not active until every step succeeded. A repair tool that can list half-provisioned tenants is not optional — it is the difference between a self-service product and a support queue.
The same applies to deletion. Suspend before delete, keep the data for a defined window, make deletion an asynchronous workflow with its own retries, and make the purge auditable. Deleting a customer’s data on a mistimed cron job is unrecoverable.
Failure stories worth testing
Run a query without the tenant predicate
With RLS on, confirm it returns nothing and logs. With RLS off, this test is how you find out that a leak is one careless commit away.
Make a pooler connection reuse leak the tenant setting
The classic multi-tenant data breach: a pooled connection carries app.tenant_id from the previous request. Test that connection reuse across tenants cannot happen, including under pooler misconfiguration.
Create 1,000 tenants and run a migration
Time it. This is the number that determines whether tier 2 or tier 3 is viable, and it is much worse than the theoretical estimate.
Let one tenant run a sequential scan on a large table
Confirm per-tenant statement_timeout catches it, and that other tenants are unaffected. This is the noisy-neighbour test that matters most.
Burst 10,000 requests from one tenant while others run normally
Confirm the fair queue holds and the other tenants’ latency does not move. If it does, you have a FIFO queue pretending to be fair.
Fill a tenant’s disk
Confirm the blast radius is one tenant. If it is the whole database, you are on tier 1 with shared storage and you need per-tenant quotas on the storage layer too.
Fail a provisioning step halfway and retry
Confirm the retry resumes and the tenant reaches a correct state, and that a half-provisioned tenant is visible to your repair tooling.
Restore one tenant’s backup
If this takes a full-database restore, your architecture is telling you something about tier 2. The test is the point.
Run an analytics query across all tenants at tier 2
Time it, with 500 and then 5,000 tenants. The fan-out pattern’s cost is superlinear in ways that are not obvious until you measure.
Delete a tenant and confirm the purge
Is it complete, is it auditable, and is a suspended tenant’s data still there after the window? Confirm all three, because each has been an incident somewhere.
A production-ready architecture
request
|
v
+---------------------------+
| control plane | auth, plan lookup, tenant
| tenant resolution | context, quotas
+------------+--------------+
|
per-tenant controls
- token bucket rate limit
- concurrency semaphore
- statement_timeout from plan
- fair queue scheduling
|
v
+---------------------------+
| data access layer | ONE interface, four backends
+------------+--------------+
|
+-----------+-----------+------------+------------+
| | | |
v v v v
shared +------------+ schema +--------+ +---------+
schema | shared db | per | tenant | | dedicated
RLS | + 1 schema | tenant | db | | instance
| per | | | |
| tenant | | | |
+------------+ +--------+ +---------+
lifecycle service
requested -> creating -> seeding -> active
-> failed (resumable, visible)
active -> suspending -> deleting -> purged
per tenant, everywhere:
quotas, statement_timeout, resource class,
audit log, own encryption key for the strict tier,
export, and per-tier backup/restore
A sensible delivery checklist:
- If PostgreSQL RLS is your isolation boundary, add it before shared production data arrives and test it through the production application role; do not treat an owner or
BYPASSRLSrole as proof. - Put
tenant_idin the primary key, not just the table. Ambiguity is a design tool. - Set
tenant_idper connection inside a transaction, and verify the pooler cannot carry it across requests. - Instrument a per-tenant cost and usage view from the beginning. Migration and provisioning decisions need data you will not have otherwise.
- Implement per-tenant rate limits, concurrency limits, and
statement_timeoutbefore you need them. - Build fair queueing if any tier-3 or larger tenant exists, and do not call a shared FIFO queue fair.
- Make provisioning resumable, idempotent, and observable, and ship a repair tool with it.
- Define suspend, retention, and purge as distinct states with distinct timelines, and audit every purge.
- Build the cross-tenant analytics path separately from the request path, and do not run it on the operational database.
- Provide per-tenant export and per-tenant restore from the start. Customers will ask, and doing it under deadline is worse.
- Design the dedicated tier as a deployment variant behind the same data access interface, so it is a configuration change rather than a fork.
- Write a per-tenant audit log. When a customer asks who accessed their data, the answer needs to exist in machine-readable form.
Common mistakes
| Mistake | What actually happens | Better decision |
|---|---|---|
tenant_id only in the WHERE clause |
One forgotten predicate is a cross-tenant breach | Composite keys plus row-level security |
| Retrofitting RLS later | A whole codebase relies on discipline it does not have | Enable it before the second tenant exists |
| RLS with a shared session pool | app.tenant_id leaks between requests |
Transaction-scoped setting, verified pooler config |
| Tier chosen on average tenant size | The 99.9th percentile destroys the plan | Plan against the tail distribution |
| No noisy-neighbour controls | One tenant’s query is everyone’s outage | Per-tenant limits, timeouts, fair queues |
| Shared FIFO queue | High-throughput tenants delay small ones | Per-tenant queues, fair scheduling |
| Tier 2 with shell-script migrations | 10,000 unschedulable migrations | A resumable migration tool, tested at scale |
| Ignoring per-tenant statistics | Query plans flip at scale and nobody knows why | Per-tenant autovacuum, plan monitoring |
| Half-provisioned tenants untracked | Invisible broken customers, support-driven discovery | Explicit state machine plus repair tooling |
| No suspend state | Deletion is the only exit and it is irreversible | Suspend, retain, purge as separate steps |
| Cross-tenant analytics on the operational db | A dashboard query is an outage | Warehouse path, decoupled |
| No per-tenant export | Emergency work under a deadline | Export as a first-class feature |
| Dedicated tier as a code fork | Every feature needs four implementations | Same interface, different deployment target |
| No per-tenant audit log | “Who accessed my data” is unanswerable | Audited access, machine-readable, per tenant |
| Assuming encryption keys can be shared | Strict-tier customers cannot be onboarded | Per-tenant keys where the tier requires it |
| Testing only the happy path for isolation | Isolation bugs live in tools and admin paths | Test admin queries, support tools, exports |
The complete story in one minute
Multi-tenancy is four tiers, and the decision is where the isolation boundary is drawn: a tenant_id predicate, a schema, a database, or a whole instance. The distribution of tenant sizes picks the answer, not the average, because a service built for ten thousand equal tenants and one built for five hundred wildly uneven ones share almost no code.
For many products, start on a shared schema with row-level security and a tested path to isolate exceptional tenants later. Shared schema wins on operations — one database, one migration path, one backup fleet — and its real risk is a missing tenant boundary. Composite tenant-aware keys prevent cross-tenant references; RLS restricts rows only when the query runs as a role subject to the policy and tenant context is set correctly. Force the policy where needed, keep the application role unprivileged, and test absence or leakage of pooled context as a security property.
Database-per-tenant buys structural isolation and per-tenant backup and export, and pays for it in connection pooling, in migrations that must run ten thousand times, in per-tenant query statistics that flip plans, and in cross-tenant analytics that needs a second path entirely. Schema-per-tenant is the compromise for when you need namespace isolation but not database isolation.
Noisy neighbour is a capacity problem before it is a security one, and per-tenant rate limits, concurrency semaphores, and statement_timeout from the tenant’s plan solve more real complaints than any isolation model. Fair queueing matters as soon as tenants differ in size, because a shared FIFO queue is not fair.
And provisioning is a state machine with a failure case nobody designs: the half-created tenant. Make every step idempotent and resumable, keep failed visible, and ship a repair tool. Deletion gets the same treatment, because a suspended tenant’s data still has to exist after the customer thinks it is gone.
one data access interface, four backends
per tenant: limits, timeouts, audit, export, restore
lifecycle: requested -> creating -> seeding -> active -> suspending -> purged
The hard part was never putting a tenant ID on a table. It was making the failure of any single provisioning step recoverable without a human reading a log.


