# Multi-Tenant Architecture: Serving Many Customers from One System Without Mixing Their Data

A customer opens a support ticket saying they can see another company's invoices. The cause is usually mundane: one query in one endpoint is missing a `WHERE tenant_id = ?`. Preventing that class of bug is most of what multi-tenant architecture is about.

Multi-tenancy means serving many customers (tenants) from shared infrastructure while keeping their data, configuration and performance separate. The design question is how much to share. Share too little and you pay for it in operations. Share too much and you risk the leak above.

Where something is opinion rather than AWS or Azure guidance, it's labelled as such.

* * *

## 1\. The isolation spectrum: silo, pool, bridge

To keep this concrete, imagine a fictional invoicing product called **Ledgerly**. It has 2,000 small businesses on its standard plan and a handful of large enterprise customers, some of whom have contractual or regulatory reasons to keep their data apart. We'll come back to Ledgerly throughout.

The AWS SaaS vocabulary has three models. **Silo** gives each tenant a dedicated application stack, with a database instance per tenant. **Bridge** shares the application tier but gives each tenant its own schema or database. **Pool** has tenants share the whole stack, including the same tables, and separates them logically.

Row-level security (RLS) is the strongest way to enforce that separation in the database, but it's an enforcement mechanism, not what defines the model. Each step trades isolation against cost: silo isolates most and costs most, while pool costs least and isolates least.

Bridge is also used more broadly for any mix of silo and pool. As systems break into smaller services, you often end up mixing them by layer. For example, a shared web tier can sit in front of per-tenant application tiers and storage.

Azure uses different names for similar ideas: automated single-tenant, fully multitenant, vertically partitioned and horizontally partitioned. A vertically partitioned setup keeps most tenants on shared infrastructure but gives dedicated infrastructure to those who need more performance or data isolation. A horizontally partitioned one separates specific components per tenant, which helps if most of your load comes from a few components.

| Model | What's shared | Isolation | Cost and ops | Typical fit |
| --- | --- | --- | --- | --- |
| Pool | Compute, database, tables (tenant_id column; RLS recommended) | Logical | Lowest cost, simplest fleet | Many small and mid-size tenants |
| Bridge | Compute; separate schema or database per tenant | Data-level | Moderate; migrations multiply per tenant | Compliance-sensitive data, shared app tier |
| Silo | Nothing (dedicated stack) | Strongest | Highest cost, most automation required | Regulated or very large enterprise tenants |

Here is how Ledgerly might use all three at once:

![Mermaid Diagram](https://mermaid.ink/img/pako:eNp1UmGP0kAQ_SvjfNKwRDR-IuaSAndKxNBQNBBqzLSdwobtbrM7FfHCfzdt78S76H6aedn33rzZvcfcFYxjLI075QfyAotVagEAkvVsl2IiZAvyBdSG7PvMv755q0ajEYSKjIGsCdpyCBxS_NbTpptdirdW2NdeB4a8CeIq9rDp2AUJQdUEgYzhyLVA4Jo8CV8Vtv9R2HYK7mTB8147q6Ctc2fFUy4g7Kt-jl4nNNneU32AcCDPRRumK4DqGkSz_-PYniiOdykuuNizN2eI4jkM4OT8kf01HNviUT1eLhezye5lim3VTdb7QBsxo8AdJmzJynddwABWiyTFVw9Sm548Wc1nH2771XChc5K_FKB0HjY953kmbVybaL5YjqFsjDnDlR-E8uOTdNtkHU0_7VKMHrLD4KnL9l8Zk_UMhsObdjePj_us33Z9r95DURx3WL-fp9hmNkGFFfuKdIHje5QDV-3vK7ikxgiqHvlKXlNmOLR3SmfljiptzjjGIdW14WE4B-FKwcRoe_xMedL1d86KghQT3juGL_MUFaxc5sQp-MjmB4vOSUHkNRkFgWwYBva6RNWZJPpXO8ubd_VPvFwUZvupM87jGF-cDloYL78BRxsEOQ?type=png)

## 2\. How to choose

Ask these questions in order:

1.  Do contracts or regulations require physical separation? Data residency rules, security reviews and large enterprise contracts can push some customers toward bridge or silo, and Azure's guidance treats commercial requirements for physically isolated data as a legitimate driver.
    
2.  How uneven are your tenants? If a few customers generate most of the load, pooling them with everyone else invites trouble (see section 6).
    
3.  How many tenants will you have? Per-tenant schemas and databases multiply every migration, backup and monitoring job, so the cost grows with tenant count.
    
4.  Can your team operate what you're designing? Azure warns against an architecture that gets more complex as you scale, or one you lack the people and skills to maintain.
    

**Recommendation (opinion, not AWS or Azure guidance):** default to pooled for the standard tier, and design so a tenant can be promoted to dedicated infrastructure later. For Ledgerly, that means the 2,000 small businesses stay pooled, and customers X and Y get promoted when their contracts require it.

AWS describes a similar hybrid, with siloed infrastructure for high-traffic customers and bridge or pool for quieter ones, but "pool first, promote the exceptions" as a default is a judgment call. Promotion is only cheap if tenant identity is built in from day one. Retrofitting it is expensive.

## 3\. Two planes: control and application

Mature SaaS systems split into two planes. The **application plane** serves tenants' requests. The **control plane** handles everything about tenants themselves: onboarding, identity, tenant registry, billing and metering. In Ledgerly, creating a new customer account and provisioning their data is control plane work. Generating one of their invoices is application plane work.

AWS's reference architecture shows how the pieces fit. A workflow service such as Step Functions orchestrates tenant onboarding and provisioning as a durable process. STS issues short-lived credentials scoped to a tenant by mapping a tenant tag into the session. KMS can encrypt each tenant's data under a per-tenant key. Other clouds have equivalents, so the pattern carries over: automated onboarding, credentials scoped per tenant, and, where practical, separate encryption keys per tenant.

## 4\. Carrying tenant context through every layer

A multi-tenant system fails when some code path forgets which tenant it is serving, so tenant context has to be explicit at every layer. Resolve it once at the edge from verified identity, such as a signed token claim or an authenticated subdomain, and never from a client-supplied header or body field.

Then carry it through logs, traces, metrics, cache keys, queue messages, object-storage prefixes, background jobs and webhooks. A `withTenant(...)` wrapper (shown below) helps because developers don't have to remember a `WHERE` clause on every query.

## 5\. Working example: pooled Postgres with row-level security

RLS moves the tenant filter into the database, so every query is checked instead of relying on each one to include it.

### Schema and policy

```sql
-- Migrations run as app_owner; the application connects as app_user.
CREATE ROLE app_user LOGIN NOSUPERUSER NOBYPASSRLS;

CREATE TABLE invoices (
  id           uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  tenant_id    uuid NOT NULL,
  amount_cents bigint NOT NULL,
  created_at   timestamptz NOT NULL DEFAULT now()
);

-- Lead every index with tenant_id
CREATE INDEX invoices_tenant_created ON invoices (tenant_id, created_at DESC);

ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
ALTER TABLE invoices FORCE ROW LEVEL SECURITY;  -- owner is NOT exempt

CREATE POLICY tenant_isolation ON invoices
  USING      (tenant_id = NULLIF(current_setting('app.current_tenant', true), '')::uuid)
  WITH CHECK (tenant_id = NULLIF(current_setting('app.current_tenant', true), '')::uuid);

GRANT SELECT, INSERT, UPDATE, DELETE ON invoices TO app_user;
```

The `NULLIF(..., '')` makes the policy fail closed. If the setting is missing or empty, the comparison yields NULL and the query returns no rows, instead of raising a cast error or matching something unintended.

### Application helper (Node/TypeScript, `pg`)

```typescript
import { Pool, PoolClient } from "pg";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });

export async function withTenant<T>(
  tenantId: string,                       // from the verified token, never from user input
  fn: (db: PoolClient) => Promise<T>
): Promise<T> {
  const client = await pool.connect();
  try {
    await client.query("BEGIN");
    // third argument `true` = local to this transaction (same as SET LOCAL)
    await client.query("SELECT set_config('app.current_tenant', $1, true)", [tenantId]);
    const result = await fn(client);
    await client.query("COMMIT");
    return result;
  } catch (err) {
    await client.query("ROLLBACK");
    throw err;
  } finally {
    client.release();
  }
}
```

### Pitfalls

-   **Connection pooling can leak tenants.** Session-level state is attached to the physical connection. When the pool hands that connection to a different tenant's request, the new request inherits the old tenant's context. The fix is to scope tenant context to the transaction rather than the connection. That is also what makes it work with transaction-mode pooling.
    
-   **Some roles skip RLS.** By default the table owner bypasses RLS, and `FORCE ROW LEVEL SECURITY` closes that gap. Superusers and roles with `BYPASSRLS` are always exempt, and `FORCE` doesn't change that.
    
-   `SECURITY DEFINER` **functions use their owner's privileges.** They run with their _owner's_ privileges rather than the caller's. If the owner is a superuser or has `BYPASSRLS`, the function runs without RLS no matter who calls it, so a function that returns data can expose every tenant's rows to any caller. If the owner is an ordinary role and the table has `FORCE ROW LEVEL SECURITY`, the policies still apply inside the function. Run the app as a restricted role, give these functions a non-exempt owner, and audit them.
    
-   **Constraints can reveal rows.** A unique index isn't subject to the policy, so a colliding insert across tenants can confirm that another tenant's row exists. Either scope uniqueness with `tenant_id` or translate the error into a generic failure.
    
-   **RLS governs rows, not columns.** A user can pass every row policy and still change a column they shouldn't, such as upgrading their own plan. Put column-level restrictions and entitlement checks in separate layers.
    

## 6\. Noisy neighbors

In shared infrastructure, one tenant's spike degrades everyone else. One engineering team describes it as inevitable in a traditional multi-tenant setup, and data and storage services are especially prone to it. Seasonality makes it worse, because a subset of tenants can see load jump suddenly.

Mitigations, roughly from cheapest to most expensive:

1.  Tag every metric with `tenant_id`. You can't fix a noisy tenant you can't identify.
    
2.  Per-tenant rate limits and quotas at the API gateway.
    
3.  Fair queuing: per-tenant concurrency caps on workers, so one tenant's bulk import can't starve the queue.
    
4.  Query guardrails: statement timeouts and row limits, with read-heavy reporting on replicas.
    
5.  Promoting the outliers to dedicated resources, sometimes called the "supertenant" approach.
    

## 7\. Scaling out with stamps (cells)

Past a certain size, one big shared environment becomes a single blast radius. The Deployment Stamps pattern repeats a complete copy of the stack, and each copy serves a group of tenants.

A stamp dedicated to one tenant gives the highest isolation and avoids noisy neighbors, but it is expensive. Stamps shared by many tenants cost less, though they still need their own patterns for managing sharing inside the stamp, and noisy neighbors can still occur there. Either way, an outage or a bad release stays within the customers on that stamp, and you can keep adding stamps as you grow.

Stamps depend on automation. Azure's guidance calls for automated deployment, typically declarative infrastructure as code such as Bicep or Terraform. Stamps can also be grouped into rings, so changes reach early-adopter cohorts first.

## 8\. Plan for the tenant lifecycle

Azure's lifecycle guidance calls out trials, tenants who upgrade to a tier that requires a dedicated deployment, and rebalancing a tenant to a different deployment to even out load. If your design can't relocate a tenant, the last two become custom migrations. A tenant-to-deployment routing table in the control plane, plus a tenant export/import tool you've rehearsed before you need it, make them routine.

## 9\. Test for leaks continuously

The main test surfaces are cross-tenant data leaks, tenant-ID propagation through every code path, noisy-neighbor behavior and per-tenant quotas. A minimal leak test belongs in CI:

```typescript
test("tenant B cannot see tenant A's invoices", async () => {
  await withTenant(A, db =>
    db.query("INSERT INTO invoices (tenant_id, amount_cents) VALUES ($1, 100)", [A]));
  const res = await withTenant(B, db => db.query("SELECT * FROM invoices"));
  expect(res.rowCount).toBe(0);
});
```

Add a startup or CI check that every tenant-scoped table has RLS enabled and forced. Also call your `SECURITY DEFINER` functions as the app role with a tenant set, and assert they return only that tenant's data.

* * *

## Checklist

-   Tenant resolved from verified identity, once, at the edge
    
-   Tenant context in every log line, metric, cache key, queue message and job
    
-   RLS **enabled and forced** on every tenant table, with explicit `USING` and `WITH CHECK`
    
-   App runs as a non-owner, non-superuser role without `BYPASSRLS`; `SECURITY DEFINER` functions have a non-exempt owner and are audited
    
-   Transaction-scoped tenant context (safe with pgBouncer transaction mode)
    
-   `tenant_id` leads indexes and uniqueness constraints
    
-   Per-tenant rate limits, quotas and metering
    
-   Automated onboarding and a tested path to move a tenant to dedicated infrastructure
    
-   Cross-tenant leak tests in CI
    
-   Everything deployable by IaC, ready for stamps when you need them
    

## Conclusion

Most of the work comes down to three decisions: how much to share, how to carry tenant context through every layer, and how to move a tenant when its needs change. Starting pooled, automating the control plane early and testing for leaks in CI keeps those options open.

* * *

## Sources

-   [AWS Well-Architected SaaS Lens: The bridge model](https://docs.aws.amazon.com/wellarchitected/latest/saas-lens/bridge-model.html)
    
-   [AWS Guidance for Multi-Tenant Architectures](https://docs.aws.amazon.com/solutions/multi-tenant-architectures-on-aws/)
    
-   [AWS SaaS Multi-Tenant Architecture Guide (Hidekazu Konishi)](https://hidekazu-konishi.com/entry/aws_saas_multi_tenant_architecture_guide.html)
    
-   [Azure Architecture Center: Architectural approaches for a multitenant solution](https://learn.microsoft.com/en-us/azure/architecture/guide/multitenant/approaches/overview)
    
-   [Azure Architecture Center: Storage and data approaches](https://learn.microsoft.com/en-us/azure/architecture/guide/multitenant/approaches/storage-data)
    
-   [Azure Architecture Center: Deployment and configuration approaches](https://learn.microsoft.com/en-us/azure/architecture/guide/multitenant/approaches/deployment-configuration)
    
-   [Azure Architecture Center: Tenant lifecycle considerations](https://github.com/microsoftdocs/architecture-center/blob/main/docs/guide/multitenant/considerations/tenant-lifecycle.md)
    
-   [Parloa: Scaling Parloa, when the platform becomes the product](https://www.parloa.com/labs/insights/scaling-parloa-when-the-platform-becomes-the-product/)
    
-   [DEV Community: Why tenant context must be scoped per transaction](https://dev.to/m_zinger_2fc60eb3f3897908/why-tenant-context-must-be-scoped-per-transaction-3aop)
    
-   [DEV Community: Postgres RLS multi-tenancy, two leaks that survive correct policies](https://dev.to/wenceslaudev/postgres-rls-multi-tenancy-two-leaks-that-survive-correct-policies-kb2)
    
-   [Harrison Songolo: Multi-Tenant Isolation in PostgreSQL](https://contra.com/p/2mF2UZY2-multi-tenant-isolation-in-postgre-sql)
    
-   [StackHarbor: Row-level security for multi-tenant PostgreSQL](https://stackharbor.com/en/knowledge-base/pg-row-level-security-multi-tenant/)

---

*Published via [ZyVOP](https://zyvop.com/multi-tenant-architecture-serving-many-customers-from-one-system-without-mixing-their-data-u7cq8?utm_source=hashnode&utm_medium=crosspost&utm_campaign=syndication) — Write once in Markdown, auto-backup to GitHub, and syndicate to Dev.to, Medium & Hashnode in 1 click.*
