Building a Scalable SaaS Architecture
Contents
When a SaaS product starts creaking, the instinct is to reach for more infrastructure. In practice, most early SaaS products are nowhere near a hardware limit. They are hitting a design limit — usually one taken in the first month, when the product had two customers and the decision felt inconsequential.
This is about the decisions that are expensive to reverse, and which ones can safely wait.
Multi-tenancy: the one decision to get right early
How you separate one customer's data from another's is the hardest thing to change later, because changing it means migrating live customer data. There are three common approaches, and the trade-off is genuinely a trade-off — none of them is simply correct.
- Shared schema, tenant ID column. Every table carries a tenant identifier and every query filters on it. Cheapest to operate, simplest to migrate, scales to large customer counts. The risk is that one missing filter leaks data across tenants — a serious class of bug that testing does not reliably catch.
- Schema per tenant. One database, separate schemas. Stronger isolation and per-tenant backup and restore become straightforward. Migrations now have to run across every schema, which becomes slow somewhere in the hundreds.
- Database per tenant. Strongest isolation, easiest story for data-residency and enterprise security review, and simple to bill by cost. Operationally the heaviest by a wide margin, and it makes cross-tenant reporting genuinely hard.
For most products the shared schema is the right default. But if you take it, take the safety measure that goes with it: enforce tenant scoping in one place — a base query layer, a repository, or row-level security in the database — rather than trusting every future developer to remember the filter on every query. This is the single highest-value hour of architecture work in a young SaaS product.
Where SaaS products actually break first
Almost never on the web tier. Stateless application servers are easy to add. The failures cluster in four less obvious places.
- Background jobs. Report generation, imports, email batches and scheduled syncs are where one large customer starves everyone else. A single tenant uploading a 200,000-row file will hold the queue while every other customer waits. Separate queues by job class, and cap per-tenant concurrency.
- The database's slowest query under real data shapes. A query that is instant against 500 rows can be unusable at 5 million. The shape of the data matters more than the volume — one tenant with 90% of the rows breaks assumptions that averages hide.
- Anything that runs per tenant on a timer. A nightly job that takes 4 seconds per tenant is fine at 50 tenants and takes over six hours at 5,000. This one arrives suddenly.
- Third-party rate limits. Your payment, email or messaging provider limits you globally while your customers experience it individually. One tenant's bulk operation consumes the allowance everyone shares.
What not to build early
Resist splitting into microservices before the domain boundaries are actually known. A well-structured single application — clear module boundaries, no cross-module database access — gives you most of the organisational benefit and none of the distributed-systems cost. You can extract a service later from a clean module; you cannot easily merge four services that were drawn in the wrong places.
The same applies to caching layers, custom orchestration and premature sharding. Every one of these adds a failure mode and an operational burden. Add them when a measurement demands it, not when a diagram suggests it.
What is worth building in from the start
A short list, because each of these is cheap now and disproportionately expensive to add once you have customers depending on the current behaviour.
- Tenant scoping enforced centrally, as above.
- Idempotency on anything that charges money or sends a message. Networks retry; without idempotency keys, so do your side effects.
- An audit trail of who changed what, when. Your first enterprise customer's security review will ask, and reconstructing history retrospectively is impossible.
- Database migrations under version control, applied the same way in every environment. Manual schema changes are the most common source of "works in staging" incidents.
- Per-tenant metrics, not just aggregate ones. Averages conceal the customer who is having a bad time, and that customer is the one who cancels.
Scale is a measurement, not a feeling
The practical discipline is to instrument before optimising. Know your slowest endpoint at the 95th percentile, your longest-running background job, and your largest tenant by row count. Those three numbers tell you what will break next far more reliably than any architectural intuition.
Most SaaS products do not need a distributed system. They need one well-structured application, correct tenant isolation, background work that cannot be monopolised, and enough visibility to see trouble a month before it becomes an outage.