Skip to content
System Design Capacity Simulator

System Design Capacity Calculator — systemdesigncalculator.online

Replace the capacity meeting with a ten-second preset

Turn daily active users and a fixed two-topology family into peak RPS, health, a monthly cost estimate, and a keepable link or PDF. Founders, interview candidates, and consultants stop rebuilding the same spreadsheet for every sketch—presets like SaaS 50k, marketplace 1M, and content 10M (golden default ≈11,574 peak RPS and ≈$2,705/month) give you a defensible baseline in one pass.

From this calculator to an agent on your stand

The public site runs one capacity model. The path ahead keeps that model and lets an agent on your own server fill it.

  1. Live

    Public calculator

    Two topologies: classic, and a queue with workers. Daily active users become peak RPS, incidents, and a monthly estimate from one of four price packs. The formulas stay on screen.

  2. Live

    A decision on the same chart

    Apply a recommendation patch, compare before and after, copy a share link, and download the report. The chart does not grow new node types.

  3. Your server

    The same engine on your stand

    Run this simulator on a server you control. The public origin and the donation wallet stay in that server’s environment.

  4. On your stand

    An agent that only fills SystemState

    The agent calls the existing simulation query, writes only SystemState fields, and applies recommendation patchJson. It tunes load, hardware, and the two supported topologies.

The agent does not add Kafka, object storage, a WAF, autoscale, or a free-form architecture graph. A new topology would be a later product, not something the agent invents.

Product definition

System Design Capacity Simulator sizes classic Users → Nginx → CDN → App → Redis → Postgres or queue-workers stacks from daily active users into peak RPS, utilization-driven incidents, a monthly cost estimate, and shareable /app#s= links or PDF exports—within a fixed two-topology model, not a drag-and-drop microservice editor.

Who this is for

System design interviewee

Defend a 1M–10M DAU sketch with named incidents (APP_THROTTLING, DB_IOPS_EXCEEDED, REDIS_NEAR_OOM, and the rest), warning-band pressure, and a growth envelope that shows how many DAU the hardware holds—not empty boxes on a whiteboard.

Founder sizing first production

See first-month spend from vCPU, RAM, SSD, and origin egress unit rates—not an AWS shopping list—and know where classic Users → Nginx → CDN → App → Redis → Postgres breaks before you commit.

Consultant delivering a one-pager

Hand the client a PDF or `/app#s=` link they can reopen without your spreadsheet. Pin A compare and recommendation patches stay tied to the same transparent formulas in the report.

Problems this saves time on

Blank architecture boxes

Teams draw Users, CDN, and database icons then stall on peak RPS, cache working set, and monthly cost—often a full hour before numbers appear.

Seven presets (SaaS 50k, marketplace 1M, content 10M, write-heavy, async-queue, media-object, edge-secure) load a complete classic or queue-workers world with DAU, CDN ratio, pool sizes, and Redis tier already set.

Saves roughly 45–60 minutes versus building a spreadsheet from scratch.

Instance count debates without load

Engineers argue +4 or +8 app instances from habit while CPU, IOPS, and Redis memory stay unmodeled against actual peak RPS.

Recommendations emit SystemState patches sized from current utilization—APP_THROTTLING and APP_PRESSURE counts follow live CPU, not a fixed increment.

Saves 20–40 minutes per sizing review.

Yellow nodes mistaken for outages

Dashboards treat 70–85% utilization like CRITICAL pages; on-call gets paged for pressure that is not yet an incident.

The warning band is explicit pressure: collectIncidents unchanged, but recommendation cards can address APP_PRESSURE and DB_IOPS_PRESSURE before thresholds cross into CRITICAL codes.

Saves 15–30 minutes clarifying severity in cross-team meetings.

Overpaying on idle replicas and cache

Finance sees large monthly numbers while REPLICA_IDLE and oversized Redis tiers stay invisible in the narrative.

Incidents surface REPLICA_IDLE and REDIS_NEAR_OOM; COST recommendations prefer enabling CDN and right-sizing before blind scale-out.

Saves 30–45 minutes reconciling cost with utilization.

Comparing two plans without a diff

Two spreadsheet tabs or two meetings to see whether adding workers or replicas actually fixes SLO_MISS.

Pin A on the compare strip freezes a baseline while you apply patches; health, peak RPS, and monthly estimate update on the live plan.

Saves 25–40 minutes per A/B architecture discussion.

“How far does this scale?” without autoscale fantasy

Stakeholders assume horizontal autoscale; nobody states the DAU ceiling with hardware held fixed.

Growth envelope holds instances, cores, Redis tier, and replicas constant and reports how many DAU the design holds—honest capacity, not elastic cloud magic.

Saves 20–35 minutes on roadmap capacity questions.

Interview sketches that leak implementation detail

Candidates over-specify vCPU counts when interviewers wanted traffic and bottleneck story only.

Interview mode hides hardware detail while keeping incidents, topology, and health narrative visible for a clean whiteboard story.

Saves 10–20 minutes resetting scope in mock interviews.

Redis memory surprises at higher DAU

Teams forget working set grows with users; Redis OOM appears late in load tests.

Working set uses DAU × 2 objects × object size (fixed objects-per-user constant); REDIS_NEAR_OOM fires before you pretend cache is free.

Saves 15–25 minutes debugging mystery cache growth.

Sync path vs queue-workers confusion

Kafka-shaped traffic gets drawn on classic topology; queue depth and WORKER_SATURATED never enter the model.

async-queue preset switches to queue-workers topology with queue publish RPS, depth estimate, worker utilization, QUEUE_BACKLOG, and WORKER_SATURATED—only where that topology applies.

Saves 30–50 minutes aligning async ingest with the right graph.

Plans trapped in one person’s laptop

PDF exports are stale the moment someone tweaks DAU; share links break when pasted on the wrong path.

Share encodes state in `/app#s=`; legacy `/#s=` redirects to `/app` with the same hash. Download report rasterizes the live report to PDF after a finished run.

Saves 15–30 minutes per client handoff or study group.

Overselling what a calculator can prove

Slides promise multi-region failover, WAF SKUs, and reserved-instance discounts the model never computed.

Model limits are on the landing and in the report: no Kafka/S3/search/WAF product SKUs, no autoscale, Users and Nginx always healthy, four unit-rate packs—not AWS list prices.

Saves 20–40 minutes defending credibility with technical buyers.

How calculations work

The simulator report follows the same order as this summary: traffic first, then peak RPS, utilization and pool sizing, incidents and the warning band, then the monthly estimate. Open any finished run for full formulas and LaTeX in the report—not a substitute for that detail.

Average RPS

average RPS = DAU × requestsPerUser / 86400

10,000,000 × 20 / 86400 ≈2315

Peak RPS

peak RPS = average RPS × peakMultiplier

≈2315 × 5 ≈11574

Redis object working set

Redis object working set = DAU × 2 × objectSizeKb

10,000,000 × 2 × 1 = 20,000,000 Kb

Worked example: content 10M golden default

These numbers come from the content 10M preset with default knobs—the same golden vector validated in domain tests, not a fictitious startup profile.

Presetcontent-10m (golden default)
Peak RPS≈11,574
Estimated monthly cost≈$2,705
Topologyclassic

Open the simulator, load content 10M, and run with default settings to reproduce peak RPS and monthly estimate in the live report.

Preset catalog: default inputs and computed load

Each row applies the same preset defaults as the simulator, then runs computeSystemCapacity on that state. Figures are illustrative capacity math—not a ranking, not a cloud vendor quote, and not list pricing.

SaaS 50k

DAU
50,000
Peak RPS
≈69
Monthly estimate
≈$207

Incidents: None on this default

No critical incident codes on this default.

Open in the simulator

Marketplace 1M

DAU
1,000,000
Peak RPS
≈2083
Monthly estimate
≈$9349

Incidents: None on this default

No critical incident codes on this default.

Open in the simulator

Content 10M

DAU
10,000,000
Peak RPS
≈11574
Monthly estimate
≈$2705

Incidents: REPLICA_IDLE

Breaks on APP_THROTTLING; this hardware holds until 19,584,001 DAU.

Open in the simulator

Write-heavy

DAU
2,000,000
Peak RPS
≈2894
Monthly estimate
≈$68546

Incidents: REPLICA_IDLE

Breaks on DB_IOPS_EXCEEDED; this hardware holds until 5,004,163 DAU.

Open in the simulator

Queue / Kafka

DAU
1,500,000
Peak RPS
≈5208
Monthly estimate
≈$60320

Incidents: None on this default

No critical incident codes on this default.

Open in the simulator

Objects + search

DAU
8,000,000
Peak RPS
≈4444
Monthly estimate
≈$3875

Incidents: REDIS_NEAR_OOM

Breaks on APP_THROTTLING; this hardware holds until 12,240,001 DAU.

Open in the simulator

WAF / regions

DAU
2,000,000
Peak RPS
≈2894
Monthly estimate
≈$12773

Incidents: None on this default

No critical incident codes on this default.

Open in the simulator

Two presets, same capacity model

Content 10M

DAU
10,000,000
Peak RPS
≈11574
Monthly estimate
≈$2705

Incidents: REPLICA_IDLE

Breaks on APP_THROTTLING; this hardware holds until 19,584,001 DAU.

Open in the simulator

Write-heavy

DAU
2,000,000
Peak RPS
≈2894
Monthly estimate
≈$68546

Incidents: REPLICA_IDLE

Breaks on DB_IOPS_EXCEEDED; this hardware holds until 5,004,163 DAU.

Open in the simulator

Top monthly-cost movers on content 10M (+10% per knob)

  • Static ratio-$875
  • Peak multiplier+$194
  • App instances+$152

Interview latency orders of magnitude

  • Memory~100 ns
  • SSD~150 μs
  • Datacenter~0.5 ms
  • Continent~150 ms

These round-trip budgets are interview reference points only. The capacity simulator does not compute or simulate them.

A session, start to artifact

  1. Preset

    Open the simulator from this page, pick a preset (or start from the golden content 10M vector), and confirm topology classic or queue-workers matches your story.

  2. Diagnose

    Read peak RPS, monthly estimate, graph health, and incidents such as APP_THROTTLING, DB_IOPS_EXCEEDED, SLO_MISS, REDIS_NEAR_OOM, REPLICA_IDLE, WORKER_SATURATED, and QUEUE_BACKLOG where applicable.

  3. Apply

    Open Recommendations on the compare strip and apply SystemState patches sized from current load—incidents first, then warning-band pressure, then cost and headroom cards.

  4. Compare

    Pin A to freeze a baseline plan while you iterate on Now; contrast health, cost, and envelope without a second spreadsheet.

  5. Share / PDF

    When the run reaches a finished job (share, PDF, or CRITICAL → OPTIMAL path), copy the `/app#s=` link or download the PDF—optional donate appears after value, never as a paywall.

What it computes — and what it refuses to fake

Included

  • Two topologies: classic Users → Nginx → CDN → App → Redis → Postgres and queue-workers with queue + workers
  • Peak RPS from DAU, requests per user, and peak multiplier
  • Incidents: APP_THROTTLING, DB_IOPS_EXCEEDED, SLO_MISS, REDIS_NEAR_OOM, REPLICA_IDLE, WORKER_SATURATED, QUEUE_BACKLOG
  • Monthly estimate from vCPU, RAM (app + Redis), SSD TB, and origin Gbps using generic, AWS, Hetzner, or Yandex unit-rate packs
  • Recommendations as Partial<SystemState> patches sized from current utilization
  • Pin A compare strip and growth envelope with hardware held fixed
  • Interview mode and share state via `/app#s=` plus PDF export

Explicitly excluded

  • Kafka, SQS, object storage, search, CDN origin-shield, auth, WAF, or multi-region failover product SKUs
  • Reserved-instance or committed-use discounts
  • Autoscale or elastic instance pickers—envelope scales DAU against fixed hardware
  • Failure models for Users and Nginx (always healthy in the graph)
  • Configurable Redis objects-per-user—the model fixes 2 objects per DAU

Query-oriented answers

What is a system design capacity calculator?

It is a fixed two-topology simulator—classic web stack or queue-workers—that turns DAU and request mix into peak RPS, health, and cost. System Design Capacity Simulator is not a free-form architecture builder; you tune presets inside that family.

How is peak RPS calculated from DAU?

Average RPS = DAU × requests per user per day / 86,400. Peak RPS = average RPS × peak multiplier—the same formulas the capacity report prints.

Can I size capacity without autoscaling?

Yes. The growth envelope holds instance counts, cores, Redis tier, and replicas fixed while DAU rises—it shows how much traffic the current chart holds, not elastic fleet autoscale.

How does Redis memory relate to DAU?

Working set size follows DAU × 2 cached objects × object size in kilobytes. Objects per user is fixed at two in the model, so higher DAU increases Redis memory even when mix stays similar.

What drives the monthly cost estimate?

Four components—vCPU, RAM (app plus Redis), SSD terabytes, and origin egress Gbps—multiplied by one of four unit-rate packs. There are no reserved-instance discounts or guaranteed AWS list prices.

Can I use this for interview or whiteboard sizing?

Interview mode hides hardware detail while keeping topology, incidents, and health narrative visible. Presets from 50k to 10M DAU give defensible peak RPS and incident names for mock interviews and consulting sketches.

Questions

Is this a free-form architecture builder?

No. You choose between two fixed topologies—classic web stack or queue-workers—and tune presets and knobs inside that family. It is capacity math for system design interviews and early production sizing, not a drag-and-drop microservice editor.

Why are Users and Nginx always green?

By design they have no failure model in this simulator. Bottlenecks appear on CDN, app pool, Redis, Postgres, queue, and workers where capacity formulas apply.

Can I share a plan with a colleague?

Yes. Finished state lives in `/app#s=` on the simulator path. Older links that used `/#s=` on the site root redirect to `/app` with the same hash so bookmarks keep working.

Are the prices AWS list prices?

No. You pick one of four unit-rate packs—generic, AWS-shaped, Hetzner, or Yandex—and the report multiplies vCPU, RAM, SSD, and egress. There are no reserved-instance or committed-use discounts in the model.

Why does Redis memory jump when I raise DAU?

Working set size follows DAU × 2 cached objects × object size in kilobytes (objects per user is fixed at 2). More daily active users mean more keys even if request mix stays the same.

Does “holds until N DAU” mean autoscale?

No. The growth envelope increases DAU while keeping instance counts, cores, Redis tier, and replicas fixed. It answers how much traffic the current hardware chart holds—not elastic scale-out.

What languages are supported?

Marketing pages live at `/en` and `/ru`. Inside the simulator you can toggle English or Russian, and `/app?locale=en` or `/app?locale=ru` sets the initial locale unless a share hash `#s=` overrides it.

Do I have to pay to use the tool?

The simulator is free. After a finished job—PDF, share link, or improving CRITICAL toward OPTIMAL—an optional donate link appears. It never blocks presets, incidents, or export.

Trust and updates

Formulas stay visible in the simulator report; model includes and excludes on this page state what the tool computes and what it refuses to fake. We do not publish download counts, user totals, review scores, or ranking guarantees.

Page content last updated: .