System Design Capacity Calculator — systemdesigncalculator.online
Replace the capacity meeting with a ten-second preset
Turn daily active users and a fixed two-topology family into peak RPS, health, a monthly cost estimate, and a keepable link or PDF. Founders, interview candidates, and consultants stop rebuilding the same spreadsheet for every sketch—presets like SaaS 50k, marketplace 1M, and content 10M (golden default ≈11,574 peak RPS and ≈$2,705/month) give you a defensible baseline in one pass.
Problems this saves time on
Blank architecture boxes
Teams draw Users, CDN, and database icons then stall on peak RPS, cache working set, and monthly cost—often a full hour before numbers appear.
Seven presets (SaaS 50k, marketplace 1M, content 10M, write-heavy, async-queue, media-object, edge-secure) load a complete classic or queue-workers world with DAU, CDN ratio, pool sizes, and Redis tier already set.
Saves roughly 45–60 minutes versus building a spreadsheet from scratch.
Instance count debates without load
Engineers argue +4 or +8 app instances from habit while CPU, IOPS, and Redis memory stay unmodeled against actual peak RPS.
Recommendations emit SystemState patches sized from current utilization—APP_THROTTLING and APP_PRESSURE counts follow live CPU, not a fixed increment.
Saves 20–40 minutes per sizing review.
Yellow nodes mistaken for outages
Dashboards treat 70–85% utilization like CRITICAL pages; on-call gets paged for pressure that is not yet an incident.
The warning band is explicit pressure: collectIncidents unchanged, but recommendation cards can address APP_PRESSURE and DB_IOPS_PRESSURE before thresholds cross into CRITICAL codes.
Saves 15–30 minutes clarifying severity in cross-team meetings.
Overpaying on idle replicas and cache
Finance sees large monthly numbers while REPLICA_IDLE and oversized Redis tiers stay invisible in the narrative.
Incidents surface REPLICA_IDLE and REDIS_NEAR_OOM; COST recommendations prefer enabling CDN and right-sizing before blind scale-out.
Saves 30–45 minutes reconciling cost with utilization.
Comparing two plans without a diff
Two spreadsheet tabs or two meetings to see whether adding workers or replicas actually fixes SLO_MISS.
Pin A on the compare strip freezes a baseline while you apply patches; health, peak RPS, and monthly estimate update on the live plan.
Saves 25–40 minutes per A/B architecture discussion.
“How far does this scale?” without autoscale fantasy
Stakeholders assume horizontal autoscale; nobody states the DAU ceiling with hardware held fixed.
Growth envelope holds instances, cores, Redis tier, and replicas constant and reports how many DAU the design holds—honest capacity, not elastic cloud magic.
Saves 20–35 minutes on roadmap capacity questions.
Interview sketches that leak implementation detail
Candidates over-specify vCPU counts when interviewers wanted traffic and bottleneck story only.
Interview mode hides hardware detail while keeping incidents, topology, and health narrative visible for a clean whiteboard story.
Saves 10–20 minutes resetting scope in mock interviews.
Redis memory surprises at higher DAU
Teams forget working set grows with users; Redis OOM appears late in load tests.
Working set uses DAU × 2 objects × object size (fixed objects-per-user constant); REDIS_NEAR_OOM fires before you pretend cache is free.
Saves 15–25 minutes debugging mystery cache growth.
Sync path vs queue-workers confusion
Kafka-shaped traffic gets drawn on classic topology; queue depth and WORKER_SATURATED never enter the model.
async-queue preset switches to queue-workers topology with queue publish RPS, depth estimate, worker utilization, QUEUE_BACKLOG, and WORKER_SATURATED—only where that topology applies.
Saves 30–50 minutes aligning async ingest with the right graph.
Plans trapped in one person’s laptop
PDF exports are stale the moment someone tweaks DAU; share links break when pasted on the wrong path.
Share encodes state in `/app#s=`; legacy `/#s=` redirects to `/app` with the same hash. Download report rasterizes the live report to PDF after a finished run.
Saves 15–30 minutes per client handoff or study group.
Overselling what a calculator can prove
Slides promise multi-region failover, WAF SKUs, and reserved-instance discounts the model never computed.
Model limits are on the landing and in the report: no Kafka/S3/search/WAF product SKUs, no autoscale, Users and Nginx always healthy, four unit-rate packs—not AWS list prices.
Saves 20–40 minutes defending credibility with technical buyers.
Query-oriented answers
What is a system design capacity calculator?
It is a fixed two-topology simulator—classic web stack or queue-workers—that turns DAU and request mix into peak RPS, health, and cost. System Design Capacity Simulator is not a free-form architecture builder; you tune presets inside that family.
How is peak RPS calculated from DAU?
Average RPS = DAU × requests per user per day / 86,400. Peak RPS = average RPS × peak multiplier—the same formulas the capacity report prints.
Can I size capacity without autoscaling?
Yes. The growth envelope holds instance counts, cores, Redis tier, and replicas fixed while DAU rises—it shows how much traffic the current chart holds, not elastic fleet autoscale.
How does Redis memory relate to DAU?
Working set size follows DAU × 2 cached objects × object size in kilobytes. Objects per user is fixed at two in the model, so higher DAU increases Redis memory even when mix stays similar.
What drives the monthly cost estimate?
Four components—vCPU, RAM (app plus Redis), SSD terabytes, and origin egress Gbps—multiplied by one of four unit-rate packs. There are no reserved-instance discounts or guaranteed AWS list prices.
Can I use this for interview or whiteboard sizing?
Interview mode hides hardware detail while keeping topology, incidents, and health narrative visible. Presets from 50k to 10M DAU give defensible peak RPS and incident names for mock interviews and consulting sketches.
Questions
Is this a free-form architecture builder?
No. You choose between two fixed topologies—classic web stack or queue-workers—and tune presets and knobs inside that family. It is capacity math for system design interviews and early production sizing, not a drag-and-drop microservice editor.
Why are Users and Nginx always green?
By design they have no failure model in this simulator. Bottlenecks appear on CDN, app pool, Redis, Postgres, queue, and workers where capacity formulas apply.
Can I share a plan with a colleague?
Yes. Finished state lives in `/app#s=` on the simulator path. Older links that used `/#s=` on the site root redirect to `/app` with the same hash so bookmarks keep working.
Are the prices AWS list prices?
No. You pick one of four unit-rate packs—generic, AWS-shaped, Hetzner, or Yandex—and the report multiplies vCPU, RAM, SSD, and egress. There are no reserved-instance or committed-use discounts in the model.
Why does Redis memory jump when I raise DAU?
Working set size follows DAU × 2 cached objects × object size in kilobytes (objects per user is fixed at 2). More daily active users mean more keys even if request mix stays the same.
Does “holds until N DAU” mean autoscale?
No. The growth envelope increases DAU while keeping instance counts, cores, Redis tier, and replicas fixed. It answers how much traffic the current hardware chart holds—not elastic scale-out.
What languages are supported?
Marketing pages live at `/en` and `/ru`. Inside the simulator you can toggle English or Russian, and `/app?locale=en` or `/app?locale=ru` sets the initial locale unless a share hash `#s=` overrides it.
Do I have to pay to use the tool?
The simulator is free. After a finished job—PDF, share link, or improving CRITICAL toward OPTIMAL—an optional donate link appears. It never blocks presets, incidents, or export.