A simulator is only as honest as its assumptions, and most tools hide theirs. RodGrid's traffic simulator takes the opposite bet: the entire model is deterministic, the capacity numbers are published inside the product, and this post walks through all of it.
A spreadsheet over the graph
There is no AI and no server involved. simulateBoard is a closed-form, steady-state calculator: you give it a requests-per-second number and it pushes that load through the diagram in one topological pass, reporting per-node utilization, latency, and error rate plus Grid-level totals. Run it twice on the same Grid and you get the same answer — which is what lets challenges grade against it fairly.
Capacity comes from the component type
Every node resolves to a capacity profile keyed by its semantic role — load balancer, cache, application service, relational database, queue, model API, and so on. Each profile is a handful of numbers: sustainable requests per second per instance, base latency, and a baseline error rate.
The absolute numbers are deliberately round, order-of-magnitude anchors. What teaches is the ratios — a cache absorbs roughly ten times what a relational database does, which is the entire reason cache-in-front-of-DB is a pattern. The in-product “Model assumptions” table is generated live from the same registry the engine reads, so the published numbers can never drift from the ones actually used.
Diagram components: Application Service, Redis, PostgreSQL.
Turn the simulator on over a Grid like this and raise traffic: the cache soaks up reads at ~10× the database's capacity, and the misses that fall through are what the database actually has to survive.
How load propagates
- Nodes are classified by role into a small set of behaviors: entry points, splitters (load balancers and gateways), caches, stores, async channels (queues/streams), services, and external dependencies. Shapes and annotation nodes are inert — they simply don't simulate.
- A node forwards only what it can actually serve: the minimum of what arrives and its capacity times its replica count. Each edge can also state calls per upstream request, so an occasional model call does not incorrectly receive every request. No phantom throughput downstream of a bottleneck.
- Traffic above a node's capacity goes unanswered, all of it, and counts against availability — no invented split into queued, rejected and erroring shares.
- The wait in a node's queue follows a standard multi-server approximation (Sakasegawa), per replica, so the number of servers decides how much utilization hurts: a wide pool at 90% barely queues, a single-threaded service at 90% waits nine times its service time, so each request takes ten. At 100% or more the wait has no bound, and the simulator says so instead of printing a number.
- Caches absorb their hit rate and forward misses plus writes; splitters divide across equivalent targets; async channels decouple the latency chain.
Replicas are just a property on the node — set a service to 4 replicas and its capacity multiplies, exactly as you'd hope, with third-party services as the deliberate exception: you don't scale someone else's API by drawing more of it.
What it deliberately is not
- Not a discrete-event simulation. No request-level queueing model, no jitter, no tail-latency distributions — a steady-state capacity check, by design.
- Not a vendor benchmark. The capacity numbers are representative anchors for reasoning, not quotes for any specific instance type.
- Not a bill. Provisioned infrastructure uses order-of-magnitude ranges. Usage-priced services such as model APIs are excluded from a combined total until you provide active hours and relevant unit rates; otherwise they say ‘usage-based — not estimated.’ Components that do not carry request traffic, such as monitoring or CI/CD, can still carry cost—the traffic model and billing classification are separate.