Real systems don't fail politely. They fail at 2am, under a retry storm, with cold caches, while one availability zone is dark. But the way we usually evaluate architecture — in design reviews and in interviews — is by looking at the sunny-day diagram. RodGrid's chaos family exists to close that gap: break the design on purpose and see who can diagnose it.
Twelve ways to hurt a healthy system
The traffic simulator accepts a stress scenario on top of whatever load you dial in. Grid-wide presets hit a whole class of nodes at once:
- Traffic surge (×3) and Flash crowd (×10) — can it take a spike, and does ×10 need architectural change rather than headroom?
- Cache stampede — every cache misses at once and the stores behind them inherit the read load.
- Cold cache — the everyday post-deploy version: 30% of the normal hit rate.
- AZ loss — every tier you operate loses the replicas it ran in the lost zone (a single-replica tier goes down), and cross-zone routing adds latency.
- Slow third-party — your external dependencies all have a bad day at once; the lesson is isolation.
- Retry storm — failures breed retries breed load, the feedback loop that turns a blip into an outage.
Node-targeted presets pin the failure to one component you pick:
- Node loss — it goes dark entirely.
- OOM kill — a thrashing instance: capacity collapses, latency spikes, errors climb.
- Thread-pool exhaustion — requests queue behind blocked workers.
- Network partition — half its callers can't reach it, and the failing half burns timeout latency first.
- Hot partition — one skewed shard takes most of the load.
Diagram components: Web App, Load Balancer, Application Service, Redis, PostgreSQL.
A healthy read-heavy stack. Open a Grid like this in RodGrid, turn on the simulator, and toggle Cache stampede — the database inherits every read the cache used to absorb, and the design's real bottleneck introduces itself.
Diagnosis challenges
Some challenges hand you a system that is already failing: a seeded starter design, a stated SLO, and an armed stress scenario. Your job is the on-call job — figure out why it's failing under that scenario and redesign it until the SLO holds. When you submit, the server rebuilds the scenario and re-runs the simulation on your final Grid, so the grade reflects what your design actually does under fire, not what a description of it claims.