Debug the failing system

Stress presets and diagnosis challenges: how RodGrid lets you break a design on purpose and find out whether you can fix it.

Real systems don't fail politely. They fail at 2am, under a retry storm, with cold caches, while one availability zone is dark. But the way we usually evaluate architecture — in design reviews and in interviews — is by looking at the sunny-day diagram. RodGrid's chaos family exists to close that gap: break the design on purpose and see who can diagnose it.

Twelve ways to hurt a healthy system

The traffic simulator accepts a stress scenario on top of whatever load you dial in. Grid-wide presets hit a whole class of nodes at once:

  • Traffic surge (×3) and Flash crowd (×10) — can it take a spike, and does ×10 need architectural change rather than headroom?
  • Cache stampede — every cache misses at once and the stores behind them inherit the read load.
  • Cold cache — the everyday post-deploy version: 30% of the normal hit rate.
  • AZ loss — every tier you operate loses the replicas it ran in the lost zone (a single-replica tier goes down), and cross-zone routing adds latency.
  • Slow third-party — your external dependencies all have a bad day at once; the lesson is isolation.
  • Retry storm — failures breed retries breed load, the feedback loop that turns a blip into an outage.

Node-targeted presets pin the failure to one component you pick:

  • Node loss — it goes dark entirely.
  • OOM kill — a thrashing instance: capacity collapses, latency spikes, errors climb.
  • Thread-pool exhaustion — requests queue behind blocked workers.
  • Network partition — half its callers can't reach it, and the failing half burns timeout latency first.
  • Hot partition — one skewed shard takes most of the load.
Try it: what does a cache stampede do here?

Diagram components: Web App, Load Balancer, Application Service, Redis, PostgreSQL.

A healthy read-heavy stack. Open a Grid like this in RodGrid, turn on the simulator, and toggle Cache stampede — the database inherits every read the cache used to absorb, and the design's real bottleneck introduces itself.

Diagnosis challenges

Some challenges hand you a system that is already failing: a seeded starter design, a stated SLO, and an armed stress scenario. Your job is the on-call job — figure out why it's failing under that scenario and redesign it until the SLO holds. When you submit, the server rebuilds the scenario and re-runs the simulation on your final Grid, so the grade reflects what your design actually does under fire, not what a description of it claims.

Note The simulator behind all of this is a deterministic capacity model, not a packet-level emulator — the same published assumptions apply everywhere, which is exactly what makes before/after comparisons fair. More on how it works in “How the traffic simulator works”.

All blog posts