[{"id":"lmi-ecs-fargate","title":"LMI Comparison — ECS Fargate Arm","description":"Interactive arm paired with the ordinary Lambda approximation scenario: load this independent ECS Fargate graph to compare task readiness and rolling replacement with the same 2,400-RPS, 100-ms workload [cwm]. This is the ECS arm, not Lambda Managed Instances; the 1-vCPU/2-GiB task, 80-concurrency setting, and 800-RPS-per-task capacity are a bounded configuration from the research baseline [aws-doc] [cwm-est], not a universal ECS claim [cwm] [unsupported]. During quiesce, users can invoke the existing rolling-deploy action explicitly; this seed does not add automatic deployment triggers [cwm]. The source Reddit measurements are validation references only [source] and were not used to tune this workload, readiness estimate, prices, or results. The existing CWM model can expose ordinary task startup, a generic four-step application-readiness ramp, and rolling replacement while preserving ready capacity in this configured rollout [cwm] [cwm-est]. Its already-ready two-task baseline is the warm-capacity side of the comparison; the configured rollout result is specific to this setup, not a universal ECS claim [cwm-est] [unsupported]. This interactive arm does not model: (1) LMI as a distinct AWS platform, managed EC2 classes, or exact lifecycle; (2) LMI worker fan-out or multiple Node.js workers inside one Lambda/LMI environment; (3) memory-to-vCPU/worker derivation; (4) initialization CPU contention; (5) LMI runtime-ready versus application-ready internals; (6) exact HTTP 500/503 attribution; (7) RDS Proxy multiplexing, pinning, or a first-class proxy resource; or (8) exact source failure counts, including 500/503 counts, or latency/p99 reproduction [unsupported]. The PostgreSQL resource is a generous 5,000-connection pool approximation, not an RDS Proxy model [cwm-est] [unsupported].","difficulty":"advanced","resources":[{"x":0,"y":0,"id":"lmi-ecs-api","name":"Node API (ECS Fargate)","type":"compute","status":"healthy","cpuUsage":1,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"AWS ECS Fargate task fleet with 1 vCPU, 2 GiB, 80-way concurrency, and 800 RPS per task [aws-doc] [cwm-est]. Two ready tasks provide the configured warm baseline; ordinary task startup, the four-step estimated application-readiness ramp, and the existing rolling-replacement action are observable in CWM [cwm]. This configured rollout result is not a universal ECS claim, and LMI workers, memory-to-vCPU/worker derivation, initialization CPU contention, exact error classes, and source latency are unsupported [unsupported].","baseLatency":5,"concurrency":80,"containerCpu":1,"maxTaskCount":10,"minTaskCount":2,"appReadySteps":4,"serviceFamily":"ecsFargate","costMultiplier":0.1285,"desiredTaskCount":2,"containerMemoryGiB":2,"perTaskCapacityRps":800,"requestServingKind":"container","taskStartupSeconds":3,"taskScaleDownSeconds":3,"requestDurationSeconds":0.1}},{"x":200,"y":0,"id":"lmi-ecs-pg","name":"PostgreSQL (RDS-Proxy-like pool approximation)","type":"database","status":"healthy","cpuUsage":5,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"AWS PostgreSQL endpoint with a generous 5,000-connection pool standing in for the source workload's RDS-Proxy-like path [cwm-est]. This is a pool approximation only: RDS Proxy multiplexing, pinning, and a first-class proxy resource are unsupported [unsupported].","baseLatency":3,"cacheHitRate":0,"serviceFamily":"rds-postgres","costMultiplier":2,"maxConnections":5000}}],"connections":[{"id":"lmi-ecs-api-pg","sourceId":"lmi-ecs-api","targetId":"lmi-ecs-pg"}],"duration":"~10 min","tags":["AWS","ECS","Fargate","LMI Comparison","Lambda Comparison","Readiness","Rolling Deploy","Warm Capacity","Research","Source Validation"],"category":"scaling","seed":295824465,"defaultTrafficPatterns":[{"name":"Idle — establish the warm task baseline","type":"step","startTime":0,"endTime":3,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true},{"name":"Cold burst — watch tasks and application readiness","type":"step","startTime":3,"endTime":13,"parameters":{"startTraffic":2400,"endTraffic":2400},"isActive":true},{"name":"Warm hold — compare ready capacity and utilization","type":"step","startTime":13,"endTime":28,"parameters":{"startTraffic":2400,"endTraffic":2400},"isActive":true},{"name":"Quiesce — invoke rolling deploy here if desired","type":"step","startTime":28,"endTime":31,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true},{"name":"Re-burst — observe replacement and readiness","type":"step","startTime":31,"endTime":46,"parameters":{"startTraffic":2400,"endTraffic":2400},"isActive":true},{"name":"Post-deployment observation — hold the 2,400-RPS workload","type":"step","startTime":46,"endTime":60,"parameters":{"startTraffic":2400,"endTraffic":2400},"isActive":true},{"name":"Traffic Recovery — enable after step 60","type":"ramp","startTime":60,"endTime":90,"parameters":{"startTraffic":2400,"endTraffic":0,"duration":30},"isActive":false}]},{"id":"ecs-fargate-task-scaling","title":"ECS Fargate Task Scaling","description":"Watch a Fargate service hold its baseline task, start additional tasks during a traffic ramp, cap at its maximum fleet, then scale to zero after traffic recovers. The model includes task startup wait, bounded queueing, and per-task vCPU and memory billing.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"fargate-alb-1","name":"Application Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Routes HTTP traffic to the ECS Fargate service.","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":250,"y":220,"id":"fargate-service-1","name":"ECS Fargate Service","type":"compute","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"0.25 vCPU / 0.5 GB","systemRole":"Fargate task fleet; allocated vCPU and memory bill while tasks run. At three tasks, excess traffic queues rather than creating VM clones.","baseLatency":12,"containerCpu":0.25,"maxInstances":3,"maxTaskCount":3,"minInstances":0,"minTaskCount":0,"serviceFamily":"ecsFargate","costMultiplier":0.1285,"coldStartLatency":1200,"desiredTaskCount":1,"containerMemoryGiB":0.5,"perTaskCapacityRps":100,"requestServingKind":"container","taskStartupSeconds":2,"taskScaleDownSeconds":3}}],"connections":[{"id":"conn-fargate-alb-service","sourceId":"fargate-alb-1","targetId":"fargate-service-1"}],"duration":"~18 min","tags":["AWS","ECS","Fargate","Tasks","Scale to Zero"],"category":"scaling","defaultTrafficPatterns":[{"name":"Traffic Ramp — watch Fargate tasks start","type":"ramp","startTime":0,"parameters":{"startTraffic":30,"endTraffic":500,"duration":30},"isActive":true},{"name":"Traffic Recovery — watch Fargate scale to zero","type":"ramp","startTime":60,"parameters":{"startTraffic":500,"endTraffic":0,"duration":30},"isActive":false}]},{"id":"pubsub-queue-saturation","title":"GCP Pub/Sub Topic Saturation","description":"Simulate an event burst that overwhelms a GCP Pub/Sub subscription: the unacknowledged message backlog surges, subscriber VMs can't keep pace, and message delivery latency climbs well past acknowledgement deadlines triggering redelivery storms. Explore how adding more GCE subscriber instances and tuning flow-control settings restore normal throughput.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"pss-lb","name":"Cloud Load Balancing","type":"network","status":"healthy","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Global load balancer distributing inbound HTTP traffic to the event producer fleet","baseLatency":2,"maxThroughput":10000,"serviceFamily":"load_balancer","costMultiplier":1}},{"x":250,"y":200,"id":"pss-producer","name":"Event Producer (GCE)","type":"compute","status":"healthy","cpuUsage":72,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"n2-standard-2","systemRole":"Publishes events to the Pub/Sub topic faster than the subscriber pool can acknowledge them","baseLatency":2,"maxThroughput":20000,"serviceFamily":"gce","costMultiplier":1}},{"x":250,"y":360,"id":"pss-pubsub","name":"Cloud Pub/Sub (saturated)","type":"queue","status":"warning","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"e2-medium","systemRole":"Subscription backlog growing — unacknowledged messages are expiring their deadlines and being redelivered, amplifying subscriber load","baseLatency":140,"maxThroughput":100000,"costMultiplier":1}},{"x":100,"y":500,"id":"pss-subscriber-1","name":"Subscriber VM 1","type":"compute","status":"critical","cpuUsage":96,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"n2-standard-2","systemRole":"Subscriber falling behind — CPU saturated and flow-control buffer full, causing nack storms","baseLatency":3,"maxThroughput":2000,"serviceFamily":"gce","costMultiplier":1}},{"x":400,"y":500,"id":"pss-subscriber-2","name":"Subscriber VM 2","type":"compute","status":"critical","cpuUsage":94,"location":{"zoneKey":"us-central1-b","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-b (Iowa)"},"provider":"gcp","characteristics":{"size":"n2-standard-2","systemRole":"Second subscriber also maxed out — combined throughput still insufficient to drain the growing backlog","baseLatency":3,"maxThroughput":2000,"serviceFamily":"gce","costMultiplier":1}},{"x":250,"y":650,"id":"pss-spanner","name":"Cloud Spanner","type":"database","status":"healthy","cpuUsage":40,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"1000pu-regional","systemRole":"Downstream database — currently healthy but at risk if redelivery storms produce duplicate-write spikes","baseLatency":6,"serviceFamily":"cloud-spanner","costMultiplier":3,"maxConnections":1000}}],"connections":[{"id":"conn-psslb-producer","sourceId":"pss-lb","targetId":"pss-producer"},{"id":"conn-producer-pubsub","sourceId":"pss-producer","targetId":"pss-pubsub"},{"id":"conn-pubsub-sub1","sourceId":"pss-pubsub","targetId":"pss-subscriber-1"},{"id":"conn-pubsub-sub2","sourceId":"pss-pubsub","targetId":"pss-subscriber-2"},{"id":"conn-sub1-spanner","sourceId":"pss-subscriber-1","targetId":"pss-spanner"},{"id":"conn-sub2-spanner","sourceId":"pss-subscriber-2","targetId":"pss-spanner"}],"duration":"~15 min","tags":["GCP","Pub/Sub","Queue","Failure","Recovery","Backlog","Subscribers"],"category":"failure","defaultTrafficPatterns":[{"name":"Steady Publisher Load","type":"step","startTime":0,"parameters":{"startTraffic":1200,"endTraffic":1200},"isActive":true},{"name":"Message Burst — overwhelm subscribers","type":"ramp","startTime":0,"parameters":{"startTraffic":1200,"endTraffic":5000,"duration":20},"isActive":false}]},{"id":"servicebus-queue-saturation","title":"Azure Service Bus Queue Saturation","description":"Simulate a burst of inbound orders that overwhelms an Azure Service Bus queue: the active message count surges past the queue's throughput unit capacity, consumer VMs fall behind, and end-to-end processing latency explodes as messages approach their lock expiry and are abandoned back to the queue. Explore how scaling out consumer instances, enabling auto-forwarding to a dead-letter queue, and upgrading to a Premium namespace restore normal throughput.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"sbqs-alb","name":"Azure Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Layer 4 load balancer distributing inbound HTTP traffic to the event producer fleet","baseLatency":2,"maxThroughput":9000,"costMultiplier":1}},{"x":250,"y":200,"id":"sbqs-producer","name":"Event Producer VM","type":"compute","status":"healthy","cpuUsage":68,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","systemRole":"Sends messages to the Service Bus queue at a rate that currently exceeds consumer throughput","baseLatency":2,"maxThroughput":15000,"serviceFamily":"vm","costMultiplier":1}},{"x":250,"y":360,"id":"sbqs-servicebus","name":"Azure Service Bus (saturated)","type":"queue","status":"warning","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_B2s","systemRole":"Queue backlog surging past throughput unit capacity — messages are expiring their lock timeouts and being abandoned back to the queue, amplifying consumer load","baseLatency":130,"maxThroughput":100000,"costMultiplier":1}},{"x":100,"y":500,"id":"sbqs-consumer-1","name":"Consumer VM 1","type":"compute","status":"critical","cpuUsage":97,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","systemRole":"Struggling to drain the queue — CPU pinned at max and lock renewals failing as processing exceeds the message lock duration","baseLatency":3,"maxThroughput":2000,"serviceFamily":"vm","costMultiplier":1}},{"x":400,"y":500,"id":"sbqs-consumer-2","name":"Consumer VM 2","type":"compute","status":"critical","cpuUsage":95,"location":{"zoneKey":"eus-zone-2","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 2)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","systemRole":"Second consumer also saturated — combined throughput still insufficient to drain the growing backlog despite both VMs at full capacity","baseLatency":3,"maxThroughput":2000,"serviceFamily":"vm","costMultiplier":1}},{"x":250,"y":650,"id":"sbqs-sql","name":"Azure SQL S3","type":"database","status":"healthy","cpuUsage":42,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"S2 Standard","systemRole":"Downstream database — currently healthy but at risk if consumers begin retrying abandoned messages in bulk, causing duplicate-write spikes","baseLatency":7,"costMultiplier":2,"maxConnections":300}}],"connections":[{"id":"conn-sbqsalb-producer","sourceId":"sbqs-alb","targetId":"sbqs-producer"},{"id":"conn-sbqsproducer-bus","sourceId":"sbqs-producer","targetId":"sbqs-servicebus"},{"id":"conn-bus-consumer1","sourceId":"sbqs-servicebus","targetId":"sbqs-consumer-1"},{"id":"conn-bus-consumer2","sourceId":"sbqs-servicebus","targetId":"sbqs-consumer-2"},{"id":"conn-consumer1-sql","sourceId":"sbqs-consumer-1","targetId":"sbqs-sql"},{"id":"conn-consumer2-sql","sourceId":"sbqs-consumer-2","targetId":"sbqs-sql"}],"duration":"~15 min","tags":["Azure","Service Bus","Queue","Failure","Recovery","Backlog","Consumers"],"category":"failure","defaultTrafficPatterns":[{"name":"Steady Order Volume","type":"step","startTime":0,"parameters":{"startTraffic":1200,"endTraffic":1200},"isActive":true},{"name":"Order Burst — overwhelm consumers","type":"ramp","startTime":0,"parameters":{"startTraffic":1200,"endTraffic":5000,"duration":20},"isActive":false}]},{"id":"azure-saas-web-app","title":"Azure SaaS Web App","description":"A production SaaS stack on Azure: Front Door caches 75% of requests at the edge so the App Service P1v3 plan climbs into the warning zone at peak without ever going critical, while PostgreSQL Flexible Server stores tenant data and Azure Monitor + Defender for Cloud add observability and security. Watch the ramp push App Service toward its limit — and compare this stack's hourly cost with the AWS and GCP SaaS scenarios.","difficulty":"beginner","resources":[{"x":250,"y":60,"id":"azsaas-fd","name":"Azure Front Door","type":"network","status":"healthy","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Azure Front Door Standard — global CDN + WAF + load balancing; caches static content at the edge so only misses reach App Service","baseLatency":-25,"cacheHitRate":0.75,"maxThroughput":50000,"costMultiplier":2}},{"x":250,"y":220,"id":"azsaas-app","name":"App Service P1v3","type":"compute","status":"healthy","cpuUsage":31,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Azure App Service Premium v3 P1v3 — 2 vCPU, 8 GB RAM, NVMe-backed; hosts the SaaS web application; $0.173/hr","autoscaling":false,"baseLatency":8,"maxThroughput":3000,"costMultiplier":1.8}},{"x":250,"y":400,"id":"azsaas-pg","name":"PostgreSQL Flexible D4s","type":"database","status":"healthy","cpuUsage":20,"location":{"zoneKey":"eus-zone-2","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 2)"},"provider":"azure","characteristics":{"systemRole":"Azure Database for PostgreSQL Flexible Server D4s_v3 — 4 vCore, 16 GB RAM; tenant data store; $0.337/hr","baseLatency":5,"costMultiplier":4.49,"maxConnections":500}},{"x":80,"y":220,"id":"azsaas-monitor","name":"Azure Monitor","type":"security","status":"healthy","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Azure Monitor + Log Analytics — collects metrics, logs, and traces from the whole stack; $0.25/GB ingested beyond 5 GB/mo free","baseLatency":0,"maxThroughput":0,"costMultiplier":0.6}},{"x":420,"y":220,"id":"azsaas-defender","name":"Defender for Cloud","type":"security","status":"healthy","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Microsoft Defender for Cloud (Server Plan 2) — CSPM + threat detection + compliance; $15/server/month","baseLatency":0,"maxThroughput":0,"costMultiplier":0.84}}],"connections":[{"id":"azsaas-c-fd-app","sourceId":"azsaas-fd","targetId":"azsaas-app"},{"id":"azsaas-c-app-pg","sourceId":"azsaas-app","targetId":"azsaas-pg"},{"id":"azsaas-c-app-monitor","sourceId":"azsaas-app","targetId":"azsaas-monitor"},{"id":"azsaas-c-app-defender","sourceId":"azsaas-app","targetId":"azsaas-defender"}],"duration":"~10 min","tags":["Azure","SaaS","App Service","Front Door","PostgreSQL","Cost"],"category":"cost","defaultTrafficPatterns":[{"name":"Business-hours ramp — watch App Service climb into the warning zone","type":"ramp","startTime":0,"parameters":{"startTraffic":200,"endTraffic":1550,"duration":60},"isActive":true},{"name":"Traffic Recovery — see the stack settle back down","type":"ramp","startTime":60,"parameters":{"startTraffic":1550,"endTraffic":200,"duration":30},"isActive":false}]},{"id":"aws-multi-region-failover","title":"AWS Multi-Region Failover — Route 53 Health Checks","description":"On July 24, 2026, the us-west-2 (Oregon) regional outage took down hundreds of single-region deployments for ~80 minutes — disrupting Apple Pay, DoorDash, Reddit, Hulu, and PlayStation Network. This companion scenario shows the resilience pattern that kept multi-region architectures standing: Route 53 Latency Routing with health checks. The scenario opens mid-failover — us-west-2 ALB is already degraded and Route 53 has begun draining DNS toward us-east-1 (N. Virginia). Watch how a 30-second health-check TTL shapes the failover window: every DNS resolver that hasn't refreshed yet still sends users west, where they see errors. Understand why a 60-second TTL (not the default 300 s) and aggressive health-check intervals (10 s, 3 failures) are the difference between a 2-minute and a 12-minute outage. The us-east-1 EC2 fleet is autoscaling — observe it absorb the redirected load as west drains.","difficulty":"intermediate","resources":[{"x":250,"y":40,"id":"amrf-r53","name":"Route 53 (Latency Routing)","type":"network","status":"warning","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"Global DNS"},"provider":"aws","characteristics":{"systemRole":"Global DNS with latency-based routing and health checks — mid-failover right now; us-west-2 health checks are failing and DNS is re-propagating to us-east-1","baseLatency":1,"maxThroughput":100000,"serviceFamily":"route53","costMultiplier":1}},{"x":100,"y":170,"id":"amrf-alb-east","name":"ALB — us-west-2 (Degraded)","type":"network","status":"warning","location":{"zoneKey":"us-west-2a","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2a (Oregon)"},"provider":"aws","characteristics":{"systemRole":"us-west-2 ALB — Route 53 health checks are failing; still receiving traffic from resolvers with cached DNS; packet loss and connection timeouts visible","baseLatency":8,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":400,"y":170,"id":"amrf-alb-west","name":"ALB — us-east-1 (Active)","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"us-east-1 ALB — receiving all new DNS lookups; load is climbing as west DNS TTLs expire and resolvers switch over","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":40,"y":310,"id":"amrf-ec2-east-1","name":"Web Server — West 1","type":"compute","status":"warning","cpuUsage":72,"location":{"zoneKey":"us-west-2a","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2a (Oregon)"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"Oregon origin server in us-west-2a — still receiving retrying clients with stale DNS; elevated CPU from retry storms and connection-reset overhead","baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":2}},{"x":160,"y":310,"id":"amrf-ec2-east-2","name":"Web Server — West 2","type":"compute","status":"warning","cpuUsage":68,"location":{"zoneKey":"us-west-2b","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2b (Oregon)"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"Oregon origin server in us-west-2b — same regional degradation; traffic will drain to zero as DNS TTLs expire and Route 53 completes failover","baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":2}},{"x":330,"y":310,"id":"amrf-ec2-west-1","name":"Web Server — East 1","type":"compute","status":"healthy","cpuUsage":30,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"East origin server in us-east-1a — absorbing the redirected load; autoscaling is armed and will scale out as DNS TTLs expire and traffic ramps up","autoscaling":true,"baseLatency":3,"maxThroughput":4000,"serviceFamily":"ec2","costMultiplier":2}},{"x":460,"y":310,"id":"amrf-ec2-west-2","name":"Web Server — East 2","type":"compute","status":"healthy","cpuUsage":30,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"East origin server in us-east-1b — paired with east-1 to spread the incoming load; will scale in symmetry once the full failover completes","autoscaling":true,"baseLatency":3,"maxThroughput":4000,"serviceFamily":"ec2","costMultiplier":2}},{"x":100,"y":460,"id":"amrf-rds-east","name":"RDS Primary — us-west-2","type":"database","status":"warning","cpuUsage":55,"location":{"zoneKey":"us-west-2a","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2a (Oregon)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Oregon RDS primary — still accepting writes; connection pool is under pressure from retrying west EC2 servers; streaming replication to the east replica is healthy","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}},{"x":220,"y":460,"id":"amrf-rds-east-standby","name":"RDS Standby — us-west-2","type":"database","status":"healthy","cpuUsage":15,"location":{"zoneKey":"us-west-2b","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2b (Oregon)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Oregon RDS standby (Multi-AZ) — in us-west-2b; healthy and ready to take over as primary if us-west-2a degrades further","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}},{"x":340,"y":460,"id":"amrf-rds-west","name":"RDS Primary Replica — us-east-1","type":"database","status":"healthy","cpuUsage":22,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Cross-region read replica in us-east-1a — serving east EC2 reads; replication lag is the key metric to watch; can be promoted to writable primary if the west region goes fully dark","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}},{"x":480,"y":460,"id":"amrf-rds-west-standby","name":"RDS Standby — us-east-1","type":"database","status":"healthy","cpuUsage":10,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Multi-AZ standby for the east replica in us-east-1b — ready for a within-region failover if us-east-1a degrades after the west-to-east promotion","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-amrf-r53-alb-east","sourceId":"amrf-r53","targetId":"amrf-alb-east"},{"id":"conn-amrf-r53-alb-west","sourceId":"amrf-r53","targetId":"amrf-alb-west"},{"id":"conn-amrf-alb-east-ec2-1","sourceId":"amrf-alb-east","targetId":"amrf-ec2-east-1"},{"id":"conn-amrf-alb-east-ec2-2","sourceId":"amrf-alb-east","targetId":"amrf-ec2-east-2"},{"id":"conn-amrf-alb-west-ec2-1","sourceId":"amrf-alb-west","targetId":"amrf-ec2-west-1"},{"id":"conn-amrf-alb-west-ec2-2","sourceId":"amrf-alb-west","targetId":"amrf-ec2-west-2"},{"id":"conn-amrf-ec2-east-1-rds","sourceId":"amrf-ec2-east-1","targetId":"amrf-rds-east"},{"id":"conn-amrf-ec2-east-2-rds","sourceId":"amrf-ec2-east-2","targetId":"amrf-rds-east"},{"id":"conn-amrf-ec2-west-1-rds","sourceId":"amrf-ec2-west-1","targetId":"amrf-rds-west"},{"id":"conn-amrf-ec2-west-2-rds","sourceId":"amrf-ec2-west-2","targetId":"amrf-rds-west"},{"id":"conn-amrf-rds-east-standby","sourceId":"amrf-rds-east","targetId":"amrf-rds-east-standby"},{"id":"conn-amrf-rds-east-west","sourceId":"amrf-rds-east","targetId":"amrf-rds-west"},{"id":"conn-amrf-rds-west-standby","sourceId":"amrf-rds-west","targetId":"amrf-rds-west-standby"}],"duration":"~15 min","tags":["AWS","Route 53","Multi-Region","Failover","DNS","Resilience","us-west-2","us-east-1"],"category":"reliability","defaultTrafficPatterns":[{"name":"Mid-failover baseline — Route 53 draining west, loading east","type":"step","startTime":0,"parameters":{"startTraffic":120,"endTraffic":120},"isActive":true},{"name":"East Scale-Out — DNS TTLs expire, us-east-1 absorbs us-west-2 drain","type":"ramp","startTime":10,"parameters":{"startTraffic":120,"endTraffic":1500,"duration":30},"isActive":true},{"name":"Traffic Recovery — watch us-east-1 scale in as us-west-2 recovers","type":"ramp","startTime":60,"parameters":{"startTraffic":1500,"endTraffic":120,"duration":30},"isActive":false}],"defaultFailureInjections":[{"name":"us-west-2 Regional Outage — us-west-2a Offline","type":"az_outage","targetZone":"usw2-az1","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}},{"name":"us-west-2 Regional Outage — us-west-2b Offline","type":"az_outage","targetZone":"usw2-az2","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}}],"realWorldIncident":{"date":"2026-07-24","provider":"AWS","summary":"On July 24, 2026, a routing failure in us-west-2 (Oregon) disrupted services for ~80 minutes. Apple Pay, DoorDash, Reddit, Hulu, and PlayStation Network were among the consumer services affected. Multi-region deployments with Route 53 health-check failover recovered in minutes; single-region deployments in us-west-2 were unavailable for the full outage window. This companion scenario runs the resilience side of the same outage — DNS TTL math plays out in real time so you can see exactly what a 30-second health-check interval and a 60-second record TTL buys you compared to the default 300-second settings when a live region goes dark.","references":[{"label":"AWS Health Dashboard","url":"https://health.aws.amazon.com/health/status"},{"label":"TechTimes · AWS us-west-2 outage report","url":"https://www.techtimes.com/articles/321567/20260725/aws-knocks-out-apple-pay-reddit-hulu-80-minutes-third-outage-since-may.htm"}]}},{"id":"web-autoscaling","title":"Web App Autoscaling","description":"Learn how autoscaling groups respond to traffic spikes and CPU thresholds","difficulty":"beginner","resources":[{"x":200,"y":100,"id":"alb-1","name":"Application Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Distributes incoming traffic across multiple targets","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":100,"y":400,"id":"ec2-1","name":"web-server-01","type":"compute","status":"healthy","cpuUsage":33,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Balanced compute for production workloads","baseLatency":3,"maxThroughput":30000,"serviceFamily":"ec2","costMultiplier":1}},{"x":300,"y":400,"id":"ec2-2","name":"web-server-02","type":"compute","status":"healthy","cpuUsage":33,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Balanced compute for production workloads","baseLatency":3,"maxThroughput":30000,"serviceFamily":"ec2","costMultiplier":1}},{"x":500,"y":300,"id":"rds-1","name":"MySQL RDS","type":"database","status":"healthy","cpuUsage":20,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Large managed relational database with high connection capacity","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-alb1-ec2-1","sourceId":"alb-1","targetId":"ec2-1"},{"id":"conn-alb1-ec2-2","sourceId":"alb-1","targetId":"ec2-2"},{"id":"conn-ec2-1-rds-1","sourceId":"ec2-1","targetId":"rds-1"},{"id":"conn-ec2-2-rds-1","sourceId":"ec2-2","targetId":"rds-1"}],"duration":"~10 min","tags":["AWS","Autoscaling","EC2"],"category":"scaling","defaultTrafficPatterns":[{"name":"Traffic Ramp — watch autoscaling fire","type":"ramp","startTime":0,"parameters":{"startTraffic":800,"endTraffic":4000,"duration":40},"isActive":true},{"name":"Traffic Recovery — see scale-in","type":"ramp","startTime":60,"parameters":{"startTraffic":4000,"endTraffic":800,"duration":40},"isActive":false}]},{"id":"db-failover","title":"Database Failover","description":"Watch a live Multi-AZ failover unfold step by step. Traffic ramps across two app servers behind an ALB, pushing connection pressure on RDS Primary until it overloads at step 15. The replica auto-promotes, latency spikes, error rates climb — then watch the system recover. A real AWS RDS Multi-AZ failure arc in under 60 steps.","difficulty":"intermediate","resources":[{"x":300,"y":50,"id":"dbf-alb-1","name":"Application Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Distributes incoming traffic across app servers in two Availability Zones","baseLatency":1,"maxThroughput":50000,"serviceFamily":"alb","costMultiplier":1}},{"x":100,"y":230,"id":"app-1","name":"App Server 1","type":"compute","status":"healthy","cpuUsage":28,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"High-performance compute for demanding applications","baseLatency":2,"maxThroughput":150000,"serviceFamily":"ec2","costMultiplier":2}},{"x":500,"y":230,"id":"app-2","name":"App Server 2","type":"compute","status":"healthy","cpuUsage":28,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"High-performance compute for demanding applications","baseLatency":2,"maxThroughput":150000,"serviceFamily":"ec2","costMultiplier":2}},{"x":100,"y":430,"id":"rds-primary","name":"RDS Primary","type":"database","status":"healthy","cpuUsage":38,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.xlarge","systemRole":"High-capacity managed relational database — primary write endpoint for all app traffic","baseLatency":5,"serviceFamily":"rds","costMultiplier":3,"maxConnections":2000}},{"x":500,"y":430,"id":"rds-replica","name":"RDS Replica","type":"database","status":"healthy","cpuUsage":15,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.xlarge","systemRole":"Standby replica in a separate AZ — auto-promotes on primary failure","baseLatency":5,"serviceFamily":"rds","costMultiplier":3,"maxConnections":2000}},{"x":300,"y":590,"id":"s3-1","name":"Backup Bucket","type":"storage","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"capacityGB":5000,"systemRole":"Object storage for files and static assets","baseLatency":20,"serviceFamily":"s3","costMultiplier":1}}],"connections":[{"id":"conn-dbf-alb-app1","sourceId":"dbf-alb-1","targetId":"app-1"},{"id":"conn-dbf-alb-app2","sourceId":"dbf-alb-1","targetId":"app-2"},{"id":"conn-app1-rds-primary","sourceId":"app-1","targetId":"rds-primary"},{"id":"conn-app2-rds-primary","sourceId":"app-2","targetId":"rds-primary"},{"id":"conn-rds-primary-replica","sourceId":"rds-primary","targetId":"rds-replica"},{"id":"conn-rds-primary-s3","sourceId":"rds-primary","targetId":"s3-1"},{"id":"conn-rds-replica-s3","sourceId":"rds-replica","targetId":"s3-1"}],"duration":"~30 min","tags":["RDS","High Availability","Multi-AZ"],"category":"failure","defaultTrafficPatterns":[{"name":"Traffic Ramp — watch RDS Primary reach its limit","type":"ramp","startTime":0,"parameters":{"startTraffic":3000,"endTraffic":8000,"duration":25},"isActive":true},{"name":"Traffic Recovery","type":"ramp","startTime":60,"parameters":{"startTraffic":8000,"endTraffic":2500,"duration":20},"isActive":false}],"defaultFailureInjections":[{"name":"RDS Primary Overload","type":"database_overload","targetResourceId":"rds-primary","severity":"severe","startTime":15,"endTime":50,"isActive":true,"parameters":{}}]},{"id":"launch-day-spike","title":"Launch Day Spike","description":"Simulate a Product Hunt-style launch surge on a single server. Find the exact upload concurrency where your transcoding queue backs up — before it happens in production.","difficulty":"beginner","resources":[{"x":150,"y":220,"id":"lds-vps-1","name":"Upload Server (VPS)","type":"compute","status":"healthy","cpuUsage":25,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-1 (Ashburn)"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (2 OCPU)","systemRole":"Single VPS that receives upload requests and enqueues transcoding jobs. No autoscaling — this is what you're stress-testing. Watch CPU climb as concurrent uploads pile up.","autoscaling":false,"baseLatency":4,"maxThroughput":1500,"serviceFamily":"oci-compute","costMultiplier":1}},{"x":400,"y":220,"id":"lds-queue-1","name":"Upload Queue","type":"queue","status":"healthy","location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-1 (Ashburn)"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (2 OCPU)","systemRole":"Buffers incoming upload jobs. Queue depth grows when the transcoding worker can't drain it fast enough — the key signal that your single server is saturated.","baseLatency":10,"maxThroughput":5000,"serviceFamily":"oci-queue","costMultiplier":0.5,"maxConnections":400}},{"x":650,"y":220,"id":"lds-worker-1","name":"Transcoding Worker","type":"compute","status":"healthy","cpuUsage":36,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-1 (Ashburn)"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (2 OCPU)","systemRole":"CPU-intensive transcoding process that drains the upload queue. Throughput is intentionally limited — this is the bottleneck you are trying to find.","autoscaling":false,"baseLatency":80,"maxThroughput":800,"serviceFamily":"oci-compute","costMultiplier":1}}],"connections":[{"id":"conn-lds-vps-queue","sourceId":"lds-vps-1","targetId":"lds-queue-1"},{"id":"conn-lds-queue-worker","sourceId":"lds-queue-1","targetId":"lds-worker-1"}],"duration":"~10 min","tags":["single-node","launch","queue","upload","OCI","Beginner"],"category":"scaling","defaultTrafficPatterns":[{"name":"Launch Day Surge (10×)","type":"ramp","startTime":0,"parameters":{"startTraffic":25,"endTraffic":250,"duration":30},"isActive":true}]},{"id":"aws-regional-outage-july-2026","title":"AWS Regional Outage — July 2026","description":"On July 24, 2026, a regional failure in us-west-2 (Oregon) degraded routing for dozens of popular services for ~80 minutes — ALBs began dropping packets, EC2 fleets lost internet connectivity, and retrying application servers flooded RDS with connection attempts. Apple Pay, DoorDash, Reddit, Hulu, and PlayStation Network were among the consumer services disrupted. This scenario opens mid-crisis: three EC2 web servers are already critical, ElastiCache has lost coherence, and RDS is absorbing a connection flood. Watch the error rate climb as EC2 loses egress, see the cache miss storm after ElastiCache degrades, and observe how the RDS read replica auto-promotes under primary pressure. The crisis is structural — not traffic-driven — so adding more servers cannot help. Around step 45, AWS restores the regional routing fabric; watch the error rate drop and servers work through their recovery cooldowns in the final stretch.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"aro-alb-1","name":"Application Load Balancer","type":"network","status":"warning","location":{"zoneKey":"us-west-2-regional","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Regional ALB — routing latency is elevated as the us-west-2 (Oregon) network fabric degrades; some health checks are timing out","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":100,"y":220,"id":"aro-ec2-1","name":"Web Server 1","type":"compute","status":"critical","cpuUsage":88,"location":{"zoneKey":"us-west-2-compute","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2 (Oregon) — Compute"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"Web tier in us-west-2a — regional routing degradation is causing packet loss and connection resets; retries are driving CPU to critical","autoscaling":true,"baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":2}},{"x":250,"y":220,"id":"aro-ec2-2","name":"Web Server 2","type":"compute","status":"critical","cpuUsage":85,"location":{"zoneKey":"us-west-2-compute","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2 (Oregon) — Compute"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"Web tier in us-west-2b — same regional event; retry storms from application clients are keeping CPU elevated","autoscaling":true,"baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":2}},{"x":400,"y":220,"id":"aro-ec2-3","name":"Web Server 3","type":"compute","status":"critical","cpuUsage":82,"location":{"zoneKey":"us-west-2-compute","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2 (Oregon) — Compute"},"provider":"aws","characteristics":{"size":"m5.xlarge","systemRole":"Web tier in us-west-2c — every AZ is affected; autoscaling cannot help when the problem is network-level, not capacity","autoscaling":true,"baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":2}},{"x":150,"y":400,"id":"aro-rds-1","name":"RDS MySQL Primary","type":"database","status":"warning","cpuUsage":62,"location":{"zoneKey":"us-west-2a","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2a (Oregon)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"RDS primary — retrying application servers have flooded the connection pool; CPU and connection pressure are building toward saturation","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}},{"x":350,"y":400,"id":"aro-rds-2","name":"RDS MySQL Standby","type":"database","status":"healthy","cpuUsage":18,"location":{"zoneKey":"us-west-2b","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2b (Oregon)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"RDS read replica in a separate AZ — currently healthy and positioned to auto-promote if the primary fails; observe promotion behavior as primary pressure rises","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}},{"x":250,"y":400,"id":"aro-cache-1","name":"ElastiCache for Redis","type":"cache","status":"critical","location":{"zoneKey":"us-west-2-compute","regionKey":"us-west-2","localityType":"zone","providerLabel":"us-west-2 (Oregon) — Compute"},"provider":"aws","recoveryPolicy":{"warningSteps":3,"criticalSteps":5,"warningCpuThreshold":70,"criticalCpuThreshold":80},"characteristics":{"systemRole":"ElastiCache cluster has lost replication coherence due to the regional routing failure — effective cache hit rate is 0 while critical, so every application request falls through to RDS, compounding the connection flood; when routing is restored the cluster warms back up over several steps","baseLatency":1,"cacheHitRate":0.8,"maxThroughput":50000,"serviceFamily":"elasticache","costMultiplier":1}}],"connections":[{"id":"conn-aro-alb-ec2-1","sourceId":"aro-alb-1","targetId":"aro-ec2-1"},{"id":"conn-aro-alb-ec2-2","sourceId":"aro-alb-1","targetId":"aro-ec2-2"},{"id":"conn-aro-alb-ec2-3","sourceId":"aro-alb-1","targetId":"aro-ec2-3"},{"id":"conn-aro-ec2-1-cache","sourceId":"aro-ec2-1","targetId":"aro-cache-1"},{"id":"conn-aro-ec2-2-cache","sourceId":"aro-ec2-2","targetId":"aro-cache-1"},{"id":"conn-aro-ec2-3-cache","sourceId":"aro-ec2-3","targetId":"aro-cache-1"},{"id":"conn-aro-ec2-1-rds","sourceId":"aro-ec2-1","targetId":"aro-rds-1"},{"id":"conn-aro-ec2-2-rds","sourceId":"aro-ec2-2","targetId":"aro-rds-1"},{"id":"conn-aro-ec2-3-rds","sourceId":"aro-ec2-3","targetId":"aro-rds-1"},{"id":"conn-aro-rds-1-rds-2","sourceId":"aro-rds-1","targetId":"aro-rds-2"}],"duration":"~15 min","tags":["AWS","Outage","Failure","us-west-2","RDS","ElastiCache"],"category":"failure","defaultTrafficPatterns":[{"name":"Steady outage load — structural crisis at low baseline","type":"step","startTime":0,"parameters":{"startTraffic":300,"endTraffic":300},"isActive":true},{"name":"Traffic Recovery — watch services stabilize","type":"ramp","startTime":60,"parameters":{"startTraffic":300,"endTraffic":2000,"duration":30},"isActive":false}],"defaultFailureInjections":[{"name":"Regional Network Routing Degradation — ALB Offline","type":"az_outage","targetZone":"us-west-2-regional","severity":"severe","startTime":0,"endTime":45,"isActive":true,"parameters":{}},{"name":"Regional Network Routing Degradation — Compute and Cache Offline","type":"az_outage","targetZone":"us-west-2-compute","severity":"severe","startTime":0,"endTime":45,"isActive":true,"parameters":{}},{"name":"RDS Connection Flood — retry storm from EC2","type":"database_overload","targetResourceId":"aro-rds-1","severity":"severe","startTime":0,"endTime":40,"isActive":true,"parameters":{}}],"realWorldIncident":{"date":"2026-07-24","provider":"AWS","summary":"On July 24, 2026, a routing failure in us-west-2 (Oregon) disrupted services for ~80 minutes. Apple Pay, DoorDash, Reddit, Hulu, and PlayStation Network were among the consumer services affected. Multi-region deployments with Route 53 health-check failover recovered in minutes; single-region deployments in us-west-2 were unavailable for the full outage window. This scenario opens mid-crisis so you can see what a network-layer structural failure looks like from the inside — three EC2 servers already critical, ElastiCache incoherent, RDS flooded — and understand why autoscaling and adding instances cannot fix a problem that lives in the routing fabric.","references":[{"label":"AWS Health Dashboard","url":"https://health.aws.amazon.com/health/status"},{"label":"TechTimes · AWS us-west-2 outage report","url":"https://www.techtimes.com/articles/321567/20260725/aws-knocks-out-apple-pay-reddit-hulu-80-minutes-third-outage-since-may.htm"}]}},{"id":"gcp-cloud-spanner-ha","title":"Cloud Spanner Outage — When Paxos Can't Help (July 4, 2023)","description":"On July 4, 2023 at 18:47 US/Pacific, Cloud Spanner went down globally for 2 hours and 16 minutes. This was not a zone failure. All three replicas — the leader in us-central1, the read-write replica in us-east1, and the witness in us-east4 — failed simultaneously because of an internal software bug in Spanner's serving layer. Paxos was intact and working correctly. The Paxos replication logs were consistent. No data was lost. But none of that mattered, because the bug was in the code that accepts client RPCs and routes them to the correct Paxos group — and that code ran identically on all three replicas. There was no healthy replica to elect as a new leader. Adding more replicas would not have helped. Switching regions would not have helped. The only fix was a rollback of the bad binary deployment, which Google engineers completed after 2h16m. This scenario opens at the moment the outage began: all three Spanner replicas are returning UNAVAILABLE, GKE pods are alive but every database call is failing with 503s, and the Global LB is routing correctly — it just has nowhere healthy to route writes to. Your learning goals: (1) Distinguish 'zone/hardware failure that Paxos handles automatically' from 'software bug that runs on all replicas identically and Paxos cannot fix'; (2) Learn the correct incident response for this class of failure — circuit breakers to stop retry storms, local caching to preserve read availability, write-queue to replay mutations on recovery, SLO alerting on 5xx rate not just p99 latency; (3) Understand why multi-region Spanner reduces but does not eliminate the need for application-level resilience patterns.","difficulty":"advanced","resources":[{"x":250,"y":60,"id":"gcs-cloud-dns","name":"Cloud DNS","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Cloud DNS was completely unaffected during the July 4, 2023 incident — all DNS zones, health-check routing, and record TTLs operated normally; the failure lived entirely in Spanner's internal serving layer, which is below the DNS and network tier; this is an important diagnostic: if DNS and the LB are healthy but writes are failing globally, suspect the database serving layer or a shared control-plane dependency, not a network partition","baseLatency":1,"maxThroughput":50000,"serviceFamily":"cloud-dns","costMultiplier":1}},{"x":250,"y":180,"id":"gcs-glb","name":"Global HTTP Load Balancer","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"The Global HTTP LB continued routing requests correctly throughout the incident — GKE pods were alive, responding to HTTP health probes, and the LB saw no backend failures from its perspective; the 5xx responses were coming from inside the application after Spanner RPCs failed, not from the LB tier; this explains why auto-mitigation (spinning up more backends, rerouting to a different region) had no effect — the bottleneck was not compute capacity or routing, it was the database","baseLatency":2,"maxThroughput":30000,"serviceFamily":"global-lb","costMultiplier":1}},{"x":400,"y":280,"id":"gcs-gke","name":"GKE Cluster — us-east1","type":"kubernetes","status":"warning","cpuUsage":35,"location":{"zoneKey":"us-east1-b","regionKey":"us-east1","localityType":"zone","providerLabel":"us-east1-b (South Carolina)"},"provider":"gcp","characteristics":{"size":"n2-standard-4","systemRole":"GKE pods are alive and accepting HTTP requests, but every call to Cloud Spanner is returning UNAVAILABLE — the application is propagating those errors as HTTP 503 to users; CPU is low because request processing terminates immediately on the Spanner error rather than doing any real work; the correct response is: (1) open circuit breakers on the Spanner client to stop flooding a broken service with retries, (2) serve reads from a local cache where staleness is acceptable, (3) queue writes in Cloud Pub/Sub or Bigtable for replay when Spanner recovers; HPA cannot help here — adding pods doesn't fix a database","autoscaling":true,"baseLatency":4,"maxThroughput":6000,"costMultiplier":2}},{"x":80,"y":420,"id":"gcs-spanner-leader","name":"Cloud Spanner — us-central1 (Leader)","type":"database","status":"critical","cpuUsage":0,"location":{"zoneKey":"us-central1-b","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-b (Iowa)"},"provider":"gcp","characteristics":{"size":"1000pu-multi","systemRole":"The Paxos leader replica is returning UNAVAILABLE on all RPCs — not because the zone is down, but because the serving-layer binary that accepts client connections has a bug; the Paxos log on this replica is intact, no transactions have been lost, and the replica is still participating in background replication — but it cannot accept new client reads or writes; this is the key distinction: Paxos protects the DATA layer, not the SERVING layer; a bad software deployment can break the serving layer while leaving Paxos intact, and that is exactly what happened on July 4, 2023","baseLatency":6,"serviceFamily":"cloud-spanner","costMultiplier":3.33,"maxConnections":1000}},{"x":380,"y":420,"id":"gcs-spanner-rw","name":"Cloud Spanner — us-east1 (Read-Write)","type":"database","status":"critical","cpuUsage":0,"location":{"zoneKey":"us-east1-c","regionKey":"us-east1","localityType":"zone","providerLabel":"us-east1-c (South Carolina)"},"provider":"gcp","characteristics":{"size":"1000pu-multi","systemRole":"The us-east1 read-write replica is also returning UNAVAILABLE — it runs the same serving-layer binary as us-central1, so the same bug affects it identically; this is the diagnostic signature of a shared-binary outage vs. a zone failure: a zone failure takes down exactly one replica and Paxos elects a new leader from the healthy ones; a software bug that ships to all replicas simultaneously takes down all replicas at once, and Paxos has no healthy candidate to elect; the election cannot complete, so no writes are accepted anywhere","baseLatency":7,"serviceFamily":"cloud-spanner","costMultiplier":3.33,"maxConnections":1000}},{"x":560,"y":520,"id":"gcs-spanner-witness","name":"Cloud Spanner — us-east4 (Witness)","type":"database","status":"critical","cpuUsage":0,"location":{"zoneKey":"us-east4-a","regionKey":"us-east4","localityType":"zone","providerLabel":"us-east4-a (N. Virginia)"},"provider":"gcp","characteristics":{"size":"1000pu-multi","systemRole":"The witness replica in us-east4 is also affected — witnesses do not serve client reads or writes (they only participate in Paxos quorum voting), so in a zone failure the witness being down would still allow us-central1 + us-east1 to form a majority; but here the witness being down alongside both read-write replicas confirms this is a global serving-layer failure, not a geographic one; the witness provides no additional resilience against same-binary bugs because it runs the same code as the read-write replicas","baseLatency":10,"serviceFamily":"cloud-spanner","costMultiplier":1,"maxConnections":1000}}],"connections":[{"id":"conn-gcs-dns-glb","sourceId":"gcs-cloud-dns","targetId":"gcs-glb"},{"id":"conn-gcs-glb-gke","sourceId":"gcs-glb","targetId":"gcs-gke"},{"id":"conn-gcs-gke-leader","sourceId":"gcs-gke","targetId":"gcs-spanner-leader"},{"id":"conn-gcs-gke-rw","sourceId":"gcs-gke","targetId":"gcs-spanner-rw"},{"id":"conn-gcs-leader-rw","sourceId":"gcs-spanner-leader","targetId":"gcs-spanner-rw"},{"id":"conn-gcs-leader-witness","sourceId":"gcs-spanner-leader","targetId":"gcs-spanner-witness"}],"duration":"~15 min","tags":["GCP","Cloud Spanner","Paxos","Multi-Region","Outage","Software Bug","Circuit Breaker","Resilience"],"category":"reliability","defaultTrafficPatterns":[{"name":"Pre-outage baseline — 500 RPS hitting Spanner hard","type":"step","startTime":0,"parameters":{"startTraffic":500,"endTraffic":500},"isActive":true}],"defaultFailureInjections":[{"name":"Serving-layer bug — us-central1-b Spanner leader offline","type":"az_outage","targetZone":"usc1-zone-b","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}},{"name":"Serving-layer bug — us-east1-c Spanner read-write replica offline","type":"az_outage","targetZone":"use1-zone-c","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}},{"name":"Serving-layer bug — us-east4-a Spanner witness offline","type":"az_outage","targetZone":"use4-zone-a","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}}],"realWorldIncident":{"date":"2023-07-04","provider":"GCP","summary":"On July 4, 2023, a software bug in Cloud Spanner's internal serving layer caused 2 hours and 16 minutes of global unavailability affecting reads and writes across all multi-region configurations. Unlike a zone failure — where Paxos automatically elects a new leader from healthy replicas — this bug ran identically on all replicas simultaneously, leaving no healthy candidate for election. The Paxos replication log was intact and no committed data was lost, but the serving layer that accepts client RPCs was broken on every node. Recovery required Google engineers to roll back the bad binary deployment across all regions. The incident is a real-world demonstration that Paxos guarantees apply to the data-replication layer, not to the serving-layer software that runs on top of it.","references":[{"label":"GCP · Cloud Spanner incident Sen4ACpGsFmquzb8FxRj","url":"https://status.cloud.google.com/incidents/Sen4ACpGsFmquzb8FxRj"}]}},{"id":"oci-multi-region-failover","title":"OCI Multi-Region Failover — Traffic Management Steering + Autonomous Data Guard","description":"A network degradation event has hit your us-ashburn-1 region, taking down your primary Load Balancer and putting your Autonomous Database primary under stress. OCI Traffic Management has already activated its FAILOVER steering policy and is routing new connections to your us-phoenix-1 backend — but there are three OCI-specific lessons worth understanding before you accept the recovery as complete. First: how OCI Traffic Management steering policies differ from Route 53 health checks. Traffic Management uses an ordered answer pool: it tries each pool in priority order and skips pools whose health-check endpoint is returning failures. The health monitor polls on a configurable interval (minimum 10 seconds for HTTP, 30 seconds for HTTPS); compare this to AWS Route 53 which polls every 10 seconds (standard health checks) or 30 seconds (basic). Both products are DNS-based, so clients must wait for TTL expiry before they resolve the failover IP — you can lower the DNS TTL on your OCI Traffic Management FQDN to reduce this window. Second: Autonomous Data Guard sync and how to promote the standby. The standby Autonomous Database in us-phoenix-1 continuously receives redo log shipments from the primary. You have two promotion paths: (1) SWITCHOVER — use this when the primary is reachable; it drains all in-flight transactions, promotes the standby to primary, and demotes the old primary to standby in a single consistent operation with zero data loss; invoke it from the OCI Console (Autonomous Database → More Actions → Switchover) or via CLI (`oci db autonomous-database switchover`). (2) FAILOVER — use this when the primary is completely unreachable; the standby applies all received redo and opens read-write; any redo not yet shipped from the primary (the replication lag at the moment of failure) represents your RPO; after failover the old primary is a 'former primary' that must be reinstated as a new standby before you can switchover back. Third: the application connection string problem. Unlike AWS RDS Multi-AZ (single DNS endpoint that transparently updates) or Azure SQL (automatic listener), OCI Autonomous Database standby has its own wallet and connection string. Your application must be pre-configured to use the standby endpoint, or you must update the connection string in your application configuration after promotion. Fast-Start Failover (FSFO) can automate the promotion decision when the primary is unreachable for a configurable observer timeout, but the application connection string update is still a manual step unless you build it into your deployment configuration.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"omrf-traffic-mgmt","name":"OCI Traffic Management — FAILOVER Steering Policy","type":"network","status":"warning","cpuUsage":0,"location":{"zoneKey":"iad-global","regionKey":"iad","localityType":"availability_domain","providerLabel":"Global DNS"},"provider":"oci","characteristics":{"systemRole":"OCI Traffic Management with a FAILOVER steering policy — the ordered answer pool has Ashburn as pool 1 and Phoenix as pool 2; the HTTP health monitor detected failure on the Ashburn Load Balancer endpoint and has removed pool 1 from DNS responses; DNS TTL for this FQDN is currently 60 seconds (lowered from the default 300 s before the incident window to minimize switchover time); note that OCI Traffic Management is DNS-based, like Route 53, so clients that already have the Ashburn IP cached must wait for TTL expiry before they resolve the Phoenix IP","baseLatency":1,"maxThroughput":100000,"costMultiplier":1}},{"x":80,"y":220,"id":"omrf-lb-ashburn","name":"OCI Load Balancer — us-ashburn-1 (Degraded)","type":"network","status":"warning","cpuUsage":0,"location":{"zoneKey":"iad-ad-1","regionKey":"iad","localityType":"availability_domain","providerLabel":"AD-1"},"provider":"oci","characteristics":{"systemRole":"Flexible Load Balancer in us-ashburn-1 — health check endpoint returning 5xx due to the regional network event; Traffic Management health monitor has marked this pool unhealthy and is directing new DNS resolutions to the Phoenix endpoint; existing connections from clients with stale DNS cache are still arriving here and failing; the Load Balancer itself is operational but cannot reach healthy backend compute instances in the affected AD","baseLatency":4,"maxThroughput":10000,"costMultiplier":1.5}},{"x":80,"y":380,"id":"omrf-compute-ashburn","name":"OCI Compute — us-ashburn-1 (Degraded)","type":"compute","status":"warning","cpuUsage":85,"location":{"zoneKey":"iad-ad-2","regionKey":"iad","localityType":"availability_domain","providerLabel":"AD-2"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (4 OCPU)","systemRole":"E4.Flex compute instances in AD-2 of us-ashburn-1 — spread across AD-2 to isolate from the AD-1 network fault; still handling a trickle of stale-TTL client connections; CPU elevated because the remaining traffic is unbalanced across fewer healthy instances; once Traffic Management TTL propagates globally these instances will drain naturally","baseLatency":8,"maxThroughput":3000,"serviceFamily":"vm","costMultiplier":1}},{"x":420,"y":380,"id":"omrf-compute-phoenix","name":"OCI Compute — us-phoenix-1 (Active)","type":"compute","status":"healthy","cpuUsage":46,"location":{"zoneKey":"phx-ad-1","regionKey":"phx","localityType":"availability_domain","providerLabel":"AD-1"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (4 OCPU)","systemRole":"E4.Flex compute instances in us-phoenix-1 — now receiving all routed traffic from Traffic Management; autoscaling is configured via an Instance Pool with a scale-out rule at 70% average CPU for 5 minutes; writes are crossing the OCI backbone (FastConnect DRG peering) to the Autonomous DB primary in Ashburn — expect ~12 ms added latency on write-heavy transactions; watch CPU climb as the full load transfers from Ashburn","autoscaling":true,"baseLatency":5,"maxThroughput":3000,"serviceFamily":"vm","costMultiplier":1}},{"x":80,"y":540,"id":"omrf-adb-primary","name":"Autonomous DB — Primary (us-ashburn-1)","type":"database","status":"warning","cpuUsage":68,"location":{"zoneKey":"iad-ad-2","regionKey":"iad","localityType":"availability_domain","providerLabel":"AD-2"},"provider":"oci","characteristics":{"size":"Autonomous DB (8 OCPU)","systemRole":"Autonomous Database primary in us-ashburn-1 — still accepting writes from Phoenix compute over the DRG cross-region link; Autonomous Data Guard is shipping redo to the Phoenix standby in near-real-time (typical lag < 1 s under normal load); to initiate a planned switchover: OCI Console → Autonomous Database → More Actions → Switchover, or CLI: oci db autonomous-database switchover --autonomous-database-id <ocid>; for an unplanned failover when primary is unreachable: More Actions → Failover (or oci db autonomous-database failover)","baseLatency":6,"costMultiplier":51.821,"maxConnections":2000}},{"x":420,"y":540,"id":"omrf-adb-standby","name":"Autonomous DB — Data Guard Standby (us-phoenix-1)","type":"database","status":"healthy","cpuUsage":24,"location":{"zoneKey":"phx-ad-2","regionKey":"phx","localityType":"availability_domain","providerLabel":"AD-2"},"provider":"oci","characteristics":{"size":"Autonomous DB (8 OCPU)","systemRole":"Autonomous Data Guard standby in us-phoenix-1 — continuously receiving redo log shipments from the Ashburn primary; replication lag is visible in the OCI Console under Autonomous Data Guard details; the standby is in a MOUNTED state and is not directly accessible for queries (unlike Azure SQL Active Geo-Replication read replicas or Amazon RDS read replicas — ADB standby is not a read replica by default); after SWITCHOVER or FAILOVER the standby opens read-write and the application must connect using the standby wallet/connection string or the new primary endpoint","baseLatency":5,"costMultiplier":51.821,"maxConnections":2000}}],"connections":[{"id":"conn-omrf-tm-lb","sourceId":"omrf-traffic-mgmt","targetId":"omrf-lb-ashburn"},{"id":"conn-omrf-tm-phoenix","sourceId":"omrf-traffic-mgmt","targetId":"omrf-compute-phoenix"},{"id":"conn-omrf-lb-compute","sourceId":"omrf-lb-ashburn","targetId":"omrf-compute-ashburn"},{"id":"conn-omrf-compute-ash-adb","sourceId":"omrf-compute-ashburn","targetId":"omrf-adb-primary"},{"id":"conn-omrf-compute-phx-adb","sourceId":"omrf-compute-phoenix","targetId":"omrf-adb-primary"},{"id":"conn-omrf-adb-dg","sourceId":"omrf-adb-primary","targetId":"omrf-adb-standby"}],"duration":"~14 min","tags":["OCI","Traffic Management","Autonomous Database","Data Guard","Multi-Region","HA","Resilience"],"category":"reliability","defaultTrafficPatterns":[{"name":"us-ashburn-1 degraded — Traffic Management routing to us-phoenix-1","type":"step","startTime":0,"parameters":{"startTraffic":800,"endTraffic":800},"isActive":true},{"name":"Scale test — confirm Phoenix absorbs full production load","type":"ramp","startTime":60,"parameters":{"startTraffic":800,"endTraffic":4000,"duration":30},"isActive":false}],"defaultFailureInjections":[{"name":"us-ashburn-1 Network Event — AD-1 Load Balancer Offline","type":"az_outage","targetZone":"iad-ad-1","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}}],"realWorldIncident":{"date":"2026-03-03","provider":"OCI","summary":"On March 3–4, 2026, a network infrastructure failure at OCI's US East (Ashburn) data center disrupted service for ~20 hours (13:24 UTC Mar 3 to 09:18 UTC Mar 4). TikTok US users and thousands of other OCI Ashburn customers were unable to reach their applications for much of that window — Oracle's own status page described intermittent connection timeouts and elevated latency across OCI service operations in the region. Customers with a prepared Traffic Management FAILOVER steering policy and an Autonomous Data Guard standby in us-phoenix-1 were able to redirect traffic within one DNS TTL; those without a standby region faced the full 20-hour outage. This scenario compresses that failover into a 14-minute practice run so you can walk through the SWITCHOVER and FAILOVER decision before the real thing.","references":[{"label":"The Register · Oracle outage & TikTok","url":"https://www.theregister.com/2026/03/04/oracle_cloud_outage_tiktok/"},{"label":"OCI Status page","url":"https://ocistatus.oraclecloud.com/"}]}},{"id":"digitalocean-multi-region-ha","title":"DigitalOcean Multi-Region HA — Floating IP Failover and the Cross-Region Database Gap","description":"Your NYC3 datacenter is experiencing a network degradation event. The Load Balancer health check is failing, and your primary Managed PostgreSQL cluster is under stress. You need to reroute traffic to your SFO3 standby Droplet — but this is where DigitalOcean's HA model diverges from AWS, GCP, and Azure in ways that matter for your architecture decisions. This scenario has three lessons. Lesson 1: DigitalOcean Floating IPs are region-scoped, not global-anycast. A Floating IP is a public IP tied to a single datacenter (nyc3). You can reassign it instantly to any Droplet in the same datacenter via API (`doctl compute floating-ip-action assign <floating-ip> <droplet-id>`) — but you cannot point a nyc3 Floating IP at an sfo3 Droplet. Cross-region compute failover on DigitalOcean therefore requires a DNS change: update your domain to point from the nyc3 Floating IP (or Load Balancer IP) to the sfo3 Droplet IP. DNS propagation is TTL-bounded — if your TTL was 300 seconds before the incident, clients can be unreachable for up to 5 minutes. Compare this to AWS Global Accelerator or Azure Front Door (both anycast — no DNS change, sub-second failover) or GCP Global Load Balancer (also anycast). If you need instant cross-region compute failover on DigitalOcean, the workaround is to set a very low DNS TTL (60 s or lower) before an incident and use a third-party DNS provider like Cloudflare with sub-second propagation. Lesson 2: DigitalOcean Managed PostgreSQL has excellent within-region HA — automatic primary/standby failover, promoted in under 60 seconds with zero manual intervention. But there is no built-in cross-region standby. The 'SFO3 standby' PostgreSQL cluster in this scenario is maintained via manual pg_logical replication that you configured yourself: a publication on the nyc3 primary and a subscription on the sfo3 cluster. pg_logical is asynchronous, so the sfo3 cluster lags the nyc3 primary by the replication delay. To promote the sfo3 cluster to writable primary you must: (1) confirm replication lag is acceptable, (2) disable the subscription on sfo3 (`ALTER SUBSCRIPTION sub_name DISABLE`), (3) update your application connection string to point to the sfo3 endpoint, (4) optionally set up reverse replication from sfo3 back to nyc3 for when you want to fail back. Compare this to AWS RDS Multi-AZ (automatic DNS failover, same endpoint), Azure SQL Active Geo-Replication (`az sql db replica set-primary`, ~30 s), and OCI Autonomous Data Guard (SWITCHOVER in the console). Lesson 3: When to choose DigitalOcean despite these limitations. DO's pricing is genuinely competitive — a 2 vCPU / 4 GB Droplet costs $26/mo vs $60–80/mo for equivalent AWS/GCP/Azure instances. DO's developer experience is simpler, its App Platform handles single-region PaaS deployments with built-in horizontal scaling, and for applications where a single region is sufficient (most internal tools, early-stage startups, read-heavy workloads with acceptable RPO), DO's within-region HA is robust. The cross-region gap only matters when you need sub-minute global failover with automatic database promotion — and that is a legitimate requirement that pushes you toward the big three.","difficulty":"intermediate","resources":[{"x":250,"y":80,"id":"domr-lb-nyc3","name":"Load Balancer — nyc3 (Degraded)","type":"network","status":"warning","cpuUsage":0,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"az","providerLabel":"nyc3 (New York 3)"},"provider":"digitalocean","characteristics":{"systemRole":"DigitalOcean Load Balancer in nyc3 — health check polling the backend Droplet HTTP endpoint every 10 seconds; NYC3 network event is causing health checks to fail; the Load Balancer IP is not a Floating IP — it is a static IP assigned at creation; to redirect incoming traffic to SFO3 you must update your DNS A record to point to the SFO3 Droplet IP (or a Floating IP assigned to it); DNS TTL is currently 60 seconds (lowered proactively); full propagation takes approximately 60–120 seconds with a modern DNS resolver","baseLatency":3,"maxThroughput":10000,"costMultiplier":1.5}},{"x":80,"y":260,"id":"domr-droplet-nyc3","name":"Droplet — nyc3 (Degraded)","type":"compute","status":"warning","cpuUsage":78,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"az","providerLabel":"nyc3 (New York 3)"},"provider":"digitalocean","characteristics":{"size":"s-2vcpu-4gb","systemRole":"2 vCPU / 4 GB Droplet in NYC3 running your application — degraded due to the regional network event; a Floating IP is assigned to this Droplet within nyc3, which lets you instantly reassign it to another nyc3 Droplet (`doctl compute floating-ip-action assign <ip> <droplet-id>`); but because Floating IPs are region-scoped, you cannot reassign this nyc3 Floating IP to the SFO3 Droplet — DNS update is the only cross-region path","baseLatency":7,"maxThroughput":1800,"serviceFamily":"droplet","costMultiplier":1}},{"x":420,"y":260,"id":"domr-droplet-sfo3","name":"Droplet — sfo3 (Active)","type":"compute","status":"healthy","cpuUsage":41,"location":{"zoneKey":"sfo3-az1","regionKey":"sfo3","localityType":"az","providerLabel":"sfo3 (San Francisco 3)"},"provider":"digitalocean","characteristics":{"size":"s-2vcpu-4gb","systemRole":"2 vCPU / 4 GB standby Droplet in SFO3 — now receiving traffic after the DNS update; DigitalOcean does not have a managed autoscaling service for individual Droplets; horizontal scaling is achieved via the App Platform (managed PaaS) or by manually creating additional Droplets and adding them to a new Load Balancer pool; for this failover scenario, this single Droplet is the failover target and is sized to handle the current traffic without scaling","autoscaling":true,"baseLatency":5,"maxThroughput":1800,"serviceFamily":"droplet","costMultiplier":1}},{"x":80,"y":440,"id":"domr-pg-nyc3","name":"Managed PostgreSQL — nyc3 Primary (Degraded)","type":"database","status":"warning","cpuUsage":71,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"az","providerLabel":"nyc3 (New York 3)"},"provider":"digitalocean","characteristics":{"size":"db-s-2vcpu-4gb","systemRole":"DigitalOcean Managed PostgreSQL primary cluster in NYC3 — has an automatic within-region standby node that DigitalOcean manages; if this primary becomes unreachable, DO promotes the standby automatically in under 60 seconds at the same connection endpoint (no DNS change required for within-region failover); the cross-region SFO3 cluster is maintained separately via pg_logical replication that you configured; the nyc3 primary is still reachable and accepting writes from the SFO3 Droplet over the public internet — but connection latency is elevated (~45 ms NYC-to-SFO) vs the ~1 ms local connection that existed before the failover","baseLatency":8,"costMultiplier":2.3,"maxConnections":200}},{"x":420,"y":440,"id":"domr-pg-sfo3","name":"Managed PostgreSQL — sfo3 Replica (pg_logical)","type":"database","status":"healthy","cpuUsage":22,"location":{"zoneKey":"sfo3-az1","regionKey":"sfo3","localityType":"az","providerLabel":"sfo3 (San Francisco 3)"},"provider":"digitalocean","characteristics":{"size":"db-s-2vcpu-4gb","systemRole":"DigitalOcean Managed PostgreSQL cluster in SFO3 — this is not a native cross-region standby; it is a separate cluster with a pg_logical subscription receiving changes from the nyc3 publication; replication lag is typically 100–500 ms under normal NYC3 load; to promote this cluster to writable primary: (1) check lag: SELECT now() - pg_last_xact_replay_timestamp() AS lag; (2) disable subscription: ALTER SUBSCRIPTION nyc3_sub DISABLE; (3) update application DATABASE_URL to point to the SFO3 connection string; DigitalOcean does not provide automatic cross-region failover — this promotion is entirely manual","baseLatency":5,"costMultiplier":2.3,"maxConnections":200}}],"connections":[{"id":"conn-domr-lb-droplet-nyc3","sourceId":"domr-lb-nyc3","targetId":"domr-droplet-nyc3"},{"id":"conn-domr-lb-droplet-sfo3","sourceId":"domr-lb-nyc3","targetId":"domr-droplet-sfo3"},{"id":"conn-domr-droplet-nyc3-pg","sourceId":"domr-droplet-nyc3","targetId":"domr-pg-nyc3"},{"id":"conn-domr-droplet-sfo3-pg-nyc3","sourceId":"domr-droplet-sfo3","targetId":"domr-pg-nyc3"},{"id":"conn-domr-pg-replication","sourceId":"domr-pg-nyc3","targetId":"domr-pg-sfo3"}],"duration":"~10 min","tags":["DigitalOcean","Floating IP","Multi-Region","PostgreSQL","pg_logical","HA","Resilience"],"category":"reliability","defaultTrafficPatterns":[{"name":"NYC3 degraded — DNS update in progress, traffic routing to SFO3","type":"step","startTime":0,"parameters":{"startTraffic":600,"endTraffic":600},"isActive":true},{"name":"Scale test — confirm SFO3 Droplet handles full NYC3 production load","type":"ramp","startTime":60,"parameters":{"startTraffic":600,"endTraffic":2500,"duration":30},"isActive":false}],"defaultFailureInjections":[{"name":"NYC3 Network Degradation — Full Datacenter Offline","type":"az_outage","targetZone":"nyc3-az1","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}}]},{"id":"multi-cloud-hybrid","title":"Multi-Cloud Hybrid Architecture","description":"Experience a hybrid architecture spanning AWS and OCI with cross-cloud connectivity","difficulty":"intermediate","resources":[{"x":100,"y":100,"id":"aws-alb","name":"AWS Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Distributes incoming traffic across multiple targets","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":100,"y":250,"id":"aws-ec2","name":"AWS Web Server","type":"compute","status":"healthy","cpuUsage":32,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Balanced compute for production workloads","baseLatency":3,"maxThroughput":50000,"serviceFamily":"ec2","costMultiplier":1}},{"x":400,"y":100,"id":"oci-lb","name":"OCI Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-1 (Ashburn)"},"provider":"oci","characteristics":{"systemRole":"Highly available load balancing service","baseLatency":1,"maxThroughput":11000,"serviceFamily":"oci-lb","costMultiplier":0.8}},{"x":400,"y":250,"id":"oci-compute","name":"OCI API Server","type":"compute","status":"healthy","cpuUsage":27,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-1 (Ashburn)","faultDomainKey":"FD-1"},"provider":"oci","characteristics":{"size":"VM.Standard3.Flex (4 OCPU)","systemRole":"Flexible compute with customizable cores","baseLatency":2,"maxThroughput":30000,"serviceFamily":"oci-compute","costMultiplier":1}},{"x":250,"y":400,"id":"oci-autonomous","name":"Autonomous DB","type":"database","status":"healthy","cpuUsage":25,"location":{"zoneKey":"us-ashburn-1-ad-2","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-2 (Ashburn)"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (2 OCPU)","systemRole":"Self-tuning, self-patching database","baseLatency":4,"serviceFamily":"autonomous-db","costMultiplier":1.7,"maxConnections":600}},{"x":100,"y":550,"id":"s3-backup","name":"S3 Backup","type":"storage","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"capacityGB":5000,"systemRole":"Object storage for files and static assets","baseLatency":20,"serviceFamily":"s3","costMultiplier":1}},{"x":400,"y":550,"id":"oci-object","name":"OCI Object Storage","type":"storage","status":"healthy","location":{"zoneKey":"us-ashburn-1-regional","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 (Regional)"},"provider":"oci","characteristics":{"capacityGB":5000,"systemRole":"Durable object storage with lifecycle management","baseLatency":16,"serviceFamily":"oci-object-storage","costMultiplier":0.8}}],"connections":[{"id":"conn-awsalb-awsec2","sourceId":"aws-alb","targetId":"aws-ec2"},{"id":"conn-ocilb-ocicompute","sourceId":"oci-lb","targetId":"oci-compute"},{"id":"conn-awsec2-ociautonomous","sourceId":"aws-ec2","targetId":"oci-autonomous"},{"id":"conn-ocicompute-ociautonomous","sourceId":"oci-compute","targetId":"oci-autonomous"},{"id":"conn-awsec2-s3backup","sourceId":"aws-ec2","targetId":"s3-backup"},{"id":"conn-ocicompute-ociobject","sourceId":"oci-compute","targetId":"oci-object"}],"duration":"~20 min","tags":["Multi-Cloud","AWS","OCI","Hybrid"],"category":"scaling","defaultTrafficPatterns":[{"name":"Steady Baseline","type":"step","startTime":0,"parameters":{"startTraffic":1200,"endTraffic":1200},"isActive":true}]},{"id":"cdn-accelerated-web-app","title":"CDN-Accelerated Web App","description":"CloudFront sits in front of your application and caches static content at edge locations worldwide. Only cache misses and dynamic API requests reach the Application Load Balancer, reducing backend load by 60–80% during traffic spikes. EC2 handles the dynamic compute tier while RDS stores persistent data.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"cdn-cf-1","name":"CloudFront CDN","type":"network","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"CDN that caches static assets at 400+ global edge locations — cache hits never reach your servers, dramatically reducing origin load and improving response times for end users","baseLatency":5,"cacheHitRate":0.75,"maxThroughput":50000,"serviceFamily":"cloudfront","costMultiplier":1.5}},{"x":250,"y":200,"id":"cdn-alb-1","name":"Application Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Receives only cache-miss and dynamic requests forwarded by CloudFront","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":130,"y":360,"id":"cdn-ec2-1","name":"Web Server","type":"compute","status":"healthy","cpuUsage":28,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Handles dynamic page rendering and API logic for requests that CloudFront cannot serve from cache","baseLatency":3,"maxThroughput":20000,"serviceFamily":"ec2","costMultiplier":1}},{"x":370,"y":360,"id":"cdn-ec2-2","name":"Web Server 2","type":"compute","status":"healthy","cpuUsage":25,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Second origin server in a separate AZ for high availability behind the ALB","baseLatency":3,"maxThroughput":20000,"serviceFamily":"ec2","costMultiplier":1}},{"x":250,"y":520,"id":"cdn-rds-1","name":"RDS MySQL","type":"database","status":"healthy","cpuUsage":20,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Managed relational database that only sees queries from dynamic origin requests — CDN caching keeps connection count well below capacity","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-cdncf1-alb1","sourceId":"cdn-cf-1","targetId":"cdn-alb-1"},{"id":"conn-cdnalb1-ec2-1","sourceId":"cdn-alb-1","targetId":"cdn-ec2-1"},{"id":"conn-cdnalb1-ec2-2","sourceId":"cdn-alb-1","targetId":"cdn-ec2-2"},{"id":"conn-cdnec2-1-rds1","sourceId":"cdn-ec2-1","targetId":"cdn-rds-1"},{"id":"conn-cdnec2-2-rds1","sourceId":"cdn-ec2-2","targetId":"cdn-rds-1"}],"duration":"~15 min","tags":["AWS","CloudFront","CDN","EC2","RDS","Caching"],"category":"scaling","defaultTrafficPatterns":[{"name":"Traffic Spike — CDN absorbs 75% at the edge","type":"ramp","startTime":0,"parameters":{"startTraffic":2000,"endTraffic":8000,"duration":25},"isActive":true},{"name":"Traffic Recovery","type":"ramp","startTime":60,"parameters":{"startTraffic":8000,"endTraffic":2000,"duration":20},"isActive":false}]},{"id":"oci-web-app","title":"OCI Web Application","description":"Deploy a simple web application on Oracle Cloud Infrastructure with load balancing","difficulty":"beginner","resources":[{"x":200,"y":100,"id":"oci-lb-1","name":"OCI Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-1 (Ashburn)"},"provider":"oci","characteristics":{"systemRole":"Highly available load balancing service","baseLatency":1,"maxThroughput":11000,"serviceFamily":"oci-lb","costMultiplier":0.8}},{"x":100,"y":300,"id":"oci-compute-1","name":"Web Server 1","type":"compute","status":"healthy","cpuUsage":30,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-1 (Ashburn)","faultDomainKey":"FD-1"},"provider":"oci","characteristics":{"size":"VM.Standard3.Flex (4 OCPU)","systemRole":"Flexible compute with customizable cores","baseLatency":2,"maxThroughput":30000,"serviceFamily":"oci-compute","costMultiplier":1}},{"x":300,"y":300,"id":"oci-compute-2","name":"Web Server 2","type":"compute","status":"healthy","cpuUsage":30,"location":{"zoneKey":"us-ashburn-1-ad-2","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-2 (Ashburn)","faultDomainKey":"FD-2"},"provider":"oci","characteristics":{"size":"VM.Standard3.Flex (4 OCPU)","systemRole":"Flexible compute with customizable cores","baseLatency":2,"maxThroughput":30000,"serviceFamily":"oci-compute","costMultiplier":1}},{"x":500,"y":300,"id":"autonomous-db-1","name":"Autonomous Database","type":"database","status":"healthy","cpuUsage":20,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"availability_domain","providerLabel":"AD-1 (Ashburn)"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (2 OCPU)","systemRole":"Self-tuning, self-patching database","baseLatency":4,"serviceFamily":"autonomous-db","costMultiplier":1.7,"maxConnections":600}},{"x":200,"y":500,"id":"object-storage-1","name":"Object Storage","type":"storage","status":"healthy","location":{"zoneKey":"us-ashburn-1-regional","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 (Regional)"},"provider":"oci","characteristics":{"capacityGB":5000,"systemRole":"Durable object storage with lifecycle management","baseLatency":16,"serviceFamily":"oci-object-storage","costMultiplier":0.8}}],"connections":[{"id":"conn-ocilb1-ocicomp1","sourceId":"oci-lb-1","targetId":"oci-compute-1"},{"id":"conn-ocilb1-ocicomp2","sourceId":"oci-lb-1","targetId":"oci-compute-2"},{"id":"conn-ocicomp1-autodb1","sourceId":"oci-compute-1","targetId":"autonomous-db-1"},{"id":"conn-ocicomp2-autodb1","sourceId":"oci-compute-2","targetId":"autonomous-db-1"},{"id":"conn-ocicomp1-objstore1","sourceId":"oci-compute-1","targetId":"object-storage-1"},{"id":"conn-ocicomp2-objstore1","sourceId":"oci-compute-2","targetId":"object-storage-1"}],"duration":"~12 min","tags":["OCI","Web App","Autonomous DB"],"category":"scaling","defaultTrafficPatterns":[{"name":"Steady Baseline","type":"step","startTime":0,"parameters":{"startTraffic":1400,"endTraffic":1400},"isActive":true}]},{"id":"github-actions-pages-db-failover","title":"GitHub Actions & Pages Database Failover (Unofficial)","description":"Independent, non-affiliated simulation inspired by a GitHub Status update posted under 'Incident with Actions.' This scenario opens mid-crisis: the primary metadata database has already begun failing over to its standby replica, and the dependent Actions job-assignment path and Pages publish/build path both start out visibly degraded — mirroring the reported 'degraded availability' (Actions) and 'degraded performance' (Pages). Watch both paths recover automatically as the standby replica takes over, with no user action required.","difficulty":"intermediate","resources":[{"x":80,"y":200,"id":"ghdb-webhook","name":"Actions Webhook Intake","type":"network","status":"healthy","location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Accepts inbound webhook events that trigger new Actions workflow runs and Pages deployment requests.","baseLatency":3,"maxThroughput":50000,"serviceFamily":"application-gateway","costMultiplier":1}},{"x":320,"y":80,"id":"ghdb-actions-queue","name":"Actions Job Queue","type":"queue","status":"healthy","location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Buffers queued Actions jobs waiting for the controller to assign them to a runner.","baseLatency":5,"maxThroughput":20000,"serviceFamily":"service-bus","costMultiplier":1}},{"x":560,"y":80,"id":"ghdb-actions-controller","name":"Actions Controller & Scheduler","type":"compute","status":"critical","cpuUsage":84,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Assigns queued jobs to available runners by reading run and queue state from the primary metadata database on every scheduling decision — the first place a database failover surfaces as degraded job-assignment availability.","autoscaling":false,"baseLatency":8,"maxThroughput":600,"serviceFamily":"app-service","costMultiplier":1}},{"x":800,"y":80,"id":"ghdb-runner-fleet","name":"Actions Runner Fleet","type":"kubernetes","status":"healthy","cpuUsage":28,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US · Zone 1"},"provider":"azure","characteristics":{"maxNodes":12,"minNodes":4,"nodeCount":4,"systemRole":"Executes assigned Actions jobs — starved of new work while the controller's job-assignment path is degraded, but not itself failing.","autoscaling":false,"baseLatency":10,"maxThroughput":20000,"serviceFamily":"aks","costMultiplier":1}},{"x":560,"y":320,"id":"ghdb-pages-build","name":"Pages Build Service","type":"compute","status":"warning","cpuUsage":76,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Builds and publishes GitHub Pages sites, writing deployment metadata to the primary database on every publish — the first place a database failover surfaces as degraded Pages performance.","autoscaling":false,"baseLatency":12,"maxThroughput":95,"serviceFamily":"app-service","costMultiplier":1}},{"x":800,"y":320,"id":"ghdb-pages-edge","name":"Pages Edge Network (Front Door)","type":"network","status":"healthy","location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Serves already-published Pages sites from edge cache — resilient to the metadata database failover because it does not need a live database read to serve cached content.","baseLatency":2,"cacheHitRate":0.85,"maxThroughput":50000,"serviceFamily":"azure-front-door","costMultiplier":1}},{"x":320,"y":500,"id":"ghdb-metadata-primary","name":"Primary Metadata Database","type":"database","status":"warning","cpuUsage":45,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Primary write endpoint for repository, workflow-run, and Pages deployment metadata; degrades every dependent write path when it fails over.","baseLatency":6,"serviceFamily":"azure-sql","costMultiplier":2,"maxConnections":6000}},{"x":560,"y":500,"id":"ghdb-metadata-replica","name":"Standby Metadata Replica","type":"database","status":"healthy","cpuUsage":18,"location":{"zoneKey":"eastus-2","regionKey":"eastus","localityType":"zone","providerLabel":"East US · Zone 2"},"provider":"azure","characteristics":{"systemRole":"Standby replica in a separate zone — auto-promotes to primary when the metadata database fails over.","baseLatency":6,"serviceFamily":"azure-sql","costMultiplier":2,"maxConnections":6000}}],"connections":[{"id":"conn-ghdb-webhook-queue","sourceId":"ghdb-webhook","targetId":"ghdb-actions-queue"},{"id":"conn-ghdb-webhook-pagesbuild","sourceId":"ghdb-webhook","targetId":"ghdb-pages-build"},{"id":"conn-ghdb-queue-controller","sourceId":"ghdb-actions-queue","targetId":"ghdb-actions-controller"},{"id":"conn-ghdb-controller-runners","sourceId":"ghdb-actions-controller","targetId":"ghdb-runner-fleet"},{"id":"conn-ghdb-controller-db","sourceId":"ghdb-actions-controller","targetId":"ghdb-metadata-primary"},{"id":"conn-ghdb-pagesbuild-edge","sourceId":"ghdb-pages-build","targetId":"ghdb-pages-edge"},{"id":"conn-ghdb-pagesbuild-db","sourceId":"ghdb-pages-build","targetId":"ghdb-metadata-primary"},{"id":"conn-ghdb-db-replica","sourceId":"ghdb-metadata-primary","targetId":"ghdb-metadata-replica"}],"duration":"~10 min","tags":["Azure","GitHub","Database Failover","Real Incident","Actions","Pages"],"category":"failure","defaultTrafficPatterns":[{"name":"Cutover Retry Storm — failover already underway","type":"step","startTime":0,"endTime":8,"parameters":{"startTraffic":620,"endTraffic":620},"isActive":true},{"name":"Backlog Drains — replica finishes taking over","type":"ramp","startTime":8,"parameters":{"startTraffic":620,"endTraffic":120,"duration":12},"isActive":true}],"defaultFailureInjections":[{"name":"Primary Metadata Database Failover","type":"database_overload","targetResourceId":"ghdb-metadata-primary","severity":"severe","startTime":0,"endTime":30,"isActive":true,"parameters":{}}],"realWorldIncident":{"provider":"GitHub-inspired (unaffiliated)","date":"2026-08-26","summary":"Independent educational simulation inspired by a GitHub Status update posted under 'Incident with Actions.' As of this writing (August 26, 2026), GitHub's posted updates state only that the team identified an issue with a database primary and failed over to a replica, and that GitHub Actions was reporting degraded availability while GitHub Pages was reporting degraded performance during the cutover. The incident was still marked 'Investigating' at the time of writing — GitHub had not published a final root-cause writeup, an incident duration, or error-rate/impact figures. This simulation therefore models the general mechanism the status page describes (a metadata database primary failing over to a standby replica, temporarily degrading dependent services) using illustrative, clearly-labeled failover characteristics rather than numbers GitHub has published. It is not an official GitHub product, is unaffiliated with GitHub, and is not an exact forensic reconstruction.","references":[{"label":"GitHub Status — Incident with Actions (in progress at time of writing)","url":"https://www.githubstatus.com/incidents/y1t7p9fzrlj2"}]}},{"id":"microservices-cache-queue","title":"Microservices with Redis and SQS","description":"See how a Redis cache and SQS message queue reduce database load and smooth out traffic spikes in a microservices architecture","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"msvc-alb","name":"Application Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Distributes incoming traffic across multiple targets","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":100,"y":200,"id":"msvc-api-1","name":"API Server 1","type":"compute","status":"healthy","cpuUsage":32,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Balanced compute for production workloads","baseLatency":3,"maxThroughput":40000,"serviceFamily":"ec2","costMultiplier":1}},{"x":400,"y":200,"id":"msvc-api-2","name":"API Server 2","type":"compute","status":"healthy","cpuUsage":32,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Balanced compute for production workloads","baseLatency":3,"maxThroughput":40000,"serviceFamily":"ec2","costMultiplier":1}},{"x":100,"y":380,"id":"msvc-redis","name":"ElastiCache for Redis","type":"cache","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"cache.t3.micro","systemRole":"Managed Redis cache that reduces database load by serving frequent queries from memory","baseLatency":1,"cacheHitRate":0.8,"maxThroughput":50000,"costMultiplier":1}},{"x":400,"y":380,"id":"msvc-sqs","name":"Amazon SQS","type":"queue","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"standard","systemRole":"Fully managed message queue that decouples and buffers asynchronous workloads","baseLatency":5,"maxThroughput":100000,"costMultiplier":1}},{"x":250,"y":540,"id":"msvc-rds","name":"RDS MySQL","type":"database","status":"healthy","cpuUsage":22,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Large managed relational database with high connection capacity","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-msvcalb-api1","sourceId":"msvc-alb","targetId":"msvc-api-1"},{"id":"conn-msvcalb-api2","sourceId":"msvc-alb","targetId":"msvc-api-2"},{"id":"conn-api1-redis","sourceId":"msvc-api-1","targetId":"msvc-redis"},{"id":"conn-api2-redis","sourceId":"msvc-api-2","targetId":"msvc-redis"},{"id":"conn-api1-sqs","sourceId":"msvc-api-1","targetId":"msvc-sqs"},{"id":"conn-api2-sqs","sourceId":"msvc-api-2","targetId":"msvc-sqs"},{"id":"conn-sqs-rds","sourceId":"msvc-sqs","targetId":"msvc-rds"},{"id":"conn-api1-rds","sourceId":"msvc-api-1","targetId":"msvc-rds"},{"id":"conn-api2-rds","sourceId":"msvc-api-2","targetId":"msvc-rds"}],"duration":"~15 min","tags":["AWS","Redis","SQS","Microservices","Cache","Queue"],"category":"scaling","defaultTrafficPatterns":[{"name":"Traffic Spike — watch cache absorb the burst","type":"ramp","startTime":0,"parameters":{"startTraffic":1000,"endTraffic":5000,"duration":20},"isActive":true},{"name":"Traffic Recovery","type":"ramp","startTime":60,"parameters":{"startTraffic":5000,"endTraffic":1000,"duration":20},"isActive":false}]},{"id":"digitalocean-starter","title":"DigitalOcean Starter Web App","description":"A simple, budget-friendly web app on DigitalOcean: a Load Balancer routes visitors to a single Droplet, which stores data in a Managed PostgreSQL database. A great starting point for developers new to cloud infrastructure who want predictable, low monthly bills.","difficulty":"beginner","resources":[{"x":250,"y":60,"id":"do-lb-1","name":"DO Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"nyc3","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3 (New York)"},"provider":"digitalocean","characteristics":{"systemRole":"Distributes incoming HTTP traffic to backend Droplets with health checks","baseLatency":2,"maxThroughput":10000,"serviceFamily":"load-balancer","costMultiplier":1}},{"x":250,"y":220,"id":"do-droplet-1","name":"Web Droplet","type":"compute","status":"healthy","cpuUsage":34,"location":{"zoneKey":"nyc3","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3 (New York)"},"provider":"digitalocean","characteristics":{"size":"s-2vcpu-4gb","systemRole":"General-purpose Droplet (2 vCPUs, 4 GB RAM) — a simple virtual machine that runs your web app","baseLatency":3,"maxThroughput":9000,"serviceFamily":"droplet","costMultiplier":1}},{"x":250,"y":390,"id":"do-pg-1","name":"Managed PostgreSQL","type":"database","status":"healthy","cpuUsage":18,"location":{"zoneKey":"nyc3","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3 (New York)"},"provider":"digitalocean","characteristics":{"size":"db-s-2vcpu-4gb","systemRole":"Fully managed PostgreSQL database with automated backups and point-in-time recovery","baseLatency":5,"serviceFamily":"managed-db","costMultiplier":2,"maxConnections":250}}],"connections":[{"id":"conn-dolb1-droplet1","sourceId":"do-lb-1","targetId":"do-droplet-1"},{"id":"conn-droplet1-pg1","sourceId":"do-droplet-1","targetId":"do-pg-1"}],"duration":"~10 min","tags":["DigitalOcean","Beginner","Web App","Droplet","PostgreSQL"],"category":"scaling","defaultTrafficPatterns":[{"name":"Steady Baseline","type":"step","startTime":0,"parameters":{"startTraffic":400,"endTraffic":400},"isActive":true}]},{"id":"api-subdomain-split-spa-cdn","title":"API Subdomain Split — SPA on CDN, API on ALB","description":"The clean fix for the CloudFront 404-trap: stop routing API traffic through CloudFront at all. The React SPA is still on S3 behind CloudFront — its error-page rewrite rule (4xx → index.html 200) now only applies to SPA routes, which is exactly what you want. API traffic goes to api.example.com, a separate ALB subdomain with no CloudFront in front of it. Real 404s and 403s from the application travel directly from the ALB to the browser with their status codes intact. Explore this architecture to understand why the subdomain split is the most reliable solution, and note what the alternative — per-behavior CloudFront error pages — requires if you ever need to keep everything on a single distribution.","difficulty":"beginner","resources":[{"x":110,"y":60,"id":"split-cf-1","name":"CloudFront Distribution (SPA only)","type":"network","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Serves only the React SPA from S3 — no API traffic ever passes through this distribution. The 4xx → index.html error-page rule applies here, and it is safe: the only 4xx responses that can come from this distribution are SPA routes like /dashboard/missing-page that the browser should handle client-side. Because no API traffic flows through CloudFront, the rule can never intercept a real API 404 or 403.","baseLatency":5,"cacheHitRate":0.8,"maxThroughput":50000,"serviceFamily":"cloudfront","costMultiplier":1.5}},{"x":110,"y":220,"id":"split-s3-1","name":"S3 SPA Bucket","type":"storage","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Static SPA assets — index.html, JS bundles, CSS, images. With CloudFront in front serving 80% cache-hit rate, only new deploys or cache misses reach S3. The 4xx → index.html rewrite rule is correct and safe here because this bucket only serves the SPA; it has no knowledge of API routes and will never produce a false-positive 200 for an API error.","baseLatency":10,"maxThroughput":100000,"serviceFamily":"s3","costMultiplier":0.5}},{"x":430,"y":60,"id":"split-alb-1","name":"Application Load Balancer (API)","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Receives API traffic directly at api.example.com — no CloudFront in front of it. When the application returns HTTP 404 (resource not found) or HTTP 403 (forbidden), the ALB passes those status codes directly to the browser without any intermediate layer that could rewrite them. This is the entire fix: the path from ALB to browser is free of any error-page transformation rules.","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":310,"y":230,"id":"split-ec2-1","name":"App Server 1","type":"compute","status":"healthy","cpuUsage":28,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Handles API business logic — authentication, data fetching, mutations. Running at healthy load. When the application returns a 404 or 403, that status code travels through the ALB directly to the browser unchanged. Developers can now rely on HTTP semantics: 404 means not found, 403 means forbidden, and the browser's fetch() receives the correct status every time.","baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":1}},{"x":550,"y":230,"id":"split-ec2-2","name":"App Server 2","type":"compute","status":"healthy","cpuUsage":25,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Second app server in a separate AZ for high availability. Also healthy. The subdomain split architecture is straightforward to reason about: SPA users hit CloudFront, API clients hit the ALB subdomain, and the two paths share only the EC2 fleet and database — no shared CloudFront distribution whose error-page rules could leak between them.","baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":1}},{"x":430,"y":400,"id":"split-rds-1","name":"RDS PostgreSQL","type":"database","status":"healthy","cpuUsage":18,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Managed PostgreSQL database — healthy and well within connection limits. Shared by both SPA-originated and API-originated requests. Because the CDN absorbs 80% of SPA asset traffic at the edge, the database only sees dynamic query load from the application tier, keeping connection pressure low.","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-split-cf1-s3","sourceId":"split-cf-1","targetId":"split-s3-1"},{"id":"conn-split-alb-ec2-1","sourceId":"split-alb-1","targetId":"split-ec2-1"},{"id":"conn-split-alb-ec2-2","sourceId":"split-alb-1","targetId":"split-ec2-2"},{"id":"conn-split-ec2-1-rds","sourceId":"split-ec2-1","targetId":"split-rds-1"},{"id":"conn-split-ec2-2-rds","sourceId":"split-ec2-2","targetId":"split-rds-1"}],"duration":"~10 min","tags":["AWS","CloudFront","ALB","S3","SPA","API Design","Architecture"],"category":"networking","defaultTrafficPatterns":[{"name":"Steady API baseline — all resources healthy, HTTP semantics intact","type":"step","startTime":0,"parameters":{"startTraffic":150,"endTraffic":150},"isActive":true}]},{"id":"aks-multi-fault-cascade","title":"AKS Multi-Fault Cascade","description":"A two-pool AKS cluster — general (2–8 nodes) and api (1–6 nodes) — runs in East US Zone 1 behind an Azure App Gateway, backed by Azure SQL in Zone 2 and Azure Cache for Redis (85% hit rate). All resources start healthy under a steady wave baseline of ≈900–2100 RPS. Nothing self-triggers: simulate a multi-fault cascade yourself by injecting faults via the chaos API or the /failures endpoint. Note the naming difference — the chaos API uses `zone_outage` while /failures uses `az_outage` for the same event type. Injecting a Zone 1 outage simultaneously impacts the App Gateway, both AKS pools, and the Redis cache while the SQL database in Zone 2 survives — observe how zone failures dominate cascade grades and how connection pressure spikes as all cache reads fall through to the database.","difficulty":"advanced","resources":[{"x":250,"y":60,"id":"amf-appgw","name":"Azure App Gateway","type":"network","status":"healthy","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Layer 7 load balancing and SSL termination — routes traffic across both AKS node pools","baseLatency":2,"maxThroughput":9000,"serviceFamily":"appgw","costMultiplier":1}},{"x":100,"y":240,"id":"amf-aks-general","name":"AKS General Pool","type":"kubernetes","status":"healthy","cpuUsage":30,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","nodePools":[{"name":"general","maxNodes":8,"minNodes":2,"nodeCount":2}],"systemRole":"General-purpose node pool (2–8 nodes) hosting web and background workloads","autoscaling":true,"baseLatency":4,"maxThroughput":48000,"costMultiplier":1}},{"x":400,"y":240,"id":"amf-aks-api","name":"AKS API Pool","type":"kubernetes","status":"healthy","cpuUsage":30,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","nodePools":[{"name":"api","maxNodes":6,"minNodes":1,"nodeCount":1}],"systemRole":"API-tier node pool (1–6 nodes) handling ingress and service mesh traffic","autoscaling":true,"baseLatency":4,"maxThroughput":48000,"costMultiplier":1}},{"x":500,"y":400,"id":"amf-redis","name":"Azure Cache for Redis","type":"cache","status":"healthy","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_B2s","systemRole":"Managed Redis cache with 85% hit rate — a Zone 1 fault clears the cache, forcing all reads through to SQL","baseLatency":1,"cacheHitRate":0.85,"maxThroughput":50000,"costMultiplier":1}},{"x":250,"y":540,"id":"amf-sql","name":"Azure SQL","type":"database","status":"healthy","cpuUsage":25,"location":{"zoneKey":"eus-zone-2","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 2)"},"provider":"azure","characteristics":{"size":"S2 Standard","systemRole":"Managed relational database in Zone 2 — survives a Zone 1 outage but absorbs the full read load when Redis goes down","baseLatency":7,"costMultiplier":2,"maxConnections":300}}],"connections":[{"id":"conn-amf-gw-general","sourceId":"amf-appgw","targetId":"amf-aks-general"},{"id":"conn-amf-gw-api","sourceId":"amf-appgw","targetId":"amf-aks-api"},{"id":"conn-amf-general-redis","sourceId":"amf-aks-general","targetId":"amf-redis"},{"id":"conn-amf-api-redis","sourceId":"amf-aks-api","targetId":"amf-redis"},{"id":"conn-amf-general-sql","sourceId":"amf-aks-general","targetId":"amf-sql"},{"id":"conn-amf-api-sql","sourceId":"amf-aks-api","targetId":"amf-sql"}],"duration":"~30 min","tags":["Azure","AKS","Kubernetes","Chaos","Multi-Fault"],"category":"failure","defaultTrafficPatterns":[{"name":"Wave Baseline — correlated faults cascade at steady load","type":"wave","startTime":0,"parameters":{"baseline":1500,"amplitude":600,"period":40},"isActive":true},{"name":"Ramp to Autoscale — observe recovery behavior after fault injection","type":"ramp","startTime":60,"parameters":{"startTraffic":1500,"endTraffic":8000,"duration":40},"isActive":false}]},{"id":"gke-kubernetes-app","title":"Kubernetes App on GKE","description":"Run a containerized app on GKE Autopilot, accelerated by a Memorystore cache and buffered by Cloud Pub/Sub for async processing","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"gke-lb","name":"Cloud Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Global load balancing across regions","baseLatency":1,"maxThroughput":12000,"costMultiplier":1}},{"x":250,"y":200,"id":"gke-cluster","name":"GKE Autopilot","type":"kubernetes","status":"healthy","cpuUsage":32,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"n2-standard-2","systemRole":"Fully managed Kubernetes with fast pod autoscaling; GKE Autopilot mode with no node management","autoscaling":true,"baseLatency":3,"maxThroughput":200000,"costMultiplier":1}},{"x":80,"y":380,"id":"gke-redis","name":"Memorystore for Redis","type":"cache","status":"healthy","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"e2-medium","systemRole":"Fully managed Redis cache with automatic failover and monitoring","baseLatency":1,"cacheHitRate":0.8,"maxThroughput":50000,"costMultiplier":1}},{"x":420,"y":380,"id":"gke-pubsub","name":"Cloud Pub/Sub","type":"queue","status":"healthy","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"e2-medium","systemRole":"Globally distributed messaging service for real-time and batch data ingestion","baseLatency":5,"maxThroughput":100000,"costMultiplier":1}},{"x":250,"y":540,"id":"gke-cloudsql","name":"Cloud SQL","type":"database","status":"healthy","cpuUsage":18,"location":{"zoneKey":"us-central1-b","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-b (Iowa)"},"provider":"gcp","characteristics":{"size":"db-n1-standard-1","systemRole":"Standard managed SQL database","baseLatency":6,"costMultiplier":2,"maxConnections":400}}],"connections":[{"id":"conn-gkelb-cluster","sourceId":"gke-lb","targetId":"gke-cluster"},{"id":"conn-cluster-redis","sourceId":"gke-cluster","targetId":"gke-redis"},{"id":"conn-cluster-pubsub","sourceId":"gke-cluster","targetId":"gke-pubsub"},{"id":"conn-pubsub-cloudsql","sourceId":"gke-pubsub","targetId":"gke-cloudsql"},{"id":"conn-cluster-cloudsql","sourceId":"gke-cluster","targetId":"gke-cloudsql"}],"duration":"~18 min","tags":["GCP","Kubernetes","GKE","Redis","Pub/Sub","Cache","Queue"],"category":"scaling","defaultTrafficPatterns":[{"name":"Traffic Ramp — watch GKE autoscale","type":"ramp","startTime":0,"parameters":{"startTraffic":2500,"endTraffic":15000,"duration":30},"isActive":true},{"name":"Traffic Recovery — see GKE scale-in","type":"ramp","startTime":60,"parameters":{"startTraffic":15000,"endTraffic":2500,"duration":30},"isActive":false}]},{"id":"cloudfront-spa-api-404-trap","title":"CloudFront SPA + API Path Routing — The 404 Trap","description":"Your React SPA is hosted on S3, served through a single CloudFront distribution. You added a second CloudFront behavior to route /api/* requests to an ALB — clean and simple, until your API started returning 200 OK for every request that should have been a 404 or 403. The culprit is CloudFront's global SPA error-page rule: when you set 4xx → index.html with 200 status to make client-side routing work, that rule applies to every behavior in the distribution, not just the SPA. API 404s and 403s from the ALB are silently rewritten to 200 before the browser ever sees them. The crisis is already live — CloudFront and the ALB are flagged warning not because of load but because of this routing logic conflict. The fix is a subdomain split (a separate API subdomain that bypasses CloudFront entirely) or per-behavior CloudFront error pages, but the latter is not yet supported for ALB origins. Explore the resources to understand the exact failure path.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"cf404-cf-1","name":"CloudFront Distribution","type":"network","status":"warning","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Single CloudFront distribution with two behaviors: /* routes to S3 (SPA) and /api/* routes to the ALB. The global SPA error-page rule — which rewrites any 4xx response to index.html with HTTP 200 — applies across the entire distribution, not per behavior. When the ALB returns a real 404 or 403 for an API call, CloudFront intercepts it and sends the browser a 200 with the SPA HTML instead. The browser's fetch() sees a 200, parses HTML as JSON, and throws a parse error — the actual status code is gone.","baseLatency":5,"cacheHitRate":0.7,"maxThroughput":50000,"serviceFamily":"cloudfront","costMultiplier":1.5}},{"x":60,"y":220,"id":"cf404-s3-1","name":"S3 SPA Bucket","type":"storage","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Hosts the React SPA — index.html, JS bundles, CSS. The S3 bucket has a CloudFront error-page rule attached: any 4xx from S3 is rewritten to index.html with HTTP 200 so that client-side routes like /dashboard or /settings load correctly on hard refresh. This rule is correct and intentional for the SPA. The problem is that CloudFront applies this same rule to the ALB behavior too, because error-page customizations are distribution-wide, not per-behavior.","baseLatency":10,"maxThroughput":100000,"serviceFamily":"s3","costMultiplier":0.5}},{"x":430,"y":220,"id":"cf404-alb-1","name":"Application Load Balancer","type":"network","status":"warning","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Routes /api/* requests forwarded by CloudFront to the EC2 app servers. When the application returns HTTP 404 (resource not found) or HTTP 403 (permission denied), the ALB faithfully passes those status codes upstream to CloudFront — but CloudFront's 4xx error-page rule catches them before they reach the browser and replaces them with 200 OK + index.html. From the client's perspective, every API error looks like a successful HTML response.","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":310,"y":390,"id":"cf404-ec2-1","name":"App Server 1","type":"compute","status":"healthy","cpuUsage":28,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Handles API business logic — authentication, data fetching, mutations. Runs at healthy load; the routing bug does not increase CPU because requests still arrive and are processed normally. The 404/403 responses the app emits are correct — they are rewritten downstream by CloudFront before the client sees them.","baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":1}},{"x":550,"y":390,"id":"cf404-ec2-2","name":"App Server 2","type":"compute","status":"healthy","cpuUsage":25,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Second app server in a separate AZ for high availability. Also running at healthy load — the routing bug is invisible at the compute layer. Requests arrive, are processed, and correct HTTP status codes are returned. The problem exists entirely in the CloudFront layer above.","baseLatency":3,"maxThroughput":5000,"serviceFamily":"ec2","costMultiplier":1}},{"x":430,"y":560,"id":"cf404-rds-1","name":"RDS PostgreSQL","type":"database","status":"healthy","cpuUsage":18,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Managed PostgreSQL database — healthy and well within connection limits. The routing bug at the CloudFront layer has no effect on database load: queries still execute correctly and return accurate results. The application is sound; only the HTTP response codes are being silently replaced before they reach the browser.","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-cf404-cf1-s3","sourceId":"cf404-cf-1","targetId":"cf404-s3-1"},{"id":"conn-cf404-cf1-alb","sourceId":"cf404-cf-1","targetId":"cf404-alb-1"},{"id":"conn-cf404-alb-ec2-1","sourceId":"cf404-alb-1","targetId":"cf404-ec2-1"},{"id":"conn-cf404-alb-ec2-2","sourceId":"cf404-alb-1","targetId":"cf404-ec2-2"},{"id":"conn-cf404-ec2-1-rds","sourceId":"cf404-ec2-1","targetId":"cf404-rds-1"},{"id":"conn-cf404-ec2-2-rds","sourceId":"cf404-ec2-2","targetId":"cf404-rds-1"}],"duration":"~10 min","tags":["AWS","CloudFront","ALB","S3","SPA","API Design","Architecture"],"category":"reliability","defaultTrafficPatterns":[{"name":"Steady API load — routing bug is live, not a load problem","type":"step","startTime":0,"parameters":{"startTraffic":500,"endTraffic":500},"isActive":true}]},{"id":"redis-cache-crash-recovery","title":"Redis Cache Crash & Recovery","description":"Observe what happens when ElastiCache for Redis crashes mid-traffic: the cache eviction storm forces every request through to the database, connection pools saturate, and latency spikes until the cache warms back up. Explore resilience patterns like circuit breakers and staggered cache warm-up.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"rcc-alb","name":"Application Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Distributes incoming traffic across multiple targets","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":100,"y":200,"id":"rcc-api-1","name":"API Server 1","type":"compute","status":"healthy","cpuUsage":32,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Balanced compute for production workloads","baseLatency":3,"maxThroughput":40000,"serviceFamily":"ec2","costMultiplier":1}},{"x":400,"y":200,"id":"rcc-api-2","name":"API Server 2","type":"compute","status":"healthy","cpuUsage":32,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Balanced compute for production workloads","baseLatency":3,"maxThroughput":40000,"serviceFamily":"ec2","costMultiplier":1}},{"x":250,"y":360,"id":"rcc-redis","name":"ElastiCache for Redis","type":"cache","status":"critical","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","recoveryPolicy":{"warningSteps":999,"criticalSteps":999,"warningCpuThreshold":70,"criticalCpuThreshold":80},"characteristics":{"size":"cache.r6g.large","systemRole":"Crashed Redis cache — all reads fall through to the database, triggering a thundering-herd event on the connection pool","baseLatency":1,"cacheHitRate":0,"maxThroughput":50000,"costMultiplier":1}},{"x":250,"y":520,"id":"rcc-rds","name":"RDS PostgreSQL","type":"database","status":"warning","cpuUsage":88,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Primary database now absorbing 100% of read traffic — connection pool nearing saturation","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-rccalb-api1","sourceId":"rcc-alb","targetId":"rcc-api-1"},{"id":"conn-rccalb-api2","sourceId":"rcc-alb","targetId":"rcc-api-2"},{"id":"conn-rccapi1-redis","sourceId":"rcc-api-1","targetId":"rcc-redis"},{"id":"conn-rccapi2-redis","sourceId":"rcc-api-2","targetId":"rcc-redis"},{"id":"conn-rccapi1-rds","sourceId":"rcc-api-1","targetId":"rcc-rds"},{"id":"conn-rccapi2-rds","sourceId":"rcc-api-2","targetId":"rcc-rds"},{"id":"conn-rccredis-rds","sourceId":"rcc-redis","targetId":"rcc-rds"}],"duration":"~15 min","tags":["AWS","Redis","Cache","Failure","Recovery","Cache Stampede","ElastiCache"],"category":"failure","defaultTrafficPatterns":[{"name":"Moderate Baseline Traffic","type":"step","startTime":0,"parameters":{"startTraffic":1000,"endTraffic":1000},"isActive":true}]},{"id":"sqs-queue-backlog-saturation","title":"SQS Queue Backlog Saturation","description":"Simulate a burst of inbound events that overwhelms the worker fleet: the SQS queue depth climbs into the thousands, workers fall behind, and end-to-end processing latency skyrockets. Watch how adding consumers and applying dead-letter queues restore stability.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"qbs-alb","name":"Application Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Distributes incoming traffic across multiple targets","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":250,"y":200,"id":"qbs-producer","name":"Event Producer","type":"compute","status":"healthy","cpuUsage":70,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Produces events at a rate that currently exceeds consumer throughput","baseLatency":2,"maxThroughput":20000,"serviceFamily":"ec2","costMultiplier":1}},{"x":250,"y":360,"id":"qbs-sqs","name":"Amazon SQS (saturated)","type":"queue","status":"warning","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"standard","systemRole":"Queue backlog growing — messages are aging past their visibility timeout and reappearing, amplifying processing pressure","baseLatency":120,"maxThroughput":100000,"costMultiplier":1}},{"x":100,"y":500,"id":"qbs-worker-1","name":"Worker 1","type":"compute","status":"critical","cpuUsage":95,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Struggling to drain the queue — CPU pinned at max and memory pressure rising","baseLatency":3,"maxThroughput":2000,"serviceFamily":"ec2","costMultiplier":1}},{"x":400,"y":500,"id":"qbs-worker-2","name":"Worker 2","type":"compute","status":"critical","cpuUsage":93,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Second worker also saturated — queue depth continues growing despite both workers running at full capacity","baseLatency":3,"maxThroughput":2000,"serviceFamily":"ec2","costMultiplier":1}},{"x":250,"y":650,"id":"qbs-rds","name":"RDS MySQL","type":"database","status":"healthy","cpuUsage":45,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Downstream database — currently healthy but at risk if workers begin retrying failed messages in bulk","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-qbsalb-producer","sourceId":"qbs-alb","targetId":"qbs-producer"},{"id":"conn-producer-sqs","sourceId":"qbs-producer","targetId":"qbs-sqs"},{"id":"conn-sqs-worker1","sourceId":"qbs-sqs","targetId":"qbs-worker-1"},{"id":"conn-sqs-worker2","sourceId":"qbs-sqs","targetId":"qbs-worker-2"},{"id":"conn-worker1-rds","sourceId":"qbs-worker-1","targetId":"qbs-rds"},{"id":"conn-worker2-rds","sourceId":"qbs-worker-2","targetId":"qbs-rds"}],"duration":"~15 min","tags":["AWS","SQS","Queue","Failure","Recovery","Backlog","Workers"],"category":"failure","defaultTrafficPatterns":[{"name":"Steady Producer Load","type":"step","startTime":0,"parameters":{"startTraffic":1200,"endTraffic":1200},"isActive":true},{"name":"Event Burst — overwhelm workers","type":"ramp","startTime":0,"parameters":{"startTraffic":1200,"endTraffic":5000,"duration":20},"isActive":false}]},{"id":"digitalocean-kubernetes","title":"DigitalOcean Kubernetes (DOKS)","description":"A container-native stack on DigitalOcean: a Load Balancer fronts a DOKS cluster that runs your app pods. A Managed Redis cache absorbs read traffic, and a Managed Kafka queue decouples background jobs from the web tier — all with zero infrastructure to patch.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"doks-lb-1","name":"DO Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"nyc3","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3 (New York)"},"provider":"digitalocean","characteristics":{"systemRole":"Exposes the DOKS cluster to the internet and routes traffic to healthy pods","baseLatency":2,"maxThroughput":10000,"serviceFamily":"load-balancer","costMultiplier":1}},{"x":250,"y":200,"id":"doks-cluster-1","name":"DOKS Cluster","type":"kubernetes","status":"healthy","cpuUsage":32,"location":{"zoneKey":"nyc3","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3 (New York)"},"provider":"digitalocean","characteristics":{"size":"doks","systemRole":"Managed Kubernetes cluster — DigitalOcean handles the control plane, you manage your app pods and worker nodes","autoscaling":true,"baseLatency":4,"maxThroughput":50000,"serviceFamily":"doks","costMultiplier":1}},{"x":80,"y":380,"id":"doks-redis-1","name":"Managed Redis","type":"cache","status":"healthy","location":{"zoneKey":"nyc3","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3 (New York)"},"provider":"digitalocean","characteristics":{"size":"db-s-1vcpu-1gb","systemRole":"Managed Redis cache that reduces database load by serving frequent reads from memory","baseLatency":1,"cacheHitRate":0.8,"maxThroughput":40000,"serviceFamily":"doManagedRedis","costMultiplier":1}},{"x":420,"y":380,"id":"doks-kafka-1","name":"Managed Kafka","type":"queue","status":"healthy","location":{"zoneKey":"nyc3","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3 (New York)"},"provider":"digitalocean","characteristics":{"size":"kafka-s","systemRole":"Managed Kafka topic that decouples background processing from the web request path","baseLatency":5,"maxThroughput":50000,"serviceFamily":"doManagedKafka","costMultiplier":1}}],"connections":[{"id":"conn-dokslb1-cluster1","sourceId":"doks-lb-1","targetId":"doks-cluster-1"},{"id":"conn-cluster1-redis1","sourceId":"doks-cluster-1","targetId":"doks-redis-1"},{"id":"conn-cluster1-kafka1","sourceId":"doks-cluster-1","targetId":"doks-kafka-1"}],"duration":"~18 min","tags":["DigitalOcean","Kubernetes","DOKS","Cache","Queue","Redis","Kafka"],"category":"scaling","defaultTrafficPatterns":[{"name":"Traffic Ramp — watch DOKS autoscale","type":"ramp","startTime":0,"parameters":{"startTraffic":2000,"endTraffic":15000,"duration":30},"isActive":true},{"name":"Traffic Recovery — see DOKS scale-in","type":"ramp","startTime":60,"parameters":{"startTraffic":15000,"endTraffic":2000,"duration":30},"isActive":false}]},{"id":"azure-event-driven","title":"Event-Driven Azure Microservices","description":"Explore an AKS-hosted microservices app that offloads read traffic to Azure Cache for Redis and fans out async work through Azure Service Bus","difficulty":"advanced","resources":[{"x":250,"y":60,"id":"az-lb","name":"Azure Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Layer 4 load balancing for Azure resources","baseLatency":2,"maxThroughput":9000,"costMultiplier":1}},{"x":250,"y":200,"id":"az-aks","name":"Azure AKS","type":"kubernetes","status":"healthy","cpuUsage":33,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","systemRole":"Managed Kubernetes with free control plane; scales compute nodes based on pod demand","autoscaling":true,"baseLatency":4,"maxThroughput":150000,"costMultiplier":1}},{"x":80,"y":380,"id":"az-redis","name":"Azure Cache for Redis","type":"cache","status":"healthy","location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_B2s","systemRole":"Managed Redis cache that reduces latency and database connection pressure","baseLatency":1,"cacheHitRate":0.8,"maxThroughput":50000,"costMultiplier":1}},{"x":420,"y":380,"id":"az-servicebus","name":"Azure Service Bus","type":"queue","status":"healthy","location":{"zoneKey":"eus-zone-2","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 2)"},"provider":"azure","characteristics":{"size":"Standard_B2s","systemRole":"Fully managed enterprise messaging for decoupling application components","baseLatency":5,"maxThroughput":100000,"costMultiplier":1}},{"x":250,"y":540,"id":"az-sql","name":"Azure SQL S3","type":"database","status":"healthy","cpuUsage":28,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"S2 Standard","systemRole":"Standard tier managed SQL database","baseLatency":7,"costMultiplier":2,"maxConnections":300}}],"connections":[{"id":"conn-azlb-aks","sourceId":"az-lb","targetId":"az-aks"},{"id":"conn-aks-redis","sourceId":"az-aks","targetId":"az-redis"},{"id":"conn-aks-servicebus","sourceId":"az-aks","targetId":"az-servicebus"},{"id":"conn-servicebus-sql","sourceId":"az-servicebus","targetId":"az-sql"},{"id":"conn-aks-sql","sourceId":"az-aks","targetId":"az-sql"}],"duration":"~20 min","tags":["Azure","Kubernetes","AKS","Redis","Service Bus","Cache","Queue"],"category":"scaling","defaultTrafficPatterns":[{"name":"Traffic Ramp — watch AKS autoscale","type":"ramp","startTime":0,"parameters":{"startTraffic":2500,"endTraffic":15000,"duration":30},"isActive":true},{"name":"Traffic Recovery — see AKS scale-in","type":"ramp","startTime":60,"parameters":{"startTraffic":15000,"endTraffic":2500,"duration":30},"isActive":false}]},{"id":"serverless-api","title":"Serverless API (Lambda + DynamoDB)","description":"A fully serverless API on AWS: an Application Load Balancer triggers Lambda functions on every request — no servers to manage and you only pay per invocation. DynamoDB handles storage at any scale. Watch for cold-start latency spikes when Lambda containers are not yet warm after a period of low traffic.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"sls-alb-1","name":"Application Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Routes incoming API requests to Lambda function targets","baseLatency":2,"maxThroughput":10000,"serviceFamily":"alb","costMultiplier":1}},{"x":250,"y":220,"id":"sls-lambda-1","name":"Lambda Function","type":"compute","status":"healthy","cpuUsage":20,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"1GB","systemRole":"Serverless function that runs on demand — scales to zero when idle but suffers cold-start delays (~300 ms) when a new container must be initialized","baseLatency":8,"maxThroughput":5000,"serviceFamily":"lambda","costMultiplier":0.3,"coldStartLatency":300}},{"x":250,"y":390,"id":"sls-dynamo-1","name":"DynamoDB","type":"database","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Serverless NoSQL database billed per read/write request unit — pairs naturally with Lambda as both scale to zero and require no capacity planning","baseLatency":3,"maxThroughput":100000,"serviceFamily":"dynamodb","costMultiplier":1.2}}],"connections":[{"id":"conn-slsalb1-lambda1","sourceId":"sls-alb-1","targetId":"sls-lambda-1"},{"id":"conn-lambda1-dynamo1","sourceId":"sls-lambda-1","targetId":"sls-dynamo-1"}],"duration":"~15 min","tags":["AWS","Lambda","Serverless","DynamoDB","Cold Start"],"category":"scaling","defaultTrafficPatterns":[{"name":"Quiet period — Lambda idles","type":"step","startTime":0,"parameters":{"startTraffic":100,"endTraffic":100},"isActive":true},{"name":"Traffic burst — cold starts fire","type":"ramp","startTime":20,"parameters":{"startTraffic":100,"endTraffic":2000,"duration":15},"isActive":true}]},{"id":"cloudfront-edge-outage","title":"CloudFront Edge Outage","description":"On July 16 2026, AWS CloudFront began returning errors globally — every edge location stopped serving cached responses. This scenario reproduces the origin overload that followed: with a 75% cache-hit rate suddenly gone, ALB and EC2 web servers absorb 4× their normal request volume and climb into the warning zone. RDS connection pressure rises as every user request now requires a full database round-trip. The crisis is already live — explore what protections (rate limiting, static failover pages, graceful degradation) would have kept your origin standing.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"cfo-cf-1","name":"CloudFront CDN","type":"network","status":"critical","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"CloudFront edge is actively failing — cache-hit rate is 0 during the outage (normally 75%), so every request bypasses the edge and hits origin directly at 4× the normal origin load","baseLatency":5,"maxThroughput":50000,"serviceFamily":"cloudfront","costMultiplier":1.5}},{"x":250,"y":200,"id":"cfo-alb-1","name":"Application Load Balancer","type":"network","status":"warning","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Receiving 4× its normal request volume directly — sized for ~25% cache-miss bypass traffic but now absorbing 100% of origin requests; maxThroughput reflects origin-only sizing, not CDN-assisted capacity","baseLatency":2,"maxThroughput":2200,"serviceFamily":"alb","costMultiplier":1}},{"x":130,"y":360,"id":"cfo-ec2-1","name":"Web Server 1","type":"compute","status":"warning","cpuUsage":65,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Origin server sized for ~25% bypass traffic (~500 RPS normal) — now handling 100% of requests and running at ~65% CPU (4× normal load); maxThroughput of 1500 reflects bypass-only sizing","baseLatency":3,"maxThroughput":1500,"serviceFamily":"ec2","costMultiplier":1}},{"x":370,"y":360,"id":"cfo-ec2-2","name":"Web Server 2","type":"compute","status":"warning","cpuUsage":65,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","systemRole":"Second origin server in a separate AZ — also running at 65% CPU as it bears its share of the full traffic volume; maxThroughput of 1500 reflects bypass-only sizing","baseLatency":3,"maxThroughput":1500,"serviceFamily":"ec2","costMultiplier":1}},{"x":250,"y":520,"id":"cfo-rds-1","name":"RDS MySQL","type":"database","status":"warning","cpuUsage":48,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","systemRole":"Connection pool under elevated pressure — cache bypass means every bypassed request now hits the database, pushing connection utilisation to ~80%","baseLatency":5,"serviceFamily":"rds","costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-cfocf1-alb1","sourceId":"cfo-cf-1","targetId":"cfo-alb-1"},{"id":"conn-cfoalb1-ec2-1","sourceId":"cfo-alb-1","targetId":"cfo-ec2-1"},{"id":"conn-cfoalb1-ec2-2","sourceId":"cfo-alb-1","targetId":"cfo-ec2-2"},{"id":"conn-cfoec2-1-rds1","sourceId":"cfo-ec2-1","targetId":"cfo-rds-1"},{"id":"conn-cfoec2-2-rds1","sourceId":"cfo-ec2-2","targetId":"cfo-rds-1"}],"duration":"~15 min","tags":["AWS","CloudFront","CDN","Outage","Incident Response","EC2","RDS"],"category":"reliability","defaultTrafficPatterns":[{"name":"Steady bypass load — origin absorbing 100% of traffic","type":"step","startTime":0,"parameters":{"startTraffic":2000,"endTraffic":2000},"isActive":true},{"name":"Traffic Drain — reduce load during incident response","type":"ramp","startTime":60,"parameters":{"startTraffic":2000,"endTraffic":500,"duration":20},"isActive":false}],"defaultFailureInjections":[{"name":"CloudFront Edge Failure — Edge Layer Offline","type":"az_outage","targetZone":"us-east-1-regional","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}}],"realWorldIncident":{"date":"2026-07-16","provider":"AWS","summary":"On July 16 2026, a configuration error disabled cache serving across all CloudFront edge locations for ~3h33m. Origin servers absorbed 4× their normal request volume — VPC Origins customers and custom-origin sites including Canvas/Blackboard, HuggingFace Spaces, and Tailscale experienced sustained overload for the duration of the outage. This scenario lets you experience the origin overload in slow motion — inject traffic, watch CPU and RDS connection pressure climb, and test whether rate limiting, static fallback pages, or a larger origin fleet would have kept you standing through the real outage.","references":[{"label":"The Register · CloudFront outage report","url":"https://www.theregister.com/off-prem/2026/07/16/aws-cloudfront-outage-serves-errors-instead-of-websites/5272421"},{"label":"AWS Health Dashboard","url":"https://health.aws.amazon.com/health/status"}]}},{"id":"zombie-infra-oci","title":"Zombie Infrastructure — OCI","description":"An idle OCI compartment still billing ~$90/month: reserved public IPs, block volume backups, a 10 Mbps load balancer minimum, and boot volumes under stopped compute and database instances. Notably cheaper than the other clouds — OCI's NAT Gateway and Service Gateway are free. Remove idle resources to watch the residual bill drop.","difficulty":"beginner","resources":[{"x":80,"y":620,"id":"zoci-ip-1","name":"Reserved Public IP (unattached)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 AD-1"},"provider":"oci","characteristics":{"systemRole":"Reserved public IP charged while not attached (OCI NAT Gateway is free)","billingState":"detached","serviceFamily":"oci-reserved-ip","costMultiplier":1}},{"x":280,"y":620,"id":"zoci-ip-2","name":"Reserved Public IP (unattached)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 AD-1"},"provider":"oci","characteristics":{"systemRole":"Reserved public IP charged while not attached (OCI NAT Gateway is free)","billingState":"detached","serviceFamily":"oci-reserved-ip","costMultiplier":1}},{"x":480,"y":620,"id":"zoci-ip-3","name":"Reserved Public IP (unattached)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-ashburn-1-ad-2","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 AD-2"},"provider":"oci","characteristics":{"systemRole":"Reserved public IP charged while not attached (OCI NAT Gateway is free)","billingState":"detached","serviceFamily":"oci-reserved-ip","costMultiplier":1}},{"x":400,"y":80,"id":"zoci-lb-1","name":"Load Balancer 10 Mbps (idle)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-ashburn-1-regional","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 (Regional)"},"provider":"oci","characteristics":{"systemRole":"Idle load balancer — 10 Mbps shape minimum hourly fee","baseLatency":2,"billingState":"idle","maxThroughput":10000,"serviceFamily":"oci-lb-idle","costMultiplier":1}},{"x":250,"y":260,"id":"zoci-vm-1","name":"app-instance-01 (stopped)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 AD-1"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (2 OCPU)","capacityGB":100,"systemRole":"Stopped compute instance — its 100 GB boot volume keeps billing","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"compute","costMultiplier":1}},{"x":550,"y":260,"id":"zoci-vm-2","name":"app-instance-02 (stopped)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-ashburn-1-ad-2","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 AD-2"},"provider":"oci","characteristics":{"size":"VM.Standard.E4.Flex (2 OCPU)","capacityGB":100,"systemRole":"Stopped compute instance — its 100 GB boot volume keeps billing","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"compute","costMultiplier":1}},{"x":250,"y":440,"id":"zoci-db-1","name":"MySQL DB System (stopped)","type":"database","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-ashburn-1-ad-1","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 AD-1"},"provider":"oci","characteristics":{"size":"MySQL.VM.Standard.E4.Flex (2 OCPU / 16 GB)","capacityGB":500,"systemRole":"Stopped DB system — 500 GB allocated block storage keeps billing","baseLatency":5,"billingState":"stopped","serviceFamily":"mysql-heatwave","costMultiplier":2,"maxConnections":500}},{"x":550,"y":440,"id":"zoci-bak-1","name":"Block Volume Backups (2 TB)","type":"storage","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-ashburn-1-regional","regionKey":"us-ashburn-1","localityType":"zone","providerLabel":"us-ashburn-1 (Regional)"},"provider":"oci","characteristics":{"capacityGB":2000,"systemRole":"Block volume backups — $0.025/GB-mo billed continuously","billingState":"idle","serviceFamily":"block-backup","costMultiplier":1}}],"connections":[{"id":"zoci-c-lb-vm-1","sourceId":"zoci-lb-1","targetId":"zoci-vm-1"},{"id":"zoci-c-lb-vm-2","sourceId":"zoci-lb-1","targetId":"zoci-vm-2"},{"id":"zoci-c-vm-1-db","sourceId":"zoci-vm-1","targetId":"zoci-db-1"},{"id":"zoci-c-vm-2-db","sourceId":"zoci-vm-2","targetId":"zoci-db-1"},{"id":"zoci-c-vm-1-bak","sourceId":"zoci-vm-1","targetId":"zoci-bak-1"},{"id":"zoci-c-vm-2-bak","sourceId":"zoci-vm-2","targetId":"zoci-bak-1"}],"duration":"~10 min","tags":["OCI","Cost","Idle Resources","Zombie Infrastructure","FinOps"],"category":"cost","defaultTrafficPatterns":[{"name":"Zero traffic — the bill keeps running anyway","type":"step","startTime":0,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true}],"realWorldIncident":{"date":"2026-08-04","provider":"OCI","summary":"The OCI version of the zombie bill an August 2026 r/aws thread ('It's always the networking costs') made famous. An idle OCI compartment: instances and the DB system are stopped, yet ~$90/month keeps billing — unattached reserved public IPs ($0.0021/hr each), block volume backups per GB-month, the load balancer's 10 Mbps shape minimum ($0.022/hr), and boot/block volumes under stopped instances. Note what's absent: OCI's NAT Gateway and Service Gateway are free, so the same architecture leaves a smaller zombie footprint than on AWS or Azure. Remove idle resources to watch the bill fall.","references":[{"label":"r/aws · 'It's always the networking costs' (Aug 2026)","url":"https://www.reddit.com/r/aws/comments/1vfav79/its_always_the_networking_costs/"},{"label":"OCI networking pricing (reserved IPs, gateways)","url":"https://www.oracle.com/cloud/networking/pricing/"},{"label":"OCI block volume pricing (backups)","url":"https://www.oracle.com/cloud/storage/block-volumes/pricing/"}]}},{"id":"gcp-multi-region-failover","title":"GCP Multi-Region Failover — Cloud DNS & Global Load Balancer","description":"On June 2, 2019, a misconfigured traffic-engineering update caused extreme congestion across Google's eastern US backbone — and it took down both Cloud Load Balancing AND Cloud DNS for ~4h. Health-check probes could not reach any backends, DNS propagation stalled, and customers behind a single regional ALB were completely unreachable for the full outage window. Multi-region deployments using the Global HTTP Load Balancer fared better for a specific reason: because the Global LB uses anycast (not DNS), its edge PoPs began finding alternate backbone paths as congestion partially cleared — the same VIP address started routing again through less-congested paths before DNS could recover. This scenario opens mid-incident: the Global LB and Cloud DNS are offline, us-central1 is completely unreachable, and us-east1 is only getting a trickle of traffic through partially-recovered anycast paths. Your job: understand why the Global LB's anycast architecture gave multi-region setups an earlier recovery window even when the LB itself was degraded — and what additional protections (Cloud Armor failover policy, multi-CDN, pre-warmed external endpoints) would give you a reliable cold-start path when the entire routing layer goes down.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"gmrf-cloud-dns","name":"Cloud DNS — Health-Check Routing","type":"network","status":"critical","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","recoveryPolicy":{"warningSteps":999,"criticalSteps":999,"warningCpuThreshold":70,"criticalCpuThreshold":80},"characteristics":{"systemRole":"Cloud DNS is offline — backbone congestion has severed the control-plane paths that health-check routing policies use to detect backend state and update resource records; DNS propagation has stalled, so clients that previously resolved successfully are stuck with cached (now-stale) IPs; this is the failure mode that showed why DNS-based routing alone is insufficient when the DNS service itself is degraded","baseLatency":1,"maxThroughput":50000,"serviceFamily":"cloud-dns","costMultiplier":1}},{"x":250,"y":180,"id":"gmrf-glb","name":"Global HTTP Load Balancer","type":"network","status":"critical","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","recoveryPolicy":{"warningSteps":999,"criticalSteps":999,"warningCpuThreshold":70,"criticalCpuThreshold":80},"characteristics":{"systemRole":"Global HTTP LB is severely degraded — backbone congestion disrupted the control plane that programs health checks and backend selection across edge PoPs; however, because anycast routes traffic to the nearest healthy edge PoP rather than resolving a new DNS name, some PoPs began finding alternate congestion-free backbone paths to the us-east1 backend before DNS could recover; this partial anycast recovery is why multi-region Global LB setups fared better than pure DNS-based routing during the 2019 event","baseLatency":2,"maxThroughput":30000,"serviceFamily":"global-lb","costMultiplier":1}},{"x":100,"y":320,"id":"gmrf-gke-central","name":"GKE Cluster — us-central1","type":"kubernetes","status":"critical","cpuUsage":88,"location":{"zoneKey":"us-central1-b","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-b (Iowa)"},"provider":"gcp","characteristics":{"size":"n2-standard-4","systemRole":"Primary GKE cluster in us-central1 — zone failure has pushed CPU to 88% and pod restarts are looping; the Global LB health check is marking this backend unhealthy, draining it from the serving path; autoscaling is trying to add nodes but the zone's capacity is constrained","autoscaling":true,"baseLatency":3,"maxThroughput":6000,"costMultiplier":2}},{"x":400,"y":320,"id":"gmrf-gke-east","name":"GKE Cluster — us-east1","type":"kubernetes","status":"warning","cpuUsage":62,"location":{"zoneKey":"us-east1-b","regionKey":"us-east1","localityType":"zone","providerLabel":"us-east1-b (South Carolina)"},"provider":"gcp","characteristics":{"size":"n2-standard-4","systemRole":"East GKE cluster — partially receiving traffic as Global LB anycast PoPs find alternate backbone paths; CPU is elevated because it is absorbing whatever traffic is making it through, but request volume is well below normal since the routing layer is still impaired; this cluster represents the multi-region advantage: it started recovering before DNS-only setups because anycast routing does not require DNS propagation to switch backends","autoscaling":true,"baseLatency":4,"maxThroughput":6000,"costMultiplier":2}},{"x":100,"y":480,"id":"gmrf-cloudsql-primary","name":"Cloud SQL Primary — us-central1","type":"database","status":"warning","cpuUsage":62,"location":{"zoneKey":"us-central1-c","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-c (Iowa)"},"provider":"gcp","characteristics":{"size":"db-n1-standard-2","systemRole":"Cloud SQL primary in us-central1-c — accepting writes; connection pool is under pressure from central GKE pod retries; the cross-region replica in us-east1 is healthy and can be promoted to primary if this instance degrades further","baseLatency":6,"costMultiplier":2,"maxConnections":500}},{"x":400,"y":480,"id":"gmrf-cloudsql-replica","name":"Cloud SQL Read Replica — us-east1","type":"database","status":"healthy","cpuUsage":18,"location":{"zoneKey":"us-east1-c","regionKey":"us-east1","localityType":"zone","providerLabel":"us-east1-c (South Carolina)"},"provider":"gcp","characteristics":{"size":"db-n1-standard-2","systemRole":"Cloud SQL cross-region read replica in us-east1-c — serving east GKE reads; replication lag is the key metric: if it stays under 5 s you have a clean promotion path; can be promoted to a standalone writable instance in under 60 s if the central primary fails completely","baseLatency":7,"costMultiplier":2,"maxConnections":500}}],"connections":[{"id":"conn-gmrf-dns-glb","sourceId":"gmrf-cloud-dns","targetId":"gmrf-glb"},{"id":"conn-gmrf-glb-gke-central","sourceId":"gmrf-glb","targetId":"gmrf-gke-central"},{"id":"conn-gmrf-glb-gke-east","sourceId":"gmrf-glb","targetId":"gmrf-gke-east"},{"id":"conn-gmrf-gke-central-sql","sourceId":"gmrf-gke-central","targetId":"gmrf-cloudsql-primary"},{"id":"conn-gmrf-gke-east-sql","sourceId":"gmrf-gke-east","targetId":"gmrf-cloudsql-replica"},{"id":"conn-gmrf-sql-replication","sourceId":"gmrf-cloudsql-primary","targetId":"gmrf-cloudsql-replica"}],"duration":"~15 min","tags":["GCP","Cloud DNS","Global Load Balancer","Multi-Region","Failover","GKE","Resilience"],"category":"reliability","defaultTrafficPatterns":[{"name":"Mid-failover baseline — Global LB draining central, loading east","type":"step","startTime":0,"parameters":{"startTraffic":500,"endTraffic":500},"isActive":true},{"name":"East Scale-Out — simulate full central drain to us-east1","type":"ramp","startTime":60,"parameters":{"startTraffic":500,"endTraffic":3000,"duration":30},"isActive":false}],"defaultFailureInjections":[{"name":"Backbone Congestion — Cloud DNS and Global LB offline (us-central1-a)","type":"az_outage","targetZone":"usc1-zone-a","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}}],"realWorldIncident":{"date":"2019-06-02","provider":"GCP","summary":"On June 2, 2019, a misconfigured traffic-engineering update caused extreme congestion in Google's eastern US backbone, taking Cloud Load Balancing and Cloud DNS offline for ~4h. Thousands of GCP customers lost inbound traffic as health-check probes timed out and DNS propagation stalled — workloads behind a single regional ALB were unreachable for the full window, while services using Google's Global Load Balancer with multi-region back-end pools began recovering once congestion cleared in alternative paths. This scenario puts you at the controls of a similar Cloud DNS and Global LB failover so you can watch exactly how health-check intervals and DNS TTL shape your recovery window.","references":[{"label":"GCP Networking Incident #19009","url":"https://status.cloud.google.com/incident/cloud-networking/19009"},{"label":"Google Cloud Blog · post-mortem","url":"https://cloud.google.com/blog/topics/inside-google-cloud/an-update-on-sundays-service-disruption"}]}},{"id":"azure-multi-region-failover","title":"Azure Front Door HA — Active Geo-Replication Through a Front Door Outage","description":"On October 9, 2025, a crash loop in Azure Front Door's data plane took its global edge network offline for ~8h10m (tracking ID: QNBQ-5W8). Unlike a backend failure — where Front Door detects unhealthy health probes and routes around the problem — this crash loop made Front Door itself the failure. Both the East US and West US App Service origins were completely healthy and ready to serve traffic, but the global proxy layer connecting users to them was gone. This is the failure mode that Front Door does not protect you from, because Front Door cannot route around its own crash. This scenario opens at peak impact: Front Door is offline globally, both App Service backends have near-zero incoming traffic (not because they're broken, but because the proxy layer isn't delivering requests), and Azure SQL is under minimal load. Your two learning goals: (1) Understand the difference between 'backend failure that Front Door routes around' vs 'Front Door itself is the single point of failure,' and (2) Learn why Traffic Manager pointing directly at App Service origins — bypassing Front Door entirely — is the resilience layer you need when the proxy itself goes down. Traffic Manager is DNS-based (TTL-bounded propagation) but is a different service from Front Door and would have remained available during the QNBQ-5W8 event.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"amrf-front-door","name":"Azure Front Door — Global Anycast","type":"network","status":"critical","cpuUsage":0,"location":{"zoneKey":"eus-zone-fd","regionKey":"eus","localityType":"az","providerLabel":"Global Edge (Front Door)"},"provider":"azure","characteristics":{"systemRole":"Azure Front Door is offline — a data-plane crash loop has taken all global edge PoPs down; health probes are not running, traffic is not being proxied to either App Service backend, and the anycast VIP is returning connection errors; this is the failure mode that Front Door cannot self-heal: the routing layer must be restarted by Microsoft, not rerouted by health checks; Traffic Manager pointing directly at App Service origins (with Front Door as a secondary option) would have provided a fallback that survived this event","baseLatency":2,"maxThroughput":50000,"costMultiplier":1}},{"x":80,"y":220,"id":"amrf-app-east","name":"App Service — East US (Healthy, Unreachable)","type":"compute","status":"warning","cpuUsage":18,"location":{"zoneKey":"eus-zone-1","regionKey":"eus","localityType":"az","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","systemRole":"App Service plan in East US — healthy, passing internal health checks, but receiving almost no traffic because Front Door is offline and not proxying requests to this origin; CPU is low not because the service is broken but because users cannot reach it; this illustrates the frustrating reality of the real October 2025 event: backends were fine, only the proxy layer was down","baseLatency":3,"maxThroughput":3000,"serviceFamily":"vm","costMultiplier":1}},{"x":420,"y":220,"id":"amrf-app-west","name":"App Service — West US (Healthy, Unreachable)","type":"compute","status":"warning","cpuUsage":18,"location":{"zoneKey":"wus-zone-1","regionKey":"wus","localityType":"az","providerLabel":"West US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","systemRole":"App Service plan in West US — also healthy and ready to serve, also receiving near-zero traffic for the same reason; having a second healthy region behind Front Door provided zero additional resilience during the October 2025 event because the failure was in the layer in front of both regions; Traffic Manager configured as a fallback (lower-priority DNS entry pointing here directly) would have started routing traffic to this origin once Front Door's endpoint failed its Traffic Manager health check","baseLatency":3,"maxThroughput":3000,"serviceFamily":"vm","costMultiplier":1}},{"x":80,"y":400,"id":"amrf-sql-primary","name":"Azure SQL — Primary (East US)","type":"database","status":"warning","cpuUsage":25,"location":{"zoneKey":"eus-zone-2","regionKey":"eus","localityType":"az","providerLabel":"East US (Zone 2)"},"provider":"azure","characteristics":{"size":"GP_Gen5_2","systemRole":"Azure SQL primary in East US Zone 2 — running normally; low CPU because App Service is receiving minimal traffic; Active Geo-Replication is keeping the West US replica in sync with near-zero lag under this low load; when Front Door recovers (or when Traffic Manager failover routes traffic to West US App Service), writes will resume and replica lag will be the key metric to monitor before promoting; promote command: az sql db replica set-primary --name <db> --resource-group <rg> --server <west-server>","baseLatency":7,"costMultiplier":2,"maxConnections":600}},{"x":420,"y":400,"id":"amrf-sql-replica","name":"Azure SQL — Active Geo-Replica (West US)","type":"database","status":"healthy","cpuUsage":12,"location":{"zoneKey":"wus-zone-1","regionKey":"wus","localityType":"az","providerLabel":"West US (Zone 1)"},"provider":"azure","characteristics":{"size":"GP_Gen5_2","systemRole":"Azure SQL readable geo-replica in West US — replication lag is near zero because the primary is under minimal load; this replica is in the best possible shape for promotion: it has caught up, and promoting it to primary (planned failover: az sql db replica set-primary) would complete in under 30 s with RPO=0; once West US App Service begins receiving Traffic Manager-routed traffic, this becomes the write-path primary","baseLatency":6,"costMultiplier":2,"maxConnections":600}}],"connections":[{"id":"conn-amrf-fd-app-east","sourceId":"amrf-front-door","targetId":"amrf-app-east"},{"id":"conn-amrf-fd-app-west","sourceId":"amrf-front-door","targetId":"amrf-app-west"},{"id":"conn-amrf-app-east-sql","sourceId":"amrf-app-east","targetId":"amrf-sql-primary"},{"id":"conn-amrf-app-west-sql","sourceId":"amrf-app-west","targetId":"amrf-sql-primary"},{"id":"conn-amrf-app-west-replica","sourceId":"amrf-app-west","targetId":"amrf-sql-replica"},{"id":"conn-amrf-primary-replica","sourceId":"amrf-sql-primary","targetId":"amrf-sql-replica"}],"duration":"~12 min","tags":["Azure","Front Door","Traffic Manager","Multi-Region","Active Geo-Replication","HA","Resilience"],"category":"reliability","defaultTrafficPatterns":[{"name":"Stale-DNS trickle — only clients with cached IPs reaching backends directly","type":"step","startTime":0,"parameters":{"startTraffic":50,"endTraffic":50},"isActive":true},{"name":"Traffic Manager failover — direct bypass traffic floods West US origin","type":"ramp","startTime":60,"parameters":{"startTraffic":50,"endTraffic":3000,"duration":30},"isActive":false}],"defaultFailureInjections":[{"name":"Azure Front Door Crash Loop — Global Proxy Layer Offline","type":"az_outage","targetZone":"eus-zone-fd","severity":"severe","startTime":0,"endTime":9999,"isActive":true,"parameters":{}}],"realWorldIncident":{"date":"2025-10-09","provider":"Azure","summary":"On October 9, 2025, a data-plane crash loop took Azure Front Door edge sites offline globally for ~8h10m (07:50–16:00 UTC). Both East US and West US backends were healthy — the proxy layer was the failure. Services that relied solely on Front Door for global routing had no fallback path. This scenario opens at peak impact so you can see what a proxy-layer single point of failure looks like: Front Door critical, both App Service origins at near-zero load despite being healthy, and SQL under minimal pressure. The key lesson is the one Microsoft's post-incident review reinforced: Traffic Manager pointing directly at App Service origins is the safety net that survives a Front Door crash.","references":[{"label":"Azure status history · QNBQ-5W8","url":"https://azure.status.microsoft/status/history/?trackingId=QNBQ-5W8"},{"label":"Azure Networking Blog · lessons learned","url":"https://techcommunity.microsoft.com/blog/azurenetworkingblog/azure-front-door-implementing-lessons-learned-following-october-outages/4479416/"}]}},{"id":"lmi-lambda-approximation","title":"LMI Comparison — Ordinary Lambda Approximation","description":"Interactive arm of the LMI-versus-ECS research baseline: load this ordinary AWS Lambda approximation separately from the linked ECS Fargate arm to observe the readiness-versus-warm-capacity trade-off at 2,400 RPS [cwm]. This is ordinary Lambda, not Lambda Managed Instances; it is the closest supportable substitute, never a claim that Lambda and LMI are equivalent [aws-doc] [cwm-est] [unsupported]. The source Reddit measurements are validation references only [source] and were not used to tune this workload, readiness estimate, prices, or results. The existing CWM model can expose a generic linear application-readiness ramp that rejects burst traffic while capacity initializes [cwm] [cwm-est]; the runner's already-ready provisioned/warm variants show how warm capacity lowers utilization and modeled CPU-overload error [cwm]. This interactive arm does not model: (1) LMI as a distinct AWS platform, managed EC2 classes, or exact lifecycle; (2) LMI worker fan-out or multiple Node.js workers inside one Lambda/LMI environment; (3) memory-to-vCPU/worker derivation; (4) initialization CPU contention; (5) LMI runtime-ready versus application-ready internals; (6) exact HTTP 500/503 attribution; (7) RDS Proxy multiplexing, pinning, or a first-class proxy resource; or (8) exact source failure counts, including 500/503 counts, or latency/p99 reproduction [unsupported]. The PostgreSQL resource is a generous 5,000-connection pool approximation, not an RDS Proxy model [cwm-est] [unsupported]. The configured ECS rollout result is specific to that arm and configuration, not a universal ECS claim [cwm-est] [unsupported].","difficulty":"advanced","resources":[{"x":0,"y":0,"id":"lmi-lambda-api","name":"Node API (ordinary Lambda approximation)","type":"compute","status":"healthy","cpuUsage":1,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Ordinary AWS Lambda request-serving function, explicitly used as an approximation arm rather than Lambda Managed Instances [aws-doc] [cwm-est] [unsupported]. One request per execution environment, 8 GiB configured memory, 100-ms request duration, and a four-step estimated application-readiness ramp can reject burst traffic while a cold environment becomes ready [cwm] [cwm-est]. CWM does not derive memory to vCPU or worker fan-out, model initializing-worker CPU contention, reproduce exact 500/503 attribution or p99, or provide an LMI lifecycle [unsupported].","baseLatency":5,"concurrency":1,"minInstances":0,"appReadySteps":4,"serviceFamily":"lambda","costMultiplier":0.3,"coldStartLatency":200,"containerMemoryGiB":8,"requestServingKind":"function","provisionedInstances":0,"requestDurationSeconds":0.1}},{"x":200,"y":0,"id":"lmi-lambda-pg","name":"PostgreSQL (RDS-Proxy-like pool approximation)","type":"database","status":"healthy","cpuUsage":5,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"AWS PostgreSQL endpoint with a generous 5,000-connection pool standing in for the source workload's RDS-Proxy-like path [cwm-est]. This is a pool approximation only: RDS Proxy multiplexing, pinning, and a first-class proxy resource are unsupported [unsupported].","baseLatency":3,"cacheHitRate":0,"serviceFamily":"rds-postgres","costMultiplier":2,"maxConnections":5000}}],"connections":[{"id":"lmi-lambda-api-pg","sourceId":"lmi-lambda-api","targetId":"lmi-lambda-pg"}],"duration":"~10 min","tags":["AWS","Lambda","Ordinary Lambda","LMI Approximation","ECS Comparison","Readiness","Warm Capacity","Research","Source Validation"],"category":"scaling","seed":295824465,"defaultTrafficPatterns":[{"name":"Idle — establish the cold baseline","type":"step","startTime":0,"endTime":3,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true},{"name":"Cold burst — watch application readiness reject traffic","type":"step","startTime":3,"endTime":13,"parameters":{"startTraffic":2400,"endTraffic":2400},"isActive":true},{"name":"Warm hold — compare ready capacity and utilization","type":"step","startTime":13,"endTime":28,"parameters":{"startTraffic":2400,"endTraffic":2400},"isActive":true},{"name":"Quiesce — clear traffic before the re-burst","type":"step","startTime":28,"endTime":31,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true},{"name":"Re-burst — observe readiness on a second cold start","type":"step","startTime":31,"endTime":46,"parameters":{"startTraffic":2400,"endTraffic":2400},"isActive":true},{"name":"Post-deployment observation — hold the 2,400-RPS workload","type":"step","startTime":46,"endTime":60,"parameters":{"startTraffic":2400,"endTraffic":2400},"isActive":true},{"name":"Traffic Recovery — enable after step 60","type":"ramp","startTime":60,"endTime":90,"parameters":{"startTraffic":2400,"endTraffic":0,"duration":30},"isActive":false}]},{"id":"zombie-infra-aws","title":"Zombie Infrastructure — AWS","description":"Everything is stopped, yet you're still paying ~$310/month. Explore what keeps billing on AWS when traffic is zero — a NAT Gateway, interface VPC endpoints, detached Elastic IPs, an idle ALB, EBS snapshots, and the disks under stopped EC2/RDS instances. Remove idle resources one by one and watch the residual bill drop.","difficulty":"beginner","resources":[{"x":100,"y":80,"id":"zaws-nat-1","name":"NAT Gateway","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Forgotten NAT Gateway — $0.045/hr fixed fee even with zero traffic","billingState":"idle","serviceFamily":"nat-gateway","costMultiplier":1}},{"x":80,"y":440,"id":"zaws-vpce-1","name":"VPC Endpoint (S3)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Interface endpoint billed $0.01/AZ-hr while unused","billingState":"idle","serviceFamily":"vpc-endpoint","costMultiplier":1}},{"x":280,"y":440,"id":"zaws-vpce-2","name":"VPC Endpoint (ECR)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Interface endpoint billed $0.01/AZ-hr while unused","billingState":"idle","serviceFamily":"vpc-endpoint","costMultiplier":1}},{"x":900,"y":440,"id":"zaws-vpce-3","name":"VPC Endpoint (SSM)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Interface endpoint billed $0.01/AZ-hr while unused","billingState":"idle","serviceFamily":"vpc-endpoint","costMultiplier":1}},{"x":1100,"y":440,"id":"zaws-vpce-4","name":"VPC Endpoint (Logs)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Interface endpoint billed $0.01/AZ-hr while unused","billingState":"idle","serviceFamily":"vpc-endpoint","costMultiplier":1}},{"x":80,"y":620,"id":"zaws-eip-1","name":"Elastic IP (detached)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Unassociated Elastic IP — $0.005/hr while detached","billingState":"detached","serviceFamily":"eip-detached","costMultiplier":1}},{"x":280,"y":620,"id":"zaws-eip-2","name":"Elastic IP (detached)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Unassociated Elastic IP — $0.005/hr while detached","billingState":"detached","serviceFamily":"eip-detached","costMultiplier":1}},{"x":480,"y":620,"id":"zaws-eip-3","name":"Elastic IP (detached)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Unassociated Elastic IP — $0.005/hr while detached","billingState":"detached","serviceFamily":"eip-detached","costMultiplier":1}},{"x":500,"y":80,"id":"zaws-alb-1","name":"Idle ALB","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Load balancer with zero LCU traffic — minimum hourly fee still applies","baseLatency":2,"billingState":"idle","maxThroughput":10000,"serviceFamily":"alb-idle","costMultiplier":1}},{"x":300,"y":260,"id":"zaws-ec2-1","name":"web-server-01 (stopped)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","capacityGB":100,"systemRole":"Stopped EC2 — instance-hours are free, but its 100 GB EBS root volume keeps billing","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"ec2","costMultiplier":1}},{"x":700,"y":260,"id":"zaws-ec2-2","name":"web-server-02 (stopped)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"m5.large","capacityGB":100,"systemRole":"Stopped EC2 — instance-hours are free, but its 100 GB EBS root volume keeps billing","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"ec2","costMultiplier":1}},{"x":500,"y":440,"id":"zaws-rds-1","name":"MySQL RDS (stopped)","type":"database","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r5.large","capacityGB":500,"systemRole":"Stopped RDS — 500 GB of allocated storage keeps billing (and AWS auto-restarts it after 7 days)","baseLatency":5,"billingState":"stopped","serviceFamily":"rds","costMultiplier":2,"maxConnections":500}},{"x":700,"y":440,"id":"zaws-snap-1","name":"EBS Snapshots (3.5 TB)","type":"storage","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"capacityGB":3500,"systemRole":"Years of daily AMI snapshots — $0.05/GB-mo billed continuously","billingState":"idle","serviceFamily":"ebs-snapshot","costMultiplier":1}}],"connections":[{"id":"zaws-c-alb-ec2-1","sourceId":"zaws-alb-1","targetId":"zaws-ec2-1"},{"id":"zaws-c-alb-ec2-2","sourceId":"zaws-alb-1","targetId":"zaws-ec2-2"},{"id":"zaws-c-ec2-1-rds","sourceId":"zaws-ec2-1","targetId":"zaws-rds-1"},{"id":"zaws-c-ec2-2-rds","sourceId":"zaws-ec2-2","targetId":"zaws-rds-1"},{"id":"zaws-c-nat-ec2-1","sourceId":"zaws-nat-1","targetId":"zaws-ec2-1"},{"id":"zaws-c-nat-ec2-2","sourceId":"zaws-nat-1","targetId":"zaws-ec2-2"},{"id":"zaws-c-ec2-1-vpce-1","sourceId":"zaws-ec2-1","targetId":"zaws-vpce-1"},{"id":"zaws-c-ec2-1-vpce-2","sourceId":"zaws-ec2-1","targetId":"zaws-vpce-2"},{"id":"zaws-c-ec2-2-vpce-3","sourceId":"zaws-ec2-2","targetId":"zaws-vpce-3"},{"id":"zaws-c-ec2-2-vpce-4","sourceId":"zaws-ec2-2","targetId":"zaws-vpce-4"},{"id":"zaws-c-ec2-1-snap","sourceId":"zaws-ec2-1","targetId":"zaws-snap-1"},{"id":"zaws-c-ec2-2-snap","sourceId":"zaws-ec2-2","targetId":"zaws-snap-1"}],"duration":"~10 min","tags":["AWS","Cost","Idle Resources","Zombie Infrastructure","FinOps"],"category":"cost","defaultTrafficPatterns":[{"name":"Zero traffic — the bill keeps running anyway","type":"step","startTime":0,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true}],"realWorldIncident":{"date":"2026-08-04","provider":"AWS","summary":"On August 4, 2026, an r/aws thread titled 'It's always the networking costs' hit the front page of the subreddit: a team shut a project down, stopped every instance, and still got a ~$310 bill. Stopped EC2 instances don't bill instance-hours, but their EBS volumes do; RDS storage bills while the instance is stopped (and AWS restarts stopped RDS after 7 days); NAT Gateways ($0.045/hr), interface VPC endpoints ($0.01/AZ-hr), detached Elastic IPs ($0.005/hr), idle ALBs, and old EBS snapshots all bill 24×7 regardless of traffic. Industry FinOps reports consistently attribute 25–30%+ of cloud spend to this kind of idle 'zombie' infrastructure. Use the Residual Charges panel to see each line item, then remove idle resources to watch the monthly bill fall.","references":[{"label":"r/aws · 'It's always the networking costs' (Aug 2026)","url":"https://www.reddit.com/r/aws/comments/1vfav79/its_always_the_networking_costs/"},{"label":"AWS VPC Pricing (NAT Gateway)","url":"https://aws.amazon.com/vpc/pricing/"},{"label":"AWS EBS Pricing (volumes & snapshots)","url":"https://aws.amazon.com/ebs/pricing/"}]}},{"id":"zombie-infra-gcp","title":"Zombie Infrastructure — GCP","description":"A decommissioned GCP project that still bills ~$120/month at zero traffic: Cloud NAT's fixed gateway fee, Private Service Connect endpoints, reserved-but-unused static IPs, persistent disk snapshots, an idle load balancer forwarding rule, and the disks under stopped VMs and Cloud SQL. Remove each idle resource to watch the residual bill drop.","difficulty":"beginner","resources":[{"x":100,"y":80,"id":"zgcp-nat-1","name":"Cloud NAT","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Cloud NAT gateway — fixed hourly fee regardless of traffic","billingState":"idle","serviceFamily":"cloud-nat","costMultiplier":1}},{"x":100,"y":440,"id":"zgcp-psc-1","name":"PSC Endpoint (SQL)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Private Service Connect endpoint — per-endpoint hourly fee","billingState":"idle","serviceFamily":"psc-endpoint","costMultiplier":1}},{"x":900,"y":440,"id":"zgcp-psc-2","name":"PSC Endpoint (APIs)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-b","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-b (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Private Service Connect endpoint — per-endpoint hourly fee","billingState":"idle","serviceFamily":"psc-endpoint","costMultiplier":1}},{"x":80,"y":620,"id":"zgcp-ip-1","name":"Static IP (unused)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Reserved external static IP not attached to any VM","billingState":"detached","serviceFamily":"static-ip","costMultiplier":1}},{"x":280,"y":620,"id":"zgcp-ip-2","name":"Static IP (unused)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-b","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-b (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Reserved external static IP not attached to any VM","billingState":"detached","serviceFamily":"static-ip","costMultiplier":1}},{"x":500,"y":80,"id":"zgcp-clb-1","name":"Idle Cloud LB","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-regional","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1 (Regional)"},"provider":"gcp","characteristics":{"systemRole":"Forwarding rule minimum charge on an idle load balancer","baseLatency":2,"billingState":"idle","maxThroughput":10000,"serviceFamily":"clb-idle","costMultiplier":1}},{"x":300,"y":260,"id":"zgcp-vm-1","name":"app-vm-01 (stopped)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"n2-standard-2","capacityGB":100,"systemRole":"Stopped VM — its 100 GB persistent disk keeps billing","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"compute-engine","costMultiplier":1}},{"x":700,"y":260,"id":"zgcp-vm-2","name":"app-vm-02 (stopped)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-b","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-b (Iowa)"},"provider":"gcp","characteristics":{"size":"n2-standard-2","capacityGB":100,"systemRole":"Stopped VM — its 100 GB persistent disk keeps billing","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"compute-engine","costMultiplier":1}},{"x":500,"y":440,"id":"zgcp-sql-1","name":"Cloud SQL (stopped)","type":"database","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"db-n1-standard-2","capacityGB":500,"systemRole":"Stopped Cloud SQL — 500 GB allocated storage keeps billing","baseLatency":5,"billingState":"stopped","serviceFamily":"cloud-sql","costMultiplier":2,"maxConnections":500}},{"x":300,"y":440,"id":"zgcp-snap-1","name":"PD Snapshots (2 TB)","type":"storage","status":"healthy","cpuUsage":0,"location":{"zoneKey":"us-central1-regional","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1 (Regional)"},"provider":"gcp","characteristics":{"capacityGB":2000,"systemRole":"Persistent disk snapshots — $0.026/GB-mo billed continuously","billingState":"idle","serviceFamily":"pd-snapshot","costMultiplier":1}}],"connections":[{"id":"zgcp-c-clb-vm-1","sourceId":"zgcp-clb-1","targetId":"zgcp-vm-1"},{"id":"zgcp-c-clb-vm-2","sourceId":"zgcp-clb-1","targetId":"zgcp-vm-2"},{"id":"zgcp-c-vm-1-sql","sourceId":"zgcp-vm-1","targetId":"zgcp-sql-1"},{"id":"zgcp-c-vm-2-sql","sourceId":"zgcp-vm-2","targetId":"zgcp-sql-1"},{"id":"zgcp-c-nat-vm-1","sourceId":"zgcp-nat-1","targetId":"zgcp-vm-1"},{"id":"zgcp-c-nat-vm-2","sourceId":"zgcp-nat-1","targetId":"zgcp-vm-2"},{"id":"zgcp-c-vm-1-psc-1","sourceId":"zgcp-vm-1","targetId":"zgcp-psc-1"},{"id":"zgcp-c-vm-2-psc-2","sourceId":"zgcp-vm-2","targetId":"zgcp-psc-2"},{"id":"zgcp-c-vm-1-snap","sourceId":"zgcp-vm-1","targetId":"zgcp-snap-1"}],"duration":"~10 min","tags":["GCP","Cost","Idle Resources","Zombie Infrastructure","FinOps"],"category":"cost","defaultTrafficPatterns":[{"name":"Zero traffic — the bill keeps running anyway","type":"step","startTime":0,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true}],"realWorldIncident":{"date":"2026-08-04","provider":"GCP","summary":"The same 'zombie bill' pattern from the August 2026 r/aws thread 'It's always the networking costs' plays out on GCP with different line items. A GCP project everyone forgot: VMs and Cloud SQL are stopped, yet ~$120/month keeps billing. Cloud NAT charges a fixed gateway fee, Private Service Connect endpoints bill per endpoint-hour, reserved static IPs cost more when NOT in use, persistent disk snapshots bill per GB-month, load balancer forwarding rules bill hourly, and the persistent disks under stopped VMs and stopped Cloud SQL instances keep accruing. Use the Residual Charges panel to see every line item, then remove idle resources to watch the bill fall.","references":[{"label":"r/aws · 'It's always the networking costs' (Aug 2026)","url":"https://www.reddit.com/r/aws/comments/1vfav79/its_always_the_networking_costs/"},{"label":"GCP VPC network pricing (Cloud NAT, static IPs)","url":"https://cloud.google.com/vpc/network-pricing"},{"label":"GCP disk & snapshot pricing","url":"https://cloud.google.com/compute/disks-image-pricing"}]}},{"id":"zombie-infra-azure","title":"Zombie Infrastructure — Azure","description":"An Azure resource group after 'shutting everything down' — still ~$175/month: a NAT Gateway, private endpoints, unassociated static public IPs, managed disk snapshots, a Standard Load Balancer's fixed fee, and the disks under deallocated VMs and stopped Azure SQL. Remove idle resources to watch the residual bill drop.","difficulty":"beginner","resources":[{"x":100,"y":80,"id":"zaz-nat-1","name":"NAT Gateway","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"NAT Gateway — fixed hourly resource charge regardless of traffic","billingState":"idle","serviceFamily":"azure-nat-gw","costMultiplier":1}},{"x":100,"y":440,"id":"zaz-pe-1","name":"Private Endpoint (SQL)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Private Endpoint — per-endpoint hourly fee even when unused","billingState":"idle","serviceFamily":"private-endpoint","costMultiplier":1}},{"x":900,"y":440,"id":"zaz-pe-2","name":"Private Endpoint (Storage)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-2","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Zone 2)"},"provider":"azure","characteristics":{"systemRole":"Private Endpoint — per-endpoint hourly fee even when unused","billingState":"idle","serviceFamily":"private-endpoint","costMultiplier":1}},{"x":80,"y":620,"id":"zaz-ip-1","name":"Static Public IP (unassociated)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"systemRole":"Standard static public IP charged while unassociated","billingState":"detached","serviceFamily":"azure-static-ip","costMultiplier":1}},{"x":280,"y":620,"id":"zaz-ip-2","name":"Static Public IP (unassociated)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-2","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Zone 2)"},"provider":"azure","characteristics":{"systemRole":"Standard static public IP charged while unassociated","billingState":"detached","serviceFamily":"azure-static-ip","costMultiplier":1}},{"x":500,"y":80,"id":"zaz-lb-1","name":"Standard Load Balancer (idle)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-regional","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Regional)"},"provider":"azure","characteristics":{"systemRole":"Standard Load Balancer — fixed hourly fee with zero traffic","baseLatency":2,"billingState":"idle","maxThroughput":10000,"serviceFamily":"azure-lb-idle","costMultiplier":1}},{"x":300,"y":260,"id":"zaz-vm-1","name":"app-vm-01 (deallocated)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","capacityGB":100,"systemRole":"Deallocated VM — its 100 GB managed disk keeps billing","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"virtual-machines","costMultiplier":1}},{"x":700,"y":260,"id":"zaz-vm-2","name":"app-vm-02 (deallocated)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-2","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Zone 2)"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","capacityGB":100,"systemRole":"Deallocated VM — its 100 GB managed disk keeps billing","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"virtual-machines","costMultiplier":1}},{"x":500,"y":440,"id":"zaz-sql-1","name":"Azure SQL (stopped)","type":"database","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-1","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Zone 1)"},"provider":"azure","characteristics":{"size":"GP_Gen5_2","capacityGB":500,"systemRole":"Stopped Azure SQL — 500 GB allocated storage keeps billing","baseLatency":5,"billingState":"stopped","serviceFamily":"azure-sql","costMultiplier":2,"maxConnections":500}},{"x":300,"y":440,"id":"zaz-snap-1","name":"Disk Snapshots (1.5 TB)","type":"storage","status":"healthy","cpuUsage":0,"location":{"zoneKey":"eastus-regional","regionKey":"eastus","localityType":"zone","providerLabel":"East US (Regional)"},"provider":"azure","characteristics":{"capacityGB":1500,"systemRole":"Managed disk snapshots — $0.05/GB-mo billed continuously","billingState":"idle","serviceFamily":"disk-snapshot","costMultiplier":1}}],"connections":[{"id":"zaz-c-lb-vm-1","sourceId":"zaz-lb-1","targetId":"zaz-vm-1"},{"id":"zaz-c-lb-vm-2","sourceId":"zaz-lb-1","targetId":"zaz-vm-2"},{"id":"zaz-c-vm-1-sql","sourceId":"zaz-vm-1","targetId":"zaz-sql-1"},{"id":"zaz-c-vm-2-sql","sourceId":"zaz-vm-2","targetId":"zaz-sql-1"},{"id":"zaz-c-nat-vm-1","sourceId":"zaz-nat-1","targetId":"zaz-vm-1"},{"id":"zaz-c-nat-vm-2","sourceId":"zaz-nat-1","targetId":"zaz-vm-2"},{"id":"zaz-c-vm-1-pe-1","sourceId":"zaz-vm-1","targetId":"zaz-pe-1"},{"id":"zaz-c-vm-2-pe-2","sourceId":"zaz-vm-2","targetId":"zaz-pe-2"},{"id":"zaz-c-vm-1-snap","sourceId":"zaz-vm-1","targetId":"zaz-snap-1"}],"duration":"~10 min","tags":["Azure","Cost","Idle Resources","Zombie Infrastructure","FinOps"],"category":"cost","defaultTrafficPatterns":[{"name":"Zero traffic — the bill keeps running anyway","type":"step","startTime":0,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true}],"realWorldIncident":{"date":"2026-08-04","provider":"Azure","summary":"The 'we deallocated everything' surprise — the Azure version of the zombie bill an August 2026 r/aws thread ('It's always the networking costs') made famous: deallocated VMs stop billing compute, but their managed disks keep billing; stopped Azure SQL keeps billing allocated storage; NAT Gateways ($0.045/hr), private endpoints ($0.01/hr each), unassociated static public IPs, disk snapshots, and the Standard Load Balancer's fixed fee ($0.025/hr) all bill around the clock. Use the Residual Charges panel to see every line item, then remove idle resources to watch the monthly bill fall.","references":[{"label":"r/aws · 'It's always the networking costs' (Aug 2026)","url":"https://www.reddit.com/r/aws/comments/1vfav79/its_always_the_networking_costs/"},{"label":"Azure NAT Gateway pricing","url":"https://azure.microsoft.com/pricing/details/azure-nat-gateway/"},{"label":"Azure Managed Disks pricing (snapshots)","url":"https://azure.microsoft.com/pricing/details/managed-disks/"}]}},{"id":"zombie-infra-digitalocean","title":"Zombie Infrastructure — DigitalOcean","description":"A wound-down DigitalOcean project still billing ~$90/month: reserved IPs not assigned to Droplets, volume snapshots, an idle $10/mo load balancer, and volumes under powered-off Droplets and a stopped managed database. DigitalOcean has no NAT Gateway or private endpoint product — its zombie surface is smaller but real. Remove idle resources to watch the residual bill drop.","difficulty":"beginner","resources":[{"x":80,"y":620,"id":"zdo-ip-1","name":"Reserved IP (unassigned)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3"},"provider":"digitalocean","characteristics":{"systemRole":"Reserved IP charged while not assigned to a Droplet","billingState":"detached","serviceFamily":"do-reserved-ip","costMultiplier":1}},{"x":280,"y":620,"id":"zdo-ip-2","name":"Reserved IP (unassigned)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3"},"provider":"digitalocean","characteristics":{"systemRole":"Reserved IP charged while not assigned to a Droplet","billingState":"detached","serviceFamily":"do-reserved-ip","costMultiplier":1}},{"x":400,"y":80,"id":"zdo-lb-1","name":"Load Balancer (idle)","type":"network","status":"healthy","cpuUsage":0,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3"},"provider":"digitalocean","characteristics":{"systemRole":"Idle load balancer — flat $10/mo regardless of traffic","baseLatency":2,"billingState":"idle","maxThroughput":10000,"serviceFamily":"do-lb-idle","costMultiplier":1}},{"x":250,"y":260,"id":"zdo-drop-1","name":"web-droplet-01 (powered off)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3"},"provider":"digitalocean","characteristics":{"size":"s-2vcpu-4gb","capacityGB":50,"systemRole":"Powered-off Droplet — its 50 GB attached volume keeps billing (note: real powered-off Droplets bill in full; destroy to stop charges)","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"droplet","costMultiplier":1}},{"x":550,"y":260,"id":"zdo-drop-2","name":"web-droplet-02 (powered off)","type":"compute","status":"healthy","cpuUsage":0,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3"},"provider":"digitalocean","characteristics":{"size":"s-2vcpu-4gb","capacityGB":50,"systemRole":"Powered-off Droplet — its 50 GB attached volume keeps billing (note: real powered-off Droplets bill in full; destroy to stop charges)","baseLatency":3,"billingState":"stopped","maxThroughput":30000,"serviceFamily":"droplet","costMultiplier":1}},{"x":350,"y":440,"id":"zdo-db-1","name":"Managed PostgreSQL (stopped)","type":"database","status":"healthy","cpuUsage":0,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3"},"provider":"digitalocean","characteristics":{"size":"db-s-2vcpu-4gb","capacityGB":100,"systemRole":"Stopped managed database — 100 GB allocated storage keeps billing","baseLatency":5,"billingState":"stopped","serviceFamily":"managed-database","costMultiplier":2,"maxConnections":500}},{"x":600,"y":440,"id":"zdo-snap-1","name":"Volume Snapshots (1 TB)","type":"storage","status":"healthy","cpuUsage":0,"location":{"zoneKey":"nyc3-az1","regionKey":"nyc3","localityType":"zone","providerLabel":"NYC3"},"provider":"digitalocean","characteristics":{"capacityGB":1000,"systemRole":"Volume snapshots — $0.05/GB-mo billed continuously","billingState":"idle","serviceFamily":"do-vol-snapshot","costMultiplier":1}}],"connections":[{"id":"zdo-c-lb-drop-1","sourceId":"zdo-lb-1","targetId":"zdo-drop-1"},{"id":"zdo-c-lb-drop-2","sourceId":"zdo-lb-1","targetId":"zdo-drop-2"},{"id":"zdo-c-drop-1-db","sourceId":"zdo-drop-1","targetId":"zdo-db-1"},{"id":"zdo-c-drop-2-db","sourceId":"zdo-drop-2","targetId":"zdo-db-1"},{"id":"zdo-c-drop-1-snap","sourceId":"zdo-drop-1","targetId":"zdo-snap-1"}],"duration":"~10 min","tags":["DigitalOcean","Cost","Idle Resources","Zombie Infrastructure","FinOps"],"category":"cost","defaultTrafficPatterns":[{"name":"Zero traffic — the bill keeps running anyway","type":"step","startTime":0,"parameters":{"startTraffic":0,"endTraffic":0},"isActive":true}],"realWorldIncident":{"date":"2026-08-04","provider":"DigitalOcean","summary":"The DigitalOcean version of the zombie bill an August 2026 r/aws thread ('It's always the networking costs') made famous. A side project wound down but never cleaned up: reserved IPs bill while unassigned to a Droplet, volume snapshots bill $0.05/GB-mo forever, the load balancer's flat $10/mo keeps running with zero traffic, and volumes under powered-off Droplets keep billing (in real DigitalOcean, powered-off Droplets bill in FULL — you must destroy them to stop charges). DigitalOcean has no NAT Gateway or private endpoint product, so its zombie surface is smaller than the hyperscalers' — but never zero. Remove idle resources to watch the bill fall.","references":[{"label":"r/aws · 'It's always the networking costs' (Aug 2026)","url":"https://www.reddit.com/r/aws/comments/1vfav79/its_always_the_networking_costs/"},{"label":"DigitalOcean Reserved IP pricing","url":"https://docs.digitalocean.com/products/networking/reserved-ips/details/pricing/"},{"label":"DigitalOcean Load Balancer pricing","url":"https://www.digitalocean.com/pricing/load-balancers"}]}},{"id":"gcp-saas-web-app","title":"GCP SaaS Web App","description":"A production SaaS stack on Google Cloud: Cloud CDN caches 80% of requests at the edge so the App Engine service climbs into the warning zone at peak without ever going critical, while Cloud SQL stores tenant data and Cloud Operations Suite adds observability. Watch the ramp push App Engine toward its limit — and compare this stack's hourly cost with the AWS and Azure SaaS scenarios.","difficulty":"beginner","resources":[{"x":250,"y":60,"id":"gcpsaas-cdn","name":"Cloud CDN","type":"network","status":"healthy","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Google Cloud CDN — caches content at 140+ global PoPs so only misses reach App Engine; $0.08/GB served from cache","baseLatency":-20,"cacheHitRate":0.8,"maxThroughput":40000,"costMultiplier":4}},{"x":250,"y":220,"id":"gcpsaas-appengine","name":"App Engine (F2)","type":"compute","status":"healthy","cpuUsage":30,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"App Engine Standard F2 — fully managed auto-scaling compute hosting the SaaS web application; billed per instance-hour while handling requests","autoscaling":false,"baseLatency":15,"maxThroughput":1500,"costMultiplier":1,"coldStartLatency":400}},{"x":250,"y":400,"id":"gcpsaas-cloudsql","name":"Cloud SQL","type":"database","status":"healthy","cpuUsage":20,"location":{"zoneKey":"us-central1-b","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-b (Iowa)"},"provider":"gcp","characteristics":{"size":"db-n1-standard-1","systemRole":"Cloud SQL for PostgreSQL (db-n1-standard-1) — tenant data store with automated backups and HA options","baseLatency":6,"costMultiplier":1.47,"maxConnections":400}},{"x":80,"y":220,"id":"gcpsaas-cloudops","name":"Cloud Operations Suite","type":"security","status":"healthy","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Cloud Logging + Monitoring + Trace — collects all GCP telemetry from the stack; $0.50/GB ingested logs beyond 50 GB/mo free tier","baseLatency":0,"maxThroughput":0,"costMultiplier":0.4}}],"connections":[{"id":"gcpsaas-c-cdn-appengine","sourceId":"gcpsaas-cdn","targetId":"gcpsaas-appengine"},{"id":"gcpsaas-c-appengine-cloudsql","sourceId":"gcpsaas-appengine","targetId":"gcpsaas-cloudsql"},{"id":"gcpsaas-c-appengine-cloudops","sourceId":"gcpsaas-appengine","targetId":"gcpsaas-cloudops"}],"duration":"~10 min","tags":["GCP","SaaS","App Engine","Cloud CDN","Cloud SQL","Cost"],"category":"cost","defaultTrafficPatterns":[{"name":"Business-hours ramp — watch App Engine climb into the warning zone","type":"ramp","startTime":0,"parameters":{"startTraffic":90,"endTraffic":680,"duration":60},"isActive":true},{"name":"Traffic Recovery — see the stack settle back down","type":"ramp","startTime":60,"parameters":{"startTraffic":680,"endTraffic":90,"duration":30},"isActive":false}]},{"id":"inference-autoscaling","title":"Inference Autoscaling","description":"A GKE GPU cluster (NVIDIA T4 nodes) serving LLM inference under a rising traffic ramp. Watch the core economics relationship live: as GPU utilization climbs from ~30% toward saturation, cost per million tokens falls — GPU nodes bill at full rate whether busy or idle, so every extra token spreads the same node cost thinner. Around 70% utilization the cluster autoscaler adds a T4 node, briefly raising cost per token before throughput catches up. The recovery ramp shows the reverse: falling traffic strands expensive GPU capacity and per-token cost climbs again.","difficulty":"intermediate","resources":[{"x":250,"y":60,"id":"infauto-lb","name":"Cloud Load Balancing","type":"network","status":"healthy","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"systemRole":"Global HTTP(S) load balancer routing inference requests to the GKE GPU cluster","baseLatency":2,"maxThroughput":40000,"costMultiplier":1}},{"x":250,"y":220,"id":"infauto-gke-gpu","name":"GKE GPU Cluster (T4)","type":"kubernetes","status":"healthy","cpuUsage":30,"location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"maxNodes":5,"minNodes":1,"nodeCount":2,"systemRole":"GKE GPU node pool (NVIDIA T4, $0.54/node-hr) serving LLM inference — watch cost per million tokens fall as utilization rises","accelerator":"t4","autoscaling":true,"baseLatency":35,"inferenceMode":true,"maxThroughput":16000,"serviceFamily":"gpu-kubernetes","costMultiplier":5.4,"tokensPerRequest":150,"inputTokensPerRequest":300}},{"x":250,"y":400,"id":"infauto-redis","name":"Memorystore (KV cache)","type":"cache","status":"healthy","location":{"zoneKey":"us-central1-a","regionKey":"us-central1","localityType":"zone","providerLabel":"us-central1-a (Iowa)"},"provider":"gcp","characteristics":{"size":"e2-medium","systemRole":"Prompt/response cache — repeated prompts skip GPU generation entirely","baseLatency":1,"cacheHitRate":0.3,"maxThroughput":50000,"costMultiplier":1}}],"connections":[{"id":"infauto-c-lb-gke","sourceId":"infauto-lb","targetId":"infauto-gke-gpu"},{"id":"infauto-c-gke-redis","sourceId":"infauto-gke-gpu","targetId":"infauto-redis"}],"duration":"~15 min","tags":["GCP","Kubernetes","GPU","Inference","Autoscaling","Cost"],"category":"cost","defaultTrafficPatterns":[{"name":"Inference ramp — watch cost/M tokens fall as GPU utilization rises","type":"ramp","startTime":0,"parameters":{"startTraffic":200,"endTraffic":2000,"duration":40},"isActive":true},{"name":"Traffic Recovery — see per-token cost climb as GPUs idle","type":"ramp","startTime":60,"parameters":{"startTraffic":2000,"endTraffic":200,"duration":30},"isActive":false}]},{"id":"aws-saas-web-app","title":"AWS SaaS Web App","description":"A production SaaS stack on AWS: CloudFront caches 80% of requests at the edge so the App Runner container service climbs into the warning zone at peak without ever going critical, while Aurora PostgreSQL stores tenant data and CloudWatch adds observability. Watch the ramp push App Runner toward its limit — and compare this stack's hourly cost with the Azure and GCP SaaS scenarios.","difficulty":"beginner","resources":[{"x":250,"y":60,"id":"awssaas-cf","name":"CloudFront CDN","type":"network","status":"healthy","location":{"zoneKey":"us-east-1-regional","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1 (Regional)"},"provider":"aws","characteristics":{"systemRole":"Content delivery network that caches static content at edge locations so only misses reach App Runner","baseLatency":-30,"cacheHitRate":0.8,"maxThroughput":50000,"costMultiplier":1.5}},{"x":250,"y":220,"id":"awssaas-apprunner","name":"App Runner Service","type":"compute","status":"healthy","cpuUsage":30,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"Fully managed container hosting for the SaaS web application — $0.064/vCPU-hr + $0.007/GB-hr × 2 GB = $0.078/hr while active","autoscaling":false,"baseLatency":15,"maxThroughput":2000,"costMultiplier":0.81,"coldStartLatency":500}},{"x":250,"y":400,"id":"awssaas-aurora","name":"Aurora PostgreSQL","type":"database","status":"healthy","cpuUsage":20,"location":{"zoneKey":"us-east-1b","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1b (N. Virginia)"},"provider":"aws","characteristics":{"size":"db.r6g.large","systemRole":"Amazon Aurora PostgreSQL-compatible cluster — tenant data store with automatic storage scaling; $0.26/hr","baseLatency":3,"costMultiplier":2.17,"maxConnections":1000}},{"x":80,"y":220,"id":"awssaas-cloudwatch","name":"CloudWatch","type":"security","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"systemRole":"AWS CloudWatch — collects metrics, logs, and traces from the whole stack; $0.50/GB ingested logs beyond 5 GB/mo free","baseLatency":0,"maxThroughput":0,"costMultiplier":0.44}}],"connections":[{"id":"awssaas-c-cf-apprunner","sourceId":"awssaas-cf","targetId":"awssaas-apprunner"},{"id":"awssaas-c-apprunner-aurora","sourceId":"awssaas-apprunner","targetId":"awssaas-aurora"},{"id":"awssaas-c-apprunner-cw","sourceId":"awssaas-apprunner","targetId":"awssaas-cloudwatch"}],"duration":"~10 min","tags":["AWS","SaaS","App Runner","CloudFront","Aurora","Cost"],"category":"cost","defaultTrafficPatterns":[{"name":"Business-hours ramp — watch App Runner climb into the warning zone","type":"ramp","startTime":0,"parameters":{"startTraffic":120,"endTraffic":850,"duration":60},"isActive":true},{"name":"Traffic Recovery — see the stack settle back down","type":"ramp","startTime":60,"parameters":{"startTraffic":850,"endTraffic":120,"duration":30},"isActive":false}]},{"id":"github-inspired-cascading-retry-storm","title":"GitHub-Inspired Cascading Retry Storm (Unofficial)","description":"Advanced reliability investigation based on the published August 17, 2026 GitHub.com incident. This independent, non-affiliated simulation starts at a healthy 8K RPS before a scheduled Central US capacity anomaly appears. Use dependency telemetry to determine which constraint failed first, why authentication retries amplified, and how a Copilot-like token path became saturated. Then compare the identical seeded fault against the included protected policy with bounded retries, exponential backoff with jitter, circuit breakers, rate limits, and load shedding.","difficulty":"advanced","resources":[{"x":80,"y":80,"id":"gh-ingress","name":"Public Ingress Load Balancer","type":"network","status":"healthy","location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Accepts web, API, raw-content, and AI-assistant traffic; its broad fan-in makes regional symptoms visible before the underlying bottleneck is known.","baseLatency":3,"maxThroughput":50000,"serviceFamily":"application-gateway","costMultiplier":1}},{"x":270,"y":80,"id":"gh-proxy","name":"Central US HAProxy Nodes","type":"compute","status":"healthy","cpuUsage":24,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"size":"Standard_D4s_v3","systemRole":"Host proxy nodes that forward traffic through the service mesh. Their healthy CPU is intentionally not a reliable proxy for sidecar concurrency headroom.","autoscaling":false,"baseLatency":5,"maxThroughput":100000,"serviceFamily":"virtual-machines","costMultiplier":1}},{"x":430,"y":80,"id":"gh-istio-sidecar","name":"Istio Sidecar Pods","type":"compute","status":"healthy","cpuUsage":31,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","systemRole":"Constrained service-mesh component. Its resilience concurrency budget, not its host throughput, is the initiating bottleneck during the traffic peak; the baseline scaling policy incorrectly observes HAProxy host CPU instead of this component's capacity utilization.","autoscaling":false,"baseLatency":4,"maxThroughput":100000,"serviceFamily":"virtual-machines","costMultiplier":1}},{"x":570,"y":80,"id":"gh-gateway-auth","name":"Gateway Authentication","type":"compute","status":"healthy","cpuUsage":26,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"size":"Standard_D4s_v3","systemRole":"Validates web/API sessions and requests short-lived AI tokens. A delayed upstream reply can turn normal authentication into repeated attempts.","autoscaling":false,"baseLatency":8,"maxThroughput":100000,"serviceFamily":"virtual-machines","costMultiplier":1}},{"x":730,"y":10,"id":"gh-web","name":"Web Frontend Services","type":"compute","status":"healthy","cpuUsage":22,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Serves interactive repository, issue, and pull-request experiences after gateway authentication.","autoscaling":false,"baseLatency":12,"maxThroughput":100000,"serviceFamily":"app-service","costMultiplier":1}},{"x":730,"y":145,"id":"gh-api","name":"Core API Services","type":"compute","status":"healthy","cpuUsage":25,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Handles authenticated API operations and fans out to database, cache, and asynchronous work dependencies.","autoscaling":false,"baseLatency":10,"maxThroughput":100000,"serviceFamily":"app-service","costMultiplier":1}},{"x":500,"y":250,"id":"gh-token-service","name":"Copilot-Like Token Service","type":"compute","status":"healthy","cpuUsage":24,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Issues AI-assistant access tokens. Its request rate is generated by gateway attempts, so a retry loop can raise a normal 7–9K RPS workload toward incident-scale load.","autoscaling":false,"baseLatency":18,"maxThroughput":100000,"serviceFamily":"app-service","costMultiplier":1}},{"x":730,"y":280,"id":"gh-copilot-workload","name":"Copilot-Like AI Workload","type":"kubernetes","status":"healthy","cpuUsage":20,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"maxNodes":8,"minNodes":3,"nodeCount":3,"systemRole":"Independent AI request path gated by token issuance; it can remain degraded after core web/API traffic begins to recover.","autoscaling":false,"baseLatency":35,"maxThroughput":100000,"serviceFamily":"aks","costMultiplier":1}},{"x":960,"y":90,"id":"gh-database","name":"Repository Metadata Database","type":"database","status":"healthy","cpuUsage":18,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Durable repository and account metadata store; connection pressure helps distinguish primary from downstream symptoms.","baseLatency":7,"maxThroughput":18000,"serviceFamily":"azure-sql","costMultiplier":1,"maxConnections":7500}},{"x":960,"y":180,"id":"gh-cache","name":"Session and Metadata Cache","type":"cache","status":"healthy","cpuUsage":15,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Caches sessions and hot repository metadata; rising misses or saturation would indicate a different initiating failure.","baseLatency":2,"maxThroughput":40000,"serviceFamily":"azure-cache-redis","costMultiplier":1,"maxConnections":15000}},{"x":960,"y":270,"id":"gh-queue","name":"Background Work Queue","type":"queue","status":"healthy","location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"systemRole":"Buffers asynchronous repository, webhook, and automation jobs while synchronous paths are under stress.","baseLatency":6,"maxThroughput":18000,"serviceFamily":"service-bus","costMultiplier":1}},{"x":1180,"y":270,"id":"gh-workers","name":"Background Worker Fleet","type":"compute","status":"healthy","cpuUsage":18,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"size":"Standard_D4s_v3","systemRole":"Drains queued jobs; queue lag exposes broader degradation without making workers the presumed root cause.","autoscaling":false,"baseLatency":15,"maxThroughput":100000,"serviceFamily":"virtual-machines","costMultiplier":1}},{"x":960,"y":360,"id":"gh-raw-content","name":"Archive and Raw Content Store","type":"storage","status":"healthy","location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"capacityGB":5000,"systemRole":"Serves repository archives and raw content, a high-error user-facing surface in the referenced incident.","baseLatency":14,"maxThroughput":16000,"serviceFamily":"blob-storage","costMultiplier":1}},{"x":280,"y":285,"id":"gh-monitoring","name":"Capacity and Flow Monitoring","type":"compute","status":"healthy","cpuUsage":12,"location":{"zoneKey":"centralus-1","regionKey":"centralus","localityType":"zone","providerLabel":"Central US · Zone 1"},"provider":"azure","characteristics":{"size":"Standard_D2s_v3","systemRole":"Collects sidecar concurrency, proxy flow, retry, token RPS, queue, and error telemetry needed to identify the original capacity mismatch.","autoscaling":false,"baseLatency":4,"maxThroughput":100000,"serviceFamily":"virtual-machines","costMultiplier":1}}],"connections":[{"id":"gh-c-ingress-proxy","sourceId":"gh-ingress","targetId":"gh-proxy"},{"id":"gh-c-proxy-sidecar","sourceId":"gh-proxy","targetId":"gh-istio-sidecar"},{"id":"gh-c-sidecar-gateway","sourceId":"gh-istio-sidecar","targetId":"gh-gateway-auth"},{"id":"gh-c-gateway-web","sourceId":"gh-gateway-auth","targetId":"gh-web"},{"id":"gh-c-gateway-api","sourceId":"gh-gateway-auth","targetId":"gh-api"},{"id":"gh-c-gateway-token","sourceId":"gh-gateway-auth","targetId":"gh-token-service"},{"id":"gh-c-token-copilot","sourceId":"gh-token-service","targetId":"gh-copilot-workload"},{"id":"gh-c-api-db","sourceId":"gh-api","targetId":"gh-database"},{"id":"gh-c-api-cache","sourceId":"gh-api","targetId":"gh-cache"},{"id":"gh-c-api-queue","sourceId":"gh-api","targetId":"gh-queue"},{"id":"gh-c-queue-workers","sourceId":"gh-queue","targetId":"gh-workers"},{"id":"gh-c-web-raw","sourceId":"gh-web","targetId":"gh-raw-content"},{"id":"gh-c-proxy-monitoring","sourceId":"gh-proxy","targetId":"gh-monitoring"}],"duration":"~20 min","tags":["Azure","Advanced","Reliability","Retry Storm","Authentication","AI","Real Incident"],"category":"reliability","seed":8172026,"resilienceConfig":{"version":1,"enabled":true,"dependencies":[{"id":"gh-ingress-proxy","sourceId":"gh-ingress","targetId":"gh-proxy","requestRatio":1,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":0,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":500,"retryBudgetRatio":0},"capacity":{"maxRps":50000,"maxConcurrent":20000,"meanServiceTimeMs":12}},{"id":"gh-proxy-sidecar","sourceId":"gh-proxy","targetId":"gh-istio-sidecar","requestRatio":1,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":0,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":500,"retryBudgetRatio":0},"capacity":{"maxRps":50000,"maxConcurrent":20000,"meanServiceTimeMs":12}},{"id":"gh-sidecar-gateway","sourceId":"gh-istio-sidecar","targetId":"gh-gateway-auth","requestRatio":1,"authRequestsPerAttempt":1,"authDependencyId":"gh-gateway-token","retryPolicy":{"maxRetries":8,"backoffMs":0,"backoffMultiplier":1,"jitterRatio":0,"timeoutMs":250,"retryBudgetRatio":10,"retryActorId":"gateway_optimistic_retries"},"capacity":{"maxRps":10000,"maxConcurrent":200,"meanServiceTimeMs":20}},{"id":"gh-gateway-web","sourceId":"gh-gateway-auth","targetId":"gh-web","requestRatio":0.65,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":800,"retryBudgetRatio":0.25},"capacity":{"maxRps":30000,"maxConcurrent":12000,"meanServiceTimeMs":30}},{"id":"gh-gateway-api","sourceId":"gh-gateway-auth","targetId":"gh-api","requestRatio":0.35,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":800,"retryBudgetRatio":0.25},"capacity":{"maxRps":30000,"maxConcurrent":12000,"meanServiceTimeMs":25}},{"id":"gh-gateway-token","sourceId":"gh-gateway-auth","targetId":"gh-token-service","requestRatio":0,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":0,"backoffMultiplier":1,"jitterRatio":0,"timeoutMs":400,"retryBudgetRatio":1,"retryActorId":"vs_code_client_retries"},"capacity":{"maxRps":20000,"maxConcurrent":10000,"meanServiceTimeMs":35}},{"id":"gh-token-copilot","sourceId":"gh-token-service","targetId":"gh-copilot-workload","requestRatio":1,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":150,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":1200,"retryBudgetRatio":0.2},"capacity":{"maxRps":25000,"maxConcurrent":8000,"meanServiceTimeMs":80}},{"id":"gh-api-db","sourceId":"gh-api","targetId":"gh-database","requestRatio":0.8,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":2,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":900,"retryBudgetRatio":0.5},"capacity":{"maxRps":18000,"maxConcurrent":7500,"meanServiceTimeMs":20}},{"id":"gh-api-cache","sourceId":"gh-api","targetId":"gh-cache","requestRatio":0.7,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":50,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":300,"retryBudgetRatio":0.25},"capacity":{"maxRps":40000,"maxConcurrent":15000,"meanServiceTimeMs":5}},{"id":"gh-api-queue","sourceId":"gh-api","targetId":"gh-queue","requestRatio":0.4,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":2,"backoffMs":200,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":1000,"retryBudgetRatio":0.5},"capacity":{"maxRps":18000,"maxConcurrent":10000,"meanServiceTimeMs":15}},{"id":"gh-queue-workers","sourceId":"gh-queue","targetId":"gh-workers","requestRatio":1,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":250,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":1500,"retryBudgetRatio":0.25},"capacity":{"maxRps":15000,"maxConcurrent":6000,"meanServiceTimeMs":40}},{"id":"gh-web-raw","sourceId":"gh-web","targetId":"gh-raw-content","requestRatio":0.25,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":2,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":1000,"retryBudgetRatio":0.5},"capacity":{"maxRps":16000,"maxConcurrent":8000,"meanServiceTimeMs":25}},{"id":"gh-proxy-monitoring","sourceId":"gh-proxy","targetId":"gh-monitoring","requestRatio":0.02,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":0,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":500,"retryBudgetRatio":0},"capacity":{"maxRps":5000,"maxConcurrent":2000,"meanServiceTimeMs":10}}],"scheduledFaults":[{"id":"gh-centralus-traffic-peak","type":"traffic_surge","dependencyId":"gh-ingress-proxy","startStep":6,"endStep":15,"trafficMultiplier":2.25}],"scalingPolicies":[{"id":"gh-sidecar-host-cpu-policy","dependencyId":"gh-sidecar-gateway","constrainedMetric":"concurrency","observationMetric":"source_cpu","observedResourceId":"gh-proxy","scaleOutThresholdPercent":85,"scaleOutCapacityMultiplier":2.5}],"maxCascadeDepth":6,"maxGeneratedRps":150000,"maxStepWork":256,"retryGeneratedTrafficAffectsCost":true},"protectedResilienceConfig":{"version":1,"enabled":true,"dependencies":[{"id":"gh-ingress-proxy","sourceId":"gh-ingress","targetId":"gh-proxy","requestRatio":1,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":0,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":500,"retryBudgetRatio":0},"capacity":{"maxRps":50000,"maxConcurrent":20000,"meanServiceTimeMs":12}},{"id":"gh-proxy-sidecar","sourceId":"gh-proxy","targetId":"gh-istio-sidecar","requestRatio":1,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":0,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":500,"retryBudgetRatio":0},"capacity":{"maxRps":50000,"maxConcurrent":20000,"meanServiceTimeMs":12}},{"id":"gh-sidecar-gateway","sourceId":"gh-istio-sidecar","targetId":"gh-gateway-auth","requestRatio":1,"authRequestsPerAttempt":1,"authDependencyId":"gh-gateway-token","retryPolicy":{"maxRetries":2,"backoffMs":500,"backoffMultiplier":2,"jitterRatio":0.25,"timeoutMs":250,"retryBudgetRatio":0.5,"retryActorId":"gateway_optimistic_retries"},"capacity":{"maxRps":10000,"maxConcurrent":200,"meanServiceTimeMs":20},"protection":{"circuitBreaker":{"enabled":true,"failureRateThreshold":0.2,"minimumRequests":100,"openSteps":3,"halfOpenMaxRequests":100},"rateLimitRps":9000,"loadShedding":true}},{"id":"gh-gateway-web","sourceId":"gh-gateway-auth","targetId":"gh-web","requestRatio":0.65,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":800,"retryBudgetRatio":0.25},"capacity":{"maxRps":30000,"maxConcurrent":12000,"meanServiceTimeMs":30}},{"id":"gh-gateway-api","sourceId":"gh-gateway-auth","targetId":"gh-api","requestRatio":0.35,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":800,"retryBudgetRatio":0.25},"capacity":{"maxRps":30000,"maxConcurrent":12000,"meanServiceTimeMs":25}},{"id":"gh-gateway-token","sourceId":"gh-gateway-auth","targetId":"gh-token-service","requestRatio":0,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":750,"backoffMultiplier":2,"jitterRatio":0.25,"timeoutMs":400,"retryBudgetRatio":0.2,"retryActorId":"vs_code_client_retries"},"capacity":{"maxRps":20000,"maxConcurrent":10000,"meanServiceTimeMs":35},"protection":{"circuitBreaker":{"enabled":true,"failureRateThreshold":0.25,"minimumRequests":100,"openSteps":3,"halfOpenMaxRequests":100},"rateLimitRps":20000,"loadShedding":true}},{"id":"gh-token-copilot","sourceId":"gh-token-service","targetId":"gh-copilot-workload","requestRatio":1,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":150,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":1200,"retryBudgetRatio":0.2},"capacity":{"maxRps":25000,"maxConcurrent":8000,"meanServiceTimeMs":80}},{"id":"gh-api-db","sourceId":"gh-api","targetId":"gh-database","requestRatio":0.8,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":2,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":900,"retryBudgetRatio":0.5},"capacity":{"maxRps":18000,"maxConcurrent":7500,"meanServiceTimeMs":20}},{"id":"gh-api-cache","sourceId":"gh-api","targetId":"gh-cache","requestRatio":0.7,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":50,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":300,"retryBudgetRatio":0.25},"capacity":{"maxRps":40000,"maxConcurrent":15000,"meanServiceTimeMs":5}},{"id":"gh-api-queue","sourceId":"gh-api","targetId":"gh-queue","requestRatio":0.4,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":2,"backoffMs":200,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":1000,"retryBudgetRatio":0.5},"capacity":{"maxRps":18000,"maxConcurrent":10000,"meanServiceTimeMs":15}},{"id":"gh-queue-workers","sourceId":"gh-queue","targetId":"gh-workers","requestRatio":1,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":1,"backoffMs":250,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":1500,"retryBudgetRatio":0.25},"capacity":{"maxRps":15000,"maxConcurrent":6000,"meanServiceTimeMs":40}},{"id":"gh-web-raw","sourceId":"gh-web","targetId":"gh-raw-content","requestRatio":0.25,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":2,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":1000,"retryBudgetRatio":0.5},"capacity":{"maxRps":16000,"maxConcurrent":8000,"meanServiceTimeMs":25}},{"id":"gh-proxy-monitoring","sourceId":"gh-proxy","targetId":"gh-monitoring","requestRatio":0.02,"authRequestsPerAttempt":0,"retryPolicy":{"maxRetries":0,"backoffMs":100,"backoffMultiplier":2,"jitterRatio":0.1,"timeoutMs":500,"retryBudgetRatio":0},"capacity":{"maxRps":5000,"maxConcurrent":2000,"meanServiceTimeMs":10}}],"scheduledFaults":[{"id":"gh-centralus-traffic-peak","type":"traffic_surge","dependencyId":"gh-ingress-proxy","startStep":6,"endStep":15,"trafficMultiplier":2.25}],"scalingPolicies":[{"id":"gh-sidecar-host-cpu-policy","dependencyId":"gh-sidecar-gateway","constrainedMetric":"concurrency","observationMetric":"capacity_utilization","observedResourceId":"gh-proxy","scaleOutThresholdPercent":85,"scaleOutCapacityMultiplier":2.5}],"maxCascadeDepth":6,"maxGeneratedRps":150000,"maxStepWork":256,"retryGeneratedTrafficAffectsCost":true},"defaultTrafficPatterns":[{"name":"Healthy 8K RPS baseline and recovery hold","type":"step","startTime":0,"parameters":{"startTraffic":8000,"endTraffic":8000,"duration":60},"isActive":true}],"realWorldIncident":{"provider":"GitHub-inspired (unaffiliated)","date":"2026-08-17","summary":"Independent educational simulation inspired by GitHub's published incident report; it is not an official GitHub product or an exact forensic reconstruction. The source incident lasted 7h47m, reached approximately 20% web/API errors and approximately 50% archive/raw-content errors, and saw Copilot Token Service traffic rise from a normal 7–9K RPS to 70–100K RPS. GitHub attributed the cascade to Central US load-balancer saturation after an Istio sidecar reached concurrency limits and failed to autoscale under a policy watching the host but not sidecar limits. Gateway authentication delays and optimistic retries amplified load, including a latent client retry loop. Documented remediation themes included sidecar-aware autoscaling, retry limits and backoff, load-balancer monitoring, regional failover safeguards, reduced gateway retries, and temporarily blocking retry-triggering token requests before a gradual ramp-up.","references":[{"label":"GitHub Status — Incident with GitHub.com (official post-incident report)","url":"https://www.githubstatus.com/incidents/zkxwbgr0cnmx"},{"label":"DevOps.com — GitHub Hit by Widespread Outage, Halting Work for Global Developers","url":"https://devops.com/github-hit-by-widespread-outage-halting-work-for-global-developers/"}]}},{"id":"idle-gpu-infrastructure","title":"Idle GPU Infrastructure","description":"An EKS GPU node group (NVIDIA A10G g5 instances) that nobody is sending traffic to — the AI-infrastructure version of a zombie resource. With zero requests the cluster generates zero tokens, so cost per million tokens reads ∞ while the full GPU node cost keeps accruing every hour. Watch idle-GPU detection fire after a few steps of sub-15% utilization, surfacing the monthly waste in the Residual Charges panel, and ask the AI insight what to do about the stranded capacity.","difficulty":"beginner","resources":[{"x":250,"y":200,"id":"idlegpu-eks","name":"EKS GPU Cluster (A10G)","type":"kubernetes","status":"healthy","cpuUsage":3,"location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"size":"g5.xlarge","maxNodes":4,"minNodes":1,"nodeCount":2,"systemRole":"Forgotten EKS g5 GPU node group (2× A10G nodes @ $1.006/node-hr) — billing around the clock with no inference traffic","accelerator":"a10g","autoscaling":false,"baseLatency":35,"inferenceMode":true,"maxThroughput":16000,"serviceFamily":"gpu-kubernetes","costMultiplier":10.06,"tokensPerRequest":150,"inputTokensPerRequest":300}},{"x":250,"y":400,"id":"idlegpu-s3","name":"S3 Model Artifacts","type":"storage","status":"healthy","location":{"zoneKey":"us-east-1a","regionKey":"us-east-1","localityType":"zone","providerLabel":"us-east-1a (N. Virginia)"},"provider":"aws","characteristics":{"capacityGB":200,"systemRole":"Model weights and checkpoints — cheap next to the idle GPU nodes they were meant to serve","baseLatency":20,"maxThroughput":5000,"serviceFamily":"s3","costMultiplier":1}}],"connections":[{"id":"idlegpu-c-eks-s3","sourceId":"idlegpu-eks","targetId":"idlegpu-s3"}],"duration":"~10 min","tags":["AWS","Kubernetes","GPU","Inference","Idle","Zombie","Cost"],"category":"cost","defaultTrafficPatterns":[{"name":"Zero traffic — the GPUs bill anyway","type":"ramp","startTime":0,"parameters":{"startTraffic":0,"endTraffic":0,"duration":60},"isActive":true}]}]