Skip to main content
Operation MirageMedium
Cache Meltdown · conversational lab
incident0m 00s
SEV-1Operation Mirage · Medium

It's 14:47 UTC on a Friday. Your e-commerce platform, ShopStream, is running a flash sale that started at 14:00 UTC. Traffic ramped up as expected — your team pre-scaled the application pods to handle 3x normal load.

But 25 minutes in, PagerDuty fires:

🔴CRITICAL: p99 API latency > 12s (threshold: 2s)
🔴CRITICAL: Error rate > 15% on /api/products and /api/cart
🟡WARNING: Database connection pool pressure detected

Customers are tweeting about timeouts. Your VP of Engineering just messaged the #incident channel. The on-call SRE escalated to you: 'We can't tell if it's the DB or the app layer — everything is slow.'

You have access to Grafana dashboards, application logs, and the deployment pipeline. Your team is standing by. What do you do first?

All dashboard and log timestamps in this lab are shown in UTC.

View system architecture
ShopStream is a standard 3-tier e-commerce architecture. The application layer i…
⏎ send · ⇧⏎ newline0/2000
Send for ad-hoc queries and actions (e.g. "Pull Redis INFO").·Submit theory when you're ready to lock in your root cause for grading.