It's 14:47 UTC on a Friday. Your e-commerce platform, ShopStream, is running a flash sale that started at 14:00 UTC. Traffic ramped up as expected — your team pre-scaled the application pods to handle 3x normal load.
But 25 minutes in, PagerDuty fires:
Customers are tweeting about timeouts. Your VP of Engineering just messaged the #incident channel. The on-call SRE escalated to you: 'We can't tell if it's the DB or the app layer — everything is slow.'
You have access to Grafana dashboards, application logs, and the deployment pipeline. Your team is standing by. What do you do first?
All dashboard and log timestamps in this lab are shown in UTC.