AI Anomaly Detection
Build an anomaly detection system that monitors time-series metrics, detects unusual patterns using statistical and ML methods, triggers alerts, and provides AI-powered root cause analysis.
"API latency spiked 3x at 14:32 — the root cause is a slow database query from the new deployment"
6 Interactive Sessions
Short, interactive sessions — watch it work, steer it, then build it yourself. Go deeper anytime with the full code walkthrough.
- 1
The Signal in the Noise — shaping data you can trust
Before you can catch the spike, four data sources have to reconcile into one clean, time-aligned, windowed stream — because you cannot detect what you never collected cleanly.
- 2
What's Normal? — the baseline that breathes
Teach the system what normal looks like — a band that moves with the daily and weekly rhythm — so a value is judged against what it should be at THIS hour, not a flat average.
- 3
When Stats Aren't Enough — the anomaly no z-score sees
Catch the multivariate anomaly a single z-score misses — where each metric looks normal on its own but the combination has never occurred — using an ensemble of isolation forest and autoencoder.
- 4
Alert Without the Fatigue — one incident, one alert
Turn a storm of anomaly points into a single, well-ranked, actionable alert — so the 14:32 incident pages once with context, not two hundred times, and the on-call team never learns to ignore the board.
- 5
See It, Own It — the dashboard the on-call lives in
Make the whole pipeline visible and actionable: an engineer opens the dashboard mid-incident and answers in ten seconds — what spiked, when, how bad, still firing — then acknowledges, resolves, or silences the alert.
- 6
From Alert to Root Cause — the whole arc pays off
Turn 'latency spiked 3× at 14:32' into 'the cause is a slow database query from deploy CHG-2207' — by correlating the anomaly with recent deploys, related metrics and past incidents — then track SLAs and error budgets so the system improves over time.
Production patterns you'll master
Synthetic data included
- API latency metrics (30 days)
- Error rate logs (JSON)
- Deployment events
- Infrastructure metrics (CPU/memory)
- Incident history
What you walk away with
Shareable portfolio
A public URL showing your module timeline, patterns mastered, and completion status.
All the code
Download everything as a ZIP — pipelines, guardrails, deployment configs. Yours forever.
Module walkthrough
Each module documented with deliverables and the production pattern you implemented.
Ready to build your ai anomaly detection?
First course free. $20 per course after that.