← All courses
$20

AI Anomaly Detection

Build an anomaly detection system that monitors time-series metrics, detects unusual patterns using statistical and ML methods, triggers alerts, and provides AI-powered root cause analysis.

"API latency spiked 3x at 14:32 — the root cause is a slow database query from the new deployment"

6 Interactive Sessions

Short, interactive sessions — watch it work, steer it, then build it yourself. Go deeper anytime with the full code walkthrough.

  1. 1

    The Signal in the Noise — shaping data you can trust

    Before you can catch the spike, four data sources have to reconcile into one clean, time-aligned, windowed stream — because you cannot detect what you never collected cleanly.

  2. 2

    What's Normal? — the baseline that breathes

    Teach the system what normal looks like — a band that moves with the daily and weekly rhythm — so a value is judged against what it should be at THIS hour, not a flat average.

  3. 3

    When Stats Aren't Enough — the anomaly no z-score sees

    Catch the multivariate anomaly a single z-score misses — where each metric looks normal on its own but the combination has never occurred — using an ensemble of isolation forest and autoencoder.

  4. 4

    Alert Without the Fatigue — one incident, one alert

    Turn a storm of anomaly points into a single, well-ranked, actionable alert — so the 14:32 incident pages once with context, not two hundred times, and the on-call team never learns to ignore the board.

  5. 5

    See It, Own It — the dashboard the on-call lives in

    Make the whole pipeline visible and actionable: an engineer opens the dashboard mid-incident and answers in ten seconds — what spiked, when, how bad, still firing — then acknowledges, resolves, or silences the alert.

  6. 6

    From Alert to Root Cause — the whole arc pays off

    Turn 'latency spiked 3× at 14:32' into 'the cause is a slow database query from deploy CHG-2207' — by correlating the anomaly with recent deploys, related metrics and past incidents — then track SLAs and error budgets so the system improves over time.

Production patterns you'll master

Statistical BaselinesML DetectionAlert RoutingRoot Cause AnalysisObservability

Synthetic data included

  • API latency metrics (30 days)
  • Error rate logs (JSON)
  • Deployment events
  • Infrastructure metrics (CPU/memory)
  • Incident history

What you walk away with

Shareable portfolio

A public URL showing your module timeline, patterns mastered, and completion status.

All the code

Download everything as a ZIP — pipelines, guardrails, deployment configs. Yours forever.

Module walkthrough

Each module documented with deliverables and the production pattern you implemented.

Ready to build your ai anomaly detection?

First course free. $20 per course after that.