← All courses
$20

AI Specs & Evals for PMs

The product manager's lens on AI: write a feature spec an engineer can build from, then define and read the eval that says whether it's good enough to ship. Runs a real, deterministic eval over a support-triage assistant — and teaches the skill that matters most: a single accuracy number will lie to you. Learn to find the false pass the headline metric hides.

"87.5% accuracy looked shippable — the eval caught the tax question it hallucinated. That's the failure that ships."

6 Interactive Sessions

Short, interactive sessions — watch it work, steer it, then build it yourself. Go deeper anytime with the full code walkthrough.

  1. 1

    The spec is mostly its non-goals

    See why framing an AI feature is less about what it does and more about pinning what it must NOT do — the non-goals are the safety surface.

  2. 2

    A refusal rule is a testable promise

    Turn the frame into a spec an engineer can build and an eval can check — a structured output contract plus explicit refusal rules.

  3. 3

    Some dimensions are averages; one is a floor

    Turn 'is it good?' into measurable dimensions with thresholds set before you see results — and recognise the one dimension that can't be a weighted average.

  4. 4

    The case you don't write is the failure you won't catch

    Author an eval set that spans happy, adversarial, and must-refuse cases — because coverage is the product, and the must-refuse cases are the ones that matter.

  5. 5

    The headline number is hiding a case

    Read an eval report well enough to find the false pass — the dangerous case a reassuring headline accuracy number waves through.

  6. 6

    The eval doesn't ship anything — you do

    Turn the eval report into an accountable decision: the call, the evidence, the condition to ship, and the signals to watch once it's live.

Production patterns you'll master

AI Feature SpecsOutput Contracts & GuardrailsEval RubricsEval Sets (adversarial/must-refuse)Reading an Eval ReportEval-Gated Go/No-Go

Synthetic data included

  • AI feature spec
  • Eval rubric (dimensions + gates)
  • Eval case set (happy/adversarial/must-refuse)
  • Pre-recorded assistant outputs
  • Go/no-go memo

What you walk away with

Shareable portfolio

A public URL showing your module timeline, patterns mastered, and completion status.

All the code

Download everything as a ZIP — pipelines, guardrails, deployment configs. Yours forever.

Module walkthrough

Each module documented with deliverables and the production pattern you implemented.

Ready to build your ai specs & evals for pms?

First course free. $20 per course after that.