Specify
Turn product intent into representative tasks, critical slices, observable invariants, and a versioned failure taxonomy.
LLMs from First Principles · Class 1
A three-hour, lifecycle-driven class mapping predictive ML deployment to LLM applications and agents—from versioned evidence and offline gates through staged release, online evaluation, drift, and incident-driven iteration.
Turn product intent into representative tasks, critical slices, observable invariants, and a versioned failure taxonomy.
Choose graders deliberately, preserve paired comparisons, report uncertainty, and inspect regressions before averages.
Stage exposure, join online outcomes, confirm drift, contain failures, and promote incidents into future release gates.
BugBoard is a fictional issue-triage agent. Version B improves overall success from 82% to 86%, yet security containment falls from 90% to 65% while latency and cost rise. Students must discover why the headline is insufficient and turn product judgment into a no-go gate.
Students commit to a release decision before the security slice is revealed.
Pairs interrogate stored A/B results without needing credentials or a model API.
Each pair writes thresholds and explains what evidence could change its decision.