The most expensive AI programs I see are not the ones that fail loudly. They are the ones that never get measured, so they never get killed.
Leadership asks for ROI. Teams show screenshots. That is not a conversation. That is theater with a budget code.
Pick metrics that survive a board packet
Good AI metrics are boring and defensible:
- Cycle time from brief to approved asset
- Revision count per piece
- Cost per asset (including human review time)
- Conversion or engagement against a human-only control
- Error or compliance incident rate
Bad AI metrics are vibes:
- "Employees are excited"
- "We generated 10x more drafts"
- "We are using the platform"
More drafts is not progress if review capacity and quality collapse.
Design the kill criterion first
Before a pilot starts, write the exit rule:
- If cycle time does not improve by X% in 60 days with equal quality, stop.
- If compliance flags rise above baseline, stop.
- If only one power user can run the workflow, stop and redesign.
Without a kill criterion, every pilot becomes a pet.
Measure programs the way you measure anything else that costs money.
A simple measurement stack
You do not need a data science team to start:
- Choose one workflow and one primary metric
- Capture two weeks of baseline (human-only)
- Run AI-assisted with the same quality bar
- Review weekly with the owner, not monthly with a steering committee
- Publish the result internally, win or lose
Honesty about a failed pilot is cheaper than funding it for a year.
Related reading
- MIT Sloan Management Review on AI and productivity (search current AI ROI coverage)
- NIST AI RMF for mapping measurement to risk controls
