Same Scores, Different Decisions: What AI Evaluation Can Miss
Why aggregate model performance can hide downstream differences in who receives what.
This is where I write through the questions behind my research and operating work: what AI evaluation misses, what institutions reveal about strategy, how healthcare technology behaves in the real world, and what education should prepare people to do.
Two systems can look equally good on an aggregate metric and still produce meaningfully different downstream allocations. That gap between scoring and action is where consequential evaluation gets interesting.
Read the essay →The categories overlap by design. Most consequential problems do.
A small archive by design. The goal is cumulative thinking, not content volume.
Why aggregate model performance can hide downstream differences in who receives what.
Institutional constraints change what “good evidence” means when decisions cannot wait for perfect information.
Authentic problems create the reason to care about the mechanism.
What changes when the system is real, the user is a patient, and failure has operational consequences.