The decision starts before the ranking

I start with a simple hypothetical: an institution has two hundred eligible applicants and twenty interview places. A model ranks the applicants. I cannot evaluate the resulting interview list responsibly by examining prediction quality alone, because the institution has already made a consequential choice about how many people can proceed. I want that choice visible in the account of the system.

I regard capacity as a policy variable even when a budget makes it difficult to change. Calling twenty places a constraint explains a limit; it does not explain why that limit is appropriate, whether alternatives were considered, or how its consequences will be reviewed. My concern is accountability for the complete decision, including the choices made before a model receives its data.

What I would require from an evaluation

I would ask a team to examine several operationally plausible capacities and show which selections change at each boundary. I would inspect predictive performance and relevant group measures alongside that account. I would also ask whether ties, incomplete records or review rules determine who crosses the boundary. Those details help me distinguish a model difference from a policy difference; they do not tell me which choice is automatically fair.

I would keep such analysis proportionate to the decision. In a hypothetical hiring review, I might examine how an interview list changes when capacity increases from twenty to thirty. I would not publish identifiable applicant records to make that comparison persuasive. I want authorized reviewers to understand the changes, while public reporting describes the method and patterns at a level that protects the people involved.

A better score does not settle the choice

I do not infer that a more stable selection is necessarily a better selection. A correction may properly replace people selected under a flawed rule. Equally, identical selections can preserve a mistake. I use allocation identity to ask sharper questions: what changed, why did it change, which evidence supports the change, and who has authority to accept it? Stability is information to interpret within that reasoning.

I connect this position to the allocation questions in my RIDI research agenda. I want the relationship between a reported audit and an actual action to be inspectable. I would resist using a compelling demonstration to skip the harder evaluation of a particular institution: its population, available resources, decision rules, uncertainty and consequences. Each setting needs an argument supported by evidence appropriate to that setting.

The accountability I want

My practical test is whether an institution can explain the decision without pointing only to a model score. I want a named owner for the capacity choice, a reason for the selection rule, and a way to reconsider decisions when evidence or resources change. I would accept a lighter process where the consequences are small and easily reversible. Where an opportunity is scarce and consequential, I believe this explanation belongs in the ordinary standard of responsible decision-making.