Same Scores, Different Decisions: What AI Evaluation Can Miss
I argue that evaluation should follow a score into the decision it produces, including who receives an opportunity when capacity is limited.
Read the essay →I believe a system should be judged by the decisions it makes possible, the people those decisions affect, and the evidence an institution can defend.
I return to one question across AI, institutions and education: what becomes invisible when a complex decision is compressed into a score? I set out my positions here, connect them to work that can be examined, and make clear what would change my mind.
Read my position →I develop my positions through concrete decisions in AI, institutional design and education.
I argue that evaluation should follow a score into the decision it produces, including who receives an opportunity when capacity is limited.
Read the essay →I believe the number of opportunities a system can allocate belongs in the explanation of its decisions. A model cannot take responsibility for a limit the institution chose.
Read the essay →I believe institutions need a defensible description of the work before they ask AI to classify a role, recommend a requirement or assess a candidate.
Read the essay →I want to know what a student can explain, test and defend when the conditions change. A polished submission gives me only part of that answer.
Read the essay →I use RIDI to investigate what an evaluation leaves unspecified about a downstream decision. I use the public demonstration as a replay of one SciFact case: the displayed retrieval metrics remain unchanged while the recorded answer changes. I treat that case as a reason to examine the relationship between an audit and an action; it is not a universal estimate of harm or proof that every deployment behaves the same way.
Implementation & project stages →I connect the identity question in IMAM to who defines the purpose of learning, selects the knowledge and judges the evidence of readiness. I believe a university should be able to explain and revise those choices in its own language and disciplinary context, even when it uses external AI tools. I would judge that control by visible authority over curriculum, assessment and revision; an Arabic interface alone would not establish it.
Read the education argument →I aim to publish one substantial essay every two months. This is an editorial intention from here forward. I will date new essays and explain substantive revisions, so readers can follow how an argument changes.