“In the back room, the team’s minds were blown. No one was thinking this woman was representative, or that the behavior was very common. That wasn’t the point. But there’s one thing a single data point can give you each and every time, and that’s an existence proof.” (Judd Antin, 2017; emphasis mine)

Do you need to assess whether users find an AI model fine-tuned for helpfulness actually helpful? Or perhaps whether two AI agents are able to coordinate to complete a series of tasks? Or maybe whether a guardrail was indeed effective at preventing a model from generating medical advice in multi-turn interactions? Or do you need to stress test or red team a system to see whether it fails in unexpected ways? Or understand how users interact with AI systems that generate outputs suggestive of a capacity for intention or morality? Or perhaps measure how prevalent a certain type of system behavior or failure is? Or check whether a model's performance degrades over time? Or determine which of several AI models performs better?

These are all evaluation questions. In my experience, however, regardless of whether one is assessing something about humans or human phenomena, or about AI and other computational systems, or about how humans interact with AI systems, and so on, they are likely asking one or a mix of the following questions:1

  1. 1.

    What is the phenomenon2 being observed

  2. 2.

    Whether a phenomenon of interest is being observed

  3. 3.

    How the phenomenon of interest is being observed (such as with respect to prevalence, frequency, saliency, severity)

  4. 4.

    When or under what circumstances a phenomenon of interest is observed 

  5. 5.

    How does this phenomenon relate or compare with other phenomena

  6. 6.

    What causes or is caused by a phenomenon of interest

In other words, what about a phenomenon of interest is the evaluation intended to help us learn. The reason why understanding which one of these questions someone is trying to answer is useful, is because they can help clarify what type of prior knowledge the evaluation assumes and what type of evidence each question requires in order to be answered. For instance, both #1 and #2which are common questions in red teaming exercisescould be answered by providing empirical evidence that the phenomenon was observedor by providing some existence proofbut only #2 assumes that the phenomenon under investigation is known, while #1 instead requires additional conceptual or theoretical work to recognize or characterize that phenomenon.

It’s worth noting, however, that none of these questions provide cues about whether acquiring the evidence needed to answer them is easy or feasible, which very much depends on what the phenomenon is and what can be observed about it. These questions are thus orthogonal to questions such as about conceptual clarity, observability, or construct or ecological validity—which are critical to trusting any evidence the evaluations rely on.

I wanted to articulate these prototypical questions as discussions about AI evaluations sometimes seem more focused on the evaluation methodse.g., LLM-as-a-judge, benchmark-driven evaluations, simulated users—rather than on what the evaluations should answer. Instead of treating evaluation methods as hammers in search of nails, focusing first on what type of evidence one needs can help practitioners consider other, more appropriate methodological approaches. For instance, qualitative methods often shine when what one needs is an existence proof