“[...] one can still observe among ‘tough-minded’ psychologists the use of words such as ‘unobservable’ and ‘hypothetical’ in an essentially derogatory manner, and an almost compulsive fear of passing beyond the direct colligation of observable data” (MacCorquodale and Meehl, 1948)
Why is it so challenging to appropriately measure and computationally represent unobservable constructs such as those related to desirable or undesirable AI system behaviors—e.g., fairness, toxicity, sycophancy, helpfulness—their impacts—e.g., emotional dependency, overreliance, anthropomorphic deception—or other related phenomena—e.g., companionship, alignment, coordination—concerning how systems interact with each other or with humans?
Such unobservable constructs are now commonplace in AI work. Though, even when clearly conceptualized—which remains rare—these constructs can be difficult to operationalize. I want to unpack a bit why that is the case or what about unobservable constructs makes them difficult to measure or computationally represent, and why this limits the claims one can make about them.
A bit of background: unobservable constructs have been broadly categorized based on whether their indicators1 are reflective—when they are caused by or understood as a ‘symptom’ of the construct—or whether they are formative—when they cause or are considered constitutive of the construct.2 In practice, one reason this distinction matters is that it tells us whether we might be able to drop an indicator when we operationalize a construct. If the indicator is considered to be formative, dropping it will essentially change the meaning of the construct—since the construct is understood to be at least partially constituted by that indicator. And as a result, we are no longer operationalizing the same construct we set out to operationalize. This affects the type of claims we can make. This isn’t necessarily a concern for reflective indicators which are just ‘symptoms’—that is, a change in the construct is expected to lead to a change in the values observed for a reflective indicator, but not vice-versa.3
Compounding such issues related to misconstruing the direction of the relationship between the constructs and what can be observed about them—the observable indicators—is what is referred to as ‘surplus meaning’ or the gap between the theoretical meaning the constructs are intended to capture and what can be observed about them in practice. Thus, regardless of whether an unobservable construct’s indicators are reflective or formative, an unobservable construct will carry ‘surplus meaning’ beyond what an indicator or any combination of indicators can capture about it. Collapsing what we can observe about an unobservable construct onto what the construct means erases its ‘surplus meaning’ and thus also essentially changes the overall meaning of that construct.
This ‘surplus meaning’ is an important reason for why one should not conflate their metrics or measurements with their constructs. Overlooking this ‘surplus meaning’ is a common source for misconceptions about what one can or cannot claim about the construct based on observable indicators. The larger the ‘surplus meaning,’ the less we can learn from the indicators about the construct, and thus the more it limits what one can claim based only on what they observe about the indicators.4 One should, thus, at least also be very skeptical about any claims that ignore either the relationship between a construct and its indicators, or the ‘surplus meaning’ that unobservable constructs carry.
As noted in my previous post, a lot has been written on this topic. If interested in reading more, I recommend starting with these papers (alas no short papers this time, but these are worth your time!):
Lee J. Cronbach and Paul E. Meehl (1955). Construct validity in psychological tests. Psychological Bulletin.
Jeffrey R. Edwards and Richard P. Bagozzi (2000). On the Nature and Direction of Relationships Between Constructs and Measures. Psychological Methods.