Validity is not a property of an instrument Chapter 9: Measurement and Construct Validation — Video 5 https://ghrbook.com/videos/validity-is-not-a-property-of-an-instrument/ [Slide 1] The previous videos focused on planning: identifying what to measure and how indicators are constructed. Now we turn to validation. [Slide 2] How do you know if an instrument actually measures what it claims to measure? Whether you're developing a new instrument, adapting an existing one, or simply selecting from available options, understanding the validation process helps you critically appraise measurement quality. [Slide 3] Construct validation is the process of establishing that the numbers we generate with a method of measurement actually represent the idea, the construct, that we wish to measure. Take the depression inventory, where sum scale scores can range from 0 to 63. Validating this construct means establishing that higher numbers correspond with greater depression severity, or that scores above a certain threshold, such as 29 out of 63, correctly classify someone as having severe depression. [Slide 4] If our numbers don't mean what we think they mean, our analyses don't either. You might be thinking that construct validation is not a top concern in your work, because you, a principled scientist, are using validated scales. If so, you'd be wrong. Validity is not a property of an instrument. [Slide 5] One methods paper makes the point clearly. Validity is not a binary property of an instrument, but instead a judgment made about the score interpretation based on a body of accumulating evidence that should continue to amass whenever the instrument is in use. Ongoing validation is necessary because the same instrument can be used in different contexts or for different purposes, and evidence that the interpretation of scores generalizes to those new contexts is needed. [Slide 6] I won't argue that you must always start from scratch to validate the instruments you select. But it's important to think critically about why you believe an instrument will produce valid results in your context. If you are using an instrument originally validated with a sample of 200 white women in one small city in America, what gives you confidence that the numbers produced carry the same meaning in rural India? [Slide 7] Whether you're developing a new instrument or evaluating an existing one for your context, it helps to think about validation as a journey through three phases. Each phase answers a different question. The substantive phase asks whether the instrument captures the right content: what topics should be included, and how should items be worded so your target population understands them? The structural phase asks whether the numbers behave. And the external phase asks whether the scores match reality. [Slide 8] The first phase of construct validation asks whether your instrument captures the right content. Whether you're building something new or adapting an existing tool for a new context, there are two questions to answer: does the instrument cover the important aspects of the construct? And are the questions understood by the target population? [Slide 9] An instrument has evidence of content validity when it assesses all of the conceptually relevant domains of a construct and excludes unrelated content. The process of establishing content validity typically starts with a review of the literature, conversations with experts, and, critically, talking to members of your target population to learn how they understand and describe the experience you're trying to measure. This last step matters more than you might think. Constructs don't always travel well across cultures. [Slide 10] My colleagues and I saw this firsthand in a study in rural Kenya, where we examined how women understood depression in the context of pregnancy and childbirth. A screening tool built somewhere else carries somebody else's account of what depression looks like, so we started from the other end. We convened focus groups and asked women to describe what observable features characterize depression, huzuni in Swahili, during the perinatal period. We had also gone through 17 existing depression screening tools and put each of their symptom terms onto a card, in Swahili and in English. Then we co-examined the overlap, and the lack of overlap, between what the women described and what was on those cards. [Slide 11] This is what that looked like. Each card carries one term from those existing tools: hopeless, crying, worried, slowed down, blame self, loss of sleep. Once we had matched the terms the women listed against the cards, the ones left over went on the table, and the group sorted each into four columns: yes, this is a characteristic of huzuni, scored three; maybe, scored two; no, scored minus one; and no opinion, scored zero. What you're looking at is a group of women deciding, term by term, which of somebody else's symptoms belong in their setting. [Slide 12] Some symptoms on standard screening tools aligned with what women described. Others didn't resonate at all, and the women identified experiences that no existing tool captured. A group of Kenyan mental health professionals then reviewed the results and offered feedback based on their local clinical expertise. Without this kind of ground-up process, we would have used an instrument that missed important aspects of the construct and included items that didn't belong. [Slide 13] Even when an instrument covers the right domains, the items themselves might not work for your population. An item like, I feel I am being punished, might be interpreted very differently depending on religious and cultural context. An item about changes in sleep patterns might not discriminate well in a population where sleep disruption is near-universal due to living conditions rather than depression. [Slide 14] Cognitive interviewing is a structured way to catch these problems. You sit down with members of your target population, ask them to read each item aloud, and have them describe in their own words what they think the item is asking. You'd be surprised how often an item that seems perfectly clear to the research team means something entirely different to a respondent, or means nothing at all. [Slide 15] The HAP team's decision to drop one item about sex for cultural reasons is a small example of what this phase produces. The item wasn't misunderstood; it was culturally inappropriate for the context. Phase 1 is about examining, questioning, and adapting an instrument so that its content fits the population and setting in which it will be used. [Slide 16] You've collected pilot data. Two hundred people in your target population responded to every item on your questionnaire. Now what? Before you can claim these items measure depression, or whatever construct you're after, you need to answer three questions: are the items working? Do they hang together? And are they consistent? [Slide 17] A common initial practice is to plot the response distributions of each item. If you ask people to rate their agreement with a statement like, I feel sad, and 100 percent of people in your sample respond strongly agree, the item has zero variance. When all or nearly all of your sample responds the same way to an item, that item tells you nothing useful. The appropriate next step is to drop the item, or conduct additional cognitive interviewing to modify it in a way that will elicit variation in responses. [Slide 18] If you have data that let you plot response distributions by group, you can also examine the extent to which items distinguish between groups. Here is a hypothetical example where 100 people responded to three items, each measured on a four-point scale from never to often. [Slide 19] The first item has very little variability. Almost everyone responded never. This item does not tell us much, and you might decide to drop or improve it. The second item has more variability, but it does not distinguish between cases and non-cases. In combination with other variables it might be useful, so you might decide to keep it unless you need to trim the overall length of the questionnaire. The third item looks the most promising. It elicits variability in responses, and a larger proportion of the cases group endorsed the item. [Slide 20] Phase 1 asks about content, and Phase 2 asks about the numbers. Item analysis tells you whether individual items are working. But knowing that each item works on its own doesn't tell you whether they work together, and that is the question that comes next.