Validity is not a property of an instrument
Slide 1The previous videos focused on planning: identifying what to measure and how indicators are constructed. Now we turn to validation.
How do you know an instrument measures what it claims to measure?
Whether you are developing one, adapting one, or selecting from what exists
Understanding validation helps you critically appraise measurement quality
Slide 2How do you know if an instrument actually measures what it claims to measure? Whether you're developing a new instrument, adapting an existing one, or simply selecting from available options, understanding the validation process helps you critically appraise measurement quality.
Construct validation
Establishing that the numbers represent the idea
Sum scores on the BDI-II can range from 0 to 63
Higher numbers should correspond with greater depression severity
Scores above 29 should correctly classify someone as severely depressed
Slide 3Construct validation is the process of establishing that the numbers we generate with a method of measurement actually represent the idea, the construct, that we wish to measure. Take the depression inventory, where sum scale scores can range from 0 to 63. Validating this construct means establishing that higher numbers correspond with greater depression severity, or that scores above a certain threshold, such as 29 out of 63, correctly classify someone as having severe depression.
If our numbers do not mean what we think they mean, our analyses do not either
You may feel safe because you are using validated scales
Slide 4If our numbers don't mean what we think they mean, our analyses don't either. You might be thinking that construct validation is not a top concern in your work, because you, a principled scientist, are using validated scales. If so, you'd be wrong. Validity is not a property of an instrument.
Flake, 2022
Validity is a judgment, and it has to keep being made
"validity is not a binary property of an instrument, but instead a judgment made about the score interpretation based on a body of accumulating evidence that should continue to amass whenever the instrument is in use"
Slide 5One methods paper makes the point clearly. Validity is not a binary property of an instrument, but instead a judgment made about the score interpretation based on a body of accumulating evidence that should continue to amass whenever the instrument is in use. Ongoing validation is necessary because the same instrument can be used in different contexts or for different purposes, and evidence that the interpretation of scores generalizes to those new contexts is needed.
An instrument validated in one place carries no guarantee in another
Originally validated with 200 white women in one small city in America
What gives you confidence the numbers carry the same meaning in rural India?
Think critically about why you believe it will produce valid results in your context
Slide 6I won't argue that you must always start from scratch to validate the instruments you select. But it's important to think critically about why you believe an instrument will produce valid results in your context. If you are using an instrument originally validated with a sample of 200 white women in one small city in America, what gives you confidence that the numbers produced carry the same meaning in rural India?
Loevinger, 1957
Validation is a journey through three phases
Substantive: does it ask the right questions?
Structural: do the numbers behave?
External: does it match reality?
Slide 7Whether you're developing a new instrument or evaluating an existing one for your context, it helps to think about validation as a journey through three phases. Each phase answers a different question. The substantive phase asks whether the instrument captures the right content: what topics should be included, and how should items be worded so your target population understands them? The structural phase asks whether the numbers behave. And the external phase asks whether the scores match reality.
Phase 1: Substantive
Does the instrument capture the right content?
Does it cover the important aspects of the construct?
Are the questions understood by the target population?
Slide 8The first phase of construct validation asks whether your instrument captures the right content. Whether you're building something new or adapting an existing tool for a new context, there are two questions to answer: does the instrument cover the important aspects of the construct? And are the questions understood by the target population?
Content validity
All the relevant domains, and nothing unrelated
Start with a review of the literature and conversations with experts
And, critically, talk to members of your target population
Learn how they understand and describe the experience you want to measure
Slide 9An instrument has evidence of content validity when it assesses all of the conceptually relevant domains of a construct and excludes unrelated content. The process of establishing content validity typically starts with a review of the literature, conversations with experts, and, critically, talking to members of your target population to learn how they understand and describe the experience you're trying to measure. This last step matters more than you might think. Constructs don't always travel well across cultures.
Green et al., 2018
Asking women in rural Kenya to describe depression
Focus groups on what observable features characterize huzuni in Swahili
During pregnancy and after childbirth
Then co-examining the overlap with existing depression screening tools
Slide 10My colleagues and I saw this firsthand in a study in rural Kenya, where we examined how women understood depression in the context of pregnancy and childbirth. A screening tool built somewhere else carries somebody else's account of what depression looks like, so we started from the other end. We convened focus groups and asked women to describe what observable features characterize depression, huzuni in Swahili, during the perinatal period. We had also gone through 17 existing depression screening tools and put each of their symptom terms onto a card, in Swahili and in English. Then we co-examined the overlap, and the lack of overlap, between what the women described and what was on those cards.
Sorting the cover terms: yes, maybe, no, no opinion
Illustrative sorting of depression cover terms by focus groups, from Green et al. (2018). Reproduced from Chapter 9.
Slide 11This is what that looked like. Each card carries one term from those existing tools: hopeless, crying, worried, slowed down, blame self, loss of sleep. Once we had matched the terms the women listed against the cards, the ones left over went on the table, and the group sorted each into four columns: yes, this is a characteristic of huzuni, scored three; maybe, scored two; no, scored minus one; and no opinion, scored zero. What you're looking at is a group of women deciding, term by term, which of somebody else's symptoms belong in their setting.
Some terms landed, and some did not
Symptoms on standard screening tools that aligned with what women described
Others that did not resonate at all
And experiences the women named that no existing tool captured
Slide 12Some symptoms on standard screening tools aligned with what women described. Others didn't resonate at all, and the women identified experiences that no existing tool captured. A group of Kenyan mental health professionals then reviewed the results and offered feedback based on their local clinical expertise. Without this kind of ground-up process, we would have used an instrument that missed important aspects of the construct and included items that didn't belong.
An instrument can cover the right domains and still fail
"I feel I am being punished" reads differently by religious and cultural context
An item on sleep may not discriminate where sleep disruption is near-universal
Slide 13Even when an instrument covers the right domains, the items themselves might not work for your population. An item like, I feel I am being punished, might be interpreted very differently depending on religious and cultural context. An item about changes in sleep patterns might not discriminate well in a population where sleep disruption is near-universal due to living conditions rather than depression.
Cognitive interviewing
Ask people to read each item aloud
Have them describe, in their own words, what they think the item is asking
An item that is clear to the research team can mean something else entirely
Slide 14Cognitive interviewing is a structured way to catch these problems. You sit down with members of your target population, ask them to read each item aloud, and have them describe in their own words what they think the item is asking. You'd be surprised how often an item that seems perfectly clear to the research team means something entirely different to a respondent, or means nothing at all.
One dropped item is what this phase looks like in practice
The HAP team dropped one BDI-II item about sex for cultural reasons
The item was culturally inappropriate for the context
Phase 1 examines, questions and adapts an instrument to fit its setting
Slide 15The HAP team's decision to drop one item about sex for cultural reasons is a small example of what this phase produces. The item wasn't misunderstood; it was culturally inappropriate for the context. Phase 1 is about examining, questioning, and adapting an instrument so that its content fits the population and setting in which it will be used.
Phase 2: Structural
You have pilot data. Now what?
Slide 16You've collected pilot data. Two hundred people in your target population responded to every item on your questionnaire. Now what? Before you can claim these items measure depression, or whatever construct you're after, you need to answer three questions: are the items working? Do they hang together? And are they consistent?
Item analysis
An item everyone answers the same way tells you nothing
Plot the response distribution of every item
If 100% respond "strongly agree", the item has zero variance
Drop it, or run cognitive interviewing to modify it so responses vary
Slide 17A common initial practice is to plot the response distributions of each item. If you ask people to rate their agreement with a statement like, I feel sad, and 100 percent of people in your sample respond strongly agree, the item has zero variance. When all or nearly all of your sample responds the same way to an item, that item tells you nothing useful. The appropriate next step is to drop the item, or conduct additional cognitive interviewing to modify it in a way that will elicit variation in responses.
Three items, 100 people, a 4-point scale
Visual example of item analysis. Responses are plotted separately for cases and non-cases. Reproduced from Chapter 9.
Slide 18If you have data that let you plot response distributions by group, you can also examine the extent to which items distinguish between groups. Here is a hypothetical example where 100 people responded to three items, each measured on a four-point scale from never to often.
Only one of the three items is doing real work
item_1 has almost no variability. item_2 varies but does not separate cases from non-cases. item_3 varies and a larger proportion of cases endorsed it. Reproduced from Chapter 9.
Slide 19The first item has very little variability. Almost everyone responded never. This item does not tell us much, and you might decide to drop or improve it. The second item has more variability, but it does not distinguish between cases and non-cases. In combination with other variables it might be useful, so you might decide to keep it unless you need to trim the overall length of the questionnaire. The third item looks the most promising. It elicits variability in responses, and a larger proportion of the cases group endorsed the item.
In Closing
Phase 1 asks about content; Phase 2 asks about the numbers
Item analysis tells you whether each item works on its own
It cannot tell you whether the items work together
Slide 20Phase 1 asks about content, and Phase 2 asks about the numbers. Item analysis tells you whether individual items are working. But knowing that each item works on its own doesn't tell you whether they work together, and that is the question that comes next.