Does it match reality, and does it travel?
Slide 1Phases 1 and 2 look inward, at the content of the instrument and the behavior of the items. Phase 3 looks outward.
Phase 3: External
The question is no longer whether the items work
Compare your scores to other measures
To groups that should already differ
And to outcomes in the real world
Slide 2You compare your scores to other measures, to groups that should differ, and to real-world outcomes. The question is no longer, do the items work? It is, do the scores mean what we think they mean? There are two broad ways to test this. The first asks whether your scores relate to other measures the way theory predicts. The second asks whether your scores predict something concrete in the real world.
Criterion validity
How well your scores relate to an external standard
Predictive validity: do scores today predict something meaningful later?
A hiring test, if high scorers go on to perform well on the job
A depression screener, if elevated scores predict later diagnosis or service use
Slide 3Criterion validity is how well your scores relate to an external standard or outcome. Predictive validity asks whether scores today predict something meaningful in the future. A hiring test has good predictive validity if high scorers go on to perform well on the job. A depression screener has good predictive validity if elevated scores predict future clinical diagnosis or service use.
Diagnostic accuracy
Can the instrument correctly classify individuals?
Compare your instrument, the index test, against a gold standard
Ask how well it sorts people into the right categories
Slide 4Diagnostic accuracy asks a more specific question: can the instrument correctly classify individuals? You compare your instrument, the index test, against a gold standard, the criterion, and ask how well it sorts people into the right categories.
Green et al., 2018
A perinatal depression screener, checked against clinical interviews
Developed in Kenya, evaluated against blinded clinical interviews
Of 193 women screened, interviewers identified 10 who met diagnostic criteria
Slide 5For example, colleagues and I developed a perinatal depression screening questionnaire in Kenya and evaluated its diagnostic accuracy by comparing questionnaire scores to blinded clinical interviews. Of 193 women screened, clinical interviewers identified 10 who met diagnostic criteria for a major depressive episode.
The confusion matrix for that study
Example confusion matrix. PDEPS score of 13 or greater against a clinical interview as the gold standard. Reproduced from Chapter 9.
Slide 6This is the confusion matrix for the study. Down the side is what the screening questionnaire said. Across the top is what the clinical interview said. The four cells are the true positives, the false positives, the false negatives, and the true negatives.
A cutoff of 13 maximizes both
Sensitivity
90% of true cases identified
Specificity
90% of non-cases identified
Slide 7A score of 13 or greater correctly identified 90 percent of true cases. It only missed 1 out of 10. This is known as sensitivity, or the true positive rate. The same cutoff also correctly identified 90 percent of non-cases, known as specificity, or the true negative rate. A score of 13 is the optimal cutoff that maximizes both.
Cross-cultural validity
The three phases apply anywhere; borders raise new problems
Does the construct itself mean the same thing in a different context?
Does the language carry the same weight?
Does the response format make sense to this population?
Slide 8The three phases apply to any validation effort. But much of global health research involves using instruments across cultural and linguistic boundaries, and that raises a distinct set of challenges. The question is whether the construct itself means the same thing in a different context, whether the language carries the same weight, and whether the response format makes sense to someone whose daily life looks nothing like the population the instrument was originally designed for.
Kohrt et al., 2011
Six questions for validating an instrument across cultures
Illustrated by adapting two scales for conflict-affected youth in Nepal
The Depression Self-Rating Scale, and the Child PTSD Symptom Scale
Slide 9Kohrt and colleagues propose six questions to ask when validating or selecting an instrument for cross-cultural use. They illustrate each one through their work adapting the Depression Self-Rating Scale and the Child PTSD Symptom Scale for conflict-affected youth, and the examples are worth learning from. I want to show you one of the six.
Technical equivalence
The format of questions and responses has to work here too
The response scale ran from 0 ("mostly") to 2 ("never")
Every child who was asked said the opposite order made more sense
The team reversed the presentation and kept the original numeric scoring
Slide 10Technical equivalence means that the format of questions and response options works the same way across settings. This sounds mechanical, but the Kohrt study shows how easily it can go wrong. Children in the focus groups found the response scale confusing because it ran from zero, mostly, to two, never, starting with the most frequent and ending with the least. Every child who was asked said the opposite order made more sense. The team reversed the presentation while keeping the original numeric scoring.
Three pictographic response scales, tested with children
Pictographic response scales tested for use with conflict-affected youth in Nepal. Water glasses, an abacus, and a dhoko-basket scale. Reproduced from Chapter 9.
Slide 11Even more striking: the team tested three pictographic response scales. Water glasses, an abacus, and a dhoko-basket scale showing men carrying baskets of bricks. The water glasses and the abacus were generally understood.
The researchers meant burden; the children read prosperity
The dhoko-basket scale was discarded because children interpreted the images in terms of earning potential rather than symptom burden. Reproduced from Chapter 9.
Slide 12The dhoko-basket scale backfired completely. Children consistently identified the empty basket, intended as not at all, with sadness and laziness. The boy had no bricks and would earn no money. The full basket, intended as extremely or always, was associated with happiness, because of the earning potential of carrying more bricks. The researchers intended the scale to represent burden; the children read it as prosperity. The scale was discarded.
What this chapter covered
Counting deaths takes judgment calls, modeling assumptions, imperfect data
Planning: conceptual models, DREAMY indicators, and how composites get built
Validation: the right content, items that behave, and scores that match reality
Slide 13This chapter covered a lot of ground. We started with a seemingly simple question, how many people died from COVID-19, and found that even counting deaths requires judgment calls, modeling assumptions, and imperfect data systems. From there we moved to constructs that are far harder to pin down: depression, empowerment, quality of life. Measuring these constructs well requires careful planning, and then it requires validation: whether the instrument captures the right content for your population, whether the items behave the way they should statistically, and whether the scores correspond to reality when compared against external criteria.
Flake and Fried, 2020
"Measurement Schmeasurement"
Subtitle: "Questionable Measurement Practices and How to Avoid Them"
Measurement is a critical part of science
Questionable practices undermine validity and slow the progress of science
Slide 14Measurement, Schmeasurement is the title of a great paper. The subtitle is, Questionable Measurement Practices and How to Avoid Them. The authors' thesis is that measurement is a critical part of science, but questionable measurement practices undermine the validity of many studies and ultimately slow the progress of science.
Measures thrown in at the end of a study design
Weeks spent on research design, then a bunch of measures added at the end
Why not, we're going to the trouble of doing the study, let's measure everything
The validity of measurement is never at the forefront
Slide 15I'm persuaded by this argument, having been in the room when investigators have spent weeks thinking about research design only to uncritically throw in a bunch of measures at the end. The thinking is often, why not, we're going to the trouble of doing the study, let's measure everything. In situations like this, the validity of measurement is never at the forefront. Doing it right is hard, and hey, measurement, schmeasurement, right?
In Closing
Measurement is not the last thing to figure out
Your design can be flawless, your sample enormous, your analysis pre-registered
If your measure does not capture the construct, none of it matters
It is one of the first things to figure out
Slide 16Wrong. Your study design can be flawless, your sample size enormous, your analysis plan pre-registered and bulletproof, but if your measure doesn't capture the construct you think it captures, none of it matters. Measurement isn't the last thing to figure out. It's one of the first.