Does it match reality, and does it travel? Chapter 9: Measurement and Construct Validation — Video 7 https://ghrbook.com/videos/does-it-match-reality-and-does-it-travel/ [Slide 1] Phases 1 and 2 look inward, at the content of the instrument and the behavior of the items. Phase 3 looks outward. [Slide 2] You compare your scores to other measures, to groups that should differ, and to real-world outcomes. The question is no longer, do the items work? It is, do the scores mean what we think they mean? There are two broad ways to test this. The first asks whether your scores relate to other measures the way theory predicts. The second asks whether your scores predict something concrete in the real world. [Slide 3] Criterion validity is how well your scores relate to an external standard or outcome. Predictive validity asks whether scores today predict something meaningful in the future. A hiring test has good predictive validity if high scorers go on to perform well on the job. A depression screener has good predictive validity if elevated scores predict future clinical diagnosis or service use. [Slide 4] Diagnostic accuracy asks a more specific question: can the instrument correctly classify individuals? You compare your instrument, the index test, against a gold standard, the criterion, and ask how well it sorts people into the right categories. [Slide 5] For example, colleagues and I developed a perinatal depression screening questionnaire in Kenya and evaluated its diagnostic accuracy by comparing questionnaire scores to blinded clinical interviews. Of 193 women screened, clinical interviewers identified 10 who met diagnostic criteria for a major depressive episode. [Slide 6] This is the confusion matrix for the study. Down the side is what the screening questionnaire said. Across the top is what the clinical interview said. The four cells are the true positives, the false positives, the false negatives, and the true negatives. [Slide 7] A score of 13 or greater correctly identified 90 percent of true cases. It only missed 1 out of 10. This is known as sensitivity, or the true positive rate. The same cutoff also correctly identified 90 percent of non-cases, known as specificity, or the true negative rate. A score of 13 is the optimal cutoff that maximizes both. [Slide 8] The three phases apply to any validation effort. But much of global health research involves using instruments across cultural and linguistic boundaries, and that raises a distinct set of challenges. The question is whether the construct itself means the same thing in a different context, whether the language carries the same weight, and whether the response format makes sense to someone whose daily life looks nothing like the population the instrument was originally designed for. [Slide 9] Kohrt and colleagues propose six questions to ask when validating or selecting an instrument for cross-cultural use. They illustrate each one through their work adapting the Depression Self-Rating Scale and the Child PTSD Symptom Scale for conflict-affected youth, and the examples are worth learning from. I want to show you one of the six. [Slide 10] Technical equivalence means that the format of questions and response options works the same way across settings. This sounds mechanical, but the Kohrt study shows how easily it can go wrong. Children in the focus groups found the response scale confusing because it ran from zero, mostly, to two, never, starting with the most frequent and ending with the least. Every child who was asked said the opposite order made more sense. The team reversed the presentation while keeping the original numeric scoring. [Slide 11] Even more striking: the team tested three pictographic response scales. Water glasses, an abacus, and a dhoko-basket scale showing men carrying baskets of bricks. The water glasses and the abacus were generally understood. [Slide 12] The dhoko-basket scale backfired completely. Children consistently identified the empty basket, intended as not at all, with sadness and laziness. The boy had no bricks and would earn no money. The full basket, intended as extremely or always, was associated with happiness, because of the earning potential of carrying more bricks. The researchers intended the scale to represent burden; the children read it as prosperity. The scale was discarded. [Slide 13] This chapter covered a lot of ground. We started with a seemingly simple question, how many people died from COVID-19, and found that even counting deaths requires judgment calls, modeling assumptions, and imperfect data systems. From there we moved to constructs that are far harder to pin down: depression, empowerment, quality of life. Measuring these constructs well requires careful planning, and then it requires validation: whether the instrument captures the right content for your population, whether the items behave the way they should statistically, and whether the scores correspond to reality when compared against external criteria. [Slide 14] Measurement, Schmeasurement is the title of a great paper. The subtitle is, Questionable Measurement Practices and How to Avoid Them. The authors' thesis is that measurement is a critical part of science, but questionable measurement practices undermine the validity of many studies and ultimately slow the progress of science. [Slide 15] I'm persuaded by this argument, having been in the room when investigators have spent weeks thinking about research design only to uncritically throw in a bunch of measures at the end. The thinking is often, why not, we're going to the trouble of doing the study, let's measure everything. In situations like this, the validity of measurement is never at the forefront. Doing it right is hard, and hey, measurement, schmeasurement, right? [Slide 16] Wrong. Your study design can be flawless, your sample size enormous, your analysis plan pre-registered and bulletproof, but if your measure doesn't capture the construct you think it captures, none of it matters. Measurement isn't the last thing to figure out. It's one of the first.