Even counting the dead takes judgment
Slide 1Measurement is one of the hardest parts of science, and when it is done poorly it has implications for the validity of our inferences. The more you learn about a discipline or an applied problem, the deeper you'll travel down the measurement rabbit hole. Debates about definitions and data are unavoidable. Even something as seemingly black and white as death is hard to measure completely, consistently, and correctly.
Measurement is one of the hardest parts of science
Done poorly, it has implications for the validity of our inferences
The deeper you go in a field, the deeper you travel down the rabbit hole
Debates about definitions and data are unavoidable
Slide 2Take the COVID-19 pandemic as an example. There are several challenges to counting the number of deaths from a disease like this one. First, some countries don't have robust civil registration and vital statistics systems that register all births, deaths, and causes of deaths. Many people die without any record of having lived. Second, an unknown number of deaths that should be attributed to COVID-19 are not, which means we can only count confirmed deaths.
Every death certificate asks for an immediate cause and any underlying causes.
Did someone die with COVID-19, or from it?
Cause of death is a judgment
The certifier makes the call
Slide 3There is also the challenge of defining what counts as a death from this disease. Did someone die with COVID-19, or from COVID-19? All cause of death determinations are judgment calls, and these are no different. When someone dies in the U.S., a physician, medical examiner, or coroner completes a death certificate and reports the death to a national registry. States are encouraged to use the Standardized Certificate of Death, which asks the certifier to list an immediate cause of death and any underlying causes that contributed to a person's death. COVID-19 is often an underlying cause, leading to pneumonia, and it's up to the certifier to make this determination. In many cases that determination is hard, because the disease can lead to a cascade of health problems in the short and long term.
The proposed fix
Excess mortality has challenges of its own
Deaths from any cause above and beyond historical trends for the period
Counts confirmed deaths, missing deaths, and deaths from other causes
Estimated with statistical models, so different inputs give different estimates
Slide 4Given those challenges, some scholars and policymakers prefer the metric of excess mortality, which represents the number of deaths from any cause that are above and beyond historical mortality trends for the period. Excess mortality counts confirmed deaths from the disease, missing deaths from the disease, and deaths from any other causes that might be higher because of the pandemic. But the quantification of excess mortality has its own challenges. Among them is that excess mortality is estimated with statistical models, so differences in input data and modeling approaches can lead to different estimates.
Measurement problems compromise the validity of your conclusions
No matter how rigorous your study design
The goal here is frameworks for thinking hard about measurement
Slide 5My goal here is not to depress you or make you question if we can ever truly measure anything. My goal is to convince you of the importance of thinking hard about measurement, and to give you frameworks for doing so. This matters because measurement problems can compromise the validity of your conclusions, no matter how rigorous your study design.
Construct validity
We rarely care about the numbers themselves
Measurement is assigning numbers to things we observe
We care about what the numbers represent: depression, poverty, health, quality of life
Slide 6At its core, measurement is about assigning numbers to things we observe. But we rarely care about the numbers themselves. We care about what the numbers represent: depression, poverty, health, quality of life. Construct validity is the extent to which your measures actually capture the concepts you intend to study.
Some constructs are closer to their measures than others
Inferred
Known only from observable indicators
Slide 7For some constructs, like height or hemoglobin levels, the connection between concept and measure is relatively direct. For others, abstract concepts like depression or empowerment that we can only infer from observable indicators, establishing construct validity requires more work.
Patel et al., 2017
The Healthy Activity Program trial
A randomized controlled trial in India
Lay counsellor-delivered psychological treatment for severe depression
Depression measured with the Beck Depression Inventory, item scores summed
Slide 8Think about the Healthy Activity Program trial, which you met in the statistical inference chapter. It was a randomized controlled trial in India, testing the efficacy of a lay counsellor-delivered psychological treatment for severe depression. To measure depression, the researchers asked participants to respond to statements from a questionnaire called the Beck Depression Inventory, assigned numbers to each response, and calculated a total score.
The researchers did not care about the scores
They cared about depression
An unobservable state, inferred from questionnaire responses
The instrument was their best attempt to quantify a condition they could not see
Slide 9But here's the thing. The researchers didn't really care about those scores. They cared about depression, an unobservable state they could only infer from the questionnaire responses. The inventory was their best attempt to quantify a complex condition they couldn't directly observe.
If it does not capture depression, the trial tells us nothing
It might measure something else
It might miss important dimensions of the experience
Either way, a perfectly executed trial says nothing about whether the treatment helps
Slide 10If it doesn't actually capture depression, if it measures something else, or misses important dimensions of the experience, then even a perfectly executed trial tells us nothing about whether the treatment helps.
Getting measurement right takes two halves
Planning
Conceptual models: what to measure
Terminology and good indicators
How composites are constructed
Validation
Do the measures capture the construct?
Evaluating an instrument you did not build
Or one you are adapting for a new setting
Slide 11The rest of this chapter is about getting measurement right, and it comes in two halves. We start with planning: using conceptual models to identify what to measure, the terminology of measurement, criteria for selecting good indicators, and how composite measures are constructed. Then we turn to validation: how to evaluate whether your measures actually capture the constructs you intend to study.
A shared vocabulary
Four terms, each building on the one before it
Construct: the abstract concept you want to study
Outcome: the change in a construct you hope to observe
Indicator: a measurable metric representing that change
Instrument: the tool used to collect the indicator data
Slide 12Before diving into the specifics of measurement, we need a shared vocabulary. Four terms form a conceptual hierarchy, each building on the one before it. A construct is the abstract concept you want to study. An outcome is the change in a construct you hope to observe. An indicator is a specific, measurable metric representing that change. And an instrument is the tool used to collect the indicator data.
Glennerster and Takavarasha, 2013
The same four terms, on one worked example
Outcome: decreased depression severity
Indicator: depression severity score
Instrument: a questionnaire of items about symptoms, scored to a total
Slide 13Think of it as moving from the abstract to the concrete. You start with a concept in your head, define what change you're looking for, identify a measurable proxy for that change, and select or develop a tool to capture it. In this trial, the construct was depression. The outcome was decreased depression severity. The indicator was a depression severity score. And the instrument was a depression scale, made up of questions about symptoms of depression, used to calculate that score.
Instruments take many forms
Surveys, questionnaires, clinical assessments
Environmental sensors, anthropometric measures, blood tests
Imaging, satellite imagery, and the list goes on
Slide 14Once you've identified your indicators, you need a way to measure them. Instruments are the tools used to collect indicator data, and they can take many forms: surveys, questionnaires, clinical assessments, environmental sensors, anthropometric measures, blood tests, imaging, satellite imagery, and the list goes on.
The two instruments used in the trial
Instruments used in Patel et al. (2017). Source: ghr.link/ins. Reproduced from Chapter 9.
Slide 15The authors measured depression with two instruments: the twenty-one item Beck Depression Inventory version II, and the nine-item Patient Health Questionnaire. Responses to each item on the inventory are scored on a scale of zero to three and summed to create an overall depression severity score that can range from zero to sixty-three, where higher scores indicate more severe depression.
In Closing
Construct validity is the question the rest of the chapter answers
Do your measures actually capture the concepts you intend to study?
Planning comes first, and it starts with a conceptual model
Slide 16Construct validity is the extent to which your measures actually capture the concepts you intend to study, and everything that follows in this chapter is an answer to that question. Planning comes first, and it starts with a conceptual model that tells you what to measure.