Indexes and scales Chapter 9: Measurement and Construct Validation — Video 4 https://ghrbook.com/videos/indexes-and-scales/ [Slide 1] Indexes and scales combine multiple items into a single score. The key difference between them lies in the relationship between items and constructs. [Slide 2] In an index, the items collectively define the construct. Add or remove an item and the construct itself changes. In a scale, the items reflect an underlying construct that exists independently of any single item. [Slide 3] That difference is easiest to see as arrows. On the left, the arrows run from the items up to the construct: the indicators establish it. On the right, they run from the construct down to the items: the indicators exist because of it. [Slide 4] An index combines items that collectively define the construct. The items don't reflect some underlying thing that exists independently; they constitute it. Wealth is a good example. A household's wealth is the sum of its assets: livestock, housing quality, appliances. Add more assets, and wealth increases by definition. The DHS wealth index uses household survey data on assets as a measure of household economic status. Asset variables include phones, televisions and cars, land ownership, and dwelling characteristics such as water and sanitation facilities, housing materials, persons sleeping per room, and cooking facilities. A household's wealth score is the weighted sum of the assets it owns, and in this case the weights are derived from the data using principal component analysis, though simpler approaches like equal weighting are also common. [Slide 5] Scales work differently. They combine items that reflect an underlying construct rather than define it. Depression manifests in symptoms, sadness, fatigue, loss of appetite, but the symptoms don't define depression; they're observable expressions of it. Two people with the same underlying depression severity might show different symptom patterns. Because scale items stem from a common, latent cause, they should be intercorrelated. [Slide 6] The modified inventory used in the HAP trial in India was designed to measure depression with 20 related items asking people to report on their experience of common symptoms. This is how those items were correlated in the trial data. Every square is one pair of items. [Slide 7] If these 20 items all reflect the same underlying construct, depression, then we'd expect them to be positively correlated with each other. And mostly, they are. The darker squares show that most item pairs correlate in the 0.2 to 0.3 range. That's a good sign: when someone endorses one symptom, they tend to endorse others too, which is what we'd expect if a common cause is driving the responses. But look at the appetite item. It's essentially uncorrelated with the tiredness item and only weakly correlated with everything else. That's a clue worth paying attention to. It suggests that in this population, changes in appetite may not travel with the other symptoms of depression the way the instrument's designers assumed. Keep it in mind. It's going to keep showing up as we dig deeper into this instrument's structure. [Slide 8] Once you've identified the items that will make up your scale, a key decision is how to combine them into a single score. Should feeling sad count the same as feeling suicidal? The most common approach is equal weighting. Responses to each item are scored 0 to 3 and simply summed to create a depression severity score ranging from 0 to 63. [Slide 9] Here is data from four people in the study. Each row is a participant, each column is an item, and the last column is the total. Each item contributes equally to that sum. No fancy math required. [Slide 10] But scales like this one could use optimal weighting instead. One alternative approach uses a factor model, where items more closely related to the construct are weighted more heavily. The model is called congeneric, which just means that each item is free to relate to the underlying construct with its own unique strength, rather than being forced to contribute equally. [Slide 11] This is the path diagram for that model. Depression severity sits at the top as a latent variable, and the loadings running down to each of the 20 items are uniquely estimated. [Slide 12] The item about feeling like a failure and the item about indecisiveness have the highest loading, 0.76, meaning these items are most closely related to the construct of depression severity. They contribute the most to the overall factor score. Changes in appetite has the weakest relationship to the construct, and contributes the least. [Slide 13] Does it matter which construction method you choose? This figure plots every participant's sum score, equal weighting, just add up the responses, on the horizontal axis, against their factor score, optimal weighting, on the vertical axis. If the two methods told the same story, every point would fall on a straight line. Many do, but some don't. Look at the two orange points. Person A and Person B both have a sum score of 11, but their factor scores diverge. Person A scores minus 0.63, while Person B scores minus 1.35. Same raw total, two different pictures of depression severity. Which one is right? [Slide 14] Why do they differ? This shows each person's item-level responses, and the answer jumps out. Person A endorsed several items with high factor loadings, the items that carry the most weight in the factor model. Person B put the same eleven points into only four items, and two of those points came from tiredness, which carries one of the lowest loadings in the model. Same sum score, but the pattern of endorsement matters when items contribute unequally. This is a validity issue. How you calculate the score changes what the number means. [Slide 15] So do these two people have the same level of depression severity, as their identical sum scores suggest? Or are they experiencing different levels, as the factor scores imply? There's no universally correct answer. For a long time, simplicity was a necessity. Practitioners administered questionnaires on paper and hand-scored them, so no fancy math was possible. Sum scores were not just convenient; they were the only practical option. And there's still value in that simplicity: sum scores are transparent, easy to calculate, and comparable across studies. Factor scores, by contrast, are often derived from sample-specific weights, making cross-study comparisons more complicated. [Slide 16] The Person A and Person B example illustrates a broader point. Two identical numbers can represent meaningfully different experiences. How you construct an indicator shapes what it captures, and what it obscures. Which brings us back to the core questions you'll face when measuring any construct: what items are essential to capture the construct? How should you combine multiple items into a single score? And how do you know that the resulting number actually represents the construct you intended to measure? That last question, how do you know, is the focus of everything that follows.