Video 6 of 7
Factor analysis and reliability
Two questions about a set of items, asked in order. Factor analysis asks whether they measure one thing or several, and reliability asks whether they measure it consistently — test-retest, internal consistency, inter-rater agreement, and responsiveness to change.
10:27 · 28 slides · printable slides · transcript
Slides
Printable deck →▶Transcript28 sections
Generated from the narration script. Plain text version.
1Item analysis tells you whether individual items are working. But knowing that each item works on its own doesn't tell you whether they work together to measure the construct of interest.
2Suppose you have 20 items that all show good variability and discriminate between groups. Are those 20 items measuring one thing, depression, or are some of them clustering around sadness while others cluster around physical symptoms? Are you measuring one construct or two?
3This is what factor analysis helps you figure out. It looks at patterns of correlation among your items and asks: can these be explained by a smaller number of underlying constructs?
4There are two main types. Exploratory factor analysis is used when you don't yet know the underlying structure, and you're asking the data to reveal it. Confirmatory factor analysis is used when you already have a hypothesis about the structure and want to test whether the data support it. As one methods paper puts it: researchers developing an entirely new scale should use exploratory factor analysis; confirmatory factor analysis should be used when researchers have strong a priori hypotheses about the factor pattern.
5To get a glimpse of what this means, imagine that, as members of the study team, we created the inventory items from scratch and wanted to examine the dimensionality of the items. Fit a two-factor model and you get a table of factor loadings. A factor loading tells you how closely each item is tied to a particular factor. Think of it as a correlation between the item and the underlying construct. Higher loadings mean the item is a better indicator of that factor. Lower loadings mean the item is doing its own thing, measuring something the factor doesn't capture well.
6What we see is a cluster of items that load strongly on the first factor, a cluster of items that load strongly on the second, and a few items, such as the appetite item, with no strong association to either factor.
7Remember that cold square in the correlation heatmap? Here it is again. The appetite item doesn't correlate well with the other items, and now the exploratory factor analysis confirms it doesn't load cleanly on either factor. The data are telling us something: changes in appetite may not be a strong marker of depression in this population.
8The software labels these factors with numbers, and it's up to us to interpret what they mean. My sense is that the first captures the affective dimension of depression, things like sadness and crying, whereas the second is about negative cognition, things like guilty feelings and self-dislike.
9Exploratory factor analysis is useful when you're exploring. But when prior research or theory gives you a reason to expect a particular structure, confirmatory factor analysis lets you test that expectation against your data. You specify the model in advance, how many factors and which items load on which factor, and then ask: does this structure fit? This inventory has decades of research behind it. While many research groups have proposed multi-factor solutions, the instrument is typically scored as a single factor, consistent with our exploratory results.
10The path diagram you saw earlier is the output of a confirmatory factor analysis: a one-factor congeneric model where each item's loading is freely estimated. The fact that this model fits the data well is evidence that a single depression severity factor adequately explains the pattern of correlations among the 20 items. If it didn't fit, we'd need to reconsider the structure. Perhaps depression in this population is better captured by two or more factors.
11Factor analysis asks about structure: do the items cluster together in the pattern you expect? Reliability asks a different question: how consistently do they measure? Every observed score consists of two parts, the true score, a person's actual level of the construct, and measurement error. Error can be random or systematic. Random error is noise, unpredictable variations that make measurements inconsistent, and reliability refers to this consistency. Systematic error is bias that pushes scores in a particular direction, and its main victim is validity.
12A bathroom scale is valid if it correctly measures your weight, and reliable if it gives the same reading when you step off and back on. A scale that consistently tells you 80 kilograms when you actually weigh 65 kilograms is reliable but not valid. Consistency is reliability, regardless of being right or wrong.
13There is no one test of an instrument's reliability, because variation in measurement can come from many different sources: items, time, raters, form. So we assess different aspects of reliability, including test-retest reliability, internal consistency reliability, and inter-rater reliability, to name a few.
14An instrument exhibits good test-retest reliability if it maintains roughly the same ordering between people when administered repeatedly under the same conditions, often benchmarked as a correlation of at least 0.70. But perfect reliability doesn't require identical scores at time 1 and time 2. Reliability is about preserving the rank order between people, not about exact agreement. Scores can be perfectly reliable even if every person's score shifts by the same amount.
15Each panel plots mock scores for 10 people measured four days apart. Points that fall on the diagonal line have identical scores at both times. Start here. Every point sits on the diagonal. The person who scored highest on day 1 scored highest on day 4. The person who scored lowest stayed lowest. The scores are perfectly reliable, the ordering is preserved, and they perfectly agree: the actual numbers are the same.
16Now look at this one. No point sits on the diagonal. Everyone's score went up. Maybe something happened between day 1 and day 4 that affected the whole group. But look at the pattern: the person who scored highest on day 1 still scored highest on day 4. The rank order is perfectly preserved, so reliability is unchanged. The scores don't agree at all, and they're still perfectly reliable. This is the key distinction: reliability is about who scores higher than whom, not about the exact numbers.
17This is what real data typically look like. The points don't fall on the diagonal, and the ordering isn't perfect, but the general pattern holds. People who scored higher on day 1 tend to score higher on day 4. This is acceptable reliability.
18And this is the one that should worry you. There's no pattern. Knowing someone's score on day 1 tells you almost nothing about their score on day 4.
19If your instrument looks like that last panel over a short, stable period, it's not measuring a consistent trait. It's mostly measuring noise. Test-retest is one aspect of reliability, and there are several more to work through.
20A key decision in assessing test-retest reliability is how long to wait in between administrations. If your instrument assesses depression symptoms in the past 2 weeks, you can't wait more than 2 weeks, because your respondent's frame of reference will change too much and symptoms of depression can come and go. On the other hand, you shouldn't re-administer the instrument the same day or even the next day, because your respondent will likely recall what they said in an effort to appear consistent.
21Repeated administrations aren't always feasible, so internal consistency reliability evaluates how closely scale items relate to each other in a single administration. If items aren't highly correlated, they probably aren't measuring the same latent construct. The most common index is Cronbach's alpha, which ranges from 0 to 1, higher being better. Values above 0.70 are generally considered acceptable, though this threshold is somewhat arbitrary. In this trial, the modified inventory showed strong internal consistency, well above that threshold.
22You might be thinking: didn't factor analysis already tell us the items hang together? It did, but the two tools play different roles. Factor analysis is diagnostic. It tells you which items belong together and how strongly each one relates to the construct, item by item. Alpha is a report card, a single number summarizing how consistently the set of items works as a unit. How much of the variability in total scores reflects the construct, and how much is noise?
23Many measurement researchers argue we should move beyond Cronbach's alpha. Alternatives like omega can provide more accurate estimates, especially when items don't contribute equally to the construct.
24Test-retest and internal consistency apply when the instrument is a questionnaire: a person answers questions, and you evaluate the consistency of their responses. But not all measurement in global health research involves questionnaires. Sometimes the instrument is a human observer, a clinician diagnosing a condition, a researcher coding interview transcripts, or a supervisor rating the quality of a counseling session. When measurement depends on human judgment, you need to ask: do different observers see the same thing? Inter-rater reliability measures the extent to which they agree.
25In this trial, researchers needed to assess whether lay counselors delivered therapy sessions with fidelity to the program design. They trained observers to rate audio recordings using a standardized instrument called the Q-HAP scale. The key question: when different observers listen to the same session, do they rate it the same way?
26For intervention studies, reliability isn't enough. You also need responsiveness, or sensitivity to change: the ability of an instrument to detect meaningful change over time when change has actually occurred. An instrument can be perfectly reliable yet insensitive to change if it suffers from floor effects, where scores cluster at the minimum, leaving no room to detect decline, or ceiling effects, where scores cluster at the maximum, leaving no room to detect improvement.
27Consider a depression intervention targeting people with mild symptoms. If your instrument was validated on severely depressed populations, participants may score near the floor at baseline, making it impossible to detect improvement. When selecting instruments, check whether they've demonstrated responsiveness in populations similar to yours. Effect sizes from prior intervention studies using the instrument can provide evidence of responsiveness.
28Let's take stock. At this point in the validation journey, we know the modified inventory asks the right questions for this population, that its items hang together in a coherent factor structure, and that it produces consistent scores. But we still haven't answered the harder question: do those scores actually mean what Patel and colleagues think they mean? A score of 35, does that really correspond to severe depression, or could it reflect something else entirely? That's what the third phase is about.