Factor analysis and reliability
Slide 1Item analysis tells you whether individual items are working. But knowing that each item works on its own doesn't tell you whether they work together to measure the construct of interest.
Are you measuring one construct, or two?
Suppose 20 items all show good variability and discriminate between groups
Some may cluster around sadness, others around physical symptoms
Knowing each item works alone says nothing about whether they work together
Slide 2Suppose you have 20 items that all show good variability and discriminate between groups. Are those 20 items measuring one thing, depression, or are some of them clustering around sadness while others cluster around physical symptoms? Are you measuring one construct or two?
Factor analysis
Can these correlations be explained by fewer underlying constructs?
It looks at the patterns of correlation among your items
And asks whether a smaller number of latent variables accounts for them
Slide 3This is what factor analysis helps you figure out. It looks at patterns of correlation among your items and asks: can these be explained by a smaller number of underlying constructs?
Flora, 2017
Two types, for two situations
Exploratory (EFA)
You do not know the structure
You are asking the data to reveal it
Use it for an entirely new scale
Confirmatory (CFA)
You have a hypothesis to test
You specify the model in advance
Use it with strong a priori hypotheses
Slide 4There are two main types. Exploratory factor analysis is used when you don't yet know the underlying structure, and you're asking the data to reveal it. Confirmatory factor analysis is used when you already have a hypothesis about the structure and want to test whether the data support it. As one methods paper puts it: researchers developing an entirely new scale should use exploratory factor analysis; confirmatory factor analysis should be used when researchers have strong a priori hypotheses about the factor pattern.
EFA example
A factor loading is how closely an item is tied to a factor
Think of it as a correlation between the item and the underlying construct
Higher loadings mean the item is a better indicator of that factor
Lower loadings mean the item is doing its own thing
Slide 5To get a glimpse of what this means, imagine that, as members of the study team, we created the inventory items from scratch and wanted to examine the dimensionality of the items. Fit a two-factor model and you get a table of factor loadings. A factor loading tells you how closely each item is tied to a particular factor. Think of it as a correlation between the item and the underlying construct. Higher loadings mean the item is a better indicator of that factor. Lower loadings mean the item is doing its own thing, measuring something the factor doesn't capture well.
Two clusters, and a few items belonging to neither
Table 9.3. Factor loadings for BDI-II items in a 2-factor exploratory factor analysis, HAP trial control arm at the 3-month endline. Shown in two columns to fit; loadings below 0.40 are not presented.
Slide 6What we see is a cluster of items that load strongly on the first factor, a cluster of items that load strongly on the second, and a few items, such as the appetite item, with no strong association to either factor.
Remember that cold square in the correlation heatmap
BDI-II correlations in the HAP trial. The appetite item correlates weakly with everything else, and the EFA confirms it does not load cleanly on either factor. Reproduced from Chapter 9.
Slide 7Remember that cold square in the correlation heatmap? Here it is again. The appetite item doesn't correlate well with the other items, and now the exploratory factor analysis confirms it doesn't load cleanly on either factor. The data are telling us something: changes in appetite may not be a strong marker of depression in this population.
The software names the factors; you have to interpret them
Table 9.3 again. Factor 1 gathers low energy, sadness, loss of interest, tiredness. Factor 2 gathers failure, guilt, suicidal thoughts, worthlessness.
Slide 8The software labels these factors with numbers, and it's up to us to interpret what they mean. My sense is that the first captures the affective dimension of depression, things like sadness and crying, whereas the second is about negative cognition, things like guilty feelings and self-dislike.
CFA
Does the structure you expected actually fit?
Specify in advance how many factors, and which items load on which
Many groups have proposed multi-factor solutions for the BDI-II
The instrument is typically scored as a single factor
Slide 9Exploratory factor analysis is useful when you're exploring. But when prior research or theory gives you a reason to expect a particular structure, confirmatory factor analysis lets you test that expectation against your data. You specify the model in advance, how many factors and which items load on which factor, and then ask: does this structure fit? This inventory has decades of research behind it. While many research groups have proposed multi-factor solutions, the instrument is typically scored as a single factor, consistent with our exploratory results.
This path diagram is the output of a CFA
A one-factor congeneric model where each item's loading is freely estimated. Fit to data from the HAP trial control arm at the 3-month endline. Reproduced from Chapter 9.
Slide 10The path diagram you saw earlier is the output of a confirmatory factor analysis: a one-factor congeneric model where each item's loading is freely estimated. The fact that this model fits the data well is evidence that a single depression severity factor adequately explains the pattern of correlations among the 20 items. If it didn't fit, we'd need to reconsider the structure. Perhaps depression in this population is better captured by two or more factors.
Reliability asks a different question: how consistently do they measure?
Every observed score has two parts: the true score, and measurement error
Random error is noise, and reliability refers to this consistency
Systematic error is bias, and its main victim is validity
Slide 11Factor analysis asks about structure: do the items cluster together in the pattern you expect? Reliability asks a different question: how consistently do they measure? Every observed score consists of two parts, the true score, a person's actual level of the construct, and measurement error. Error can be random or systematic. Random error is noise, unpredictable variations that make measurements inconsistent, and reliability refers to this consistency. Systematic error is bias that pushes scores in a particular direction, and its main victim is validity.
A bathroom scale can be one without the other
Valid
It correctly measures your weight
Reliable
Same reading when you step off and back on
Even if it reads 80 kg when you weigh 65 kg
Slide 12A bathroom scale is valid if it correctly measures your weight, and reliable if it gives the same reading when you step off and back on. A scale that consistently tells you 80 kilograms when you actually weigh 65 kilograms is reliable but not valid. Consistency is reliability, regardless of being right or wrong.
There is no one test of an instrument's reliability
Variation in measurement comes from items, time, raters, form
So we assess different aspects: test-retest, internal consistency, inter-rater
Slide 13There is no one test of an instrument's reliability, because variation in measurement can come from many different sources: items, time, raters, form. So we assess different aspects of reliability, including test-retest reliability, internal consistency reliability, and inter-rater reliability, to name a few.
Test-retest
The same ordering between people, on repeated administration
Often benchmarked as a correlation of at least 0.70
Perfect reliability does not require identical scores at time 1 and time 2
Scores can be perfectly reliable even if everyone's score shifts by the same amount
Slide 14An instrument exhibits good test-retest reliability if it maintains roughly the same ordering between people when administered repeatedly under the same conditions, often benchmarked as a correlation of at least 0.70. But perfect reliability doesn't require identical scores at time 1 and time 2. Reliability is about preserving the rank order between people, not about exact agreement. Scores can be perfectly reliable even if every person's score shifts by the same amount.
Panel 1
Perfect reliability, perfect agreement
Mock BDI-II scores for 10 people measured four days apart. Points on the diagonal have identical scores at both times. Reproduced from Chapter 9.
Slide 15Each panel plots mock scores for 10 people measured four days apart. Points that fall on the diagonal line have identical scores at both times. Start here. Every point sits on the diagonal. The person who scored highest on day 1 scored highest on day 4. The person who scored lowest stayed lowest. The scores are perfectly reliable, the ordering is preserved, and they perfectly agree: the actual numbers are the same.
Panel 2
Perfect reliability, no agreement
Every score went up between day 1 and day 4, and the rank order is perfectly preserved. Reproduced from Chapter 9.
Slide 16Now look at this one. No point sits on the diagonal. Everyone's score went up. Maybe something happened between day 1 and day 4 that affected the whole group. But look at the pattern: the person who scored highest on day 1 still scored highest on day 4. The rank order is perfectly preserved, so reliability is unchanged. The scores don't agree at all, and they're still perfectly reliable. This is the key distinction: reliability is about who scores higher than whom, not about the exact numbers.
Panel 3
Acceptable reliability, low agreement
What real data typically look like: the ordering is imperfect, but the general pattern holds. Reproduced from Chapter 9.
Slide 17This is what real data typically look like. The points don't fall on the diagonal, and the ordering isn't perfect, but the general pattern holds. People who scored higher on day 1 tend to score higher on day 4. This is acceptable reliability.
Panel 4
Chance reliability and agreement
Knowing someone's score on day 1 tells you almost nothing about their score on day 4. Reproduced from Chapter 9.
Slide 18And this is the one that should worry you. There's no pattern. Knowing someone's score on day 1 tells you almost nothing about their score on day 4.
Reliability is about who scores higher than whom
Not about whether the exact numbers agree
An instrument that looks like Panel 4 over a short, stable period is measuring noise
Test-retest is one aspect of reliability; there are several more
Slide 19If your instrument looks like that last panel over a short, stable period, it's not measuring a consistent trait. It's mostly measuring noise. Test-retest is one aspect of reliability, and there are several more to work through.
Test-retest period
How long should you wait between administrations?
If the instrument asks about the past 2 weeks, you cannot wait more than 2 weeks
The respondent's frame of reference changes, and symptoms come and go
Wait too little and they recall what they said, trying to appear consistent
Slide 20A key decision in assessing test-retest reliability is how long to wait in between administrations. If your instrument assesses depression symptoms in the past 2 weeks, you can't wait more than 2 weeks, because your respondent's frame of reference will change too much and symptoms of depression can come and go. On the other hand, you shouldn't re-administer the instrument the same day or even the next day, because your respondent will likely recall what they said in an effort to appear consistent.
Internal consistency
How closely items relate to each other in one administration
Repeated administrations are not always feasible
Items with weak correlations probably measure different constructs
Cronbach's alpha runs from 0 to 1, and above 0.70 is generally acceptable
Slide 21Repeated administrations aren't always feasible, so internal consistency reliability evaluates how closely scale items relate to each other in a single administration. If items aren't highly correlated, they probably aren't measuring the same latent construct. The most common index is Cronbach's alpha, which ranges from 0 to 1, higher being better. Values above 0.70 are generally considered acceptable, though this threshold is somewhat arbitrary. In this trial, the modified inventory showed strong internal consistency, well above that threshold.
Two tools, two different jobs
Factor analysis
Which items belong together
And how strongly each relates
Cronbach's alpha
One number for the set as a unit
How much is construct, how much is noise
Slide 22You might be thinking: didn't factor analysis already tell us the items hang together? It did, but the two tools play different roles. Factor analysis is diagnostic. It tells you which items belong together and how strongly each one relates to the construct, item by item. Alpha is a report card, a single number summarizing how consistently the set of items works as a unit. How much of the variability in total scores reflects the construct, and how much is noise?
McNeish, 2018
Many measurement researchers argue for moving beyond alpha
The 0.70 threshold is somewhat arbitrary
Alternatives like omega can give more accurate estimates
Especially when items do not contribute equally to the construct
Slide 23Many measurement researchers argue we should move beyond Cronbach's alpha. Alternatives like omega can provide more accurate estimates, especially when items don't contribute equally to the construct.
Inter-rater reliability
Sometimes the instrument is a human observer
A clinician diagnosing, a researcher coding transcripts, a supervisor rating a session
When measurement depends on human judgment, do different observers see the same thing?
Slide 24Test-retest and internal consistency apply when the instrument is a questionnaire: a person answers questions, and you evaluate the consistency of their responses. But not all measurement in global health research involves questionnaires. Sometimes the instrument is a human observer, a clinician diagnosing a condition, a researcher coding interview transcripts, or a supervisor rating the quality of a counseling session. When measurement depends on human judgment, you need to ask: do different observers see the same thing? Inter-rater reliability measures the extent to which they agree.
Rating recorded counseling sessions against a scale
Quality of the Healthy Activity Programme instrument. Observers were trained to rate audio recordings of therapy sessions. Reproduced from Chapter 9.
Slide 25In this trial, researchers needed to assess whether lay counselors delivered therapy sessions with fidelity to the program design. They trained observers to rate audio recordings using a standardized instrument called the Q-HAP scale. The key question: when different observers listen to the same session, do they rate it the same way?
Responsiveness
An instrument can be reliable and still miss the change
Floor effects: scores cluster at the minimum, leaving no room to detect decline
Ceiling effects: scores cluster at the maximum, leaving no room to detect improvement
Responsiveness is the ability to detect change when change has actually occurred
Slide 26For intervention studies, reliability isn't enough. You also need responsiveness, or sensitivity to change: the ability of an instrument to detect meaningful change over time when change has actually occurred. An instrument can be perfectly reliable yet insensitive to change if it suffers from floor effects, where scores cluster at the minimum, leaving no room to detect decline, or ceiling effects, where scores cluster at the maximum, leaving no room to detect improvement.
Check that it has been responsive in a population like yours
An intervention for mild symptoms, an instrument validated on severe cases
Participants score near the floor at baseline, so improvement cannot show up
Effect sizes from prior intervention studies are evidence of responsiveness
Slide 27Consider a depression intervention targeting people with mild symptoms. If your instrument was validated on severely depressed populations, participants may score near the floor at baseline, making it impossible to detect improvement. When selecting instruments, check whether they've demonstrated responsiveness in populations similar to yours. Effect sizes from prior intervention studies using the instrument can provide evidence of responsiveness.
Taking stock
The harder question is still open
Phase 1: the instrument asks the right questions for this population
Phase 2: the items hang together and the scores are consistent
But a score of 35 - does it really correspond to severe depression?
Slide 28Let's take stock. At this point in the validation journey, we know the modified inventory asks the right questions for this population, that its items hang together in a coherent factor structure, and that it produces consistent scores. But we still haven't answered the harder question: do those scores actually mean what Patel and colleagues think they mean? A score of 35, does that really correspond to severe depression, or could it reflect something else entirely? That's what the third phase is about.