Question
Cognitive load theory is mostly unfalsifiable, and we treat it like gospel anyway
The consensus is that cognitive load is a real, measurable constraint on human attention and working memory, and that reducing it makes everything better: interfaces, pedagogy, communication. Fair enough on the surface. But then you notice nobody can actually agree on what it measures or how to quantify it without circular reasoning.
The framework came out of Sweller's work in the 1980s on instructional design—genuine insight, to be fair. But somewhere between the lab and practice, it became a catch-all explanation. Is your form confusing? Cognitive load. Is the onboarding failing? Cognitive load. Is the policy document unreadable? Cognitive load. The term does work as a diagnosis, but so does "bad design." The difference is we think we understand cognitive load scientifically when mostly we're just redescribing something we already knew.
The predictive problem runs deeper. Cognitive load theory predicts that stripping context and reducing variables should improve learning or performance. Sometimes it does. Often it doesn't. Desai and others found that students with *more* scaffolding sometimes learned worse, not better—maybe because the scaffolding removed the productive struggle that actually cements understanding. Or take interface design: many of the most usable systems (Unix, Twitter's early design) violate every principle of load minimization. They work because they're consistent, not because they're simple.
I suspect what we're actually measuring isn't cognitive load but something looser: does this *feel* effortful? Does it match my expectations? Am I getting signal or noise? Those are real questions. But they're not physics, and calling them cognitive load doesn't make them more predictive than just... asking whether people are confused.
3 comments
Log in to comment.
What would actually falsify this for you? Because I'm wondering if you're pointing at a real problem—that CLT gets applied too loosely—or arguing the core mechanism doesn't exist.
The reason I ask: in my research with teachers, I see something consistent. When we actually measure working memory constraints in specific ways (digit span tasks before and after a lesson, error patterns on novel problems), the predictions *do* hold. Students with lower available capacity perform worse on transfer tasks. But the moment someone says "just make it simpler" without measuring anything, it falls apart. They strip out the thing that would've helped.
Your Unix example is interesting but maybe proves the point differently—Unix isn't simple, it's *consistent*. Consistency reduces the novel load on each interaction. That's still load theory, just applied correctly instead of as "fewer elements = better."
The Desai finding on scaffolding troubles me less now that I've read their full methodology. They're measuring something real (the hazard of removing productive struggle), but that's not actually contradicting CLT so much as showing you can reduce *extraneous* load while tanking *germane* load. Sweller called this out himself in the refinements.
So genuinely: do you think the effect disappears when you actually operationalize it rigorously, or does it just become a less useful explanatory framework than it looks?
I'm curious whether you're drawing a distinction between the *theory itself* being unfalsifiable versus the *way it gets applied* in practice being sloppy. Because those feel like different problems to me.
Sweller's original work on worked examples and element interactivity was pretty specific—testable claims about split-attention effects in particular learning domains. The theory itself isn't unfalsifiable; people have run proper experiments, found null results, found reversals depending on learner expertise. That's messy but it's not non-falsifiable. Where I think you're right is that somewhere between "worked examples help novices learn procedural skills" and "simplify everything to reduce cognitive load" we stopped doing the hard work of replication and context-sensitivity.
But here's what I'm less sure about: when you say the productive struggle stuff and the consistency-over-simplicity examples show cognitive load theory *fails*, aren't those just showing it's incomplete or domain-specific rather than wrong? The Desai finding on scaffolding and productive struggle—I'd want to know if they're actually contradicting CLT or just adding the constraint that *some* struggle is necessary for retention. Unix and early Twitter work because they're consistent, sure, but do people experience them as lower or higher load once they've learned the model? That's a different question than whether they're "simple."
What would actually convince you the theory has predictive power, or would you count it as unfalsifiable if it can accommodate both "more structure helps" and "less structure helps" depending on context?
I taught high school English for seven years, and I saw this play out constantly. Every few years there'd be a new pedagogical push—usually from admin who'd read one book—and it was always framed as reducing cognitive load. Shorter assignments. Fewer steps. Clearer rubrics. Strip away the ambiguity.
The problem was that some of my weakest students got *worse* under this approach. They needed the productive struggle, yeah, but also—and this matters—they needed to see an adult model the struggle. When I showed them my messy thinking about a text, circled back, changed my mind, that was apparently too "cognitively demanding." But it was also the only time some of them figured out that confusion wasn't a sign they were dumb. The streamlined version just made them feel stupid faster.
I'm also skeptical of the Unix/Twitter example in your post, though I get the instinct. Those systems work for experienced users who've already paid the upfront cost. They're not actually simpler—they're just consistent in a way that lets you build a mental model. That's different from saying cognitive load theory fails. It's more that we conflate "simple" with "intuitive to beginners" when they're not the same thing. But you're right that we throw the term around without distinguishing between these cases, and that sloppiness means we end up designing for the wrong problem half the time.