> read --source=primary _ Cognitive Complacency. >_ THINKING · PRAGMA · 2026
PRAGMA July 15, 2026 12 min read

Cognitive Complacency
Four famous studies. We read all four.

The recruiters

In 2022, Fabrizio Dell’Acqua, then a postdoctoral fellow at Harvard Business School, ran the experiment everyone now cites.

He recruited 181 professional HR recruiters and had each of them evaluate 44 résumés for a software engineering role — close to eight thousand evaluations, about five thousand of which survived his attention checks. He randomly assigned each recruiter a different quality of AI assistant and, importantly, told them which one they had. One group got a perfect predictor. One got a good one, around 85% accurate. One got a bad one, around 75%. One got nothing.

He had real ground truth to score against: administrative math scores from the OECD’s adult skills survey.

The finding that made the paper famous is in its title. Give people an AI that’s good enough to trust, and they stop thinking. They click faster, deliberate less, and let the machine’s answer become their answer. Dell’Acqua called it falling asleep at the wheel, and he borrowed the image on purpose — his paper opens on the federal investigation into Tesla’s Autopilot.

His explanation is not that people are lazy. It’s that they’re rational. Effort is expensive. If the machine is usually right, checking it is a bad trade. So people stop checking. The paper puts it precisely: as AI quality rises, “human workers are more likely to free-ride,” and the AI begins to substitute for judgment rather than augment it.

That’s a genuinely important idea. It’s also where the retelling starts drifting from the paper.

FIGURE 1The retelling, and the sourceSTUDYWHAT EVERYONE SAYSWHAT THE SOURCE SAYSHBSrecruitersBad AI beats good AI. The “perfect” AI was99% accurate.That gap is p = 0.52 — noise. The perfect AI beatevery group. No 99% is in the paper.HBS/BCGconsultantsAI lifts quality 40%, and tanks you outsidethe frontier.The 40% is the pre-review draft; peer review says19%. The 19 pt drop is real — from one task.MITessaysBrain scans prove ChatGPT causes brain rot.Not peer-reviewed. EEG, not scans. 18 peoplefinished. Authors ask you not to say “brain rot.”MicrosoftCHI ’25AI measurably reduces critical thinking.It’s a survey — “Self-Reported” is in the title.Half the headline misses its own p = 0.007 bar.Science2011The Google effect: we remember where, notwhat.Failed replication twice — the second using theoriginal author’s own corrections.Sources: Dell’Acqua 2022 (unpublished working paper) · Dell’Acqua et al., Organization Science 2026 · Kosmyna et al., arXiv:2506.08872 (preprint) ·Lee et al., CHI ’25 · Sparrow et al., Science 2011; Camerer et al. 2018; Vöhringer et al. 2020. >_

What the recruiter study does not say

Harvard’s own AI Institute published an article about this research. It’s good. It’s also wrong in two specific ways, and I’d rather show you than assert it.

The Institute article says the perfect AI had “approximately 99% accuracy.”

There is no 99% in Dell’Acqua’s paper. The perfect condition predicted candidate quality correctly — it’s a full-delegation baseline, deliberately unrealistic, used to bracket the experiment. The 99% is a gloss that appeared somewhere between the paper and the write-up. It’s now loose in the world.

The second one matters more. The popular version of this study — including Harvard’s — is that recruiters with the bad AI outperformed recruiters with the good AI. Vigilance beat quality. It’s a great line. It’s the whole reason the study gets cited.

Look at the actual table. Measured as a straight right-or-wrong count, the difference between the good-AI group and the bad-AI group carries a p-value of 0.52. That is not a marginal result. That is noise. On a finer ten-point accuracy scale it reaches p=0.06, which is suggestive and nothing more, and Dell’Acqua says so himself — he notes the effect is “more precisely estimated” on the 1–10 scale than the binary one, which is a careful researcher’s way of telling you the binary version didn’t hold.

And then there’s the part almost nobody repeats. The perfect-AI group performed better than every other group. It’s in the paper in plain text: they “performed better than any other group.” Every AI group beat the no-AI group. Nobody was made worse off by having a machine.

So “better AI makes people worse” — the thing this study is famous for proving — is contradicted at the top of its own quality range.

There’s one more crack. The effort finding is real on one measure and not on another. Recruiters with the bad AI spent about ten more minutes per résumé, roughly 50% over the control group, and that’s solid at p<0.01. But the study’s other effort measure, how many times people clicked to dig into a candidate, came back at p=0.29. Nothing. When you read that this study proved people “exert less effort” with good AI, understand that one of the two effort measures fired and the other didn’t.

None of this makes the paper bad. It makes it a paper. It is, to this day, an unpublished working paper — not peer-reviewed, not replicated, with its preregistration under embargo. It is a suggestive first look at a real phenomenon, and it is being cited like settled law.

The honest version of the finding is narrower and, I think, more interesting than the myth: mid-quality AI may be the worst partner you can have, because it’s good enough to trust and not good enough to be right. That’s a real claim. It’s defensible. It’s just not as loud.

The consultants

Dell’Acqua’s other paper is stronger, and it’s the one to lean on.

He and a team including Ethan Mollick, Karim Lakhani, Edward McFowland III, and others ran a preregistered experiment with 758 BCG consultants — around 7% of BCG’s individual-contributor workforce. Everyone was measured at baseline, then randomly assigned to work with no AI, with GPT-4, or with GPT-4 plus training on how to use it.

On tasks inside the model’s competence, the results were what you’d hope. Consultants finished 12.2% more tasks, moved 25.1% faster, and produced better work. The weakest performers improved the most.

Then they gave people one task engineered to sit outside what the model does well — a business case where the data the AI could most easily read pointed at the wrong answer.

Consultants without AI got it right 84.5% of the time. Consultants with AI got it right 60% to 70.6% of the time. That’s an average drop of 19 percentage points, and it is the single most important number in this entire literature. Not because AI is dangerous, but because the same tool that lifted people on one task sank them on the next one, and the people using it could not tell which task they were on.

That paper has since been peer-reviewed and published in Organization Science. Which is exactly why I want to flag two things about it.

First, if you’ve seen “40% higher quality” attached to this study — and you have, it’s everywhere — that figure is from the 2023 working paper. The peer-reviewed version reports 19% and 17.1% improvements on its scored measure and drops the 40% from the abstract. Peer review sanded it down. The internet kept the pre-review number.

Second, if you’ve heard the “centaurs and cyborgs” framing that came out of this research, note that the words centaur and cyborg do not appear in the published version at all. They were cut. They survive in a separate paper that is still, as of now, unpublished and under revision.

And the authors themselves put a fence around the 19-point finding that most people quote it without: it rests on one task. They wrote the limitation into the paper — “a single outside-the-frontier task — an inherent limitation.” It proves degradation can happen. It does not tell you how often real work sits outside the frontier, which is the number you’d actually want.

FIGURE 219percentage pointsHow much worse BCG consultants did with AI, on the onetask built to sit outside what the model is good at.Without AI: 84.5% correct. With it: 60–70.6%.19 ptDell’Acqua et al., Organization Science, 2026. n=758.One task — the authors flag this limitation themselves. >_

The brain scans that weren’t brain scans

In June 2025, a team at the MIT Media Lab led by Nataliya Kosmyna released a study called Your Brain on ChatGPT. It got enormous coverage. You’ve seen the headlines.

They put 54 people in three groups — write an essay with ChatGPT, with a search engine, or with nothing — wired them for EEG, and ran them across several months. The core result: the more external help you had, the weaker your measured brain connectivity, and the less you felt the essay was yours. The ChatGPT group struggled to quote work they had produced minutes earlier.

Here is what the coverage left out.

The study has never been peer-reviewed. It went up on arXiv as a preprint and, as far as I can determine, remains one. The authors said at release that “all the conclusions are to be treated with caution and as preliminary.”

Fifty-four people started. Eighteen finished the fourth session — the crossover session that produced the most-quoted findings. Eighteen, split across two conditions.

There are no brain scans. EEG measures electrical activity at the scalp. The authors say plainly that its spatial resolution can’t localize deep structures, and that fMRI is what they’d need.

And then the detail that should end most of the coverage: the authors published a list of words they are asking people not to use about their work. Not paraphrasing. Their FAQ says, in bold, don’t use “stupid,” “dumb,” “brain rot,” “harm,” “damage,” “brain scans,” “brain damage,” “LLMs make you stop thinking,” or “terrifying findings.” They wrote: “It does a huge disservice to this work, as we did not use this vocabulary in the paper, especially if you are a journalist reporting on it.”

They also noted, with what I read as some despair, that a lot of the coverage was written by people who fed their paper to an LLM and published the summary.

Sit with that for a second.

There is now also a formal methodological critique of the study — Stanković and colleagues, filed in 2026 — raising the sample size, the reproducibility of the analysis, the EEG methods, inconsistent reporting, and transparency. If you cite this study in public, assume a critical reader knows that comment exists.

FIGURE 35418People who started the MIT study, and people who finishedthe session that produced its most-quoted finding — splitacross two conditions.18 / 54Kosmyna et al., arXiv:2506.08872.Preprint; not peer-reviewed. >_

The study that says “self-reported” in its own title

The fourth pillar is a 2025 CHI paper from Hank Lee at Carnegie Mellon and six co-authors at Microsoft Research: The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers.

It’s peer-reviewed. It’s a good paper. Its central finding is genuinely useful: the more confidence you have in the AI, the less critical thinking you do (β=-0.69, p<0.001). And the nature of the work shifts — from gathering information to verifying it, from solving the problem to integrating the machine’s answer, from doing the task to supervising it. They call that last one stewardship, and it’s the right word.

Note what’s in the title, though. Self-reported. This is a survey of 319 people recruited on Prolific, asked to recall three times they used AI at work. There is no behavioral measurement anywhere in it. No task, no test, no observation. The authors chose that design for a defensible reason — they say critical thinking “as a pure mental phenomenon is difficult for people to self-observe” — but it means the study measures what people believe about their own thinking.

The authors are admirably direct about the limits. “Our analysis does not establish causation.” Participants “occasionally conflated reduced effort in using GenAI with reduced effort in critical thinking.” The sample skews young and technical.

And one more thing, which I only found by opening the paper: the authors set themselves a corrected significance threshold of p=0.007 because they were testing 98 predictors. The famous second half of their headline — that higher self-confidence means more critical thinking — comes in at p=0.026. It doesn’t clear their own bar. Half of the most-quoted sentence from this paper is weaker than the paper’s own standard.

The one that fell apart

For fifteen years, the go-to citation for “technology is changing our memory” has been Sparrow, Liu and Wegner, Google Effects on Memory, published in Science in 2011. When you expect to be able to look something up, you remember where it is instead of what it is. It’s a beautiful finding. It launched a thousand op-eds about digital amnesia.

In 2018, a large replication project published in Nature Human Behaviour tried to reproduce the study’s first experiment. It failed. No effect, with adequate power.

The original author objected that the replication had design flaws. So another team ran it again in 2020, preregistered, using her own suggested corrections. It failed again — with Bayesian evidence favoring the null result by better than five to one. Those authors also documented that the original paper’s reported degrees of freedom don’t work under either possible sample size, and that its headline result was computed on a post-hoc subset of four words.

Be precise here, because this is exactly the kind of thing that gets over-claimed in the other direction: only the priming experiment has been directly replicated and failed. The more famous “we remember where, not what” experiments haven’t had the same high-powered test either way. So the honest status is contested, not debunked.

But if you’re still citing the Google effect in 2026 without mentioning any of this, you are doing the thing this essay is about.

“Please do not use the words like ‘stupid’, ‘dumb’, ‘brainrot’, ‘harm’, ‘damage’… It does a huge disservice to thiswork, as we did not use this vocabulary in the paper.”THE MIT AUTHORS’ OWN FAQ · BRAINONLLM.COM >_

Notice what just happened

Four studies. Every one of them is quoted more confidently than its own authors quote it.

Harvard’s write-up is more certain than Harvard’s paper. The MIT study is described in language its authors formally asked people to stop using — by writers who, in some cases, had a language model read it for them. The Microsoft paper’s own title contains the caveat that most citations of it omit. The Science paper has failed replication twice and is still doing the rounds.

Nobody involved is dishonest. What happened is much more ordinary, and much more interesting.

Somebody read a summary and trusted it, because the summary was good enough. Then somebody summarized the summary. At every step the person doing the summarizing had a rational reason not to open the primary source: it’s long, the summary is usually right, and checking is expensive.

That is Parasuraman’s definition, executed exactly. Monitoring dropped below the rate the situation warranted. Failures got through that a person watching properly would have caught. And the reason monitoring dropped is that the system — Harvard, Science, MIT, a well-built language model — is reliable enough that checking feels like a waste of time.

The discourse about cognitive complacency is the single best available demonstration of cognitive complacency.

That’s not irony. It’s confirmation. This is the most-studied failure mode in human factors research, it has been sitting in Human Factors since 2010 with a clean operational definition, and it caught the people writing about it. It would catch you. It caught us — we came into this expecting the “bad AI beats good AI” result to be real, because we’d read it a dozen times, and it isn’t there.

What’s actually true

Strip out everything that doesn’t survive contact with the primary sources, and you’re left with a short list. It’s smaller than the headlines. It’s also solid enough to build on.

People under-monitor reliable automation, and the more reliable it is, the less they watch. That’s established, it’s decades deep, and it’s the one thing here that isn’t in dispute.

Good AI helps on work it’s good at, and hurts on work it isn’t — by about 19 points in the one careful test we have — and people can’t reliably tell which kind of work they’re doing. That’s peer-reviewed, preregistered, and n=758.

Trusting the tool more is associated with thinking about it less. That’s peer-reviewed and it’s self-reported, which means it’s real evidence about belief and thin evidence about behavior.

And there’s a suggestion — nothing more — that the most dangerous assistant isn’t the best one or the worst one, but the one in the middle. Good enough to trust. Not good enough to be right.

Everything else you’ve read about this is louder than the evidence underneath it.

Why we spent a week reading footnotes

We’re PRAGMA. We measure how often AI systems recommend a business when a customer asks. That’s the whole job.

Which means we’re in the business of being trusted about a number, and we’re acutely aware of what this literature says about numbers people trust: you’ll stop checking ours. Not because you’re careless — because checking is expensive and we’ll usually be right. That’s the trap. It’s the same trap, and there’s no version of our business that sits outside it.

So we don’t think the answer is to ask you to trust us harder. Dell’Acqua’s own current work is exploring something called controlled error injection — deliberately keeping a small amount of imperfection visible in a system so that the human stays awake. We think the publishable version of that idea is simpler. Show your work. Publish the ruler. Report what the evidence supports and not one inch more, especially when a bigger number was available and nobody would have checked.

We could have written the version of this essay everyone else writes. Harvard proved that bad AI beats good AI. MIT scanned people’s brains and found ChatGPT is rotting them. It would have gotten more attention.

It just wouldn’t have been true, and the whole point of a measurement company is that it doesn’t say things that aren’t true when nobody’s looking.

Sources. Parasuraman, R., & Manzey, D. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, 52(3), 381–410. · Dell'Acqua, F. (2022). Falling Asleep at the Wheel: Human/AI Collaboration in a Field Experiment on HR Recruiters. Working paper, Laboratory for Innovation Science, Harvard Business School. Unpublished; not peer-reviewed. · Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. (2023/2026). Navigating the Jagged Technological Frontier. HBS Working Paper 24-013; published in Organization Science (2026), DOI 10.1287/orsc.2025.21838. · Kosmyna, N., et al. (2025). Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. arXiv:2506.08872. Preprint; not peer-reviewed. Authors' FAQ at brainonllm.com/faq. See also Stanković et al. (2026), arXiv:2601.00856. · Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The Impact of Generative AI on Critical Thinking. CHI '25, DOI 10.1145/3706598.3713778. · Sparrow, B., Liu, J., & Wegner, D. M. (2011). Google Effects on Memory. Science, 333(6043), 776–778. Replication failed: Camerer et al. (2018), Nature Human Behaviour; and Vöhringer et al. (2020), PeerJ 8:e10325.