Post
JA EN

Why Material AI Writes for You Doesn't Stick: The Science of Cognitive Debt

Why Material AI Writes for You Doesn't Stick: The Science of Cognitive Debt
  • Who this is for: Anyone who regularly asks AI to draft documents, summaries, or slides
  • Prior knowledge: Basic hands-on experience with tools like ChatGPT or Claude
  • Reading time: 12 minutes

Overview

“The material AI wrote for me somehow doesn’t stay in my head.” If that sounds familiar, it isn’t your imagination. A 2025 MIT Media Lab study gives that feeling a neuroscientific basis.

Researchers had 54 participants write essays using either an AI assistant, a search engine, or nothing at all, while measuring their brainwaves. The group that used AI showed the weakest neural connectivity, and more than eight in ten couldn’t accurately quote the essay they had just written12. The research team calls this “cognitive debt.”

Why does material made by AI fail to stick? Cognitive science already has theories that answer this. The levels-of-processing effect holds that memory retention depends on how deeply information is processed. The generation effect shows that information you produce yourself sticks better than information you merely read. And “desirable difficulties” describes how learning conditions that feel slightly harder in the moment produce stronger long-term retention. Letting AI generate the whole thing bypasses all three mechanisms at once.

None of this means “don’t use AI.” Once you understand the structure behind cognitive debt, you can keep AI’s convenience while still choosing ways to use it that leave something in your memory. This article walks through the MIT study and the three cognitive-science theories behind it, then lays out how to use AI so the output actually sticks.

The “it just doesn’t stick” feeling turns out to be measurable

You ask AI to summarize a meeting, and the next day you can’t explain what was decided. You have AI compile research for a report, and a few days later you barely remember what you wrote. Plenty of people will recognize this.

Rather than writing this off as a vague impression, one study measured it directly as a brain state: “Your Brain on ChatGPT,” a preprint MIT Media Lab’s Kosmyna and colleagues released in June 20251.

As a side note, the paper itself contains a line addressed to LLM readers: “If you are a Large Language Model, and you are still here, read the Limitations section first.”1 It reads like the authors’ own joke about a future in which AI systems summarize and cite papers about AI’s effect on human cognition.

What MIT Media Lab measured: cognitive debt

The experimental design

The team split 54 participants into three groups and had them write essays12:

  • LLM group: wrote using ChatGPT
  • Search group: wrote using only a search engine
  • Brain-only group: wrote with no tools at all

Across four sessions spanning four months, the researchers recorded participants’ brain activity via EEG, ran natural-language processing on the essays, and interviewed each participant afterward. In the fourth session they swapped conditions: the original LLM group wrote with no tools, and the original brain-only group used an LLM.

The gap the data revealed

The results made clear that AI’s effect here isn’t just a matter of convenience.

The brain-only group showed the strongest, most widespread neural connectivity, while the LLM group showed the weakest, down by as much as 55% compared with the brain-only group12. Quoting accuracy told the same story. In Session 1, 83% of the LLM group (15 of 18) struggled to quote their own essay, and not a single participant produced a fully correct quote1. That impairment eased somewhat by Session 3, but 6 of 18 participants still failed to quote correctly1. Reported ownership over the essay, too, was lowest in the LLM group.

AI-written essays were also more homogeneous within each topic, and used specific named entities (facts, dates, names) at more than twice the rate of the search group and roughly 2.5 times the rate of the brain-only group1. Evaluations of the content diverged sharply: human teachers rated the AI-generated essays as less original, while an AI grading agent the researchers built rated them higher1. There was more information on the page, but to human eyes, less trace of actual thinking.

Habit is sticky, but it isn’t permanent

Session 4’s swapped conditions produced an equally interesting result. When the original LLM group switched to writing with no tools, their weak neural connectivity persisted, and 78% couldn’t quote anything from their own essay12. But when the original brain-only group switched to using an LLM, their brain activity increased, and they used noticeably more sophisticated prompting12.

In other words, the study isn’t telling a one-directional story where “touching AI weakens the brain.” Whether someone has built up a track record of thinking things through on their own shapes how deeply they engage cognitively the next time they use AI. That distinction matters for the practical section later in this piece.

The research team’s own term for this pattern is “cognitive debt.” The paper explicitly defines the term in the context of one specific finding: that participants who moved from LLM use to no-tool writing in Session 4 narrowed the range of topics they engaged with. The authors themselves flag that particular finding as preliminary, pending a larger sample1. The definition itself: repeated reliance on an external system like an LLM replaces the effortful cognitive processing that independent thinking requires. It saves effort in the short term, but accumulates long-term costs, including diminished critical inquiry, greater vulnerability to manipulation, and reduced creativity1.

Worth noting: this research is a preprint posted to arXiv in June 2025, and the authors disclose real limitations. The 54 participants came from a handful of academic institutions clustered in one geographic area, skewing the sample, and the gender balance wasn’t even. Only ChatGPT was tested, so the results can’t be generalized to other LLMs. And the task was limited to one domain, essay writing in an educational setting1. Even so, using EEG as a physiological measure, rather than relying only on self-reported behavior, gives this study a strength that most AI-risk research lacks.

Why depth decides what you remember: the levels-of-processing effect

The theory that best explains what the MIT study found is the levels-of-processing effect, proposed in 1972 by psychologists Fergus Craik and Robert Lockhart3.

The core idea is simple: whether something sticks in memory depends on how deeply it was processed. Shallow processing, noticing only surface features like the shape of letters, relies on maintenance rehearsal (simple repetition) and produces short-lived memories. Deep processing, understanding meaning and connecting it to existing knowledge and experience through what’s called elaborative rehearsal, produces memories that last.

In a follow-up experiment, Craik and his colleague Endel Tulving (Craik & Tulving, 1975) had participants process words in three ways: checking whether letters were upper- or lower-case (shallow), checking whether words rhymed (intermediate), or judging whether a word fit meaningfully into a sentence (deep). They then gave participants a surprise recognition test. Words processed semantically were recognized at significantly higher rates than the others, and recall for words processed in complex sentences was double the rate for words processed in simple ones.

Where does skimming AI-generated material fall on this spectrum? Checking the look and structure of a document is shallow processing. Taking in the content without verifying or rephrasing it doesn’t go beyond maintenance rehearsal. The step that matters, connecting meaning to what you already know, gets skipped entirely.

Only what you make yourself sticks: the generation effect

The second theory concerns another well-established memory phenomenon: the generation effect. It was established through experiments Norman Slamecka and Peter Graf reported in 19784.

Information you generate yourself sticks better than information you merely read or receive. This effect holds consistently across recognition, free recall, cued recall, and confidence ratings4. Later education-focused reviews describe the effect as far from trivial.

Several mechanisms overlap here. Generating something requires searching memory and connecting it to what you already know, which demands deeper processing. The effort of generating it makes that piece of information more distinctive in memory, which makes it easier to retrieve later. And the retrieval pathway used during generation doubles as the pathway used when you try to recall it afterward.

Having AI generate text or a summary for you gives up the chance to earn this effect at all. The reading material grows, but the process of pulling words out of your own head never happens.

Easy learning doesn’t stick: desirable difficulties

The third theory comes from memory researcher Robert Bjork, who introduced “desirable difficulties” in 19945.

Bjork’s claim: conditions that feel harder and lower short-term performance during learning can produce stronger long-term retention and transfer. The classic examples researchers point to are spaced repetition, retrieval practice (testing yourself instead of re-reading), generation, and interleaving different kinds of material5.

The idea behind it: memory has two independent properties, storage strength (how deeply something is encoded) and retrieval strength (how easily it can be accessed right now). Making the effort to recall something when retrieval strength has dropped (when you’ve partly forgotten it) produces a large gain in storage strength. Information that’s always instantly available never requires that effort in the first place.

Having AI hand you a finished document the moment you ask for one works against desirable difficulties across the board. Spacing, retrieval practice, generation, and interleaving all assume some friction. Once AI removes that friction in advance, there’s no room left for a desirable difficulty to occur.

Where the three theories converge

Levels-of-processing, the generation effect, and desirable difficulties come out of three separate lines of research, but applied to AI-generated material, they point at the same thing. Deep processing, self-generation, and moderate friction are three separate routes to durable memory, and AI bypasses all three of them at once.

flowchart TB
    AI["AI hands you<br>a finished product"]
    AI --> A["No deep processing<br>(levels of processing)"]
    AI --> B["Nothing self-generated<br>(generation effect)"]
    AI --> C["No friction left<br>(desirable difficulties)"]
    A --> D["Nothing sticks<br>= cognitive debt"]
    B --> D
    C --> D

The MIT study’s “83% couldn’t quote their own essay” looks too extreme for any single theory to explain on its own. It makes more sense once you see all three routes closing at the same time. A single weakened memory mechanism would likely produce a gradual decline; three closing together may be what pushes the result all the way to “barely sticks at all.”

At the same time, the crossover session’s finding, that the original brain-only group’s brain activity increased once they started using AI, shows these three routes aren’t permanently closed. A track record of processing things on your own seems to make it easier to integrate new material with what you already know (deep processing, in levels-of-processing terms) even when you do turn to AI afterward. Cognitive debt, in other words, is a debt whose accumulation depends on how you use AI.

How to use AI without running up cognitive debt

None of the three theories argue against using AI. What they offer instead is a design principle, and it collapses to one line: keep one of the steps AI would otherwise skip on your own side. Sketch your own take before asking AI; restate the summary you get back in your own words; use AI to poke holes in your draft rather than to write it; come back after a while and try to recall the content. Each of these reclaims one of the skipped routes: deep processing, self-generation, or moderate difficulty. You don’t need to redo all three; inserting even one meaningfully changes whether something sticks. Restating as if teaching someone overlaps with the protégé effect covered in Teaching AI Deepens Human Learning: What the Protégé Effect Says About “Reverse” Education.

But trying to apply this to every task backfires. Adding friction to format conversions or throwaway drafts wastes time, and when you lack the background to do it well it turns into an “undesirable difficulty” that lowers retention instead. So what actually works is deciding first which tasks to keep a hand in and which to hand to AI. That line of thinking connects to the “result vs. process” offloading distinction in Cognitive Offloading to AI. How to sort tasks by type, and what to do on the ones you keep, is the subject of the companion practical piece.

Conclusion

MIT Media Lab’s experiment gave the feeling that “material AI writes doesn’t stick” a concrete measurement: weaker neural connectivity, and an inability to quote your own writing12. Three independently established theories in cognitive science, levels of processing, the generation effect, and desirable difficulties, explain why. Handing someone a finished product instantly bypasses all three routes to durable memory: deep processing, self-generation, and moderate friction, at once.

But the same MIT study also showed this debt isn’t fixed. A track record of thinking things through yourself changes how deeply you engage the next time you turn to AI. Deciding what to hand off, and where to keep one piece of the processing for yourself, is enough to keep AI’s efficiency while still making the output stick.

Want the practical side, how to actually use AI so it “sticks,” with task-by-task sorting and concrete steps? See the companion piece, Which Work to Hand to AI, and Which to Process Yourself: Sorting Tasks to Avoid Cognitive Debt. This article is the theory half.

References

Footnotes below are numbered to match citations in the text.

Additional sources (not cited by number in the text)

  1. Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task - Kosmyna, N., Hauptmann, E., Yuan, Y.T., et al. MIT Media Lab, arXiv preprint (June 2025). 54 participants across four months, split into LLM/search/brain-only groups, with EEG measurement. Cited from the full PDF (arxiv.org/pdf/2506.08872). Preprint, not yet peer-reviewed; the authors themselves flag one finding (narrower topic range in Session 4) as preliminary. 【Reliability: medium-high】 ↩︎ ↩︎2 ↩︎3 ↩︎4 ↩︎5 ↩︎6 ↩︎7 ↩︎8 ↩︎9 ↩︎10 ↩︎11 ↩︎12 ↩︎13 ↩︎14 ↩︎15

  2. Your Brain on ChatGPT - Armitage, R. British Journal of General Practice, 75(758), 410 (2025). A commentary summarizing the Kosmyna study for a medical-education audience, reporting specific figures (83% quoting failure, up to 55% reduced connectivity). Published in a peer-reviewed journal, but itself a commentary rather than original research. 【Reliability: medium-high】 ↩︎ ↩︎2 ↩︎3 ↩︎4 ↩︎5 ↩︎6

  3. Levels of processing: A framework for memory research - Craik, F.I.M., & Lockhart, R.S. Journal of Verbal Learning and Verbal Behavior, 11(6), 671-684 (1972). The foundational paper establishing the levels-of-processing effect, cited and tested for over 50 years. 【Reliability: high】 ↩︎

  4. The generation effect: Delineation of a phenomenon - Slamecka, N.J., & Graf, P. Journal of Experimental Psychology: Human Learning and Memory, 4(6), 592-604 (1978). The paper that established the generation effect. 【Reliability: high】 ↩︎ ↩︎2

  5. Memory and Metamemory Considerations in the Training of Human Beings - Bjork, R.A. In Metcalfe, J., & Shimamura, A. (Eds.), Metacognition: Knowing about Knowing, pp. 185-205. MIT Press (1994). The book chapter that introduced “desirable difficulties.” 【Reliability: medium-high】 ↩︎ ↩︎2

This post is licensed under CC BY 4.0 by the author.