In short
- A paper claiming AI erodes your sense of owning your ideas was published in April, covered widely, and retracted on 31 August. The university said it never approved or supervised the research.
- Repetition alone makes statements feel truer, with no new evidence and no argument. Knowing better does not protect you: the check does not fail, it never runs.
- Your instinct is that the surprising statistic is the risky one. The opposite holds. The eyebrow is the check running, and a claim that agrees with you raises none.
- Your feeling is data. It is not evidence. The whole distance between those two words is where thinking happens.
You have had the feeling. You finish something with AI in eleven minutes that would have taken two hours, and along with the satisfaction there is a small hollow note underneath it, like arriving somewhere without remembering the drive. You did the work. The work is good. And still there is a thin voice asking whether you were actually there for it.
So when a study landed in April saying exactly that, you did not interrogate it. You accepted it. Nearly two thousand participants, an American Psychological Association journal, a write-up in TIME on the fifteenth of April, and fifty-eight per cent of people in the study agreeing that the AI “did most of the thinking”. The headline said that leaning on AI erodes your confidence in your own reasoning and your sense of owning your ideas, and you read it and thought: yes, that is my Tuesday.7
On the fourteenth of September, Retraction Watch reported that the paper had been retracted.
01What actually happened
The paper appeared in Technology, Mind, and Behavior, published by the APA, in April 2026.4 It claimed a survey of 1,923 adults in the United States and Canada, each completing ten simulated work tasks: drafting plans with incomplete information, interpreting ambiguous data, explaining the reasoning behind strategic choices. The APA put out a press release,6 TIME covered it, Futurism covered it. It travelled a very long way in a very short time, because it was the rare kind of finding that everybody was ready to believe. (Figure 1)
Sandra Grinschgl, who studies human and technology interaction at the University of Bern, read it and started asking questions on PubPeer within weeks. She was the first. Later she brought in two colleagues, Ian Hussey and Malte Elson. Between them they listed inconsistencies in the graphs and the results, problems with how the study was designed, references that pointed to the wrong papers, and an absence of external ethics approval.5
Figure 1
Every institutional check passed. A person caught it.
Five months from press release to retraction
- April 2026 Published in Technology, Mind, and Behavior (APA).APA press release. TIME covers it on 15 April.
- Within weeks Sandra Grinschgl (University of Bern) posts concerns on PubPeer.Graphs inconsistent with results. Erroneous references.
- Ian Hussey and Malte Elson join the review.No ethics approval from any external body.
- May 2026 The three email the journal directly.No reply. They continue anyway.
- June 2026 Journal opens an investigation and requests the data.The author declines to provide it.
- 31 Aug 2026 Retraction notice published.Middlesex: not approved, conducted or supervised by us.
- 14 Sep 2026 Retraction Watch reports the story.
Peer review missed it. The press office missed it. TIME missed it. Three people reading carefully, unpaid, did not.
Retraction Watch, 14 September 2026; retraction notice, Technology, Mind, and Behavior, 31 August 2026.
The journal opened an investigation in June and asked for the data.2 The author, Sarah Baldeo, declined, saying the transfer of participant-level materials and analytic code was not authorised under the consent framework governing the study. Middlesex University, where Baldeo is a doctoral candidate, formally stated that the research had not been approved, conducted or supervised by the university and asked for its affiliation to be removed from the paper. The ethics approval had come from a body called ID Quotient Advisory Group, which is the company Baldeo founded and runs.3
Grinschgl’s summary to Retraction Watch is worth sitting with. She accepted that there can be genuine limits on releasing participant data, but pointed out that this does not extend to demonstrating the platform the study was run on, or to sharing the analytic code, neither of which contains participant data. And then: “We still doubt that the data collection platform ever existed.”1
Baldeo disputes the characterisation and says she asked for the paper to be withdrawn in June.
02Why did you believe it
The story of one bad paper is not that interesting. This next question is.
You believe that a number with credentials has been checked by someone.
Not consciously. You would never say it out loud in those words, but it is the operating assumption underneath almost every statistic you have ever repeated. A big sample size means it is robust. A journal name means peer review happened. A press release means an institution put its name behind it. A TIME article means a journalist looked. Somewhere in that chain, you assume, a competent adult ran the check so that you would not have to.
03What the research actually shows
Start somewhere smaller than a journal. Think of a piece of trivia you are completely sure about, that you have never once verified. That the Great Wall is visible from space.10 That we use ten per cent of our brains.11 That goldfish have a three-second memory.12 You did not research any of these. You heard them, and then you heard them again, and at some point they crossed a line from something you had been told into something you simply knew. None of them are true.
Lynn Hasher, David Goldstein and Thomas Toppino published the experiment that explains this in 1977, in the Journal of Verbal Learning and Verbal Behavior. They gave students at Villanova and Temple sixty plausible statements drawn from politics, sport and the arts, and asked them to rate on a scale of one to seven how certain they were that each was true. Then they brought the same people back twice more, at two-week intervals. Twenty statements appeared on all three lists; the other forty on each list were new.8 (Figure 2)
Figure 2
Repetition alone made statements truer
No new evidence. No argument. Just the same sentence, twice more.
Sixty statements, twenty of them repeated across all three sessions. The repeated ones climbed. The rest did not move.
Hasher, L., Goldstein, D. & Toppino, T. (1977), ‘Frequency and the conference of referential validity’, Journal of Verbal Learning and Verbal Behavior, 16(1), 107–112. Seen-once line drawn flat: the paper reports no reliable change, not a mean series.
One of the statements they used was this: a sari is the name of the short plaid skirt worn by Scots. You know that is wrong. You know the word. So did their participants: asked separately what the short pleated skirt worn by Scots is called, they answered kilt, correctly. Then they saw the sari sentence a second time, and some of them began to rate it as true.
Repetition alone, with no new evidence and no argument, made things truer. They called it the conference of referential validity. Everyone since has called it the illusory truth effect.
Knowing better does not protect you
For decades there was a comforting assumption around this finding: it must only work on things you do not already know. Repetition might fill an empty space, but surely it could not overwrite an occupied one. Lisa Fazio, Nadia Brashier, Keith Payne and Elizabeth Marsh tested that in 2015, in the Journal of Experimental Psychology: General. Their first experiment used forty Duke undergraduates, which is a small sample and worth saying plainly, but the design is what matters: they used statements that contradicted things people demonstrably knew, and checked each participant’s knowledge separately rather than assuming it.9 (Figure 3)
The effect held anyway. People who could correctly tell you that the Pacific is the largest ocean on Earth, who proved they knew it on the knowledge check, still rated the repeated statement “the Atlantic Ocean is the largest ocean on Earth” as more true than a statement they had not seen before. The knowledge was in there. It was accessible. It simply did not get consulted.
Figure 3
Which statistics actually get checked
Your verification runs on friction. No friction, no check.
Path A
A claim that surprises you
“Only 3% of them ever return”
Friction
The eyebrow rises. Processing slows down.
You check it
Your knowledge gets consulted
Path B
A claim that agrees with you
“58% felt the AI did the thinking”
Friction
Nothing snags. It reads as description.
You believe it
The check never ran at all
Fazio, Brashier, Payne & Marsh (2015) called Path B knowledge neglect. Knowing better does not protect you.
Fazio, L.K., Brashier, N.M., Payne, B.K. & Marsh, E.J. (2015), ‘Knowledge does not protect against illusory truth’, Journal of Experimental Psychology: General, 144(5), 993–1002.
Fazio and her colleagues named this knowledge neglect: the failure to draw on what you already know when something feels fluent. That word, fluent, is the mechanism. Your brain runs a constant background check on how easily information is going down, and easy processing feels like truth, because across most of evolutionary history the things that were familiar were the things that were safe. Familiarity was a decent proxy for verification when the only information you ever encountered came from people standing in front of you.
The check does not fail. It never runs.
It is a terrible proxy now. But it is not a proxy you can switch off, and Fazio’s finding is that being smart and informed does not switch it off either.
04Why this study in particular
Knowledge neglect happens when a claim feels fluent, and nothing on earth is more fluent than a claim that describes a feeling you already have.
The Baldeo paper did not have to persuade anyone. That was its enormous advantage. It arrived into a population of people who had all, privately, felt that hollow note after a fast piece of AI-assisted work, and who had all, privately, wondered what it meant. The study did not introduce an idea. It named one that was already sitting there unlabelled, and the sensation of a thing being named is almost indistinguishable, from the inside, from the sensation of a thing being proven.
Which is why the fifty-eight per cent went so far. Nobody was checking it, because it did not feel like a claim. It felt like a description.
The dangerous statistic is the one you liked.
Go and look at the last deck you built. Find the number you were happiest to include. That is the one to trace.
05My own position here
I run a company whose entire proposition is marketing strategy backed by behavioural science, which means what I have just described is not an industry problem I am observing from outside. It is my business model’s occupational hazard.
Here is the objection. Every agency in the world has a slide with a statistic on it. “Seventy-three per cent of consumers say they want brands to be authentic.” “It takes seven touchpoints before a purchase.” “Users decide in fifty milliseconds.” These numbers are load-bearing. They are what turns an opinion into a recommendation and a recommendation into an invoice. Most of them, if you actually chase them, dissolve into a content marketing blog citing another content marketing blog citing a survey somebody ran in 2011 on a sample of four hundred people who were incentivised to respond. The client cannot check. That is the point of hiring the service.
The distinction: citing research is not the same as borrowing its authority. A study used well changes what you do. It predicts something specific, it can be wrong, and if it turns out to be wrong then your recommendation changes. A study used badly decorates a conclusion you had already reached. (Figure 4)
Figure 4
Is that citation evidence, or decoration?
One question separates them, and it takes about six seconds.
If I deleted this number from the slide, would my recommendation change?
No
Yes
It is decoration.
You already reached the conclusion. The number is dressing it up.
Cut the number. Keep the claim.
“In our experience, buyers tend to…” is a respectable sentence. Use it.
It is load-bearing.
The advice depends on it being true. So it has to actually be true.
Trace it to the primary source.
A name, a year, a journal. The study, not an article about the study.
“In our experience” is a perfectly respectable sentence. More respectable than a decimal point you cannot defend.
The test 1&O holds itself to. You are welcome to hold us to it.
You can tell which one you are doing by asking whether removing the citation would change the advice. I have absolutely put numbers on slides that I had not traced to source. Not many and not recently, but the honest figure is not zero, and anyone in this industry who tells you their figure is zero is either new or not counting.
The pressure that produces it is real. You are building the deck at eleven at night, the point is correct, you know it is correct from a hundred client conversations, and a number would make it land. So you reach for the first one that fits. You are not lying. You are decorating something true with something you have not checked, which is a smaller sin and a real one.
06What the retraction does not take away
Most people finish a story like this feeling slightly worse than when they started, and I do not think that is the correct response at all. Something got taken away from you last week, and it was not what you think.
The study was false. The feeling was not. That hollow note when you finish in eleven minutes what should have taken two hours, that sense of having arrived without the drive, is yours. You had it before April. You had it before you read about 1,923 strangers who supposedly had it too. The retraction does not reach back and delete your Tuesday. Nothing in this story touches your experience.
What the retraction takes away is the borrowed certainty you wrapped around the experience, and you were never actually served well by that wrapper. (Figure 5)
Figure 5
Your feeling is data. It is not evidence.
Both are worth having. Only one of them closes a question.
Data
The thing you noticed.
- Arrives from your own experience
- Cannot be wrong: it happened
- Belongs only to you
- Opens a question
Evidence
The thing you tested.
- Arrives from a repeatable method
- Can be wrong, and that is the point
- Belongs to anyone who checks it
- Closes a question, for now
A statistic that agrees with you does not add to your thinking. It ends it.
The whole distance between those two words is where thinking happens.
Because look at what the number was doing for you. It was converting something open into something closed. Before the study, you had a live question: what is that feeling, is it fatigue or is it loss, is it the discomfort of a new tool or the beginning of a real erosion? That is a genuinely interesting question and you were in the middle of it. Then a statistic arrived, and the question stopped.
A statistic that agrees with you does not add to your thinking. It ends it. That is the actual cost, and it is much larger than the cost of being wrong about one paper. So take the question back. It is a better possession than the answer was.
Your feeling is data. It was always data. It is simply not evidence, and the entire distance between those two words is where thinking happens. Data is the thing you noticed. Evidence is the thing you tested. Nobody ever got to the second one without first respecting the first, and almost nobody gets to the second one at all, because the first one is so satisfying that we stop there and call it knowing.
07And one more thing, which is the good news
There is a detail in this story that I have deliberately held back. Nobody was assigned to catch this. There is no department for it. Sandra Grinschgl read a paper in Bern, in what I assume were hours that belonged to her own work, and noticed that the graphs did not agree with the results. She wrote it up on a public forum. Then she asked Ian Hussey and Malte Elson to look, and they found more. They emailed the journal directly in May. The journal did not reply to them. They kept going anyway.
Four months later the paper is gone.
None of them were paid for this. None of them get a citation for it. Hussey spent his own evening after the retraction was announced going through a related book line by line, listing digital object identifiers that resolve to the wrong papers or to nothing at all, and publishing the list underneath the article so that anyone could verify it themselves.
I find this genuinely moving, and I think you should too. We talk about science as though it were a machine, a set of procedures that grinds truth out of data whether or not anyone is paying attention. It is not. Peer review missed this. The journal missed it, the press office missed it, TIME missed it, the university put out a celebratory post about it. Every institutional check in the chain went through and came back clean.
What caught it was a person reading carefully.
That is what the system actually is, underneath the letterhead. Not a machine. A small number of people who look properly at things they had no particular reason to look at, who say something when the numbers do not sit right, and who do not stop when nobody answers the email. Science self-corrects, and the self doing the correcting is always somebody in particular, usually for free.
You can be that. On a much smaller scale, on your own deck, on a normal working day. Tracing one number to its source takes about six minutes, and it is the single most underrated act of professional integrity available to you. Nobody will notice you did it. That is rather the point.
The world does not get more honest because the institutions improve. It gets more honest because a few people keep reading carefully and refuse to pass along what they have not checked.
Be one of them. It is quite a nice thing to be.
08Sources
The retraction
- Orrall, A. (14 September 2026). CEO of AI company loses paper on cognitive impact of generative AI. Retraction Watch. ↩
- Retraction Watch (16 June 2026). Journal investigating paper on cognitive impact of generative AI. ↩
- Retraction notice, Technology, Mind, and Behavior, dated 31 August 2026. ↩
- The retracted paper: Baldeo, S. (2026). Technology, Mind, and Behavior. ↩
- PubPeer thread, concerns first raised by Sandra Grinschgl (University of Bern), later joined by Ian Hussey and Malte Elson. ↩
How the claim travelled
- American Psychological Association press release (April 2026). Overreliance on AI programs may undermine confidence at work. ↩
- TIME (15 April 2026). Letting AI Do Your Work Erodes Your Confidence, According to a New Study. ↩
The behavioural science
- Hasher, L., Goldstein, D. & Toppino, T. (1977). Frequency and the conference of referential validity. Journal of Verbal Learning and Verbal Behavior, 16(1), 107–112. ↩
- Fazio, L.K., Brashier, N.M., Payne, B.K. & Marsh, E.J. (2015). Knowledge does not protect against illusory truth. Journal of Experimental Psychology: General, 144(5), 993–1002. ↩
The three myths
- Great Wall: Britannica, Can you see the Great Wall of China from space? — and Al Jazeera (17 October 2003) on Yang Liwei, the first Chinese astronaut in orbit, reporting he could not see it. ↩
- Ten per cent of the brain: Beyerstein, B.L. (1999). Whence Cometh the Myth that We Only Use 10% of Our Brains? In S. Della Sala (ed.), Mind Myths: Exploring Popular Assumptions about the Mind and Brain (pp. 3–24). Wiley. Also BrainFacts.org, Society for Neuroscience. ↩
- Goldfish: López, J.C., Broglio, C., Rodríguez, F., Thinus-Blanc, C. & Salas, C. (1999). Multiple spatial learning strategies in goldfish (Carassius auratus). Animal Cognition, 2, 109–120. ↩
A note on that last one, since it is the sort of thing this piece is about. The figure most often quoted alongside the goldfish myth is Culum Brown’s finding that fish retained a trained escape route eleven months later. That study is real and worth reading, but the fish were crimson spotted rainbowfish rather than goldfish (Brown, C., 2001, Animal Cognition, 4(2), 109–113), so it does not debunk the goldfish claim directly. The goldfish work above does.