Skip to main content

Psychology already has what medicine is trying to build

Dr Aisha Tariq · 15 September 2026

A Nature Medicine paper argues that clinicians who lean on AI during training may never develop the reasoning safe practice needs. The argument transfers to psychology more cleanly than I would like. What we do about it comes down to what we ask in supervision.

A trainee brings me a formulation. It holds together. The maintenance cycle makes sense, the language is right, it accounts for the presenting problem. I ask why she chose that model over two others that would fit the same material, and the answer thins out. Not wrong. Thin.

Five years ago I would have read that as nerves, or as someone who had done the thinking and couldn't yet narrate it. Now I don't know what to read it as. That uncertainty is the actual problem, and it isn't about the trainee. The things I used to treat as evidence that someone had thought, a clean formulation, a fluent report, a differential that covers the ground, have become cheap to produce. My proxies stopped working and nobody told me.

The word for it

There is now a term for what I'm circling. A Perspective in Nature Medicine in May, led by Yuhe Ke and Nan Liu at Duke-NUS with David Bates and Pearse Keane among the authors, calls it never-skilling.

It sits alongside two older ideas. Deskilling is what happens to an experienced clinician who leans on a tool until the old competence dulls, and the underlying architecture is still there. Mis-skilling is absorbing the tool's errors as fact. Never-skilling is neither of those. It's the failure to build the reasoning at all, because something else did the cognitive work during the years that work was supposed to be constructing something.

The authors are more careful than most of the coverage has been. There is no direct evidence for never-skilling in clinical trainees. None. Nobody has run the longitudinal study. What they have is learning theory, mainly desirable difficulties, deliberate practice and cognitive load, and a handful of signals from elsewhere. Endoscopists who had got used to AI-assisted colonoscopy detected around six per cent fewer adenomas when they went back to working unassisted, which is deskilling rather than never-skilling. A randomised study of high school maths students found that unrestricted access to an AI tutor improved performance during assisted practice and produced a seventeen per cent relative reduction on the later closed-book exam, with the heaviest losses among the students who started weakest.

They call it a risk model rather than an established phenomenon. I'd hold it the same way, and I'd say so to anyone who quotes it at you as settled.

Why it lands harder on us

The maths study is the one I keep returning to, because of what it describes. Performance that looks fine while the tool is there and doesn't survive its withdrawal. The paper calls this false proficiency.

Medicine has a partial defence against false proficiency, which is that a lot of medical reasoning gets checked by something outside the clinician's head. The scan shows the mass or it doesn't. The potassium comes back at 6.8 whatever anyone believed.

A formulation has no potassium. We assess it on internal coherence, fit with the evidence base, and whether it explains what the person brought. A formulation that is fluent, plausible and wrong looks very like one that is fluent, plausible and right. It can run for several sessions before anything contradicts it, sometimes considerably longer, and sometimes with the client agreeing to it, because a confident account of your own life is difficult to refuse when it arrives from someone you've been told is an expert.

This is what the paper calls the calibration paradox, and we are unusually exposed to it. Knowing when to trust an output requires a standard to check it against, and that standard gets built by reasoning independently, being wrong, and noticing you were wrong. A trainee who hasn't built it cannot verify anything. They can only accept or reject, and both are guesses dressed as judgement.

The counter-argument is better than people admit

It would be easy to write this up as a warning and stop. The paper doesn't, and the honest version is more interesting.

The same authors point out that AI designed for learning rather than for answers may accelerate competence. A Socratic tutor that asks questions instead of supplying conclusions produced more reasoning engagement than a standard one. AI-assisted study has improved postgraduate exam performance, though those exams were sat under AI-enabled conditions, which is precisely the gap they're worried about.

Their distinction is between answer-delivery mode and learning mode, and it carries the argument. The risk isn't exposure. It's sequencing, and what the tool is doing in the moment it's used. Most tools built for clinical settings run in answer-delivery mode because that is what practising clinicians need from them. Trainees then meet them in exactly that mode, at the stage when they would benefit most from friction.

Documentation is not exempt

The paper explicitly puts documentation and administrative tasks outside its scope, on the grounds that they involve different cognitive demands.

That's convenient for me and I don't entirely accept it. In medicine the note largely follows the reasoning. In our work the writing often is the reasoning. I have watched trainees discover what they actually think about a case somewhere in the third paragraph of a report, and anyone who has rewritten a formulation and found that the rewriting changed their mind knows the drafting isn't clerical.

I should be straightforward about my position here. I use these models across my practice and I build software for psychologists in this space, so I have a commercial interest and I'd rather state it than have it inferred. It also means the question lands on me rather than at a comfortable distance. If an assistant psychologist's first fifty reports are drafted from a transcript and then edited, what did the fifty teach them? I don't know. Nobody does, because nobody has looked. What I'd want is for the question to be asked out loud rather than settled by default, one practice at a time, according to whatever each of us happens to do.

We already have phase three

The framework the paper proposes runs in three phases. Establish competence without AI. Introduce it through structured work on deliberately flawed reasoning, so trainees learn to detect and correct. Then integrate it under supervision.

Medicine is proposing to build supervised practice into training. That is where we start. Every assistant psychologist, every trainee on placement, every newly qualified psychologist is already sitting down for an hour a week with someone whose job is to examine how they reached their conclusions. The phase medicine finds hardest is our standing weekly appointment.

What we don't have is anything to say in the hour. The HCPC standards are silent on this. The BPS has issued nothing to training providers. So the decision is being made a few hundred times a week by individual supervisors improvising, and mostly it is being made by not asking. The reference list of that paper contains one piece of psychiatric literature, and it's about deskilling in established clinicians. Our training pipeline hasn't been looked at by anyone.

What I ask

Not "did you use AI for this". That question is useless. It gets a defensive answer, it positions me as compliance rather than supervision, and it tells me nothing I want to know.

I ask for the formulation before anything is opened, said aloud, in the room. I ask which model was ruled out and on what grounds. I ask where this account would break, and what we would expect to see if it were wrong. Those questions worked before any of this and they aren't a response to AI, which is rather the point. They test the thing itself instead of testing for the tool.

I also ask about AI use in the first supervision session of a placement, in the same register as I'd ask about anything else in how someone works, and well before they've worked out what I think of it. Ask in month four and you get whatever they've decided you want to hear.

The cost is time. Reasoning aloud in front of someone more senior is uncomfortable, and slower than reading a finished document and nodding. The discomfort is the mechanism, which is a hard thing to explain to a trainee and harder to hold on a full caseload.

The part I can't resolve

We will find out about this the slow way. The cohort currently training will qualify, and in ten years we'll have some idea. The argument for moving before then is that a deficit like this is difficult to detect and difficult to reverse once it exists at the scale of a cohort, and that waiting for proof means waiting until the thing you were worried about has already happened.

I don't find that argument entirely comfortable. That reasoning has justified a lot of unnecessary caution in our profession before now, and I don't want to be the person telling trainees to work harder for its own sake while I use every tool available to me.

What I'd want us to avoid is narrower than a ban and duller than a debate. A generation of psychologists who produce good work and cannot tell you why any of it is right.

Related insights

Join us at the AI in Psychology Summit

A one-day, 6-hour CPD summit on AI in psychology for UK mental health professionals. Online, 5 October 2026.