Posted on Leave a comment

The AI Adoption Gap in L&D: What the Data Actually Shows

Abstract bar chart in deep navy representing an adoption gap

Eighty-seven percent of L&D teams are already using AI in some form, according to Synthesia’s 2026 AI in Learning & Development report. And yet only 25 percent of U.S. employees say their organization has actually communicated a clear plan for integrating it, per Gallup’s May 2026 data. That gap between adoption and structure, not a lack of enthusiasm, is the real story in corporate training right now, and understanding it requires separating four things that get casually lumped together as “using AI.”

Four different things people mean by “AI adoption”

StageWhat it actually looks like
ExperimentationIndividuals trying AI tools informally, often without organizational awareness or policy
AdoptionAI use becomes routine for specific tasks, typically content production, but stays within individual workflows
ImplementationAI use is deliberately designed into a process, with defined checkpoints, training, and quality standards
EnablementThe organization has trained its people, set policy, and built the capability to use AI well and consistently, not just permission to use it

The 87 percent figure describes adoption. The 25 percent figure describes enablement. Nearly everyone is past experimentation; very few organizations have reached enablement. That gap between the two numbers is not a rounding error, it is the actual condition of the field in 2026.

The numbers, and what they actually mean together

  • 87% of L&D teams use AI in some form, and only 2% have no adoption plans at all (Synthesia, 2026).
  • 55% of U.S. workers regularly use AI, but only about 1 in 3 received employer-provided AI training in the past six months (The Conference Board, 2026).
  • 49% of employees believe AI is advancing faster than their company’s training programs can keep up with (TalentLMS, 2025).
  • The AI-powered corporate training market itself is valued at $7.49 billion in 2026 (Mordor Intelligence), meaning real money is being spent, just not evenly.

Put simply: people are using AI at work whether or not their organization has trained them to, and the training is not keeping pace with the use. That is not fundamentally an adoption problem. It is an enablement problem, and the two require different responses.

Where L&D teams are actually putting AI to use

Synthesia’s 2026 data breaks down where the real usage sits today: mostly in content production, not in strategy or measurement. Voice generation (63%), content and quiz drafting (60%), video creation (52%), and translation (38%) dominate current use. The most commonly cited benefits are faster production (84%) and a better learner experience (66%). This matters because it shows AI is currently being used to do L&D’s existing work faster, which is adoption, not to change what L&D measures or how it proves impact, which would be closer to implementation; only 55% cite clearer business impact as a benefit today, well behind the production-speed gains.

What the data does not tell us

It is worth being explicit about the limits of this evidence, because most coverage of these statistics is not. Synthesia’s report drew from a sample that likely overrepresents early adopters, since it circulated mainly within AI-forward professional networks; the 87% figure should be read as describing an engaged segment of the field, not the entire profession uniformly. It is also vendor-produced research, from a company that sells AI video and content tools, which does not make the data false, but does mean the questions asked and the framing of “benefit” naturally lean toward use cases that company’s product supports well.

Similarly, self-reported survey data on “using AI” does not distinguish between someone who occasionally asks a chatbot to draft an email and someone running AI-assisted workflows daily; both count as “using AI” in most surveys, but they represent very different levels of actual capability. None of the statistics above measure whether AI use is actually improving learning outcomes, only whether it is being used and whether people feel positively about it. That is a meaningfully different, and much harder, question to answer, and the data here simply does not answer it yet.

The caution behind the enthusiasm

Interest in agentic AI, tools that can complete multi-step tasks autonomously, is notably more cautious than interest in generative AI generally. Only 27% of L&D professionals describe themselves as excited about agentic AI, while 39% say they are cautious and 29% say they need to learn more before forming a view. That caution is reasonable: agent-based tools raise the same workflow and checkpoint questions we cover in our explainer on what AI agents actually change, and adopting them without a defined review process is exactly where the current enablement gap becomes a real risk rather than a missed opportunity.

What to actually do with this data

If you are an L&D professional trying to translate these numbers into action, the practical sequence is straightforward, even if it is not exciting: audit what AI use is already happening informally in your organization before introducing anything new, since 55% of employees are already using it whether or not there is a policy. Then close the training gap the data consistently points to as the actual bottleneck, the move from adoption to enablement, not the technology itself. Only after that does it make sense to evaluate specific tools using a structured process; see our framework for evaluating an AI tool for exactly that step.

Key takeaways

  • Adoption (87%) and enablement (25%) are different stages; the gap between them, not low enthusiasm, is the actual problem.
  • Current AI use in L&D skews toward content production speed, not strategy or measured business impact.
  • Treat vendor-produced survey data as directional, not definitive; sample bias and framing both matter here.
  • The practical first step is auditing existing informal use and closing the training gap, before adopting new tools.

Sources: Synthesia, AI in Learning & Development Report 2026; Gallup, May 2026; The Conference Board, 2026; TalentLMS, 2025; Mordor Intelligence.

Posted on Leave a comment

How to Evaluate an AI Tool Before Your School or Team Adopts It

Abstract illustration in warm sand representing a framework for evaluating AI tools

Most advice on choosing an AI tool focuses on features. The more useful question for a school or training team is narrower: what happens to the data, how reliable is it on your actual use case, and who is accountable when it gets something wrong? This article builds a practical evaluation framework around those questions, one a teacher, school, or L&D team could genuinely use before adopting a tool, not just a list of things to vaguely keep in mind.

The evaluation framework

Nine dimensions cover most of what actually determines whether an AI tool is safe and useful to adopt. Not every tool needs deep scrutiny on every dimension, a simple internal writing aid needs less security review than a tool handling student records, but knowing which dimensions matter for your specific case is itself part of the evaluation.

DimensionKey questionWhy it matters
Problem fitDoes this solve your exact use case, or a general version of it?A general-purpose tool asked to do a narrow job often produces plausible but inconsistent results
AccuracyHave you tested it on a case where you already know the correct answer?This surfaces failure patterns faster than any feature list or demo
Privacy and data handlingIs data used to train the vendor’s models, and is it retained after your account ends?Directly affects compliance obligations for student or employee data
SecurityDoes the vendor have a clear, specific security posture, not just a generic privacy page?A vendor that cannot answer specifically is itself the answer
ReliabilityWhat happens when the tool is wrong, and how often does that happen on your use case?Determines how much unsupervised trust the tool has earned
Pedagogical or learning valueDoes the tool support the underlying skill you are teaching, or bypass it entirely?A tool can be accurate and still undermine the learning goal it is used for
AccessibilityDoes it work for students or employees with different needs, devices, and connectivity?An otherwise excellent tool that excludes part of your population is not actually adopted equitably
Integration and costDoes it fit your existing systems and budget without heavy custom work?High integration cost often outweighs a tool’s raw quality advantage
Vendor dependence and human oversightWhat is the defined checkpoint where a human reviews output before it reaches someone, and what happens if the vendor changes terms or shuts down?The more a tool automates, the more this checkpoint and this exit plan both matter

Start with data, not features

Before evaluating what a tool can do, establish what happens to what goes into it. For any AI tool touching student or employee data, ask directly: is data used to train the vendor’s models, is it retained after your account ends, and does the vendor’s policy meet your institution’s actual data protection obligations, not just a generic privacy page. If a vendor cannot answer this clearly and specifically, that is itself the answer, and no feature is worth proceeding without it.

Match the tool to a specific use case, not a general one

“AI tool for teachers” and “AI tool for grading essays against a specific rubric” are different evaluation problems. A general-purpose tool asked to do a narrow job will often produce plausible-looking but inconsistent results, because it was not built or tuned for that specific task. Before testing anything, write down the exact task the tool needs to do, including the format of the input and the output you actually need, and evaluate against that, not against a demo video.

Test it on a case where you already know the right answer

The single most useful evaluation step is also the most skipped: run the tool on a real example where you already know what a correct or good output looks like, before trusting it on a case where you don’t. Grade a past assignment with an AI grading tool and compare it to how you actually graded it. Ask a lesson-planning tool to build a lesson you have already taught well, and see where it diverges from your judgment. This surfaces failure patterns far faster than reading a feature list, and it is the closest thing to a real accuracy test most teams can run without a formal research budget.

Pedagogical value: accurate is not the same as useful

A tool can be technically accurate and still be a poor fit if it bypasses the exact skill you are trying to build. A grammar-correction tool that silently fixes a student’s writing without explaining the error teaches the tool to write, not the student. When evaluating pedagogical fit, ask specifically what the tool does with its own output, does it explain, does it just fix, does it check the student’s reasoning or just the final answer, because that determines whether it reinforces learning or quietly substitutes for it.

Decide the human checkpoint before you adopt, not after something goes wrong

Every AI tool that touches instruction, assessment, or communication needs an explicit answer to one question: at what point does a human review the output before it reaches a student, parent, or employee? This matters even more for tools built on AI agents, which complete multi-step tasks with less prompting; the more a tool does unsupervised, the more that final checkpoint matters, not less. Decide this during evaluation, as a condition of adoption, rather than improvising it after a mistake reaches someone.

Ask what the tool is not good at

Any vendor demo will show you what a tool does well. A useful evaluation actively looks for where it breaks: ambiguous instructions, edge cases, unusual student needs, non-standard formats. If a vendor cannot or will not discuss known limitations, treat that as a gap in the evaluation, not a reason for confidence. This applies to vendor dependence too: ask what your exit plan is if the vendor changes terms, raises prices, or shuts down, before you are relying on the tool operationally.

A short practical checklist

  • Data: is data used for model training, and is it retained after the account ends?
  • Specificity: does the tool solve your exact use case, or a general version of it?
  • Testing: have you run it on a case where you already know the correct answer?
  • Pedagogical fit: does it explain and build the skill, or silently bypass it?
  • Checkpoint: is there a defined point where a human reviews the output before it reaches someone?
  • Limitations: can the vendor clearly describe where the tool performs poorly?
  • Exit plan: what happens if the vendor changes terms or the tool is discontinued?

Key takeaways

  • Evaluate data handling and accountability before comparing features.
  • Test against a specific use case with a known-correct example, not a demo.
  • Accuracy and pedagogical value are different questions; a tool can be correct and still undermine the skill you are teaching.
  • Decide the human review checkpoint and the vendor exit plan as conditions of adoption, not afterthoughts.

This framework is intentionally vendor-neutral. If you are specifically weighing agent-based tools that complete multi-step tasks, read this alongside our explainer on what AI agents actually change before applying the checklist above.

Posted on Leave a comment

What the OECD’s 2026 Report Actually Says About AI in Education

Abstract illustration in periwinkle representing an education research report

The OECD’s Digital Education Outlook 2026 is the most substantial piece of international research yet on generative AI in classrooms, and its central finding is more useful than either the optimistic or alarmist headlines it generated: generative AI can genuinely support learning, but only when it is guided by clear pedagogical intent. Used without that guidance, it tends to boost immediate task performance while producing no real learning gain at all. This article separates what the OECD actually found from what we think educators should practically take from it, because those are two different things and conflating them is exactly how good research gets flattened into a slogan.

What the report actually studied

The Outlook examines generative AI across three distinct scenarios: students using it to learn independently, students and teachers using it together as part of instruction, and teachers using it alone to support their own work. It also looks at how generative AI can improve efficiency at the institutional level, such as analyzing learning pathways or supporting study advisors. This scope matters, because most classroom conversation about AI collapses all three scenarios into one, when the report treats them as genuinely different problems with different risks and different evidence behind them.

The three scenarios the report treats separately

Before getting to the headline finding, it helps to see how differently the OECD treats each usage scenario, because the risk profile is not the same across all three.

ScenarioWhat the OECD found
Students learning independentlyHighest risk of the metacognitive lag effect; benefit depends heavily on whether the tool has pedagogical structure built in, not just general capability
Students and teacher using AI together in instructionMost consistently positive scenario in the evidence reviewed; the teacher’s real-time guidance is what converts AI use into actual learning
Teacher using AI alone for their own workGenerally lower-risk; framed mainly as a productivity and preparation question rather than a learning-outcomes question

The practical implication is that the same tool can be low-risk or high-risk depending purely on which of these three modes it is used in. A tool used well by a teacher for prep work carries little of the risk the OECD documents for unsupervised independent student use, even if it is the exact same underlying AI system.

What the OECD found: performance gains do not equal learning gains

The report’s most important distinction is between task performance and actual learning. When students outsource a task to generative AI without pedagogical structure, they tend to complete it faster and more successfully in the moment, but that improvement often does not transfer to independent skill. Researchers involved in the OECD’s accompanying conference described this as a form of metacognitive lag: efficiency gains alongside flat or declining underlying competence.

This is a finding about design, not a blanket verdict on AI in the classroom. The report frames the difference as one of intent: AI tools built or deployed with explicit pedagogical purpose, rather than general-purpose AI used as an undirected shortcut, are what actually produce learning benefit. That is the OECD’s own framing, and it is worth holding onto precisely because it resists the two lazy conclusions, that AI simply helps or simply harms, that dominate public discussion.

What the OECD found: the equity pattern

One of the report’s more concrete patterns: students in well-supported learning environments tend to use AI the way a good tutor would be used, iteratively, with guidance, treating it as a partner in the work. Students without that support more often use it as a pure shortcut. Left unaddressed, this risks widening the exact achievement gaps that additional support was supposed to close, rather than narrowing them.

What the OECD says: this is a design and implementation issue, dependent on how AI is introduced and supported, not an inherent property of the technology. Our interpretation: this places real responsibility on institutions, not just individual teachers, since the support structure the OECD describes, guided, iterative use with feedback, is difficult for a single teacher to build alone without institutional backing, time, and training.

What the OECD found: the framing of teachers’ role

The OECD frames teachers’ relationship to AI along a spectrum from replacement to complementarity to augmentation, and the meaningful distinction is not whether AI helps a teacher, but how it affects their professional judgment: whether it substitutes for their decisions, leaves them unchanged, or genuinely expands what they can do. The report is explicit that the goal is not simply improving output; it is preserving and strengthening teacher agency, not routing around it.

Our interpretation: this lines up closely with how we think about AI agents in education: the tools that genuinely help are the ones that expand what a teacher or trainer can do, with a human checkpoint still firmly in place, not the ones that quietly remove the teacher from the loop. The OECD’s augmentation category and the “final human check” principle we describe there are, practically speaking, describing the same thing from different angles.

What this means in practice for a teacher or trainer

Translated into something usable day to day, the report’s findings suggest a few concrete habits, not a policy document:

  • When introducing an AI tool to students, build in the guidance and iteration the OECD found matters, do not just grant access and assume productive use will follow.
  • Watch specifically for the gap between who uses AI well and who uses it as a shortcut; it will not be evenly distributed across a classroom or cohort.
  • Treat AI-assisted output the way the augmentation framing suggests, as expanding your options, not replacing the judgment call at the end.

What the report does not settle

It is worth being precise about the limits here. The OECD’s findings are a synthesis of emerging research, not a single definitive experiment, and the underlying evidence base is still developing. The report itself frames this as a snapshot of where the evidence currently points, with design guidance for institutions rather than a finished verdict. It also does not resolve the longer-running question of exactly how AI use affects independent thinking over time; we cover that evidence separately, including its own significant caveats, in what the research says about AI and critical thinking.

Key takeaways

  • OECD finding: generative AI supports real learning only with deliberate pedagogical structure; without it, gains are limited to task performance.
  • OECD finding: unequal AI use risks widening achievement gaps unless institutions actively support how students use it, not just whether they can access it.
  • Our interpretation: this places real responsibility on institutional support, not individual teacher effort alone.
  • This is a synthesis of emerging evidence, not a closed case; expect the picture to keep developing.

Source: OECD (2026), OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education, OECD Publishing, Paris.

Posted on Leave a comment

AI Agents, Explained for People Who Teach and Train

Abstract network diagram in cobalt blue representing AI agents

An AI agent is software that can take a goal, break it into steps, and carry those steps out on its own, searching, writing, checking its own work, and adjusting, rather than just answering one prompt at a time the way a chatbot does. For teachers and trainers, that distinction matters more than the terminology suggests: it is the difference between a tool you have to operate at every step, and one you can hand a task to and check on later. This article explains what actually separates a chatbot, an AI assistant, and an AI agent, what that difference looks like in a classroom or training program, and where human oversight still genuinely matters.

Chatbot, assistant, or agent: what is actually different

These three terms get used almost interchangeably in casual conversation, but they describe genuinely different levels of autonomy. The clearest way to separate them is by what happens between your instruction and the finished result.

TypeWhat it doesWhat you still have to do
ChatbotAnswers one message at a time; no memory of taking action, only of the conversationRe-prompt for every step; assemble the pieces yourself
AI assistantCarries context across a task, can use a tool or two (search, a document), but generally completes one bounded action per requestDirect each major step; review each output before moving on
AI agentPlans a multi-step sequence, uses several tools, checks its own output against a goal, and revises without being re-prompted at each stageSet the goal and constraints up front; review the finished result before it reaches anyone else

The practical marker of an agent is not intelligence, it is persistence: it keeps working toward a goal across multiple steps without you re-engaging at each one. A chatbot that drafts one paragraph when asked is not an agent. A tool that takes “build a two-week unit plan on this topic, aligned to these three learning objectives,” researches the topic, drafts the sequence, checks it against the objectives you gave it, and revises the weak parts on its own, is functioning as an agent, regardless of what the vendor calls it.

Anthropic’s 2026 State of AI Agents report, based on real usage data rather than survey opinion, found that agents have moved from experimental novelty to production infrastructure inside organizations: over 57 percent of surveyed organizations are now running multi-step agent workflows, not just single-action assistants. The report also found that model capability is no longer the main bottleneck to adoption; integration with existing systems is, cited by 46 percent of organizations as their top challenge. That second point is the one worth sitting with: the hard part of using agents well is not the AI, it is fitting the workflow around it.

What this actually looks like in a classroom or training program

Strip away the enterprise framing, and the practical version for education and training looks like this:

  • Research and drafting agents that can be given a topic and a set of constraints (grade level, learning objectives, length) and return a structured first draft of a lesson plan or training module, not a single paragraph you then have to assemble yourself. A trainer building a compliance module, for instance, could set the required topics and pass criteria, and let the agent draft the sequence, checking its own draft against the pass criteria before handing it over.
  • Multi-step feedback tools that check a piece of student or trainee writing against a rubric, flag specific issues, and draft suggested comments, then wait for a human to approve before anything goes back to the learner. The agent behavior here is the multi-pass checking: reading the rubric, checking the draft against each criterion, and only then producing comments, rather than a single generic pass.
  • Administrative agents that can be pointed at a messy task, such as reformatting a slide deck, compiling attendance data, or drafting a parent or manager email from bullet notes, and asked to complete it end-to-end.

None of this requires a technical background to use. It does require a workflow decision: what are you comfortable handing over completely, and what needs a human checkpoint before it reaches a student or employee? That question is worth answering deliberately, not by default; see our framework for evaluating an AI tool before adopting one for exactly this reason.

What agents actually change, and what they do not

What changes: the amount of multi-step work you can delegate without babysitting every stage. A lesson-planning agent does not just draft a paragraph when asked, it can research the topic, structure the plan, check it against a rubric you provide, and revise itself, then hand you something closer to a finished draft. That is a genuinely different capability from last year’s single-turn AI tools.

What does not change: judgment. An agent can execute a plan; it cannot tell you whether the plan serves your actual students or trainees, whether the tone is right for your specific group, or whether a shortcut it took quietly undermined the learning goal. This is the same caution that applies to generative AI in education generally; see our piece on what the research says about AI and critical thinking. Handing a multi-step task to an agent does not remove the need for a human to check the outcome; if anything, because agents complete more steps unsupervised, the final check matters more, not less.

Current limitations worth knowing before you rely on one

Agents fail differently than chatbots do, and it is worth understanding how before you hand one a real task. A chatbot that misunderstands a prompt produces one visibly wrong answer you catch immediately. An agent that misunderstands a goal can complete several plausible-looking steps built on that misunderstanding before anyone notices, because each individual step looks reasonable in isolation, it is the cumulative direction that is wrong. This is sometimes called compounding error, and it is the main practical risk of increased autonomy.

Agents also still depend heavily on how clearly a goal and its constraints are specified. Anthropic’s own reporting on integration being the top adoption barrier, ahead of raw capability, reflects this: an agent given a vague goal will confidently fill the gaps with its own assumptions, which may not match yours. And because an agent can use multiple tools and take multiple actions, the surface area for something to go wrong, a wrong data source, an outdated reference, a misapplied rule, is larger than with a single-turn tool.

None of this is an argument against using agent-based tools. It is an argument for treating the final review as a required step, not an optional one, and for being explicit and specific about goals and constraints rather than assuming the agent will infer what you actually meant.

If you are in a corporate learning and development role weighing whether agent-based tools are worth adopting yet, that decision connects directly to a broader pattern we cover in the AI adoption gap in L&D: enthusiasm for AI in L&D is high, but structured, trained, checkpoint-driven adoption is still rare, and that gap is exactly where agent-based tools can go wrong if adopted without a workflow to match.

Key takeaways

  • The functional difference between a chatbot, an assistant, and an agent is persistence across multiple steps without re-prompting, not raw intelligence.
  • The practical uses in education and training are drafting, feedback-checking, and admin work, not replacing instructional judgment.
  • Agents fail through compounding error, several plausible steps built on one wrong assumption, which is why the final human check matters more with agents, not less.
  • Integration into an existing workflow is the actual adoption bottleneck, not the technology itself.

Source: Anthropic, 2026 State of AI Agents Report.

Posted on 1 Comment

Does AI Really Hurt Critical Thinking? What the Research Shows

Abstract layered arcs in cobalt blue representing cognitive processes

The honest answer is: it depends what you offload, and what happens to the thinking time that gets freed up. A growing body of research from 2025 and 2026 links heavy AI use to weaker independent thinking, but the more recent and more careful studies are converging on a sharper point: the danger is not AI use itself, it is AI use with no structure for what stays human. Getting this right requires separating several concepts that casual discussion of this topic tends to blur together.

Six terms that are not interchangeable

TermWhat it actually means
Cognitive offloadingUsing an external aid to reduce mental demand, a neutral mechanism, the same one behind a grocery list or a calculator
DelegationA deliberate choice about which specific task to hand over, made consciously rather than by default
LearningBuilding durable skill or understanding you can access without the external aid present
Independent reasoningThe capacity to work through a problem unaided, which is what atrophies if the wrong tasks are chronically offloaded
MetacognitionAwareness of your own thinking process, including noticing when you are offloading and whether that is appropriate
VerificationActively checking an AI’s output rather than accepting it, the specific behavior most research finds declining with heavy trust in AI

Most of the public debate about “AI and critical thinking” is really only about two of these six: whether offloading is happening, and whether learning suffers as a result. The more useful and more evidence-backed version of the question involves all six, especially metacognition and verification, which is where the research below actually locates the risk.

What the evidence actually shows

A 2025 study published in the journal Societies (Gerlich, 2025) surveyed and interviewed 666 people across age groups and education levels, and found a significant negative correlation between frequent AI tool use and critical thinking scores, with cognitive offloading as the mechanism connecting the two. Younger participants showed the strongest effect.

A separate 2025 MIT study (Kosmyna et al.) used EEG to measure brain activity while participants wrote essays using ChatGPT, a search engine, or no tool. The ChatGPT group showed the weakest neural engagement while writing, and when later asked to write without any AI assistance, showed reduced brain connectivity and struggled to recall their own earlier work, a pattern the researchers termed cognitive debt.

A third study, from Microsoft and Carnegie Mellon researchers presented at CHI 2025 (Lee et al.), surveyed 319 knowledge workers about 936 real AI use cases and found that the more a person trusted the AI’s output, the less critical thinking, meaning verification, in the framework above, they reported applying to it, while trusting their own judgment correlated with more thinking, at a higher mental cost. This study is the clearest evidence that the actual mechanism at risk is verification specifically, not reasoning capacity in general.

Important caveat: correlation, not proof

None of this evidence establishes that AI use causes declining critical thinking in a strict sense. The Gerlich study is explicitly correlational: it is equally possible that people who already think less analytically reach for AI assistance more often, or that a third factor, such as time pressure, drives both patterns at once. The MIT EEG study, while methodologically distinctive, involved a small sample (54 participants) and measured a specific task, essay writing, that may not generalize to every kind of AI-assisted work. These findings are a real and statistically significant signal worth taking seriously, not a settled verdict that AI use damages cognition across the board.

Where the research moved next: not all offloading is equal

This is the part most coverage of this topic misses. Cognitive offloading itself is not new or inherently harmful, as the table above makes explicit, it is the same mechanism behind writing a grocery list. It becomes a problem specifically when it offloads work that would otherwise build or maintain a skill you need, in other words, when delegation happens without metacognitive awareness of what is being given up. A March 2026 synthesis from the University of Technology Sydney frames the real question as what happens to the freed-up mental effort, not whether offloading happens at all.

A concrete intervention study supports this distinction directly. 240 university students learning English essay writing were split into two groups over 12 weeks. One group was explicitly taught to delegate lower-order tasks to AI, such as brainstorming, grammar checking, and initial co-revision, while deliberately keeping higher-order work, analysis, evaluation, and reflection, for themselves. The other group received standard instruction. The group with the explicit offloading structure showed significantly greater critical thinking gains, not smaller ones. The intervention did not restrict AI use; it made the boundary between what to hand over and what to protect explicit, which is precisely the metacognitive step the correlational studies above found generally missing.

What educators and learners can practically do

Translating this research into practice comes down to making two things explicit that usually happen invisibly:

  • Name the boundary before the task starts. Decide in advance which parts of an assignment or project are meant to be AI-assisted (brainstorming, formatting, first-pass grammar) and which parts must stay unaided (the actual analysis, the argument, the final judgment), rather than letting the boundary drift by default.
  • Build in a verification step deliberately. Since the CHI 2025 findings point specifically to verification declining with trust, not general reasoning, the most targeted fix is requiring a check step, asking a student or trainee to identify what they would need to confirm before accepting an AI-generated answer, rather than banning AI use outright.
  • Treat the freed-up time as the actual variable. The UTS synthesis and the 12-week intervention both point the same direction: offloading is beneficial when the time it frees up gets redirected to higher-order thinking, and harmful when it is simply time saved with nothing redirected. That redirection has to be designed in, it does not happen automatically.

This is the same structural principle the OECD’s 2026 education research points to independently; we cover that in more detail in our analysis of the OECD’s 2026 report. It also directly informs how to think about AI agents, which by design take on more multi-step work with less human involvement at each stage; the more a tool offloads, the more deliberate the boundary and the verification step both need to be.

Key takeaways

  • Multiple 2025-2026 studies link heavy, unstructured AI use to weaker independent thinking; the evidence is real but correlational, not proof of causation.
  • The specific mechanism most consistently affected is verification, checking AI output, not reasoning capacity in general.
  • Cognitive offloading itself is not the problem; it becomes one when it offloads a skill you still need, without metacognitive awareness of that trade-off.
  • A 12-week intervention study found that explicitly defining what to delegate and what to protect produced greater critical thinking gains, not smaller ones.
  • The practical takeaway is structure, not restriction: name the boundary, build in verification, and make sure freed-up time gets redirected somewhere.

Sources: Gerlich, M. (2025), Societies, 15(1), 6; Kosmyna et al. (2025), MIT Media Lab; Lee et al. (2025), presented at CHI 2025; University of Technology Sydney synthesis (March 2026).