Skip to main content

7 AI Practices to Avoid in the Classroom

Seven uses that hand the tool the thinking.

· Updated August 14, 2026

These seven are one move. Something the lesson needed gets taken out, and something that resembles it goes in the gap. The work still arrives and still looks right, which is why none of them are visible in the finished product.

17%

worse than classmates who never had the tool, once it was taken away

A field experiment in Turkey with nearly 1,000 high-school students measured both halves of that at once. Working alongside a general GPT interface, students performed about 48% better during practice. Once the tool was withdrawn, they scored worse than classmates who had never had it. A safeguarded, constrained tutor produced a gain of roughly 127% and largely avoided the drop.

20 of 818

studies carried strong causal evidence, in a 2026 Stanford review

Those figures come from mathematics classrooms outside the United States. Read the numbers as a direction rather than a measurement. A 2026 Stanford review screened the field and found strong causal evidence in a small fraction of it, none of that from student-facing research in American K-12 schools.

Every one of the seven has a legitimate neighbour: a job the same tool does well, in the same lesson, on the same afternoon. So the repair is never to shut the tool off. It is to put the displaced thing back, and hand the tool what was always standing around it.

1. A chatbot cannot be the first explanation

Students cannot evaluate an answer in a subject they have not studied yet. They have no way to see what it left out, flattened, or got wrong, because seeing that is the knowledge the unit was built to produce. Open chatbot use at first exposure closes a circle: the student needs the knowledge to judge the system, and is using the system in place of acquiring it.

Take a unit on the Dust Bowl that opens with a chatbot explaining the causes. Nobody in the room can yet tell whether the answer treats it as a drought that happened to farmers, or as the consequence of how that land had been broken and farmed. Telling those apart is the unit. Whichever account comes back is what students carry forward, and you meet it again in their essays.

The repair is sequence, not prohibition. The tool has a place here, several steps later than it usually turns up.

First, teach it

Instruction, a text, a source, a demonstration, a discussion. The first encounter comes from you or from the record.

Then have them explain it back

Students put the idea in their own words while it is still new enough to be effortful.

Then apply it once

One use of the idea on a problem or a source, so there is student thinking on the page to work with.

Now bring the tool in (the earliest it belongs)

The student has something to judge the output against, and enough of the subject to notice what is missing from it.

Last, transfer without it

A task completed unaided, which is the only evidence that the learning survived the support.

2. Never automate the capability you are grading

An assignment turns unstable the moment the tool performs the thing the assignment exists to build. If the objective is to construct a thesis, the thesis is the one part the tool should not construct. The instability is hard to see from outside: what lands on your desk still looks like the assignment, while the grade has quietly moved onto how well each student directed a model.

Sorting which is which takes one pass before you write the directions. Ask what the objective is, then ask whether the tool is doing that or doing the work standing around it.

Asking questions about a draft, naming what a student left unaddressed, building the opposing case, and holding the writing against the rubric all put the student back inside the capability you meant to develop.

3. Changing the reading level is not personalization

Reading level, vocabulary, examples, language, format, sequence — a model changes any of them in seconds. All of that is adaptation, and adaptation is worth doing. Personalizing instruction is a larger claim. You need to know what this student already understands, which misconception is in the way, their language and cultural context, and the standard the work answers to. You also have to decide which difficulty to strip and which to protect, and what would count as evidence it worked. A general-purpose model has access to none of it.

What it does instead is documented. A 2026 review of 28 studies of teachers designing with AI names two weaknesses that keep recurring: generic content, and poor alignment to the curriculum and the local context. Left alone, a model strips the disciplinary vocabulary a standard names, drops the conceptual difficulty that made a reading worth assigning, or supplies a friendly analogy that quietly misleads.

So name the difficulty you are removing before you remove it. Confusing wording, an inaccessible layout, an unfamiliar term the standard does not test, directions in a language the student is still learning — those come out freely. Retrieving what they know and tracing evidence to a claim is the assignment, and it stays.

Ask a model to simplify a passage on the Fourteenth Amendment and it will do it well. “Due process of law” may come back as “fair treatment,” which is close enough for a summary and useless for the lesson, because the lesson was that phrase.

4. A plausible lesson is not a coherent one

You can read a generated plan straight through, find nothing wrong with any single point, and still watch it fail as an hour. Coherence is a property of the whole, and a complete lesson has to align all of this at once:

  • prior knowledge - the learning goal - the disciplinary content - a model or exemplar - the sequence of student thinking - anticipated misconceptions
  • checks for understanding
  • guided practice - independent practice - assessment - accommodations - local context - materials and time - the next lesson in the unit

A model assembles those from the shape of plans it has seen. Whether they hold together is something you find out in front of the class.

The trial evidence points the same way. An EEF/NFER randomized trial with 259 teachers across 68 English secondary schools studied ChatGPT-supported planning in Key Stage 3 science. A blinded expert panel did not detect an apparent reduction in resource quality. The teachers in it used AI selectively, for one or two components — quiz questions, activity ideas, examples — rather than delegating whole lessons. Bounded use is what the evidence supports. Autonomous design is not.

So hold the goals, the sequence, and the evidence of learning, and delegate components. Ask for three options rather than an answer, then verify, edit, and localize what comes back. That division is the part of the work Kindred K-12 is built around: you decide what the lesson is for, and it builds the materials underneath that decision.

5. One polished artifact is one observation

Grade a unit almost entirely on an essay written at home and you are measuring several things at once. The content knowledge you meant to assess may be the weakest of them. Access to tools moves that score. Prompting skill moves it, and so does editing ability. Adult help at the kitchen table counts, along with a read on how firmly the syllabus rule was meant, and how sophisticated a model the student could reach.

Tighter proctoring does not repair this, because the defect is in what a single artifact can tell you.

Collect affirmative evidence in more than one form instead. A short in-class response written without tools, an oral explanation of one choice the student made, and the process artifacts behind a draft each show you something a finished document cannot. Two of those agreeing tell you more than the essay did. Designing assessment for an AI-present classroom works through how to build a secure writing sample worth comparing against.

6. A detector score is a reason to look, not a finding

Turnitin states that its own AI-writing indicator can misidentify human writing, AI-generated writing, and AI-paraphrased writing, and that it should not be the sole basis for adverse action. Vanderbilt University disabled that detector after weighing false positives, the opacity of the score, privacy, and the potential effect on students who are not native English writers. Research has separately documented systematic detector bias against non-native English writing.

The cost of a false positive does not land evenly. It falls hardest on the student with the least standing to argue back, and an accusation you cannot substantiate costs you the working relationship the rest of your teaching runs on.

Treat the score as a weak investigatory signal, then go get something you can act on:

  • a conversation with the student about what they wrote
  • comparison against a secure sample you watched them produce
  • the process artifacts behind the draft
  • the document’s version history
  • their source notes
  • an oral explanation of one decision in the piece
  • a short follow-up task

Every one of those gives you a finding. A percentage by itself never does.

7. Simulation is not testimony

A chatbot is asked to “be” an enslaved person, or a Holocaust survivor, and to answer students’ questions in character.

Asked to speak as someone who lived through something, a model invents experiences it has no record of and collapses many different lives into one composite. It delivers all of it in the register of a first-hand account. Students then reason from the performance the way they would reason from a document, in the one discipline whose central skill is telling evidence apart from everything that resembles it. Simulation is not testimony, and synthetic language is not a primary source.

Start from the documentary record and give the tool a job beside it. Authentic testimony and scholarship carry the account of what happened. The tool can generate the questions students should investigate, surface the tensions among sources you supplied, organize evidence students have already verified, or argue an abstract policy position rather than impersonate someone who was harmed.

None of that puts AI-generated perspective writing off-limits. Used as an object of critique, with the learning objective stated openly, it is one of the strongest things a model can do in a history classroom. The students audit the output, and the audit is what you grade. Put a fabricated diary entry beside a real one and watch how fast someone notices that the invented narrator states their own age, trade, and county in the opening line. A real diarist has no reason to write down what they already know about themselves.

Putting the displaced thing back

All seven make the same trade, which turns them into one question you can ask before the tool is ever open. What is this lesson supposed to produce in the student, and is that the part you are handing over? Everything else in the room is fair game, and most of it the tool handles well.

Asked from the positive side, that question is five tests for any classroom AI use. Kindred K-12’s AI in K-12 library takes up the other half of it: running the good uses well.

Back to top