What ChatGPT for Teachers Still Can't Do
A free workspace can draft, differentiate, and grade fast — but it can't build trust, manage a room, or know your students. An honest look at the hard limits.
Thirty-five essays into this series, ChatGPT for Teachers can start to look like it does everything. It plans the week, differentiates one lesson across five reading levels, drafts the parent email, builds the quiz, and writes the sub plan while you finish your coffee. That's a real and useful list. It is also, deliberately, a list of tasks — discrete, describable things that begin and end. The job of teaching is not a list of tasks. It's a relationship carried out over a room of thirty individuals across nine months, and the most important parts of it never fit inside a chat window.
This is the honest limits piece, and it closes the workflows cluster on purpose. Before the next several essays wade into the real debate — the skeptics, the cheating data, the mixed reviews — it's worth being clear-eyed about what this tool genuinely cannot replace. Not "can't do well yet." Can't, structurally, because the thing being asked for isn't the kind of thing a language model produces.
Key Takeaways
- It doesn't know your students. The model has no memory of which kid's parents just separated, who reads two years below grade, or who needs to be asked twice. That context lives in you, and it drives every good decision you make.
- It can't manage a room. Behavior management is physical, relational, and real-time — presence, proximity, tone. A chatbot can suggest a strategy; it cannot notice the back corner going sideways.
- It shouldn't own a legally-required process. IEPs, 504 plans, and mandated reporting carry legal weight and human accountability. ChatGPT can draft supporting language; it cannot be the responsible party.
- The final grading judgment stays with you. A model can score against a rubric and speed the first pass, but the call that a struggling student earned a break — or didn't — is a professional judgment, not an output.
- The limits are the point, not a bug. Knowing what the tool can't do is what lets you trust it with what it can. Offload the tasks; keep the teaching.
It doesn't know your students
A language model responds to what's in front of it. It has no standing knowledge that the quiet kid in row three just lost a grandparent, that two students in your fourth period cannot be seated together, or that one of your strongest writers freezes completely on timed assessments. That information isn't in any file you can upload cleanly, and even if it were, the reason it matters is that you weigh it in the moment — softening a deadline here, pushing harder there, reading a facial expression and changing course mid-sentence.
OpenAI's own framing of ChatGPT for Teachers positions it as a workspace for preparation and administrative lift, not as a stand-in for the teacher-student relationship — and that's the correct framing. The tool is genuinely good at the generic version of a task: a differentiated lesson for "a class with a range of reading levels." It cannot produce the specific version — the one shaped by knowing that this range, in this room, includes a newcomer who arrived in March and a gifted reader who's bored into disruption. That translation from generic to specific is teaching, and it happens in your head.
It can't manage a room
Behavior management is the clearest hard limit because it's so obviously physical. Classroom management runs on presence, proximity, timing, and a hundred micro-decisions per hour that depend on watching a room breathe. You move toward the noise before it becomes a problem. You catch an eye. You change your voice. None of that is text, and none of it can be delegated to a system that isn't in the room.
You can, of course, ask ChatGPT for a de-escalation strategy or a seating plan or a script for a hard conversation with a student — and those are reasonable uses. But there's a category error waiting here. A strategy is not the same as the execution of a strategy under live conditions with a thirteen-year-old who is testing you. The broader sentiment data reflects exactly this tension: a large share of teachers report that AI can add work rather than remove it, precisely because the parts of the job that exhaust them most — managing behavior, holding attention, sustaining relationships — are the parts no tool touches.
It shouldn't own a legally-required process
Some parts of the job carry legal weight, and those are the ones where "the AI drafted it" is not an acceptable answer. Individualized Education Programs (IEPs), 504 plans, and mandated-reporting decisions sit inside statutory frameworks with named human accountability. A teacher or case manager signs their name. A district assumes liability. The process assumes a responsible professional made the call.
ChatGPT can genuinely help around these processes. It can rewrite a goal in plainer language, suggest accommodations to consider, or turn dense clinical notes into parent-friendly prose — and this series covered that use in writing IEP-friendly materials with ChatGPT. But there's a bright line between drafting supporting language and owning the decision. FERPA and IDEA don't recognize a chatbot as a responsible party, and no district should let workflow convenience blur that. The rule of thumb is simple: the tool can hold the pen, but a human holds the pen's accountability.
It shouldn't make the final grading call
Grading is where the limit gets subtle, because a model can score against a rubric, and doing so is a legitimate time-saver on a first pass. The problem is that grading isn't only measurement. It's judgment layered on top of measurement — the decision that a student who's been through something this month gets some grace, or that a capable kid coasting on minimum effort gets the grade that finally gets their attention. Those are professional calls informed by everything you know about the person, and a rubric score is only an input to them.
| The task | What ChatGPT can do | What stays with you |
|---|---|---|
| Grading | Score a first pass against a rubric, flag patterns | The final mark, and the human judgment behind it |
| Behavior | Suggest strategies and scripts | Reading and managing the live room |
| IEPs / 504s | Draft and simplify supporting language | Legal accountability and the signed decision |
| Knowing students | Handle the generic case | The specific case, shaped by real context |
Use the model to get faster to the judgment, not to replace it. A teacher who lets a first-pass AI score become the grade has handed away the one part of assessment that requires them specifically.
Where the line actually sits
None of this is an argument against the tool. It's the opposite: knowing precisely what ChatGPT can't do is what makes it safe to lean on for what it can. The pattern across all four limits is the same — the model is excellent at the task and useless at the relationship and accountability wrapped around it. Offload the drafting, the formatting, the first passes, the repetitive lift. Keep the knowing, the managing, the judging, and the signing.
That's also why the debate the next several essays engage with is worth taking seriously rather than dismissing. When EdWeek weighed the launch as a boon, a bust, or just "meh", part of what it was measuring is exactly this gap between the impressive task list and the untouched core of the job. A tool that saves you an hour of prep is genuinely valuable. A tool sold as replacing the teacher would be selling something that doesn't exist. For the parts it does serve well, purpose-built helpers like Lesson Plan Studio and structured education agent skills tend to beat a blank chat window — but even the best of them respects the same line.
Frequently Asked Questions
Can ChatGPT for Teachers write an IEP?
It can draft and simplify language that supports an IEP — rewriting goals in plainer terms, suggesting accommodations to consider, or making clinical notes parent-friendly. It cannot own the IEP. The plan carries legal weight under IDEA, and a named human professional is accountable for it. Treat the model as a drafting aid, never the responsible party.
Will it replace teachers?
No — and OpenAI doesn't market it that way. The product is built for preparation and administrative work, not for the relational, in-room, judgment-heavy core of teaching. It removes tasks; it doesn't remove the teacher. The parts of the job that most define it — knowing students, managing a room, exercising judgment — sit entirely outside what a language model produces.
Is it safe to let ChatGPT assign final grades?
It's safe to let it score a first pass against a rubric to save time, but the final mark should be a human decision. Grading blends measurement with judgment informed by everything you know about a student, and that judgment is the part only you can supply. Use the model to reach the decision faster, not to make it for you.
Can it handle classroom behavior management?
Only in the abstract. It can suggest de-escalation strategies, seating plans, or scripts for hard conversations. It cannot execute any of them, because behavior management is real-time, physical, and relational — reading a room and responding in the moment. The strategy is not the same as the live handling of it.
So what should I actually use it for?
The tasks with clear inputs and outputs: lesson drafts, differentiation, quiz generation, parent emails, sub plans, and grading first passes. Those are covered across this series' workflows cluster. Keep the relationship, the room, the legal accountability, and the final judgment for yourself.
Part 36 of 100 in the ChatGPT for Teachers series. Previously: Prompt Patterns Every Teacher Should Steal. Next: EdWeek's Verdict: Boon, Bust, or Just Meh?. Browse more builder insights or explore AI skills for education at aiskill.market.