An IT Director's Checklist for a ChatGPT Pilot
A numbered checklist for IT directors piloting OpenAI's ChatGPT for Teachers before district-wide rollout — vendor review, opt-in, metrics, and escalation.
Free access to ChatGPT for Teachers through June 2027 will land on your desk as a request: "Can we turn this on for everyone?" The right answer is almost never "yes, today." It's "yes, after a pilot" — because a pilot is how you convert a vendor's promises and a superintendent's enthusiasm into evidence you can actually defend.
This is the checklist for running that pilot properly. Not a vague "try it and see," but a structured evaluation with a vendor review up front, a defined cohort, success metrics that aren't vanity, and an escalation path for when something goes sideways. Run it well and you'll either scale with confidence or say "not yet" with reasons. Both are wins.
Key Takeaways
- Do the vendor review before the pilot, not after. Read OpenAI's actual data terms, confirm the FERPA posture in writing, and document your findings before a single teacher signs in.
- Scope the pilot to learn, not to prove. 20-50 volunteer teachers across a few schools for 4-6 weeks generates real adoption and support data without district-wide risk.
- Pick five metrics before launch. Verification success rate, weekly active usage, support-ticket volume, privacy questions raised, and teacher-reported time saved — decided up front, not backfilled to look good.
- Build the escalation path first. Know exactly who handles a suspected data-exposure incident, a verification failure, and a policy question before one happens.
- A "not yet" is a valid pilot outcome. The point is a defensible decision, so a pilot that surfaces real blockers has done its job as well as one that greenlights rollout.
Before you touch the tool: the vendor review
The pilot starts on paper. Before any teacher enrolls, work through the vendor-review items — because these are the questions you cannot answer later by watching usage graphs.
Confirm what's actually being offered: the launch, reported by CNBC and GovTech, gives verified U.S. K-12 educators a workspace with a stated default that data isn't used for model training. Your job is to move that from a headline into your own documentation. Read the real data-handling terms — OpenAI's help center is the reference point — and get the FERPA-support and data-training claims confirmed in a form your district's counsel can rely on, not a marketing summary. The companion pieces on FERPA and what's protected and whether OpenAI trains on your lesson plans map the exact questions to bring to that review.
The pilot checklist
Here's the working checklist. Treat every row as a gate — you don't advance the pilot with an open item you can't explain.
| # | Item | What "done" looks like |
|---|---|---|
| 1 | Vendor data review | Data terms read; FERPA + no-training claims confirmed in writing |
| 2 | Cohort defined | 20-50 volunteer teachers named across 2-3 schools |
| 3 | Opt-in process | Teachers know what SheerID asks for and how to get help |
| 4 | Data policy communicated | Written "what to keep out of the box" guidance sent to every pilot teacher |
| 5 | Success metrics set | Five metrics defined before launch (see next section) |
| 6 | Escalation path built | Named owner for incidents, verification failures, policy questions |
| 7 | Support channel open | A single place teachers report problems; you watch ticket volume |
| 8 | Review date scheduled | End-of-pilot decision meeting on the calendar before day one |
Items 3 and 4 are where pilots quietly leak risk. The opt-in runs through SheerID's third-party verification, so expect a handful of teachers to get flagged — plan a help queue. And "data policy communicated" means every pilot teacher has, in hand, a plain-language rule about keeping identifiable student information out of a general workspace. Give them a safe first task too — building next week's plans in something like Lesson Plan Studio touches no student records — so the pilot generates real usage instead of cautious paralysis.
Defining success: metrics that aren't vanity
The failure mode of an ed-tech pilot is measuring "logins" and declaring victory. Logins are vanity. Pick metrics that tell you whether the tool is working and safe, and pick them before launch so nobody's tempted to grade on a curve afterward:
- Verification success rate — what percentage of your cohort got through SheerID without a support ticket? A low rate here predicts a painful district-wide rollout.
- Weekly active usage — are teachers using it in week 5, or did they abandon it after week 1? Sustained use is the only adoption signal that matters.
- Support-ticket volume — normalized per teacher. This is your rollout-cost forecast.
- Privacy questions raised — track them. Frequent, similar questions tell you exactly what your district-wide communication needs to preempt.
- Teacher-reported time saved — a short survey. This is the number your superintendent actually cares about, and the one that justifies the whole exercise.
The escalation path: build it before you need it
The item most pilots skip is the one you'll be judged on if it's missing. Decide, in advance and in writing, who owns each of these:
- A suspected data-exposure incident — a teacher pastes something they shouldn't have, or worries they did. Who do they tell, and what happens next? Because a free AI workspace holds classroom materials, the stakes are real; the scenario in what a K-12 AI data breach would actually mean is worth reading so your response plan isn't improvised.
- A verification failure — a teacher can't get through SheerID. Who resolves it, and how fast?
- A policy question — "am I allowed to use this for X?" Route these to one named person so answers stay consistent.
An escalation path isn't bureaucracy; it's the thing that lets you say "yes, we're ready" honestly, because you'll know how problems get handled rather than hoping none arise.
Scope the pilot to learn, not to prove
The mindset that makes all of this work: you are running the pilot to learn the truth, not to manufacture a yes. An IT director who scopes a pilot to confirm a decision already made isn't piloting — they're rubber-stamping. Go in genuinely willing to conclude "not yet," and the pilot becomes useful.
That posture also sets up the real rollout. Everything you learn — the verification-failure rate, the privacy questions, the support load — feeds directly into the district-wide sequence covered in District Rollout 101. And your pilot should deliberately surface the integration questions too, because the gaps around single sign-on and rostering (detailed in the SSO and LMS gap analysis) are exactly the kind of thing you'd rather discover with 50 teachers than 5,000.
Frequently Asked Questions
How long should a ChatGPT for Teachers pilot run?
Four to six weeks is the sweet spot. It's long enough to see whether adoption survives past the novelty of week one, but short enough to keep momentum and reach a decision while stakeholders are still engaged.
How many teachers should be in the pilot?
Twenty to fifty volunteers across two or three schools. That's enough to generate meaningful support-load and adoption data without exposing the whole district, and volunteers give you a fairer read on real-world usefulness than conscripts.
What's the single most important thing to check before a pilot?
The vendor data review. Confirm OpenAI's FERPA-support and no-training-by-default claims against the actual terms — not the marketing — and document it. Everything else in the pilot assumes that groundwork is solid.
What if the pilot goes badly?
Then the pilot did its job. A "not yet" backed by specific blockers — poor verification rates, low sustained usage, unanswered privacy questions — is a defensible, valuable outcome that saves you a far more expensive district-wide failure.
Do we need OpenAI's involvement to run a pilot?
Not to start — verified teachers can self-enroll. But for anything beyond a small pilot you'll want to confirm managed-account and data terms with OpenAI directly, especially the parts not detailed in public launch coverage.
A good pilot is boring on purpose: a vendor review you can defend, a scoped cohort, five honest metrics, and an escalation path nobody had to invent under pressure. Do that, and the district-wide decision makes itself — you'll either scale on evidence or hold the line with reasons. The IT directors who get AI rollouts right aren't the fastest to say yes; they're the ones who made "yes" mean something.
Part 12 of 100 in the ChatGPT for Teachers series. Previously: Does OpenAI Train on Your Lesson Plans?. Next: SSO, Rostering, and the LMS Gap in ChatGPT for Teachers. Browse more builder insights or explore AI skills for education at aiskill.market.