Build Your Tutor Stack: Four Skills, Not Forty
You don't need forty education skills. A working AI tutor is four composable behaviors — question, explain-and-ground, drill, and assess — assembled from the marketplace. Here's the stack.
The temptation, once you discover education skills, is to install all of them. Every subject tutor, every flashcard tool, every essay grader, every study coach — forty skills, and the vague hope that somewhere in the pile a real tutor is hiding. It isn't. A pile of overlapping skills doesn't teach; it argues with itself, and your agent splits the difference into mush. The builders whose AI tutors actually move a learner don't have the biggest stack. They have the right four.
That's the whole argument of this closing piece. A working AI tutor is not forty behaviors — it's four, running in a loop: question, explain-and-ground, drill, and assess. Each is a distinct move. Each maps to a specific, real skill on the marketplace. And because the four don't overlap, they compose into one clean workflow instead of a tangle. This is the practical payoff of everything the series argued: teaching is a loop you can assemble, and the parts are already on the shelf.
The whole tutor is four moves
Strip tutoring down to what a good human actually does across an hour and you get a cycle, not a monologue. They ask a question to find out where you are. When you're stuck at the edge of what you can reach, they explain — but they ground the explanation in something checkable, not vibes. Then they make you practice the thing until it sticks past tomorrow. And before they let you walk away believing you've learned it, they make you prove it back under test conditions. Question, explain, drill, assess. Then around again.
Every one of those moves is a behavior you install, not a capability you wait for — which is the thesis the series opened with, that AI tutoring is a skills problem, not a model problem. The four-move loop is why. Each move is a governed restraint the raw model won't perform on its own. Left uninstructed, the model explains when it should question, asserts when it should ground, tells you "great job" when it should test, and never drills at all. The stack fixes each failure with one skill apiece.
A tutor isn't a skill you install. It's a loop you assemble — question, explain-and-ground, drill, assess — from four parts that each own one move and refuse to do the others' jobs.
The stack, in one table
Here's the whole thing in a single view: the four behaviors, the job each does in the loop, and the marketplace skill that owns it. There's an optional fifth row we'll get to — it's for the human, not the material.
| Behavior | Job in the loop | Skill(s) |
|---|---|---|
| Question | Diagnose where the learner is; withhold the answer | socratic-tutor |
| Explain & ground | Give the smallest correct explanation, checked against a source of truth | concept-explainer + wolfram-alpha-llm-mcp |
| Drill | Convert understanding into durable memory via spaced repetition | flashcard-deck + anki-connect-mcp |
| Assess | Make the learner prove it back under test conditions | exam-blueprint |
| Wellbeing (optional) | Keep the human regulated enough to learn | exam-anxiety-reset |
Four core slots, four jobs, no overlap. That last property is the point. When two skills both claim to "explain," your agent reconciles them toward the median and you lose the edge you installed them for. When each skill owns exactly one move, they reinforce instead of contradict — and the loop runs clean.
Question and ground: the front half of the loop
The front half is where most AI tutors quietly fail, because it's where the model is most tempted to just help. The fix is two skills that pull in opposite directions on purpose.
socratic-tutor owns the question. Its entire job is to withhold the answer and hand back a question that surfaces your thinking — to diagnose the misconception under the wrong answer instead of correcting the surface. This is the oldest move in teaching, and the one a fluent model breaks instantly if you don't govern it. Install this and the tutor stops being an answer key and starts being an interviewer.
But questioning alone strands a learner who genuinely can't reach the next step. That's where explain-and-ground comes in — and it's deliberately two skills, not one. concept-explainer supplies the explanation itself: the smallest correct scaffold, an analogy sized to the learner, the next rung and no more. Then wolfram-alpha-llm-mcp does the part the model can't be trusted to do alone — it grounds the explanation in a source of truth. When the tutor claims a derivative, a unit conversion, or a numeric result, that claim gets checked against a computational engine rather than confabulated. A tutor that explains fluently but sometimes explains wrong is worse than no tutor, because a confident error is stickier than a gap. Grounding is what makes the explanation safe to trust — the case the series made in full in grounding the tutor with tools, not talk.
Drill and assess: the back half that makes it stick
Explaining something well is not the same as making it last. The Ebbinghaus forgetting curve is brutal and old news: without deliberate review, most of what you understood on Tuesday is gone by Friday. Understanding is necessary and nowhere near sufficient. The back half of the loop exists to beat the curve.
flashcard-deck owns the drill: it turns the concept you just grasped into retrieval-practice cards — the format that decades of memory research keeps confirming beats re-reading, because the effortful act of recalling is what strengthens the trace. Then anki-connect-mcp pushes those cards straight into Anki, the spaced-repetition system that schedules each card to reappear exactly as you're about to forget it. This is the one piece you cannot fake with cleverness in the moment — spacing is a function of time, and time needs a scheduler that lives outside the conversation. Generate in the tutor, drill in Anki, forever.
Assess closes the loop, and it's the move every soft tutor skips. exam-blueprint makes the learner prove the concept back under test conditions — not "does this make sense?" but produce it unaided, cold, the way the real assessment will demand. Retrieval under pressure is both the truest measure of whether learning happened and, conveniently, a powerful way to cause more of it. If the assessment surfaces a gap, you've got a fresh diagnosis — which is just the question move again, and the loop turns over. Assess isn't the end of the pipeline. It's the thing that feeds the next lap.
The optional fifth slot is for the human
There's one more skill worth a slot, and it's categorically different from the other four. The core loop teaches the material. But a learner who is panicking doesn't have the working memory to learn anything — anxiety eats the exact cognitive bandwidth that questioning and drilling need. exam-anxiety-reset is the wellbeing layer: a skill that regulates the human so the pedagogy can land. Treat it as optional because not every learner needs it every session — but when the blocker is the nervous system and not the concept, no amount of Socratic questioning helps. Keep it in the stack, invoked by state, not by default.
Assemble it this afternoon
That's the whole build. Four skills for the loop, one optional layer for the human, and no forty-first skill you have to talk yourself out of installing. The reason a tight stack wins is the same reason a good tutor wins: every part knows its one job and refuses to do the others'. socratic-tutor questions. concept-explainer and wolfram-alpha-llm-mcp explain and ground. flashcard-deck and anki-connect-mcp drill. exam-blueprint assesses. Compose those and you've built the loop that Bloom's two-sigma finding said a good one-to-one tutor runs (Bloom, 1984) — not by scaling a model, but by stacking four behaviors.
So don't go collect forty education skills. Go browse the AI tutoring category, pull the four moves, add the wellbeing layer if your learner needs it, and wire them into one workflow. That's ten essays' worth of pedagogy reduced to a build you can finish before dinner — because a tutor was never a bigger brain or a longer skill list. It was always four moves in a loop, and now they're on the shelf, waiting for you to assemble them.
Part 10 of the AI Tutoring series. Previously: Grounding the Tutor. Browse the AI tutoring skills or more builder insights.