How to evaluate an AI tutor before opening it to students
The evaluation set for an educational tutor: the five case types it has to include, how correct abstention is measured, and what threshold gets agreed before the feature opens.
Short answer
RAG over educational content fails for reasons specific to the material: untranscribed video stays out of the index, blind chunking cuts examples in half, students ask using words that aren't in the text, the pedagogical sequence gets broken, and material has versions. None of them is fixed by changing model.
Retrieval-augmented generation works well over technical documentation: structured text, consistent vocabulary, each section answering one thing. Over course material it works noticeably worse, and teams arriving from the former are surprised.
The causes aren’t about the model. They’re five properties of educational content that break assumptions RAG takes for granted.
On many learning platforms, the explanation a student needs is spoken, not written. If the video isn’t transcribed, the index contains the supplementary notes and the module title, and the tutor answers from the written fraction of the course while the student knows the answer was in the class.
This gets perceived as “the AI doesn’t know the course”, when in reality the course was never shown to it.
Transcription is the first and most expensive job of the project in large catalogues. It’s also the highest-return one: without it, everything else is built on half the material.
The standard way to split a document — every so many characters, with overlap — assumes the text is homogeneous. Educational material isn’t: it’s made of units that only make sense whole.
A worked exercise has a prompt, a derivation and a result. An example has a setup and a conclusion. A proof has hypotheses and steps. If the cut lands in the middle, the index holds a fragment that answers nothing, and what gets retrieved is half a line of reasoning.
Chunking has to follow the pedagogical structure of the material, not a character count. It’s manual work and it’s the intervention with the best effort-to-improvement ratio in retrieval.
This is the domain’s central asymmetry, and the least anticipated.
The person asking is precisely the one who hasn’t learned the vocabulary yet. They write “the graph thing that goes up and then falls” when the material says “function with a local maximum”. They write “that semicolon error” when the material says syntax error. Semantic search helps considerably here, but it has a limit: it can’t connect two texts that share no expressed concept.
What works is enriching the index with student vocabulary: forum questions, support tickets, messages from the previous cohort. It’s content the platform already has and almost never indexes.
A course is designed so that certain things are learned before others. Retrieval doesn’t know that: if the best answer to a unit-two student’s question sits in unit seven, it brings it.
The result is a tutor that answers correctly and wrecks the course design, surfacing material the student isn’t equipped to use, and sometimes unlocking content from a plan they didn’t buy.
The scope filter — courses they’re enrolled in, units already unlocked — isn’t an awkward restriction bolted on afterwards. It’s part of what “correct answer” means in this domain.
Courses get updated. A regulation changes, a unit gets rewritten, an exercise gets corrected. Without a reindexing process, the index keeps the old version and the tutor cites it with complete confidence.
In education this error is worse than in other domains, because the student has no way to detect it: that’s why they’re taking the course. And the citation, which is what makes the system trustworthy, here makes it convincingly wrong.
The requirement is a reindexing process tied to content publication, and the version date visible in the answer.
When an educational tutor answers badly, the first question isn’t about the model. It’s where the chain broke:
| Symptom | Likely cause |
|---|---|
| Correct generalities, but not from the course | The passage isn’t in the index |
| Half a line of reasoning | Chunking cut the unit |
| Finds nothing for the student’s phrasing, finds it for the teacher’s | Vocabulary |
| Answers with material from a later unit | Missing scope filter |
| Cites something no longer in the course | Stale index |
None of those five rows improves by changing the model. All of them improve by working the content — the part of the project you can’t buy ready-made, and the part that decides whether the tutor is any good.
The evaluation set for an educational tutor: the five case types it has to include, how correct abstention is measured, and what threshold gets agreed before the feature opens.
Criteria for choosing a learning platform's first AI feature: what each one solves, what they cost, what risk they carry, and why the order is almost always the same.
How to calculate model cost per active student in a learning platform, what drives it up, and the four levers that bring it down without degrading the answer.