How to evaluate an AI tutor before opening it to students
The evaluation set for an educational tutor: the five case types it has to include, how correct abstention is measured, and what threshold gets agreed before the feature opens.
Short answer
An AI tutor's cost is measured per active student per month, not per query. It depends on three variables: messages per student, context retrieved per message, and the model chosen. In EdTech the unit economics are tight, so the math comes before the prototype: if it exceeds the margin of the plan selling the feature, there's no product.
In almost every sector we work in, the cost question arrives last: decide what to build, then find out what it costs. In EdTech that order doesn’t work, and it’s the mistake that sinks the most projects.
The reason is unit economics. A student plan costs what it costs, usage is high and sustained through the academic term, and an AI feature with a variable per-student cost comes straight out of a margin that was already finite.
Not the token, not the query, not the course. Cost is measured per active student per month, because that’s the unit revenue is expressed in.
The calculation has three variables:
Messages per active student per month. The hardest number to estimate before launch and the one that varies most. A reasonable proxy is current forum and support volume multiplied by a factor: an always-available tutor receives considerably more questions than a forum where you have to wait for an answer.
Context retrieved per message. How many passages of material go into each answer, and how long they are. This is the variable the team controls and the one most underestimated.
The model chosen. With an order-of-magnitude difference between the small and large models of the same family.
Multiplied out, they give the number you compare against the plan’s margin.
Almost always the second variable, and almost always through the same architectural decision.
The fastest way to build a tutor is to include a lot of context: the whole unit, the full conversation history, related material just in case. It works well, it ships in days, and it multiplies cost several times over without improving the answer proportionally.
The second cause is history. A long tutoring conversation drags every prior message into each turn, so the twentieth message costs several times what the first did. Without a summarization strategy, long conversations are the expensive ones — and they belong precisely to the most engaged students.
Ordered by return per unit of effort.
Tight retrieval. Bring the passages that answer the question, not the whole unit. It requires material split by conceptual unit and properly indexed, which is upfront work and the same work that makes the answer good. It’s the only lever that lowers cost and improves quality at the same time.
Model routing. Most course questions are comprehension over already-retrieved material and don’t need the most expensive model. A cheap classifier up front decides, and the difference on the invoice is large.
Caching the repeated. In a course, questions cluster heavily: the same twenty doubts cover a high share of volume, especially around assignment deadlines. Recognizing that and answering from cache is straightforward and very profitable.
History summarization. Compressing old turns instead of carrying them whole. This is the lever that stops long conversations from dominating the bill.
It’s arithmetic, which is why it’s better done before rather than after.
On the revenue side: what the plan leaves per student per month, after everything else. On the cost side: the number from above.
If the estimated tutor cost is a small fraction of the margin, there’s a product and the conversation becomes about quality. If it’s a large fraction, the feature has to go in a higher tier, carry a usage ceiling, or not ship.
And if the estimated cost exceeds the plan price, what you have isn’t a product but a loss-making promotion dressed as innovation. It happens more often than it seems, because the prototype is built with twenty internal users and nobody multiplies by the whole base.
Before the tutor prototype, semantic search over the same content. It costs an order of magnitude less, it’s measurable without touching assessment, and it produces exactly the missing data: what students ask, in what words, and how often.
With that data, the estimate of messages per student stops being a guess. And since search needs the same transcribed, chunked and indexed content the tutor does, none of the work is wasted: it’s the first half of the same project.
The evaluation set for an educational tutor: the five case types it has to include, how correct abstention is measured, and what threshold gets agreed before the feature opens.
Criteria for choosing a learning platform's first AI feature: what each one solves, what they cost, what risk they carry, and why the order is almost always the same.
The five reasons retrieval over course material fails more than over technical documentation, and what has to change in how the content is prepared.