How to evaluate an AI tutor before opening it to students
The evaluation set for an educational tutor: the five case types it has to include, how correct abstention is measured, and what threshold gets agreed before the feature opens.
Real costs, criteria for deciding, and the failures that only show up once the agent is handling actual customers. We publish what we wish we had read first.
The evaluation set for an educational tutor: the five case types it has to include, how correct abstention is measured, and what threshold gets agreed before the feature opens.
Criteria for choosing a learning platform's first AI feature: what each one solves, what they cost, what risk they carry, and why the order is almost always the same.
The five reasons retrieval over course material fails more than over technical documentation, and what has to change in how the content is prepared.
How to calculate model cost per active student in a learning platform, what drives it up, and the four levers that bring it down without degrading the answer.
What an evaluation set for an AI agent is, how to build one from real cases in your operation, and how to use it to catch a quality drop before a customer complains.
The six failures that show up once an AI agent leaves the demo and starts handling real customers, why none of them are visible in testing, and how each one is prevented.
Real investment ranges for automating a business process with AI in 2026, what actually drives the price up, and what it costs to keep running after month one.