Skip to content
Discuss your product
EdTech and learning platforms

AI assessment and question bank generation

Writing good questions is expensive and is the most-postponed task on a learning platform: that's why courses have had the same assessment for three years, the answers circulate over WhatsApp between cohorts, and the item bank never grows.

AI handles volume well and quality badly if left alone. A model that reads a unit produces twenty questions in a minute, and fifteen of them measure whether the student read the paragraph. The difference between a useful bank and a decorative one lies in the process around generation, not in the model.

Short answer

Generating assessments with AI works when the item comes from your own material, declares which learning objective it measures, and carries distractors that correspond to real conceptual errors. Without those three conditions you get a surface quiz that measures recent reading, not learning. Every item is reviewed before publishing.

How it works

01

Each item is tied to a learning objective

Generation starts from the course's declared objective and the cognitive level being measured, not from loose text. Without this you can't know what the assessment covers or what it misses.

02

Distractors come from real errors

They're fed by the wrong answers students already gave in previous cohorts. An invented distractor is dismissed at a glance and turns the question trivial.

03

It's filtered before a person sees it

Semantic duplicates against the existing bank, items with more than one defensible answer, and questions whose answer is literally in the stem. It's an automatic filter, and it removes half.

04

A teacher approves and the item calibrates through use

Once published, success rate and discrimination index say whether the item works. Ones everybody gets right or everybody gets wrong leave the bank.

What gets measured

  • Generated items that survive teacher review, over total generated.
  • Bank coverage per learning objective: which objectives have no items.
  • Discrimination index of new items against existing ones.
  • Teacher hours per newly published assessment.

What's needed on your side

  • Learning objectives declared per unit. If they don't exist, that's the project's first deliverable.
  • The history of wrong answers — what makes distractors good.
  • A teacher with allocated time to approve. Review is the bottleneck, not generation.
  • Item statistics in the platform: success rate and discrimination per question.

When it isn't worth it

  • If nobody is going to review the items, don't do it. An unreviewed question bank is worse than a small one: it measures badly and nobody finds out until an entire cohort appeals.
  • If the assessment is certifying or for admissions, generation can propose but the validation process is the one the regulation already demands, and that doesn't get shortened.
  • If the course has ten students per cohort, there's no item statistics to calibrate against and the bank doesn't improve with use.

Related questions

Does it work for open questions or only multiple choice?
Both, but the value sits in different places. In multiple choice the hard work is the distractors. In open questions it's the accompanying rubric, without which the question can't be graded consistently by a person or a system.
How many items get discarded in review?
In the first batches, most of them — and that's useful information: it tells you what's missing from the prompt, the material or the declared objectives. Survival rate rises when generation starts feeding on the items the teacher already approved, not when you change models.
Can it generate a different assessment per student?
Technically yes, and it deserves a second thought. If each student gets different items, grades stop being comparable unless every item is calibrated to the same difficulty — which requires a mature bank. The reasonable route is variants of already-calibrated items, not generation on the fly.

Use cases in this industry

Rubric-based AI assisted grading

How to implement AI assisted grading in a learning platform: a draft grade per criterion, evidence quoted from the student, and the teacher's signature.

Semantic search over learning content

How to implement semantic search in a learning platform: over transcribed video and your own material, citing the timestamp and respecting the student's scope.

Next step

What education product should exist next?

Tell us what you are building, what is not working yet or which AI opportunity you want to evaluate. In the first conversation, we will tell you where we would start, what it requires and what we would not build.

Message us on WhatsApp