The same tutoring score can mean different things to different AI models

arqiv✳

← All research storiesAI

PreprintAI

The same tutoring score can mean different things to different AI models

When an AI tutor decides whether a student is ready to move on, a cutoff like 0.65 doesn't mean the same thing across models.

How far alongPreprintShared publicly before peer review so others can see it early. Results may change, and experts have not formally checked it yet.
How soon could it reach you?A few years outPromising, but needs more testing or engineering first, likely 3 to 10 years.
Where it comes fromarXiv (preprint server)The University of Hong Kong · Sep 8, 2026Read the original →

What the scientists did

AI tutoring programs track what a student seems to know (this is called knowledge tracing) and decide when they’re ready for the next lesson, usually when a score passes a cutoff.

Researchers compared six knowledge-tracing models on four public education datasets, testing 12 possible cutoffs across 30 teaching scenarios.

Words to know

Threshold (cutoff)

The number the program's estimate must reach before it lets a student move ahead. The estimates run from 0 to 1, so 0.80 means 80%.

  • 0.65
  • 0.80
  • 0.95

Knowledge tracing

Software that estimates what a student knows based on their past answers.

What they found

The same number produced very different decisions depending on the model. Some models estimate whether a skill is mastered; others predict the chance of getting the next question right. Those aren’t the same thing, so the best cutoff changed from model to model. Higher cutoffs didn’t reliably shrink gaps between students, and could hold some students back unevenly.

6knowledge-tracing models compared
30teaching scenarios tested

Why it matters to you

If you use an online learning program that decides when you’re ready to move on, this is a reminder that “80% mastery” in one app may not mean the same as 80% in another. A student could be held back, or pushed ahead too soon, because of how the software reads that number.

Where this could lead

Fairer learning apps

Designers could calibrate cutoffs for each model instead of reusing a familiar number.

Few years

Better questions from teachers

Schools choosing software can ask vendors what their mastery scores actually measure.

Now

These are possibilities the research points toward, not promises. Most early findings take years of testing before they reach everyday life, and some never do.

Why scientists care

As AI tutors spread in classrooms, small design choices like cutoffs can decide who advances. This study shows those choices need checking, not assuming.

Keep in mind

This is a preprint that re-analyzes existing data; it didn’t run a new classroom trial. It shows a calibration problem but not which cutoff gives students the best results.

How far along is the science?

  • Peer-reviewedChecked by independent experts before a journal published it. The strongest level of evidence, but still not the final word.
  • PreprintShared publicly before peer review so others can see it early. Results may change, and experts have not formally checked it yet.
  • Conference talkPresented to other scientists at a meeting. Usually short and early; a full paper may come later.
  • Agency reportPublished by a science agency such as NASA, NOAA or NIH. Reviewed internally, but not by an outside journal.
  • Press releaseAn announcement from a university or company. Useful, but always check the research it describes.

How soon could it reach you?

  • Could reach you soonCould show up in products, services or advice within about 1 to 3 years.
  • A few years outPromising, but needs more testing or engineering first, likely 3 to 10 years.
  • Long-term scienceBasic research that builds knowledge. Real-world uses, if any, are far off.
  • Small stepA small update to earlier work. Useful to researchers, but not a change you would notice yet.

Every story links to its original source: the paper itself, or the conference program when no paper is out yet. Explanations are written with AI, usually from the paper’s abstract (each story says what it was written from), then checked against the source. We only use journals, recognized preprint servers, science agencies, scientific conferences and established science news outlets.