Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Please Don’t Let the AI Decide Who Your Child Is From One Test

AI design essay: This article sets out the standards we would want educational AI to meet. The examples illustrate proposed behaviour; they do not establish that a deployed BTT AI currently performs these functions. Mathematics advice should be checked against the learner’s actual work and the teacher’s evidence.

One of the easiest mistakes in education is to turn a temporary state into an identity.

“She failed this test” becomes “She is weak at Mathematics.”

“She made repeated careless errors” becomes “She is careless.”

“She froze during the exam” becomes “She cannot handle pressure.”

I do not want BTT AI doing that.

A test tells us what happened under particular conditions

That is already valuable.

It may show that a concept is unstable.

That method selection is unreliable.

That time pressure is exposing a weakness.

That the student cannot yet transfer learning to unfamiliar wording.

But “cannot yet” is very different from “is not capable”.

Describe the pattern. Do not turn the pattern into the person.

Language matters because children hear the labels too

If an AI tells a parent, “Your child is weak in algebra,” that sentence may sound efficient.

I would rather it say:

“The recent work suggests algebraic manipulation is currently less reliable than the conceptual understanding.”

That is longer.

It is also more accurate.

One describes a current relationship in the work.

The other quietly defines the child.

Patterns become more trustworthy when they repeat

One test may be enough to notice something.

It is rarely enough to freeze a conclusion.

I want the AI to ask whether the same issue appears across:

different topics;

different weeks;

schoolwork and tests;

timed and untimed conditions.

If it repeats, confidence increases.

If it disappears, the system should update.

Improvement should be allowed to change the description

This seems obvious, but educational labels are surprisingly sticky.

A child was weak at fractions in Primary school.

Years later, everyone still talks about her as though fractions are the explanation for every new difficulty.

I want BTT AI to remain willing to say:

“That was a real weakness earlier. The recent evidence no longer supports treating it as the main problem.”

That is what it means for a model of the learner to remain answerable to the learner.


I want BTT AI to be very good at describing what is happening.

I do not want it to become confident about who a child is.

The work can be weak.

The method can be unstable.

The performance can be poor.

Those are things we can examine and change.

The child should remain larger than the diagnosis.

A temporary state should become a repairable next edge

Once the pattern is described carefully, I want one practical consequence.

If the current state is “method selection is unreliable”, then repair method selection. Do not quietly expand that into “weak at Mathematics”.

Then test the repair somewhere different: another topic, another week, an unfamiliar wording, or a timed condition.

If the difficulty survives the change, the pattern deserves more weight.

If it does not, update the description.

That is how we keep diagnosis useful: specific enough to repair, provisional enough to change.

Sources and Further Reading

Discover more from Bukit Timah Tutor

Subscribe now to keep reading and get access to the full archive.

Continue reading