MATHEMATICS EXPERT · EVIDENCE & CORRECTABILITY
Mathematics Evidence and Validation Harness | Testing the BTT Framework
A teaching framework should be willing to change when evidence disagrees with it. Bukit Timah Tutor therefore treats its explanations, learning connections and teaching approaches as working models that should be checked against student performance, curriculum requirements and relevant research.
Three questions should stay separate
- Is the mathematical connection reasonable? For example, does an earlier idea genuinely support the later one?
- Is the diagnostic question informative? Does it help us understand the student’s difficulty rather than simply produce another score?
- Does the teaching response work? Does the student improve and retain the improvement when the question changes?
Research-informed does not mean infallible
External research can support principles such as formative assessment, worked examples, retrieval practice or feedback. It does not automatically prove that every local teaching sequence, diagnostic question or intervention is correct for every learner.
That is why actual student work matters. If a predicted improvement does not appear, the explanation should be reconsidered rather than protected.
What improvement should survive
- A changed-looking question.
- Less prompting from the tutor.
- A delay before the idea is used again.
- A mixture of topics rather than one repeated exercise type.
- Reasonable examination pressure where that is part of the student’s goal.
A framework should keep a memory of failure
Teaching improves when contradictions are not quietly discarded. Cases where an explanation was wrong, a question misled us or an intervention failed are useful information. They help refine future decisions and prevent confidence from becoming certainty without evidence.
Good teaching does not merely collect successes. It learns from the cases that do not behave as expected.
PHASE 4 · EVIDENCE & VALIDATION READER GUIDE
Quick Read: how should a Mathematics teaching framework earn confidence?
A teaching framework earns confidence when its mathematical claims, diagnostic interpretations and interventions survive independent checks against student work, curriculum demands, relevant research, changed questions and later retention.
Three different questions must remain separate. A mathematical connection can be valid while a diagnostic probe is uninformative. A diagnosis can be correct while a chosen intervention fails. A teaching method can produce short-term success while the learning disappears when prompts are removed. Treating these as one “evidence” question makes weak claims look stronger than they are.
One-sentence answer: mathematical validity, diagnostic validity and intervention effectiveness should be tested separately, then brought back together only when the evidence supports the whole chain.
Three claims, three kinds of evidence
| Claim | Question | Useful evidence |
|---|---|---|
| Mathematical dependency | Does the earlier idea genuinely support the later one? | Mathematical structure, curriculum progression, research, counterexamples, learner performance. |
| Diagnostic interpretation | Does this observation distinguish the proposed cause from alternatives? | Discriminating probes, changed representation, repeated observations, error-pattern comparison. |
| Intervention effectiveness | Does this teaching response improve the learner’s capability? | Reduced prompting, changed-surface success, delayed retention, reconnection to current work. |
Confidence should attach to the specific claim that the evidence supports. Strong research about retrieval practice, for example, does not prove that one particular worksheet or diagnostic sequence is appropriate for one particular learner.
Student work is not merely an outcome—it is a test of the model
Suppose a tutor believes weak distributive reasoning is causing repeated algebra errors. That explanation should generate predictions. The learner should struggle with numerical distribution, symbolic distribution or reverse factorisation in a patterned way. If those tasks are strong, the proposed explanation loses confidence.
- State the hypothesis. What do we think is causing the visible error?
- Predict another consequence. Where else should the weakness appear?
- Probe independently. Use a changed number, representation or context.
- Compare prediction with observation.
- Update the explanation. Keep, weaken or reject it.
This keeps diagnosis falsifiable. The model has to answer to the learner’s actual Mathematics.
Contradictions are valuable evidence
A framework becomes weaker when contradictory cases are explained away automatically. If a supposed prerequisite is missing but the learner performs the later task reliably, that case matters. If an intervention produces immediate improvement but no retention, that matters too.
- Unexpected success can show that a dependency is helpful rather than necessary.
- Unexpected failure can reveal a missing condition or hidden prerequisite.
- Short-term gain with later loss can reveal support dependence rather than durable learning.
- Transfer failure can reveal that practice was too surface-specific.
A robust framework keeps a memory of where its predictions failed.
Evidence has different strengths
Not all evidence answers the same question. A useful hierarchy is not “research beats classroom” or “classroom beats research.” Each source contributes something different.
| Evidence source | Strongest use | Important limit |
|---|---|---|
| Mathematical structure | Establish logical or conceptual relationships. | Does not by itself prove the best teaching sequence. |
| Official curriculum / syllabus | Clarify required content, stage expectations and assessment boundary. | Does not explain every learner’s mechanism. |
| Education research | Inform principles and likely effects across groups/settings. | May not transfer unchanged to every local context or learner. |
| Student work | Test what this learner currently understands and can deploy. | One response can be noisy or misleading. |
| Transfer and retention | Test whether learning has become portable and durable. | Requires time and changed conditions. |
The strongest decision usually comes from convergence rather than one source carrying the whole claim.
A teaching response must survive more than same-day success
- Understanding: can the learner explain why the correction works?
- Reconnection: can the repaired idea be used in the original current topic?
- Fade: can prompts, examples or tutor guidance reduce?
- Changed surface: does the idea survive a different-looking question?
- Mixed context: can the learner choose the method when the chapter is not announced?
- Delay: is the capability still available later?
- Pressure: where relevant, does it remain under realistic examination conditions?
A repair that passes only the first step is promising, not proven.
Three examples of evidence changing the teaching story
Case 1: assumed prerequisite is not actually blocking performance
A learner has imperfect mental arithmetic but solves algebraic equations accurately and efficiently with reliable written work. The evidence suggests mental speed may be helpful, but it is not the current limiting prerequisite. Repair should not automatically divert the learner away from the target topic.
Case 2: diagnosis is plausible but intervention fails
A tutor correctly identifies weak fraction magnitude and uses one visual method. The student succeeds while the model is present but fails after it is removed. The diagnosis may still be right; the intervention has not yet produced independent capability.
Case 3: immediate improvement does not retain
A student performs perfectly after a worked example but cannot reconstruct the method a week later. Same-day success overstated the strength of learning. The next cycle needs retrieval, spacing or stronger relational understanding rather than simply recording the topic as “mastered.”
What parents should expect from evidence-based teaching
- Explanations should be specific enough to be checked.
- A child should not receive a permanent label from one mistake.
- Teaching methods should change when they are not producing improvement.
- Progress should be verified on changed work, not only repeated worksheets.
- Support should generally reduce as capability improves.
- Retention and examination use should matter when they are part of the learner’s goal.
Evidence-based does not mean teaching becomes mechanical. It means claims about what a learner needs remain answerable to what the learner subsequently demonstrates.
Frequently asked questions
Does research tell us exactly how to teach every child?
No. Research can inform principles and likely effects, while individual student evidence helps determine how those principles should be applied, combined or reconsidered in a particular case.
Can a mathematically correct dependency still be a poor diagnostic tool?
Yes. A connection may be real while a particular question fails to distinguish the suspected weakness from other causes.
What is stronger evidence than a correct answer immediately after teaching?
Independent success with less support, on a changed question, after some delay and inside a mixed context provides stronger evidence of durable learning.
What should happen when the framework is wrong?
The explanation, dependency or intervention should be revised. Protecting a model from contradictory student evidence defeats the purpose of using evidence at all.
The larger idea: good teaching is correctable
Teaching requires judgement because learners are not identical systems. The safeguard is not pretending uncertainty can be removed completely. The safeguard is making educational claims testable and remaining willing to update them when reality disagrees.
That is what keeps a framework useful over time: not that it never makes a mistake, but that mistakes become evidence for a better next version.
A strong teaching framework does not ask the learner to fit the model. It asks the model to remain accountable to the learner.
When evidence needs a controlled experiment: route into the BTT Mathematical Lab. The Evidence & Validation Harness remains the canonical evidence owner; MathLab supplies the bounded probe, intervention, transfer or retention test, then returns a traceable evidence receipt instead of treating one correct answer as proof.
