A chapter of the free book Learn to Learn

Published 2026-08-31 · updated 2026-08-31

Also in: සිංහල · தமிழ்

Before You Trust Any AI With Your Exam: The Four Questions

Before You Trust Any AI With Your Exam: The Four Questions

Modern chatbots are genuinely capable — and they are trained on the whole world's material, not marked against your marking scheme. This chapter of Learn to Learn gives you a four-question test for judging ANY AI study tool, ours included, and names the two risks that fluent answers hide from exam candidates.


The short answer: Use ChatGPT freely for exploration and tutor-mode questions — but don't hand it your exam preparation unchecked, because it answers in the world's dialect rather than your marking scheme's, and it cannot mark your handwritten script against the official rubric. Judge any AI study tool by four questions: checked against what? guarantee maintained by whom? where do my confusions go? can it mark my handwriting?

Let's start by being fair, because you will not trust this chapter otherwise.

Today's ChatGPT and its rivals are excellent. They remember your past chats, and you can tell them once — "I'm a Sri Lankan A/L physics student" — and be remembered. If you use one daily, you have probably tuned it well, and a chapter that pretends otherwise would deserve your eye-roll.

So this chapter makes a narrower, more honest claim: for the specific job of preparing for a Sri Lankan national exam, a fluent answer is not the same as a mark-scoring answer — and the gap between the two hides in places you can't check by feeling. Here is where it hides, and the four questions that find it in any tool.

Risk one: the wrong dialect

Your exam has a dialect. The marking scheme rewards specific NIE terminology, specific notation conventions, specific diagram labels, specific ways of laying out working. A general model learned physics from the whole world — American AP courses, British A-Levels, Indian JEE prep, university textbooks — and its answers average across all of them.

So the failure mode that actually costs you marks is rarely "the physics is wrong." It's right physics, wrong dialect: a derivation laid out in a foreign convention, a term your scheme doesn't recognise, a definition phrased beautifully — and differently from the phrasing the examiner's checklist pays. You met this exact problem in the textbook chapter, with tuition handouts and internet summaries; a world-trained chatbot is the same drift, fluent at scale. And your tuned instructions reduce it without removing it, because the model has no copy of your marking scheme to check itself against — it has your description of one.

The honest version of the hallucination warning belongs here too: outright fabrication has become rarer on standard material, but it hasn't become detectable — a wrong step arrives in the same confident voice as a right one, and confidence is the only signal the chat gives you. Rare-but-invisible is precisely the risk profile that hurts an exam candidate, because you'll carry the error into the hall unchecked.

Risk two: the answer that changes nothing

The second gap is not about answer quality at all.

When your chat tool explains a concept tonight, what happens to the fact that you needed it explained? In a chat, nothing. The conversation may be remembered as context — but your confusion doesn't change what you revise on Thursday, doesn't resurface as a practice question next week, doesn't adjust anything about your preparation. You have received an answer; your system has learned nothing.

Compare that with what this book keeps calling the A-student's engine: every error and every confusion recorded, categorised, and scheduled — returning as practice until it dies. A study tool is wired into that engine when your questions and mistakes automatically feed your revision plan, your practice queue, and your readiness picture. A chat tool, however smart, sits outside it: brilliant in the moment, disconnected from the machine that actually manufactures your grade.

What's the one job no chat tool does?

There is a third thing, and for many students it matters more than either risk: your exam is handwritten, and a chat window cannot honestly mark your handwritten script against the official rubric. You can type an answer into a chatbot and get an opinion; you cannot photograph three pages of your Combined Maths working and get back method marks, accuracy marks, and keyword checks the way an examiner would award them. Yet that marking loop — as the whole Papers section of this book argues — is where grades are actually made. Whatever AI you keep for questions, this is the capability to go looking for.

How do you judge any AI study tool?

Fold all of that into a portable test — for any tool, this year and every year after:

  1. What is it checked against? Does the tool verify answers against the national syllabus and official sources — or does it generate from world-knowledge and leave the checking to you?
  2. Who maintains the guarantee? Is the syllabus alignment engineered in, or is it your custom instructions doing their best — a description of the scheme, not the scheme?
  3. Where do my confusions go? Into a system that schedules them back to me as practice — or into a chat history?
  4. Can it mark my handwriting? Against the real rubric, with the marks itemised — the one job the exam actually pays.

And in the open, as this book prefers: Idasara Academy was built against exactly this test. Ask AI — free to try with 5 messages a day; uncapped on the paid tiers — answers from grounding in the NIE curriculum and official textbooks — checked against the source, not just prompted toward it — and every confusion it sees feeds your Mistake Bank — your notebook of lost marks — and revision schedule. Exam Paper Marking (also paid) does the job no chat window does: reads your photographed handwritten script and marks it against the official scheme, method marks and all. That's not a claim of magic — no AI is beyond checking, ours included, and the standing rule stays the standing rule: the official textbook outranks every AI, always. It's a claim about design: the four questions were the specification.

Keep your general chatbot for what it's genuinely good at. The test simply tells you which jobs to trust it with — and which jobs your exam needs done by something that answers to your syllabus.

Do this tonight

  1. Run question one as an experiment on whatever AI you use: ask it for a definition your textbook states exactly, then compare word by word against the official book. Whatever you find, you've practised the habit that matters most: verification before trust.
  2. Run question four honestly: when your last practice answer was handwritten, what marked it? If the answer is "nothing" or "me, generously," that's the gap in your toolchain — whatever you decide to fill it with.
  3. File both results in your question log's margin — your running notebook of precise questions: which tool, which check, what you found.

At a glance

  • Modern chatbots are good — the exam risk isn't stupidity, it's the wrong dialect: right content, foreign convention, unpaid marks. Your custom instructions describe the scheme; they aren't the scheme.
  • Errors are rarer now but arrive in the same confident voice as truths — rare-but-invisible is the worst risk profile for an exam candidate.
  • A chat answer changes nothing in your system: unwired tools leave your confusions unscheduled and your revision untouched.
  • The job no chat tool does: marking your handwritten script against the official rubric — where grades are actually made.
  • Four questions for any AI tool: checked against what? · guarantee maintained by whom? · where do my confusions go? · can it mark my handwriting? The textbook outranks everything, always.
The Idasara Academy app on a phone

Start Free — Create Your Account →

FAQ

I've set custom instructions and keep my syllabus in a project. Isn't that enough? It's genuinely good practice and closes part of the gap. What it can't close: the model still has no marking scheme to check its output against — it has your summary of one — and your confusions still end at the chat's edge instead of feeding your revision. Keep the setup; just know which questions (1 and 3) it leaves open.

How would I even spot a "wrong dialect" answer? Warning signs: notation or spellings your textbook doesn't use, steps compressed where your scheme awards method marks, definitions that paraphrase rather than match. The reliable detector isn't sign-spotting, though — it's the standing rule: anything headed for your notes gets checked against the official source first.

So is free ChatGPT bad for students? No — for exploration, language practice, brainstorming, and tutor-mode questioning (the previous chapters' rules work in any AI), it's genuinely useful. The four questions are about one specific, high-stakes job: syllabus-accurate exam preparation with a feedback loop. Use the right tool per job — and the verification habit with every tool.

My child uses a chatbot for homework — should I be worried? (for parents) Ask two things: which mode (the tutor-vs-vending-machine chapter gives you the two-minute check), and which jobs (this chapter's four questions). Tutor-mode questioning on a general chatbot is studying. The combination to catch is answers being copied while nothing ever gets checked or marked — that's renting, in any tool.

Further reading

The full comparison — grounding, syllabus drift, and mistake telemetry — is written up in the Idasara Method: Part 1 — The Human Foundation & Active Recall and Part 2 — The Plan.


Sources