AI by IndustryEducation

AI for tutoring and education: feedback, planning, and safeguards

A randomized trial of 900 tutors and 1,800 students found a real learning gain when AI coached the tutor. That human-plus-AI design is the centre of this playbook.

By Adi Huric, founder of Most AI LabsAugust 20269 min read

On this page
    The rare randomized resultBuild the tutor’s copilot firstAdministration is useful, but keep the claim modestChildren change the privacy calculationMeasure learning, not output volumeThe honest bottom lineSources

The best evidence in AI tutoring does not show a bot replacing the tutor. It shows AI giving a tutor better suggestions during the lesson—especially when that tutor had less experience to draw on.

The rare randomized result

Tutor CoPilot was tested in a preregistered randomized controlled trial involving 900 tutors and 1,800 K–12 students from historically underserved communities. Students whose tutors had access to the system were four percentage points more likely to master mathematics topics. Students working with lower-rated tutors improved by nine points relative to the control group.

The researchers analysed more than 550,000 messages and found assisted tutors used more strategies that encouraged understanding, such as guiding questions, and were less likely to give away answers. The estimated usage cost was $20 per tutor per year. Tutors also reported suggestions that were not always appropriate for the student’s grade level. It was one system, one partnership, and one instructional context; the result should not be generalized to every subject or product.

Key takeaway

The design lesson: the system advised the educator in real time. The educator saw the student, chose the response, and remained accountable for instruction.

Build the tutor’s copilot first

  • Before the session: draft examples, misconception checks, vocabulary, and a lesson sequence from the centre’s curriculum.
  • During the session: suggest a guiding question, alternate explanation, worked example, or scaffold for the tutor to choose.
  • After the session: draft a progress note tied to observed work, prepare practice, and summarize the next objective.
  • Across the team: identify where tutors need coaching and build reviewed examples from strong practice.

Ground the system in the curriculum, the student’s authorized learning plan, and work produced during the session. Do not let it invent mastery, diagnoses, accommodations, or parent assurances. A progress note should distinguish what the student demonstrated from what the model recommends practising next.

AI should help the tutor ask a better next question—not hand the student a faster answer.

Administration is useful, but keep the claim modest

Scheduling, reminders, intake summaries, tutor matching from documented subject and availability, and parent-message drafts can remove routine coordination. For matching, show the criteria and allow staff to override; do not infer personality, disability, or family circumstances from language and then call the result “fit.”

Statistics Canada found 53% of workers in educational services reported using generative AI at work by March 2026. That broad category includes schools, post-secondary institutions, and other education work; it does not establish adoption among tutoring centres or learning outcomes.

Children change the privacy calculation

UNESCO’s guidance calls for a human-centred and age-appropriate approach, data-privacy protection, and limits on independent conversations between children and general generative-AI platforms. A tutoring company should tell families what tool is used, what data enters it, how long data remains, whether it trains a model, and how a student can learn without it.

Use minimum necessary data, approved accounts, role-based access, and contracts that cover deletion and subprocessors. Avoid sending names, diagnoses, school records, recordings, or identifiable work to a public tool. Keep a way for parents and students to access and correct records where applicable.

Measure learning, not output volume

Set a baseline with a comparable assessment, define the instructional period, and compare topic mastery, error patterns, attendance, completion, tutor preparation time, and parent questions. Where possible, randomize or stagger the rollout. Do not use grades alone when teachers, courses, or assessment difficulty differ.

Watch out for this

The failure mode: a student can complete more work with AI while learning less because the tool supplied the reasoning. Review process, transfer to new problems, and what the student can explain without assistance.

The honest bottom line

  • Build for tutors first. The strongest trial supports human-plus-AI instruction.
  • Ground suggestions in curriculum. Grade level and learning objective must be explicit.
  • Protect children’s data. Consent, minimization, age-appropriate use, and a no-AI path belong in the workflow.
  • Measure independent mastery. Finished worksheets and longer chats are not learning outcomes.

Our free 7-day audit maps curriculum, student data, tutor workflow, and measurement before implementation.

The free lead calculator uses the broad Education & Instruction benchmark of about $77 per Google inquiry. It does not model enrolments because no credible inquiry-to-student rate exists.

Sources