Illustration for Screenwise guide: Smartest AI Models Make Poor Math Tutors for Kids
Parent Guide

Smartest AI Models Make Poor Math Tutors for Kids

High accuracy in large language models rarely translates to effective teaching and guided hints

Updated 8/23/26
Based on researcharXiv logo

A high-performing AI might be great at solving your child’s math homework but terrible at teaching it. New research shows that a model's accuracy often has little to do with its ability to provide the hints and guidance necessary for actual learning.

Junyi Yao, Zihao Zheng, Baichuan Li (2026). arXiv (preprint)
Who was studied: Eight large language models (LLMs) evaluated using public performance data from the MathTutorBench leaderboard.
How: The researchers compared accuracy-based scores against pedagogy-based scores across eight AI models to measure the 'gap' between solving and teaching.
Read the original paper
Honest caveats
  • This is a preprint and has not yet undergone formal peer review.
  • The study analyzed a small sample of only eight AI models.
  • The findings are based on standardized benchmarks (MathTutorBench) rather than observations of real-world interactions between children and AI.

An AI model that reliably solves your middle schooler's math homework might actually be a terrible tutor. New research reveals that an AI's math accuracy has almost nothing to do with its ability to help children learn.

TL;DR

Smart AI models frequently spoil the learning process by handing over full answers instead of guiding children to figure out math problems on their own.

Why it matters

When choosing a homework assistant, parents naturally reach for the highest-performing, most "advanced" chatbot available. But assigning a genius problem-solver to assist with math often backfires. If an AI immediately hands over step-by-step solutions, it removes the friction and mental effort required for real skill acquisition.

Kids who rely on answer-focused AI risk developing a "solving dependency"—appearing high-performing on paper while failing to build the underlying conceptual understanding needed for tests and real-world reasoning.

What's driving this

AI developers typically benchmark their products on whether they output correct answers, not on whether they know how to teach. Researchers created a diagnostic tool called MathTutorBench to evaluate whether popular large language models use "non-disclosive scaffolding"—the practice of offering small hints that prompt students to solve problems without giving away the final solution.

What they're saying

An AI's ability to solve a math problem tells you remarkably little about its ability to teach it.

  • A weak connection: The correlation between an AI model's problem-solving accuracy and its pedagogical skill is only 0.421, meaning raw performance is a unreliable indicator of teaching ability.
  • Accuracy leaders fall behind: Several prominent AI models that ranked at the top of standard accuracy leaderboards dropped significantly when scored on how well they supported student learning.
  • Defaulting to answer-giving: Generic AI tools prioritize speed and precision, leading them to explain full solutions rather than asking guiding questions that preserve a child's agency.
Between the lines

Consumer AI interfaces are engineered for efficiency and user satisfaction, which directly conflicts with effective teaching. When an adult uses an AI assistant at work, a fast answer is an asset. When a child uses an AI assistant for learning, a fast answer is a shortcut that skips the cognitive labor required to rewire the brain. A model optimized to generate quick solutions will treat a student's confusion as a bug to fix rather than a necessary step in the learning process.

Grain of salt

This study is an unreviewed preprint that analyzed a small sample of eight AI models using an automated evaluation benchmark rather than real-world trials with human children. While benchmark data clearly isolates structural differences between solving and teaching, human classroom dynamics and individual student behavior may vary.

If [this], then [that]
  • If your child uses a general-purpose AI for math homework… test the tool yourself by typing "I don't know how to do this" to see if it gives away the answer or offers a small hint.
  • If you are selecting a digital math assistant… prioritize platforms designed explicitly for education rather than standard, unconstrained chatbots.
  • If your child turns in completed AI-assisted homework… ask them to explain the steps back to you without looking at the screen to verify they understand the underlying concepts.
The bottom line

Do not mistake an AI model's raw intelligence for teaching ability. A useful educational tool acts like an encouraging coach who prompts a child to think, not a classmate who hands over an answer key.

Junyi Yao, Zihao Zheng, Baichuan Li (2026). Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact. arXiv (preprint). — arxiv.org