AI models are geniuses at writing essays but toddlers at basic geometry. Even the most advanced systems struggle to solve the same visual coding puzzles taught in middle school, frequently failing to translate simple shapes into the code required to draw them.
AI models currently fail significantly when asked to translate geometric shapes into code. Success rates for leading models like GPT-4o are below 30% for basic visual programming tasks, meaning they cannot be trusted as reliable tutors for geometry or introductory computer science homework.
If your child is using ChatGPT or Claude to "check" their coding homework, they are likely getting wrong answers. This creates a false sense of security and can bake in fundamental misunderstandings of geometry and logic.
Parents who assume AI is an all-knowing math tutor are actually handing their kids a broken compass for visual learning. Until these models improve their "spatial reasoning," they are more likely to confuse a student than help them.
Large Language Models were built on text, not spatial logic. While they have gotten better at "seeing" images, the bridge between a visual pattern and the precise mathematical steps needed to recreate it is still shaky.
Researchers created a benchmark called TurtleAI to see if the latest "vision" models could handle Turtle Graphics—the classic Python tool used to introduce kids to programming by having them "drive" a cursor to draw shapes. They found a massive gap between the AI's ability to talk about code and its ability to actually solve a visual puzzle.
The most advanced AI models are surprisingly bad at "seeing" how shapes relate to each other in space.
- Over 20 top-tier models were tested on 823 educational coding tasks; most failed about seven out of ten times.
- The bottleneck isn't the Python code itself—it's the spatial reasoning. Models struggle to accurately translate a visual pattern into the specific angles and distances required.
- Specialized training can help—one model (Qwen2-VL-72B) improved its performance by 20% after being fed specialized data—but general-purpose AI is still in its infancy here.
We are currently in a "spatial gap." AI has mastered the grammar of human language but lacks the internal "mental map" that even a human novice uses to understand that a square requires four equal lines and exactly 90-degree turns.
Using AI for STEM subjects that rely on visual logic—like robotics, geometry, or graphic design—is fundamentally different and riskier than using it for history or English.
This study is an arXiv preprint, meaning it has not yet undergone formal peer review. The findings are specific to "Turtle Graphics" and might not reflect how AI performs in other areas of visual learning, like reading charts or identifying objects in photos.
Additionally, the researchers included "GPT-5" in their testing, which likely refers to a future-dated version or a specific internal iteration; parents should treat claims about unreleased models with caution until they are publicly available.
- If your child is using AI to debug visual Python or Scratch projects, tell them to treat the AI's suggestions as a "guess" rather than a "fix."
- If a student is stuck on a geometry-based coding problem, have them manually draw the shape on paper first instead of asking a chatbot for the answer.
- If you are looking for a digital tutor for STEM, prioritize specialized platforms with built-in logic checkers (like Khan Academy's Khanmigo) rather than generic, open-ended AI chatbots.
Don't outsource your child's visual logic to a machine that can't tell a 60-degree angle from a 90-degree one. AI is currently a better poet than it is a math teacher; for now, the best way to learn coding remains human trial and error.
Chao Wen, Jacqueline Staub, Adish Singla (2026). TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics. arXiv (preprint). — arxiv.org


