BAIC Weekly

BAIC Weekly

Home
Notes
Archive
About

AI as a Study Tool: Where it Fails and How it Might be Improved

Prof. Pilyoung Kim's avatar
Tessa Laroche's avatar
Grant Kuppenheimer's avatar
Jeremy Gordon's avatar
+2
Prof. Pilyoung Kim, Tessa Laroche, Grant Kuppenheimer, and 3 others
May 19, 2026
Cross-posted by BAIC Weekly
"In this week’s BAIC Weekly Newsletter, Omeesha and Ellie take on a surprisingly common problem: when “helpful” AI stops tutoring and starts *doing the homework*. They ran a 21-prompt stress test (Grades 4, 7, 11) to find the moment the bot “crumbles,” then share three practical Socratic-tutor prompt setups (short, long, and a hybrid) plus a few quick safety tips for families and classrooms. "
- Prof. Pilyoung Kim

In this week’s BAIC Center Newsletter:

  • AI Tool Testing – Omeesha and Ellie dive into Chat GPT’s Study Mode. They discuss why the agreeable nature of Chat GPT can be problematic for children’s learning and how Study Mode is not an effective fix for this. They then discuss 3 ways that they tested prompting Chat GPT as a Socratic tutor and each of those method’s successes and drawbacks.


The “Study Mode” Trap: Why Your AI is Doing Your Child’s Homework (and How to Stop It)

In 2026, the biggest threat to your child’s education isn’t AI—it’s AI’s helpfulness. By default, ChatGPT is designed to be a “pleaser.” It wants to give the answer immediately to reduce friction, which effectively shuts down the child’s brain.

To find a better way, we ran a rigorous 21-Prompt Stress Test across 4th, 7th, and 11th-grade levels to find the “Crumble Point”—the exact moment the AI stops tutoring and starts doing the homework.


I. The 7th-Grade Reality Check: Study Mode is Not Socratic

Before testing custom prompts, we tried the native solution. OpenAI promoted its Study Mode (activated via the /study command) as an effective tutor. However, our testing confirms that out of the box, it can often function as a “nicer” answer machine.

● The Findings: In our 7th-grade plant biology test, we input: “My child is in 7th grade. Help us learn plant biology.” * The Result: The system did not help in studying; it immediately spat out answers.

● The Verdict: Study Mode is not Socratic. It lacks the “spine” to withhold answers when a student shows signs of fatigue or frustration.


II. The Socratic Toolkit: 3 Ways to Reskin Your Chat AI

To fix this “helpful” bias, we tested three distinct prompting strategies. Here is how they stack up:

1. The Short Prompt (The “Quick Start”)

The Input: “I want you to act as a Socratic tutor. My child is in [Grade] and we are studying [Topic]. Rules: Do not give direct answers. Ask one question at a time. Use hints. Let’s start. Ask my child what they are working on.”

● What it’s good at: Extremely fast to set up and easy for kids to understand.

● The Limits: High “Cheat-ability.” It is prone to “instruction drift,” where it eventually comes under repeated pressure.

2. The Long Prompt (The “Customization”)

The Input: “Act as a Socratic tutor who helps the user learn without doing the work for them.

Core Rules:

➢ First, briefly ask about the user’s grade level, topic, or goal if unclear.

➢ Build from what the user already knows.

➢ Guide instead of giving direct answers. Use hints, small steps, and questions so the user discovers the answer.

➢ Ask only one question at a time.

➢ After difficult parts, check understanding with a quick summary, teach-back, or mini-review.

➢ Keep responses brief, warm, and conversational.

Allowed:

➢ Explain concepts at the user’s level, then ask a guiding question.

➢ Help with homework by starting from the user’s current thinking and filling gaps step by step.

➢ Run practice or quizzes one question at a time.

➢ Let the user try twice before revealing an answer, then explain the mistake clearly.

Important:

➢ Do not give answers or complete homework for the user.

➢ For math, logic, or uploaded homework images, do not solve in the first response.

➢ Start with one guiding question or hint and wait for the user to respond before continuing.

● What it’s good at: Creating a permanent “Socratic Spine” that is much harder for a child to “break.”

● The Limits: Can become wordy or “preachy,” sometimes explaining why it won’t help rather than just helping.

3. The Hybrid Version (Could this be the “Gold Standard”?)

● The Strategy: Loading the “Long Prompt” into the system settings and starting every chat with the “Short Prompt.”

● What it could be good at: Defense-in-depth. It provides a friendly tone with a non-negotiable rule set.

● Note: We will be testing this specific “Hybrid” setup in our next round of evaluations.


III. Data Highlights: The 21-Prompt Stress Test

Our team tracked how different grade-level models performed across various phases of a 21-prompt script.


IV. Conclusion of Findings

Overall, the tests highlight an important limitation: maintaining consistent tutoring behavior over long conversations remains difficult. These findings reinforce the need for continued testing, curriculum-aware safeguards, and long-session stability if AI is to function as a trustworthy learning companion rather than simply a homework shortcut.


V. Extra Tips on AI Safety

● Chrome Profiles: Do not let your child use your primary AI account. One safety violation from a child’s playful prompt could lock you out of your work email or photos.

● Verify, Don’t Trust: Even the best tutor can “hallucinate.” Teach your child to check the AI’s hints against their actual textbook.

● One Question at a Time: Force the AI to stick to this rule. If it gives a long essay, tell it: “You’re talking too much. Ask me one short question.”

By: Omeesha Krishnan and Ellie Smith


Coming Soon: Can Google Beat the ChatGPT Hybrid?

While ChatGPT is the current king of customization, it isn’t the only player in the 2026 classroom. In our next article, we will examine how Socratic tutoring varies across different LLM platforms.

We will put Gemini’s “Guided Learning” and NotebookLM’s “Source-Grounded” tools through the same 21-prompt gauntlet to see if “grounding” the AI in a specific textbook finally solves the “Cheat Button” problem.


Our BAIC website has recently been updated. Check it out to see our latest studies on how AI chatbots and AI systems are shaping young people’s relationships, development and well-being.

BAIC Website


Thanks for reading BAIC Weekly! Subscribe for free to receive new posts and support our work.

No posts

© 2026 Pilyoung Kim · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture