Discussion summary
A new AI tutor demonstrated a 0.71-1.30 SD effect size in Dartmouth courses, with high adoption rates. Opinions vary on its educational value and commercial potential.
What the discussion says
- Some see AI as a game changer for motivated learners.
- Concerns about hallucinations and accuracy in AI tutoring.
- High adoption rates suggest strong interest despite skepticism.
- Questions about the effectiveness compared to traditional methods.
- Discussions about the commercial viability of AI in education.
“0.71 SD effect size is impressive but I would be more convinced if they ran the control group against a TA.”
“This AI thing got 90% voluntary usage, which is more telling than just effectiveness.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Conflicted about this study. On one hand, LLMs have been incredible for my personal learnings of new concepts.
On the other, I'm sceptical of that it'll have "strong benefits" at scale; I'd be more in favor if the wording was "some"/"moderate". I reckon self-selection plays a huge part, as mentioned in the "Limitations" section of the paper.
I'd also caution against attaching the tool to grading. That means students have to put more effort into the course, which increases the chances that they will use LLMs to save time rather than make the investment.
by mmarian - Yes! Very exciting to see this.
Bloom's Two Sigma Opportunity suggests that there's another SD improvement available: https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem
by rictic - There's a famous post by Erik Hoel that calls the human version of this Aristocratic Tutoring [1] (Scott Alexander is unconvinced [2]).
In the 1980s, a researcher called Benjamin Bloom claimed a z=2.0 (that is 2σ) advantage for a combination of mastery learning (don't move on the the next topic until you've mastered the current one) and 1-on-1 tutoring. Later replications show there is definitely something going on, but the effect size is much lower, for example around z=0.7 in a 2020 paper [3].
I'm still open on AI tutoring, though the Dartmouth results look impressive. Someone please try and replicate this.
There's a saying that AI helps the best students get better, and the worst ones get worse. (Anthropic sort-of agrees [4].) It'll be interesting to see how that turns out.
[1] https://www.theintrinsicperspective.com/p/why-we-stopped-mak... [2] https://www.astralcodexten.com/p/contra-hoel-on-aristocratic... [3] https://www.nber.org/papers/w27476 [4] https://www.anthropic.com/research/AI-assistance-coding-skil...
by red_admiral - This is exciting because the effect size is so large. But as the author's acknowledged, selection bias is nearly impossible to control for in this non-randomized study:
> and lacks randomized controls. Self-selection is the central threat: students who complete more quizzes may be more motivated or higher-performing generally
But this is still a strong result. I'm excited to see more in this space.
by rusbus - The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder.
> constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria
> Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient.
They specifically call out that the "RAG chat assistant" part of Phosphor (the platform) wasn't used much.
I commend the effort here, but I don't think these results are particularly noteworthy. The conclusion is essentially that people who do practice quizzes will do better on exams.
by wxw - I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of the student - agentic tutor pairs, a referee when the student and model disagree, etc. and most importantly still being the human teacher you can just talk to in the human education process.
I'm convinced this is the future of education - models are there, we need the classroom tech to catch up. The alternative is obvious and quantified in the paper - students just use models to do their work for them and learn nothing.
by baq - I'm not an expert, but how much of this is down to novelty, ie https://en.wikipedia.org/wiki/Hawthorne_effect ?
(ie changing the environment can lead to short term productivity gains because either participants are aware they are being watch, or it breaks up the monotony and makes people work a bit harder. )
by KaiserPro - I am somewhat skeptical of this.
First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement
Second, trying to incorporate past grades into their modelling is not a substitute for a randomized trial.
Third, the headline engagement number of 90% is for "engaging with the platform, via Module Review or Lesson Quizzes, at least once". I don't know why much of that couldn't just be attributed to novelty. Or even partly a professor with all sorts of enthusiasm for the platform.
Fourth, the "full dosage" effectiveness is measured based the final exam scores. Were these exam questions produced independently from the "Phosphor" materials? (e.g. by blinding?) Were they checked for direct overlap with those materials? The 0.7 sigma shift is 3 points on a 24 point exam; if even a few of the questions on that exam were very similar to those materials it could account for almost all of it. This is not clear to me from the manuscript.
If this was the case, then it's a question less of "is AI effective" vs. "did the students look at the materials". You could still argue that the AI platform got them to read, but that is a somewhat different statement than the AI helped them learn.