Discussion summary

A new AI tutor demonstrated a 0.71-1.30 SD effect size in Dartmouth courses, with high adoption rates. Opinions vary on its educational value and commercial potential.

What the discussion says

  • Some see AI as a game changer for motivated learners.
  • Concerns about hallucinations and accuracy in AI tutoring.
  • High adoption rates suggest strong interest despite skepticism.
  • Questions about the effectiveness compared to traditional methods.
  • Discussions about the commercial viability of AI in education.
0.71 SD effect size is impressive but I would be more convinced if they ran the control group against a TA.
luciana1u
This AI thing got 90% voluntary usage, which is more telling than just effectiveness.
Rperry2174

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Interesting article, wonder where we're going with this though, I find it's very difficult to keep LLMs on track and critical enough to be useful.

    Just want to say that:

    >In our deployment, student-reported reading completion baselines for MATH 010 were approximately 15%, with instructors estimating 10%. Individual student reports of reading compliance ranged from "literally no one does that" to "is this being recorded?"

    is hilarious

  • Do you have a larger study planned for the Fall? It definitely seems promising.

    I'm curious how well you feel this worked because the subject was Statistics (objective grading) versus something more subjective like Civics or Literature.

    PS - I'd say this qualifies for Show HN, too!

    Do you

  • They were using Sonnet 4.6 for some fre form responses so that could be applied to something subjective.
  • I currently study Multivariate Calculus by using very new and nodern method: I read the text book, while solving the examples of the general for,ulas, or try to come up with my own. Then I do a s$ht-ton of exercises. I only use LLM’s to quickly clarify confusing topics or notation, but not really much else. I cancelled my Claude subscription. Now I use just Mistral and local Vibethinker-3B, but they work just fine.

    Earlier I used Claude by giving it the course material and asking it to generate me exercises (our cpurse work went way over my head) and yeah i learned to differentiate a gradient or Jacobian, but it was very shallow - I knew the formulas, but not what they meant or how to apple them correctly. After I just filled glaring holes I had in Univariate Calculus by readong and doing, I actually started to understand something.

    Lon story short, in my experience Learning with LLM’s is ok with very unfamiliar material that is not too complex (there’s obvious problems of LLM’s themselves being pretty ghastly with maths sometimes), but at least it os not better than the traditional method of just putting your nose on the grimd stone.

  • In the Matrix movies you can upload complex knowledge to a person in seconds. If LLMs could help with this even a little it would be useful. But alas, people like myself who are slow learners still get overwhelmed when trying to learn too much or too fast without rest breaks.
  • The article explicitly calls out selection bias (this is entirely based on 90% that opted into using the tutor, there was no control group), I wish the headline did as well. "Engaged students score 0.71 - 1.30 SD better in tests" sounds like a much simpler explanation.
  • "Full dosage of the Phosphor material is associated with an increase in final exam performance."

    This sentence is accurate, but inevitably leads to the confusion you see in these comments.

  • I used to TA a graduate level CS math class at Georgia Tech. We regularly saw that the students who self-organized study groups did dramatically better in the course than average. One semester they told us to put everyone in study groups to see if it helped. The effect disappeared. Turns out that it was the self-selection of the most engaged students into a small group that mattered, not the study group itself.
  • Conflicted about this study. On one hand, LLMs have been incredible for my personal learnings of new concepts.

    On the other, I'm sceptical of that it'll have "strong benefits" at scale; I'd be more in favor if the wording was "some"/"moderate". I reckon self-selection plays a huge part, as mentioned in the "Limitations" section of the paper.

    I'd also caution against attaching the tool to grading. That means students have to put more effort into the course, which increases the chances that they will use LLMs to save time rather than make the investment.

  • > LLMs have been incredible for my personal learnings of new concepts.

    Mind if I ask what did you learn and how you're using it?

    The reason I'm asking is that I repeatedly felt excitement only to realize down the line that the explanations didn't actually translate into practical skills. I'm not sure it's even an AI problem, it's a "doing versus reading" problem. Same as with reading a pop-science article and thinking to myself that I learned something about physics or medicine or mathematics.

  • Yes! Very exciting to see this.

    Bloom's Two Sigma Opportunity suggests that there's another SD improvement available: https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem

  • As far as I know, replications showed an effect size that's still pretty amazing by educational standards but nowhere near z=2.0 (I think z=0.6 is the current best guess).
  • The story around Bloom's two sigma is a bit complex https://nintil.com/bloom-sigma/
  • This is exciting because the effect size is so large. But as the author's acknowledged, selection bias is nearly impossible to control for in this non-randomized study:

    > and lacks randomized controls. Self-selection is the central threat: students who complete more quizzes may be more motivated or higher-performing generally

    But this is still a strong result. I'm excited to see more in this space.

  • They tried to control for this. It's described in the first paragraph of section 4.
  • There's a famous post by Erik Hoel that calls the human version of this Aristocratic Tutoring [1] (Scott Alexander is unconvinced [2]).

    In the 1980s, a researcher called Benjamin Bloom claimed a z=2.0 (that is 2σ) advantage for a combination of mastery learning (don't move on the the next topic until you've mastered the current one) and 1-on-1 tutoring. Later replications show there is definitely something going on, but the effect size is much lower, for example around z=0.7 in a 2020 paper [3].

    I'm still open on AI tutoring, though the Dartmouth results look impressive. Someone please try and replicate this.

    There's a saying that AI helps the best students get better, and the worst ones get worse. (Anthropic sort-of agrees [4].) It'll be interesting to see how that turns out.

    [1] https://www.theintrinsicperspective.com/p/why-we-stopped-mak... [2] https://www.astralcodexten.com/p/contra-hoel-on-aristocratic... [3] https://www.nber.org/papers/w27476 [4] https://www.anthropic.com/research/AI-assistance-coding-skil...

  • The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder.

    > constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria

    > Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient.

    They specifically call out that the "RAG chat assistant" part of Phosphor (the platform) wasn't used much.

    I commend the effort here, but I don't think these results are particularly noteworthy. The conclusion is essentially that people who do practice quizzes will do better on exams.

    by wxw
  • > a practice quiz platform with an AI autograder.

    What do you think tutoring is?

  • I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of the student - agentic tutor pairs, a referee when the student and model disagree, etc. and most importantly still being the human teacher you can just talk to in the human education process.

    I'm convinced this is the future of education - models are there, we need the classroom tech to catch up. The alternative is obvious and quantified in the paper - students just use models to do their work for them and learn nothing.

    by baq
  • Today I saw a demo of Remarkable turned into Voldemort's diary from Harry Potter - you write to it, and it writes back, in handwriting.
  • A 'smart pen' that records the student's writing in some way, maybe? My first thought was a tablet that boots straight into a writing software but students should not be subjected to any amount of latency in their writing.

    Practically, I think if you want the AI system to have a live view of what the student's doing you're going to have to replace one of either the tablet or the writing instrument. A wearable camera could work as well but there are issues with that.

  • I would add that somewhere in there should be a spaced repetition algorithm.

    Spaced repetition is very effective, but it's really really clunky to use. My unpopular opinion is that we all have Stockholm syndrome when it comes to creating "cards", and people talk about how valuable creating cards is; but I think it stucks, it takes a lot of time.

    If AI is already teaching me math (let's say), it would be nice to tell the AI/app "quiz me on this periodically", and then the AI makes up a fresh polynomial to factor (or whatever) and presents that to you according to a spaced repetition algorithm.

    Behind the scenes, the AI should have access to what has happened the last several times a specific topic has been quized, so the AI can watch to see that certain mistakes are resolved, and the AI might also know better how to correct the user if it has context about previous quizzes of that topic.

  • I work in consulting and one of my projects is piloting an AI use case for a department within one of my clients. On a discovery call someone casually brought up that they bought a reMarkable notebook themselves and were wondering if it could be integrated into the use case. It really got me thinking.

    Maybe reMarkable or something like it could help bridge a student's writing with an LLM without having to fall back to a laptop or ipad.

    https://remarkable.com/

  • I'm not an expert, but how much of this is down to novelty, ie https://en.wikipedia.org/wiki/Hawthorne_effect ?

    (ie changing the environment can lead to short term productivity gains because either participants are aware they are being watch, or it breaks up the monotony and makes people work a bit harder. )

  • I am somewhat skeptical of this.

    First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement

    Second, trying to incorporate past grades into their modelling is not a substitute for a randomized trial.

    Third, the headline engagement number of 90% is for "engaging with the platform, via Module Review or Lesson Quizzes, at least once". I don't know why much of that couldn't just be attributed to novelty. Or even partly a professor with all sorts of enthusiasm for the platform.

    Fourth, the "full dosage" effectiveness is measured based the final exam scores. Were these exam questions produced independently from the "Phosphor" materials? (e.g. by blinding?) Were they checked for direct overlap with those materials? The 0.7 sigma shift is 3 points on a 24 point exam; if even a few of the questions on that exam were very similar to those materials it could account for almost all of it. This is not clear to me from the manuscript.

    If this was the case, then it's a question less of "is AI effective" vs. "did the students look at the materials". You could still argue that the AI platform got them to read, but that is a somewhat different statement than the AI helped them learn.

  • Thank you for the feedback! Maybe the following info will be helpful when considering our results:

    1. Quiz completion is our deliberately conservative lower bound on reading compliance, and the 0.71 figure is not a claim that those 16 students each gained that much. The estimate is from a regression carried by the per-lesson slope, fit across the whole dosage distribution, and the underlying dosage-performance relationship is essentially unchanged whether or not zero-completion students are included (R² 0.091 vs 0.096). In other words, more Phosphor use is strongly associated with better performance across the whole range of usage - not just the group who completed all content.

    The numbers in Table 1 show how dosage was distributed across the course. We report that across the class, the median percentage of lessons reached on Phosphor, including both students with an account and those who never logged in, was above 90%. Among platform users in particular, it was 96%.

    We'd like to emphasize that for this pilot, the platform was presented to students as an entirely optional "study aid", and our adoption rates far exceed those reported in the past for optional interventions. It will be interesting to see how things go when we attach completion to the course grade, as we're thinking of doing in the fall. Past literature from interventions in college courses predicts that this will achieve far higher levels of engagement, bringing the high-dosage effect to a large proportion of the class.

    2. We explicitly note this in Limitations; it's an observational study. We were unable to do an RCT for this course since it raised an ethics consideration - neither we nor the instructors wanted to deprive students entirely of a course material that could have been helpful for them. We'd love to run a randomized trial at some point though - one way to do this is a crossover, where we offer the treatment to one of two groups, then switch it over to the other midway through the trial, so that both get even treatment. Another possibility is randomly selecting students to get access to MCQ-only vs. CRQ-enabled quizzes. That being said, this mechanism of conditioning on past performance is well-known and relatively robust for observational studies of educational interventions.

    3. The platform was created independently of the instructors of the course. The instructors designed their curriculum ahead of time (as had been taught for years of past offerings of the course), lectured in a conventional style, referenced the course's official textbook (Freedman, Pisani, Purves) and suggested homework problems from the textbook only. The instructional content was authored using material that every student in the course had access to, and did not feature exam questions that students were evaluated on after-the-fact.

    Phosphor was not endorsed publicly by the instructors, and was spread primarily via student word-of-mouth. In fact, one instructor in this course initially believed the project would be "a waste of time" and refused to collaborate with us for the pilot. Despite this, 97% of the students in this instructor's section used the platform!

    Engagement also persisted well across the full ten-week term, and two-thirds of Review attempts involved retries spaced a day or more apart — not a pattern typically produced by novelty effects.

    4. Instructors wrote exams independently with their long-running FPP-based curriculum. Even if we steelman and suppose that "the platform just got students to engage" rather than truly learn, this is refuted by our result that the MCQ-only Module 2 had similar engagement but no dosage relationship. This strongly suggests that the CRQ format was a driver of the results.

    As we mention in the paper, we agree that replication, especially across contexts, is a priority. For us in particular, this means not only across other courses, but across other institutions as well. And an RCT would certainly help lock in the causal claim.

  • Education research is hard hard hard. Getting clean studies, especially at large sample sizes, is extremely difficult, leading to a lot of ambiguity in results.
  • This is a helpful explanation - am not a researcher so I have little idea how to run an unbiased, meaningful experiment (except that it takes a lot of effort and thought to run one). Useful analysis
  • Yeah, calling this an "effect size" is just nonsense, and it is alarming that educational software can get away with such poor statistical practice. I'm hoping this was just a student project.
  • It feels to me that the venn diagram between "students that fully engaged with the material" and "students that learned well from the material" is going to basically be a circle for any teaching method.