education
Best AI Tutors
The best AI tutors in 2026, ranked by subject and age group and judged against the independent evidence — Khanmigo, ChatGPT Study Mode, Claude Learning Mode, Gemini Guided Learning, Synthesis, Amira, MATHia and ALEKS — including the August 2026 randomised trial that found the tutoring layer barely got used.
Quick answer: The best AI tutor for most learners in September 2026 is Khanmigo at $4 a month or $44 a year, because it is the only consumer AI tutor that has been through a large independent randomised trial and it refuses to hand over answers by design. But read the evidence before you buy anything: that trial, published on 17 August 2026, found gains of roughly 0.06 to 0.08 standard deviations over a school year and concluded the improvement came from structured practice rather than from the AI tutor, which the median student messaged on only a third of practice days. For a free option, ChatGPT Study Mode, Claude Learning Mode and Gemini Guided Learning all deliver Socratic questioning at no cost and are the sensible starting point for secondary and university students. For early reading, Amira Learning has the strongest evidence rating of any product here, and for structured maths curricula Carnegie Learning’s MATHia and ALEKS have decades of trial history that no chatbot can match. The caveat that governs the whole category: a randomised trial published in PNAS in June 2025 found that unrestricted GPT-4 access made students 17% worse on unaided exams, and that a guardrailed tutor version removed the harm without producing a benefit. What you pick matters less than whether it makes the learner do the thinking.
This page ranks one-to-one tutoring products — tools that hold a sustained teaching relationship with a learner, adapt to what they get wrong, and are built to instruct rather than to answer. That is a narrower thing than the pages around it, and the distinction is worth stating plainly because these tools overlap:
- Best AI for students covers which general assistant to use, plus the free student offers from the labs.
- Best AI study tools covers the study layer — flashcards, spaced repetition, turning your notes into revision material.
- Best AI for homework and maths covers getting tonight’s problem set done, including camera solvers.
- Best AI tools for teachers covers the teacher side — planning, differentiation and marking.
- This page covers the tutor itself: the thing that sits with a learner over weeks and teaches them.
What the evidence actually says
Almost every claim made for AI tutoring traces back to one number, and that number does not hold up.
Bloom’s “2 sigma” is not a benchmark. Benjamin Bloom’s 1984 paper reported that one-to-one tutoring moved the average student two standard deviations above a conventional class. It came from a handful of doctoral studies running three to four weeks on deliberately unfamiliar topics, the tutored group also received extra quizzing and roughly an extra hour of instruction a week, and Bloom framed the figure as a problem to solve, not a product specification. It has never replicated. The best modern meta-analysis of tutoring, Nickow, Oreopoulos and Quan in the American Educational Research Journal (2024), pools 96 studies at 0.288 SD — note that the widely-quoted 0.37 figure is from the earlier working paper, not the peer-reviewed version. VanLehn’s 2011 review put human tutoring at d = 0.79 and intelligent tutoring systems at 0.76, explicitly rejecting the 2.0 figure. For scale, Kraft (2020) puts the median education RCT effect at 0.10 SD.
The strongest independent trial of a consumer AI tutor found a small effect, and credited it to something other than the AI. Philip Oreopoulos and Nina Low ran a two-year cluster-randomised trial of Khan Academy with Khanmigo across 18 middle schools in Hamilton County, Tennessee (NBER Working Paper 35620, August 2026 — a working paper, not yet peer-reviewed). The effect was +1.3 national percentile ranks per term, about 0.05 SD, or roughly 0.06 to 0.08 SD over a school year. The usage data is the finding that matters: 96% of students messaged Khanmigo at least once, but the median student messaged it on only a third of practice days, and used it in only 17% of the exercise sessions where they made a mistake — the exact moment a tutor is supposed to earn its place. Of the messages that were sent, the authors report most were “bare answers or clicks on suggested prompts”. Their conclusion is worth quoting twice over: “These gains resemble those from Khan Academy practice without AI assistance”, and “the binding constraint appears to be engagement: realizing the promise of AI tutoring will require getting students to use it, not just giving them access.”
Unguarded chatbots measurably damage learning. Bastani and colleagues, in PNAS (June 2025), randomised 839 Turkish high-school students across four maths sessions. Students given unrestricted GPT-4 improved 48% on practice problems while the tutor was available — then scored 17% worse than the control group on an unaided exam (-0.054 against a control mean of 0.321, roughly -0.19 SD). A guardrailed “GPT Tutor” version eliminated the harm entirely but produced no measurable gain. Students in the harmed group did not perceive that they had done worse. This is the single most important result in the field and it is the reason the answer-refusing design of Khanmigo, Study Mode and Learning Mode matters more than any feature list.
A well-engineered tutor can do much better — under conditions consumers do not have. Kestin and colleagues in Scientific Reports (June 2025) ran a crossover trial with 194 Harvard physics students and found median learning gains more than double those of expert-led in-class active learning, with an effect size the paper puts between 0.73 and 1.3 SD across quantiles. But the tutor was PS2 Pal, a custom GPT-4 system built by the course’s own instructor with pre-written step-by-step answers fed in to suppress hallucination, tested on two lessons, measured on a bespoke quiz, with no retention follow-up. It is a demonstration of what careful instructional design can achieve, not a product you can buy.
Putting AI behind a human tutor works better than putting it in front of a student. Stanford’s Tutor CoPilot gave real-time AI suggestions to human tutors rather than to learners: 783 tutors and 1,013 students in grades 3 to 6, a +4 percentage point improvement on session exit tickets, and — the interesting part — the least experienced and lowest-rated tutors gained most (+9pp), with students of low-rated assisted tutors performing at or above students of high-rated unassisted ones. There was no significant effect on end-of-year test scores, and the much-quoted $20 per tutor per year covers GPT-4 API calls only, not engineering or training. It remains a preprint.
The meta-analyses are inflated, and one has been retracted. The most-cited pro-chatbot meta-analysis, Wang and Fan (2025), reporting g of about 0.867, was retracted on 22 April 2026 over discrepancies in the analysis; it is still being quoted. The better-venued reviews are more modest — Han, Peng and Liu in Educational Research Review (2025) pool 68 studies at 0.45 SD. The most honest number in the literature is from Doo and Park (2026) in IRRODL: a pooled g = 0.573 with a 95% prediction interval of [-0.47, 1.61], meaning a new study could plausibly find AI tutoring makes things worse.
And there is no official adjudicator. The What Works Clearinghouse has no rating for any AI tutor — not Khanmigo, not MATHia as currently branded, not ALEKS, not Amira. After the 2025 contract terminations, the WWC contract was reinstated for website and reviewer training but not for reviewing education research. No independent referee is coming.
The defensible summary: AI tutoring produces small positive effects when it is structured and refuses to give answers, and negative effects when it does not. Anyone quoting two sigma is selling something.
Best AI tutors ranked (September 2026)
| Rank | Tutor | Best for | Ages | Subjects | Evidence | Price |
|---|---|---|---|---|---|---|
| 1 | Khanmigo | Best overall, best evidenced | 10–18 | Maths, science, humanities, writing | Independent RCT, 0.06–0.08 SD/year | $4/mo or $44/year; free for teachers |
| 2 | ChatGPT Study Mode | Best free tutor for secondary and up | 13+ | All | Guardrailed design matches the PNAS null-harm arm | Free on all plans |
| 3 | Claude Learning Mode | Strictest “make me think” tutor | 13+ | Writing, humanities, reasoning | No independent trial | Free tier; Pro paid |
| 4 | Gemini Guided Learning | Best for auto-generated practice | 13+ | All, strong on quizzing | No independent trial | Free tier |
| 5 | Amira Learning | Best for early reading | 5–10 | Reading fluency only | ESSA Tier 2, +0.15 — but see caveat | School-sold; data not available |
| 6 | Carnegie Learning MATHia | Best structured maths curriculum | 11–18 | Maths only | RAND RCT +0.21 SD in year two | School-sold; data not available |
| 7 | ALEKS | Best mastery-based maths placement | 8–18 | Maths, chemistry | One independent null-equivalent trial | School and consumer; data not available |
| 8 | Synthesis Tutor | Best conversational tutor for young children | 5–11 | Maths | No independent trial | Data not available |
| 9 | Duolingo Max | Best for language practice | 8+ | Languages only | No independent RCT of the AI features | Data not available |
| 10 | Human tutoring with AI support | Best outcomes if you can afford it | All | All | Tutor CoPilot +4pp, largest gains for weak tutors | Market rate |
Pricing note. Khanmigo’s rates were read on its own pricing page on 2 September 2026. Several products on this page are sold through schools and districts and publish no rate card; rather than repeat figures from directories and affiliate roundups — which for Khanmigo alone we found quoted at $15, $35 and $44 a year, only one of which is right — those rows say “data not available”. Confirm current pricing with the vendor.
The tutors in depth
1. Khanmigo — the best-evidenced consumer AI tutor, with an honest caveat
Khanmigo is Khan Academy’s tutor, built on GPT-4-class models and wrapped in a pedagogy that refuses to give answers: it asks what you have tried, where you got stuck, and what the next step might be. It sits on top of Khan Academy’s existing exercise library, which is the part that appears to do the work.
It is the only consumer AI tutor to have been through a large independent randomised trial, and that trial is two weeks old at the time of writing. The effect was real but small — roughly 0.06 to 0.08 SD over a school year — and the authors attribute it primarily to structured practice rather than to the conversational tutor. Khan Academy’s own response, published on 24 August 2026, anchors on a 0.14 SD figure, which is the effect implied for a student who participates actively all year rather than the pre-registered intention-to-treat result. Sal Khan also concedes the central usage finding: “students used Khanmigo infrequently, which matches what we’ve been saying publicly for a while.”
Two further things a buyer should know. Khan Academy changed Khanmigo to open-by-default on 29 January 2026 precisely because students were not seeking it out — so the interface tested in the trial is not the one shipping now, in either direction. And the product had a documented maths problem: the Wall Street Journal reported in February 2024 that it “regularly struggled with basic computation”, corroborated by Education Week in September 2024, and Khan Academy now routes numerical work to a calculator rather than the model.
Pricing, verified at source on 2 September 2026: free for teachers, $4 a month or $44 a year for learners and families, district pricing on request.
Best for: a learner aged roughly 10 to 18 who needs a tutor that will not do the work for them, at a price that makes the decision easy. Against: the evidence says the tutoring layer is the least-used part of what you are buying.
2. ChatGPT Study Mode — the best free tutor, and the one that matches the trial evidence
Study Mode is free on every ChatGPT plan and turns the assistant into a Socratic questioner that walks a student toward an answer rather than producing one. It is the broadest free option: any subject, any level, with the strongest general model underneath.
Its significance is structural rather than promotional. The PNAS trial found that unrestricted GPT-4 harmed unaided exam performance while a guardrailed tutor build removed that harm — Study Mode is the guardrailed build, shipped free to everyone. It has not itself been independently trialled, so treat that as an argument about design rather than measured efficacy. Note the honest reading of the PNAS result: the tutor arm eliminated harm but did not produce a gain.
Best for: secondary and university students who want a capable tutor at no cost. Against: the guardrail is a mode, not a lock — a student can leave it in one click, and the same account gives unrestricted access.
3. Claude Learning Mode — the strictest of the three
Claude’s Learning Mode is the least willing of the free tutors to do the thinking for you, which makes it the best fit for writing, humanities and any subject where the goal is reasoning rather than a numeric answer. It is also the most likely to frustrate a student who wants an answer and the most useful for one who wants to understand.
No independent trial. Free tier available, with a paid Pro tier.
Best for: essay work, dense reading, argument construction. Against: unevaluated, and deliberately harder work than the alternatives.
4. Gemini Guided Learning — best at generating practice
Gemini’s Guided Learning is the strongest of the free tutors at turning a topic into practice questions and quizzes, which matters because retrieval practice is the mechanism with the best evidence behind it in the whole of learning science. It integrates with Google’s school tooling, which decides it for a lot of students.
No independent trial.
Best for: revision and self-testing, and students in Google-based schools. Against: unevaluated as a tutor.
5. Amira Learning — the strongest evidence rating here, and read the small print
Amira listens to a child read aloud, detects errors in real time and intervenes with the specific prompt that particular error calls for. It is the only product on this page with a formal evidence rating: Evidence for ESSA rates it Moderate (ESSA Tier 2) on two studies covering 15,602 students, average effect +0.15, grades K–4.
The caveat is significant and rarely reported. Of those two studies, one is Mostow, Nelson-Taylor and Beck (2013), which tested the 2000–01 version of Project LISTEN’s Reading Tutor — a 25-year-old system, 178 children, in an affluent district — and returned +0.64. The other is a vendor-authored, non-peer-reviewed matched-comparison study by Amira itself covering 15,424 Texas children, returning +0.26 in kindergarten and +0.06 in grade 1. The Florida Center for Reading Research separately lists Amira with figures that match neither. So “Amira is ESSA Tier 2” is true and, on its own, misleading.
Best for: early reading fluency, ages roughly 5 to 10, in a school that buys it. Against: one job only, sold to schools, and an evidence base that mostly predates the product.
6. Carnegie Learning MATHia — real trial history, on an older version
MATHia is the descendant of Cognitive Tutor, and the reason it is on this page is the RAND trial by Pane and colleagues (2014): 73 high schools, 11,066 students, an effect of -0.07 SD in year one (not significant) and +0.21 SD in year two (p = .04). The middle-school arm never reached significance.
Two caveats. The What Works Clearinghouse intervention report on Cognitive Tutor Algebra I records “mixed effects” with an improvement index of +4, while the live WWC database entry currently displays it as “positive effects, ESSA Tier 1” — the PDF and the database disagree, and we cite the PDF. And all of this evidence concerns Cognitive Tutor as it ran between 2007 and 2010. There is no independent randomised trial of MATHia as currently branded.
Best for: a school or district wanting a structured maths curriculum with a genuine trial record. Against: the record is old, the product has changed, and the year-one effect was negative.
7. ALEKS — mastery-based, widely used, thinly evidenced
ALEKS assesses what a student actually knows, then serves only what they are ready to learn next. It is used at enormous scale in US maths and chemistry. Its evidence base is much weaker than its footprint suggests: no WWC intervention report exists and there is no Evidence for ESSA entry. The strongest independent study, Craig and colleagues in Computers & Education (2013), found after-school ALEKS students performed at the same level as students taught by expert teachers while needing significantly less teacher assistance — a null-equivalent result rather than a positive one.
Worth knowing: a $1.54 million IES-funded RAND randomised trial of ALEKS ran to completion in 2017 and has never been published.
Best for: placement and mastery sequencing in maths. Against: an unpublished trial is not a good sign.
8. Synthesis Tutor — the conversational tutor for young children
Synthesis is a conversational maths tutor aimed at primary-age children, and the thing it does better than anything else here is hold a young child’s attention in a genuine back-and-forth. There is no independent trial, and pricing is quoted inconsistently across directories, so we have not published a figure.
Best for: ages roughly 5 to 11, maths, where engagement is the binding constraint. Against: no evidence, and a single subject.
9. Duolingo Max — practice, not tutoring, and that is fine
Duolingo’s AI features add conversational practice and explanations of your mistakes. Language learning is the one domain where high-frequency, low-stakes practice is close to the whole game, so the fit is good. There is no independent randomised trial of the AI features specifically.
Best for: conversational practice in a language you are already studying. Against: not a tutor in the sense the rest of this page uses the word.
10. Human tutoring with AI support — the option the evidence actually favours
If the budget exists, the strongest evidence on this page is not for any product a learner uses directly. It is for putting AI behind a human tutor: Stanford’s Tutor CoPilot improved session outcomes by 4 percentage points and helped the weakest tutors most. Human tutoring itself pools at roughly 0.29 SD across 96 studies, several times the effect of any AI tutor measured to date.
Best for: anyone who can afford it, and any programme deciding where to spend a fixed budget. Against: cost, which is the entire reason AI tutoring is interesting.
Best AI tutor by subject
Maths. Khanmigo for a tutor that will not hand over the answer; MATHia or ALEKS where a school wants a full structured curriculum. For solving a specific problem tonight rather than learning the method, that is a different job — see best AI for homework and maths.
Early reading. Amira Learning, on the evidence, though read the caveat above. Nothing else on this page listens to a child read.
Writing and humanities. Claude Learning Mode, which is the strictest about making the student produce the reasoning, with ChatGPT Study Mode as the more forgiving alternative.
Science. ChatGPT Study Mode or Gemini Guided Learning for explanation and practice; Khanmigo where a structured progression matters.
Languages. Duolingo Max for practice volume. Conversation frequency beats tutoring sophistication in this domain.
Exam revision. Gemini Guided Learning for generating practice questions, because retrieval practice is the best-evidenced study mechanism there is. The best AI study tools page covers this layer properly.
Best AI tutor by age group
Ages 5 to 10. Amira Learning for reading and Synthesis for maths, both designed as walled gardens for young children. General-purpose chatbots are not appropriate here: the major assistants set a minimum age of 13, and a tutor for this age group should be a closed product supervised by an adult, not a general model with a mode switched on.
Ages 11 to 14. Khanmigo is the clearest recommendation on this page — it is the age band the Hamilton County trial covers, the price is trivial, and the answer-refusing design matters most at exactly the point where students are capable of outsourcing the whole task.
Ages 15 to 18. Khanmigo plus a free lab tutor. This is also the age band the PNAS harm result should worry you about most: the students who lost 17% on unaided exams were high schoolers, and they did not notice.
University and adult. ChatGPT Study Mode, Claude Learning Mode or Gemini Guided Learning, free. At this level the constraint is discipline rather than product, and the Harvard trial suggests a well-structured AI tutor can genuinely outperform a class — provided it is structured.
A note on supervision. Every result on this page that shows a benefit involved an adult in the loop: a teacher facilitating, a course designer building the tutor, a tutor receiving the suggestion. Every result that shows harm involved a student alone with an unrestricted model. That pattern is the most consistent thing in the evidence base, and it is a better guide to buying than any feature comparison.
How we rank
We rank AI tutors on independent evidence first, product design second, and price third — in that order, because in a category this noisy the marketing and the measurement point in different directions.
Independent evidence means a randomised or quasi-experimental study the vendor did not run or fund. Where the only evidence is vendor-produced we say so and do not count it. Where a product has no evidence at all we say that too, rather than substituting a feature list for a result. Prices are read on the vendor’s own page and dated; where a product is sold through schools with no published rate card, the row reads “data not available” rather than borrowing a figure from a directory.
This page is re-scored monthly, and this category is moving fast enough that the date at the top matters more than usual — the central trial here was published seventeen days ago.
Frequently asked questions
What is the best AI tutor in 2026?
Khanmigo, at $4 a month or $44 a year, is the best-evidenced consumer AI tutor: it is the only one to have been through a large independent randomised controlled trial, and it is designed to refuse to give answers. The trial, published as NBER Working Paper 35620 on 17 August 2026, measured gains of roughly 0.06 to 0.08 standard deviations over a school year across 18 Tennessee middle schools. For a free alternative, ChatGPT Study Mode, Claude Learning Mode and Gemini Guided Learning all provide Socratic tutoring at no cost, though none has been independently evaluated.
Do AI tutors actually work?
They produce small positive effects when they are structured and refuse to give answers, and negative effects when they do not. The independent trial of Khanmigo found about 0.06 to 0.08 standard deviations over a school year, against a median education-research effect of about 0.10 standard deviations. A 2025 randomised trial in PNAS found that unrestricted GPT-4 access made high-school students 17% worse on unaided exams, while a guardrailed tutor version removed that harm without adding a benefit. The best-venued meta-analysis pools AI tutoring at about 0.45 standard deviations, but one widely cited meta-analysis reporting a much larger effect was retracted in April 2026.
What is the best free AI tutor?
ChatGPT Study Mode is the broadest, free on every plan and available for any subject. Claude Learning Mode is the strictest about making the student do the reasoning, which suits writing and humanities. Gemini Guided Learning is the best at generating practice questions and quizzes, which is the study mechanism with the strongest evidence behind it. None of the three has been independently trialled, so the case for them rests on design rather than measured results.
Is Khanmigo worth it?
At $4 a month it is the cheapest way to put a tutor that refuses to give answers in front of a learner, and it is the only consumer option with independent trial evidence. Set expectations from that evidence rather than from the marketing: the trial found the gains came mainly from structured practice, and that the median student messaged the tutor on only a third of practice days and used it in only 17% of the sessions where they made a mistake. Khan Academy’s own response to the trial concedes that students used Khanmigo infrequently. Teachers get it free.
Can an AI tutor replace a human tutor?
Not on current evidence. Human tutoring pools at about 0.288 standard deviations across 96 studies in the peer-reviewed meta-analysis by Nickow, Oreopoulos and Quan, several times the effect measured for any consumer AI tutor. The strongest AI result in this space actually comes from combining the two: Stanford’s Tutor CoPilot gave AI suggestions to human tutors and improved session outcomes by 4 percentage points, with the largest gains going to the least experienced tutors. The honest framing is that AI tutoring is cheap, not that it is equivalent.
What about Bloom’s 2 sigma?
It does not replicate and should not be used to set expectations. Bloom’s 1984 figure came from a small number of doctoral studies running three to four weeks, in which the tutored group also received extra quizzing and about an extra hour of instruction per week, and Bloom presented it as a challenge to solve rather than an achievable benchmark. Of 96 studies in the most recent tutoring meta-analysis, none reached two standard deviations. Modern estimates put human tutoring at roughly 0.29 to 0.79 standard deviations depending on design.
Are AI tutors safe for young children?
Treat age five to ten as a different product category, not a younger version of the same one. The major general-purpose assistants set a minimum age of 13, so a tutor for younger children should be a closed, single-purpose product with adult supervision — Amira Learning for reading or Synthesis for maths rather than a chatbot with a study mode enabled. Check the vendor’s stated minimum age and its data-handling terms directly before signing a child up, since these change frequently.
What is the difference between an AI tutor and an AI homework helper?
An AI tutor is built to make the learner produce the reasoning and holds a relationship across sessions; a homework helper is built to resolve tonight’s problem. The distinction is not pedantic — it is the single strongest finding in the research. The PNAS trial found that students given an answer-producing assistant improved 48% on practice while it was available and then scored 17% worse without it, while the tutoring version showed no such penalty. If you want tonight’s problem set solved, see best AI for homework and maths; if you want the learner to know it next month, use a tutor.
Do schools have evidence to justify buying an AI tutor?
Less than most procurement processes assume. The What Works Clearinghouse has no rating for any AI tutor, and after the 2025 contract changes it is not currently reviewing new education research. Amira holds an ESSA Tier 2 rating that rests on a 25-year-old predecessor system and one vendor-authored study. MATHia’s trial record is for Cognitive Tutor as it ran between 2007 and 2010. ALEKS has no WWC report at all, and a $1.54 million federally funded randomised trial of it completed in 2017 without ever being published. The Khanmigo trial is the most current evidence available and it is a working paper.
Which AI tutor has the best evidence?
For early reading, Amira Learning holds the only formal evidence rating on this page, ESSA Tier 2 at +0.15 — with the caveat that its underlying studies are a 2013 evaluation of a 25-year-old system and a vendor-authored report. For maths curricula, Carnegie Learning’s MATHia has the strongest randomised record, +0.21 standard deviations in year two of the RAND trial, though for a version of the product that no longer ships. For consumer AI tutors specifically, Khanmigo is the only one independently trialled at all.
Conclusion: how to choose in September 2026
Buy Khanmigo if you want a tutor for a school-age learner and you want to spend as little time deciding as possible: $4 a month, refuses to give answers, and the only one anyone has actually measured. Use ChatGPT Study Mode, Claude Learning Mode or Gemini Guided Learning if you would rather spend nothing, and expect the same order of benefit. For a child under about ten, buy a closed product built for children rather than switching a mode on in a general assistant.
Then set your expectations from the evidence rather than the category’s marketing. The effects that have been measured are small — around 0.06 to 0.08 standard deviations a year for the best-evidenced consumer product — and the one large trial that looked at how students actually behave found they mostly did not ask the tutor for help, especially when they got something wrong. The gains came from structured practice. The harm, when it appears, comes from unrestricted answer-giving.
Which points at the thing worth paying attention to. Every result on this page showing a benefit had an adult in the loop somewhere, and every result showing harm had a student alone with a model that would answer anything. That is a better buying guide than any feature table.
For related rankings, see best AI for students, best AI study tools, best AI for homework and maths and best AI tools for teachers.
This ranking is re-scored monthly. Effect sizes are quoted from the studies named and linked; where evidence is vendor-produced we say so. Prices in USD and read at source on 2 September 2026 where a vendor publishes them. Several products here are sold through schools without a public rate card and are marked “data not available” rather than estimated.