Language learning looks like the perfect AI use case. The task is text and speech, the feedback loop is immediate, and the alternative — a human tutor at €30 an hour, three times a week — is expensive enough that even a partial substitute has obvious value.
And yet most people who open a general-purpose chatbot with the intention of “practising Italian” abandon the habit within two weeks. The technology isn’t the problem. The problem is that language acquisition has a structure, and a blank chat box has none.
This is a practical breakdown of what AI does well in language learning, where it fails in ways that are easy to miss, and how to assemble a stack that actually produces progress rather than the feeling of progress.
Where AI genuinely outperforms the alternatives
Unlimited low-stakes output practice. The single biggest bottleneck for adult learners isn’t vocabulary or grammar knowledge — it’s production. You need thousands of hours of speaking and writing, and most learners get a few dozen. The reason is social: producing bad sentences in front of another person is uncomfortable, so people avoid it. An AI conversational partner removes the social cost entirely. You can say something wrong at 11pm, be corrected, and try again immediately, with no embarrassment budget being spent. That’s not a marginal improvement — it directly attacks the scarcest input in the whole process.
Correction at the moment of the error. Traditional courses correct you in a class on Thursday for a mistake you made on Thursday. A tutor corrects you in real time but costs money per minute. AI can do it continuously and at zero marginal cost. Research on feedback timing consistently favours immediacy, and this is one of the rare cases where the technology’s economics and the pedagogy point in the same direction.
Turning encounters into review items. Spaced repetition works, but almost nobody maintains a flashcard deck by hand for more than a month — the friction of authoring cards kills the habit. Automated card generation from the words you actually met, in the sentences you actually read, closes that gap. This is unglamorous and probably the highest-ROI AI feature in the entire category.
Level-appropriate input on demand. Comprehensible input — material slightly above your current level — is the engine of acquisition, and it’s historically been hard to source. Generating or adapting a text to a specific CEFR level is something models do reliably well.
Where it fails, including in ways you won’t notice
Confident wrong answers about grammar. Ask a general model to explain a subtle point — in Italian, say, the difference between ci vuole and ci mette, or when the subjunctive is genuinely obligatory versus merely formal — and you’ll often get a fluent, well-structured explanation that is partly invented. A beginner has no way to detect this. The error rate is low enough to be trusted and high enough to install false rules that take months to remove.
Register and regional flattening. Models are trained toward a neutral, slightly formal, written-ish standard. Real Italian is full of register: what you’d say to a colleague, a mechanic, a barista, or your partner’s grandmother are four different languages. A learner trained exclusively on model output ends up speaking correct Italian that no Italian actually speaks — a subtle failure mode, because everything is technically right.
No curriculum and no measurement. A chatbot answers what you ask. It doesn’t know what you should be asking, it doesn’t sequence material, and it will happily let you spend six months practising the same three tenses. Nor does it tell you where you are: “I’ve been chatting with an AI for a year” is not a level.
Motivation is not a software feature. Streaks and gamification substitute for accountability without providing it. The learners who finish are, overwhelmingly, the ones with a deadline, a community, or a human who notices when they disappear.
The stack that actually works
The pattern that produces results is not “AI instead of a course” or “course instead of AI.” It’s AI slotted into a structured system, where each element covers the others’ weaknesses:
- A sequenced curriculum so you know what comes next and in what order.
- Authentic input — news, podcasts, stories written for or by natives, not model-generated filler.
- An AI conversation layer for daily output practice and real-time correction.
- Automated spaced repetition fed by the vocabulary you actually encountered.
- Human contact — teachers, a community, anything that makes disappearing socially costly.
- Periodic external assessment against a real standard like the CEFR.
Most learners assemble this from five disconnected tools and lose the connective tissue: the words from Monday’s podcast never make it into Wednesday’s review, and the AI conversation never touches the grammar point the course covered.
A worked example of the integrated version: LearnAmo’s Italian Academy combines weekly podcasts and explained news, graded mini stories, downloadable exercises, an AI chat that corrects errors and answers grammar questions (by typing or speaking aloud), AI-generated flashcards with spaced repetition, and a Telegram group with teachers and other students. It’s a useful reference point because it shows the architecture clearly — the AI handles conversation and review, while curriculum, authentic material and human accountability come from elsewhere in the system. The platform’s own suggested week is instructive regardless of which tools you use: listen Monday, drill Tuesday, read Wednesday, story Thursday, AI conversation Friday, community Saturday, review-only Sunday.
How to evaluate any AI language tool
Four questions cut through most of the marketing:
- Does it correct me, or just respond to me? Many conversational tools optimise for a pleasant exchange and silently accept broken grammar.
- Where does the content come from? Model-generated practice text is cheap and slightly artificial. Authentic material is expensive and closer to the language you’ll actually meet.
- Does what I learn on Monday come back on Sunday? Without automated review, retention leaks.
- Can I find out my level from something other than the app itself? If the only measure of progress is the product’s own dashboard, treat it sceptically.
The honest summary: AI has solved the practice bottleneck and the review bottleneck, which were the two most expensive parts of learning a language. It has not solved sequencing, authenticity, or motivation — and tools that claim otherwise are describing an ambition, not a product. Build the stack accordingly.

