Skip to content
All articles
  • Product
  • Speaking practice

Why an AI language tutor is suddenly a real thing

Speaking is the part of a language nobody practises, for reasons that were economic rather than pedagogical — and those reasons stopped being true about a year ago.

DZ
Dan ZabrotskiFounder & CEO

6 min read

Almost everyone who has studied a language has the same story. Three years of school French, or two years of Duolingo Spanish, or a shelf of Japanese textbooks. You can read a menu. You can follow the gist of a podcast if the host talks slowly. And then somebody asks you a question in the street and nothing comes out.

That gap is so common that people treat it as a personal failing — a lack of talent, or confidence, or the mythical ear for languages. It is neither. It is the predictable result of practising one skill and hoping a different one shows up.

Comprehension and production are different skills

Understanding a sentence and producing one run on different machinery. When you read or listen you are doing recognition: the words arrive, and you match them against something you already hold. When you speak you are doing retrieval under time pressure — you have to find the word, inflect it, order it, and get it out of your mouth before the other person's attention moves on, all while planning the next clause.

Recognition is comfortable. Retrieval is not. And the thing about retrieval is that it only improves by being done. You can recognise the Spanish subjunctive in a thousand sentences and still not produce one when you need it, for the same reason that watching a thousand hours of tennis does not give you a serve.

So the honest version of the problem is this: the one skill that everybody wants is the one skill that almost no learning product actually trains. Not because anyone is being lazy. Because until recently it was the expensive one.

Speaking practice was rationed by economics

To practise speaking you have historically needed a second human being who is (a) fluent, (b) patient with your mistakes, (c) awake at the same time as you, and (d) willing to spend an hour on you. Each of those conditions costs something, and together they set a price.

A tutor on a marketplace runs somewhere between twenty and sixty dollars an hour depending on the language. Booked in advance, in a slot, with a cancellation policy. That is not a bad deal — a good tutor is worth every cent of it — but it means practice arrives in scheduled blocks. Two hours a week, if you are diligent. Nine minutes a day, averaged out.

The alternatives are worse. Language exchange partners are free, which is exactly what they are worth in reliability; half the conversation is in your native language anyway, and both of you spend it being polite rather than correcting each other. Conversation groups mean six people and one hour, so you speak for ten minutes and listen for fifty. Talking to yourself in the shower works better than people expect, and still has nobody in it to tell you that what you just said was wrong.

None of this is a pedagogical decision. Nobody looked at how languages are learned and concluded that speaking should be the scarcest, most expensive, hardest-to-schedule part. It is scarce because human attention is scarce.

What actually changed

Language models could hold a conversation in Spanish four years ago. That was never the blocker. Three other things were, and they resolved more or less at once.

Latency. This is the one that matters most and gets discussed least. A conversation is a real-time system. Humans take turns with gaps measured in about two hundred milliseconds, and the gap itself carries meaning — hesitation reads as doubt, an instant answer reads as confidence. Chain speech-to-text, then a model, then text-to-speech, and you get two or three seconds of dead air per turn. Technically it is a conversation. In practice it is a walkie-talkie, and after five minutes you stop trying to be spontaneous because spontaneity is not rewarded. Speech-to-speech models cut that to a few hundred milliseconds. Below roughly half a second, people stop performing for the machine and start just talking — which is the entire point.

Tolerance of bad input. A learner's speech is exactly the input that older recognition systems were worst at: accented, hesitant, half in the wrong language, full of restarts. The systems that fell over on that were useless precisely for the people who needed them. Current models handle it, and more importantly they can hear the mistake as a mistake rather than transcribing something plausible and moving on.

Memory. A tutor who does not remember you is not a tutor, it is a conversation partner. The value of a real teacher accrues over months: they know you keep dropping articles, that you are learning for a job in Berlin and not for a holiday, that you gave up on the past subjunctive twice already. That is a product problem more than a model problem — deciding what to keep, what to resurface, and when — but it is the difference between practice and chat.

Put those three together and the constraint that made speaking practice scarce stops being a constraint. Not because an AI is better than a good teacher. It is not. Because it is available at eleven at night, for the fourteen minutes you actually have, without either of you having to book anything.

What this is honestly good for

An AI tutor is very good at the boring, high-volume, slightly humiliating part of learning a language: saying the same structure forty times until it stops requiring thought. It has infinite patience for that, which humans do not, and you have no social reason to feel embarrassed in front of it, which is worth more than it sounds. A large part of the speaking gap is not knowledge but the fear of being slow in front of another person, and that fear does not attach to software.

It is good at being available at the moment the motivation exists. Most learning does not fail during the lesson, it fails in the gaps between lessons.

It is good at noticing patterns across sessions that neither of you would notice within one.

What it is not good at is being a person. It does not care whether you succeed, it cannot tell you what a sentence sounds like to a Berliner at a party, and it will not notice that you have been avoiding the language for three weeks because your job got hard. A human teacher does all of that, and if you can afford one, have one. Use the AI for the reps in between, which is where the hours are anyway.

The uncomfortable version

Here is the part that language products do not usually say out loud: none of this makes learning a language easy. It is still hundreds of hours. It is still mostly boring. The only thing that has changed is that the most useful hour — the one where you actually open your mouth and get something wrong and get corrected — went from being the hardest hour to arrange to the easiest.

That is a smaller claim than most of the marketing you will read. It also happens to be the one that matters.

HolaBonjourこんにちはCiao안녕OláHallo

Your first conversation is two minutes away

Join 20,000 learners speaking new languages with Speekl. No credit card, no fixed lessons — just talk.

Free plan forever · Cancel anytime