Skip to content
All articles
  • Product
  • AI

Why can't I just use ChatGPT?

You can, and for some things you should. Here is an honest account of what a general assistant does well, where it quietly fails as a tutor, and what has to be built around the model.

DZ
Dan ZabrotskiFounder & CEO

7 min read

It is the first question anybody sensible asks, and it deserves a real answer rather than a defensive one. ChatGPT speaks better Spanish than most Spanish teachers. It has a voice mode. It costs twenty dollars a month for everything, not just for languages. So what exactly is left for a language product to do?

Quite a lot, it turns out — but almost none of it is the part people assume.

Start with what it genuinely does well

If you want a grammar point explained, a general assistant is excellent and you should use it. It will explain the difference between ser and estar more clearly than the textbook, adapt the explanation when you say you did not follow, and generate twenty example sentences on demand. That is a real capability and it used to cost you a teacher's time.

It is also good at translation, at drafting the email you need to send in German by Thursday, at telling you whether a phrase sounds natural, and at playing a scenario if you set it up carefully. For all of that, it is the right tool and there is nothing to add.

The problems start when you try to use it as the thing you practise with, several times a week, for months.

It is trained to understand you. That is the bug.

A general assistant's whole job is to be helpful, and being helpful means figuring out what you meant. So you say something mangled — wrong gender, wrong tense, a word borrowed from English with a hopeful accent — and it understands you perfectly and answers the question.

This feels wonderful. You are being understood in a foreign language. It is also the single most efficient way to make a mistake permanent, because the error was not just tolerated, it was rewarded with a successful conversation. Linguists call the endpoint fossilisation: an error repeated enough times that it stops being a mistake and becomes your dialect. Every fluent-but-permanently-wrong speaker you have ever met got there by being understood too often.

A tutor's job includes not understanding you when it matters. That is a strange thing to ask a model that has been optimised in the opposite direction for its entire training run.

Correction is a policy, not a capability

The obvious fix is to ask for corrections. Try it and watch what happens.

Say "correct all my mistakes" and the next three turns come back as a graded essay: every article, every preposition, every hesitation marked up. You stop speaking, because nobody wants to talk to a machine that grades them. Say nothing and you drift back to being understood.

The real answer is neither, and it changes by the sentence. An error that blocks meaning gets corrected now. An error that is merely wrong gets corrected at the end of the turn, if it is the one you have made four times this week. An error above your level gets ignored entirely, because you are not ready and telling you would only cost you your fluency for the next minute. A beginner gets corrected on maybe one thing in five; someone preparing for a C1 exam wants the opposite.

That is a policy. It depends on your level, your error history, what today's session is for, and how the last thirty seconds went. You cannot put it in a prompt because it is not one decision, it is a decision per utterance, and it needs to survive across sessions.

The memory is about facts, not about you

General assistants have memory now, and it is genuinely useful: it remembers that you are learning Italian, that you have a trip in October, that you prefer short answers.

What a tutor needs is a different shape of thing. It needs to know that you have produced hätte correctly nine times out of thirty-one attempts across six weeks, that your errors cluster in verb-final subordinate clauses, that you have met the word Verspätung four times and never once said it out loud. That is not a memory feature, it is a database with a scheduler on top — a record of every word and structure you have encountered, produced, and failed, with dates, so that the right thing can come back at the right interval.

Spaced repetition has been the best-evidenced result in learning research for about forty years. It requires bookkeeping. Bookkeeping is not something you get from a better model.

Nothing keeps the difficulty in the right place

The most robust finding in second-language acquisition is roughly: you learn from input that is a little above your current level. A little. Comprehensible input, slightly stretched.

Ask an assistant to speak at A2 and it will, for a while. Then the conversation gets interesting, and it drifts back up to its own register, because its register is fluent adult native speaker and there is nothing pulling it down. You will not notice — you will just find the session tiring and conclude that you had a bad day.

Holding the input at the edge of your ability, turn after turn, session after session, is an active control problem. It needs a current estimate of your level, and something checking the output against it.

Voice is not a mode, it is the whole thing

Text and voice look like the same conversation in a different wrapper. They are not.

Speaking is real-time. The gap before you answer is part of the exchange, and a two-second pipeline delay teaches you to compose your sentence fully before opening your mouth — the exact habit that keeps people from ever becoming fluent. Fluency is the ability to start a sentence before you know how it ends.

Pronunciation feedback needs the audio. Once speech has been turned into text, the evidence is gone: the transcript says you said the word, and the transcript is wrong, because the whole question was how you said it. Anything useful about your vowels or your stress has to happen before that step.

And interruption matters. A tutor who lets you talk yourself into a wall is not helping; a tutor who cuts in every time you hesitate is unbearable. That is a turn-taking design problem, not a prompt.

It waits. A tutor arrives with a plan.

Open a chat window and it does nothing until you type. That is correct behaviour for an assistant and hopeless for a teacher, because the hardest part of language learning is not the lesson, it is deciding to have one.

A tutor knows what today is for before you do. Fifteen minutes on the past tense you have been avoiding, five minutes of vocabulary that is due, then the restaurant scenario because your trip is in three weeks. You should have to decide to show up, and nothing after that.

So what is the actual difference?

Not the model. We use frontier models too, and if a better one ships next month we will use that instead. The model is the most commoditised part of the whole stack and everyone has the same access to it.

The difference is everything around it: a level estimate that updates, an error profile that persists, a correction policy that changes by utterance, a scheduler deciding what comes back today, a voice pipeline built for turn-taking instead of transcription, and a plan for the session you have not started yet.

That is unglamorous work, and it is where the learning actually comes from.

When you should just use ChatGPT

Genuinely: when you want something explained, translated, checked, listed, or drafted. It is better at that than any language app, ours included, and pretending otherwise would be silly.

Use a tutor for the other thing — the reps, the speaking, the being corrected over months. The two are not competing for the same hour.

HolaBonjourこんにちはCiao안녕OláHallo

Your first conversation is two minutes away

Join 20,000 learners speaking new languages with Speekl. No credit card, no fixed lessons — just talk.

Free plan forever · Cancel anytime