The journal On learning

Why you understand more than you can say — and how to close the gap

Every language learner runs into it: you follow the conversation, you get the joke, and then someone turns to you and you freeze. The science behind that feeling is illuminating — and the fix is more specific than you might expect.

You have been studying the language for a year. You sit down to watch a film in it — no subtitles, as an experiment — and you are quietly amazed. You catch the gist of whole scenes. You laugh at the right moments. The sounds resolve into words and the words into meaning in a way that would have been impossible a year ago. You feel, briefly, like someone who knows what they are doing.

Then a native speaker asks you, simply and warmly, what you thought of the film. And you open your mouth and the words are gone. You know you know them. You just cannot, in this moment, find a single one.

This is one of the most disorienting experiences in language learning, and also one of the most universal. It has a name — the comprehension-production gap — and understanding what causes it turns out to be genuinely useful, because the cause points directly at the cure.

Two vocabularies in one head

Every language learner carries two overlapping vocabularies at once. The first is receptive: words you understand when you encounter them in speech or on a page. The second is productive: words you can actually summon when you need to speak or write. Researchers in applied linguistics, most notably Paul Nation at Victoria University of Wellington, have spent decades measuring these two stores, and the consistent finding is that receptive vocabulary is considerably larger — sometimes dramatically so — than productive vocabulary. The words you passively recognise vastly outnumber the words you can actively reach for.

This is not a failure of memory. It is how language acquisition works. Comprehension can happen with a partial, rough sense of a word — you hear soledad and the feeling of aloneness washes over you without your needing to pin down the exact grammar or all the word's collocations. Production demands far more: to use soledad, you need to know how it inflects, what it combines with, whether it fits this register, and you need to retrieve all of that from nothing, in real time, under social pressure. The bar for understanding is lower than the bar for speaking, and so understanding always runs ahead.

Recognition and recall are different acts

The gap is not just a matter of vocabulary size — it is a matter of cognitive architecture. Recognising a word and retrieving a word are genuinely different mental operations, and research suggests they engage overlapping but distinct neural processes. Recognition asks: does this incoming signal match something I have stored? The threshold is low, the answer comes fast, and partial knowledge is often enough to tip you over the line. Recall asks: starting from nothing, produce the word, its sound, its spelling, its meaning, its grammar. It is a harder search with no prompt to recognise and match.

Psychologists Roger Brown and David McNeill described a revealing edge case in 1966: the tip-of-the-tongue state, in which you know a word exists, you can describe its meaning, you can even recall how many syllables it has — but you cannot retrieve it. The word is, in some real sense, in your memory. Your access to it has simply broken down. Tip-of-the-tongue moments happen in your native language too, but they are far more frequent in a language you are learning, where the retrieval pathways are newer and thinner. The word is not missing. It is just harder to reach.

Understanding a word is arriving at a destination on a map you recognise. Producing it is finding the road from a standing start with no map at all.

Why input alone does not close the gap

In the 1980s, the linguist Stephen Krashen put forward an influential idea: that language acquisition happens through comprehensible input — through reading and listening to things just beyond your current level — and that this alone is sufficient for fluency. The theory is elegant, and there is genuine truth in it. Comprehensible input is powerful; immersion works; reading in a foreign language is deeply good for you.

But in 1985, the Canadian applied linguist Merrill Swain noticed something uncomfortable in the data from French immersion programmes in Canada. Students who had received years of rich, comprehensible input in French could understand the language well. They still could not produce it with anything like native-level accuracy. Swain proposed what she called the Output Hypothesis: that comprehensible input, for all its virtues, is not the whole story. You also need to be pushed to produce — to speak, to write, to put the language together yourself — and that act of production does something input cannot.

The reason, Swain argued, is that output forces you to notice the gap. When you try to say something and discover you cannot say it the way you want, you have identified a hole in your knowledge with a precision that passive listening rarely achieves. The frustration of reaching for a word and missing it is not wasted; it is a precise diagnostic. It tells you exactly what you need to learn next, and it primes you to notice that thing the next time it appears in input. Comprehension tells you what you know. Production tells you what you do not yet know well enough.

What actually pushes words from passive to active

The path from receptive to productive knowledge runs through retrieval practice — not through more reading, and not through more listening. Every time you successfully pull a word from memory without a prompt, that retrieval pathway grows a little stronger. Every time you look at the English meaning and have to produce the Spanish, or hear a definition and have to supply the word, you are doing something qualitatively different from encountering the word in a sentence and understanding it. You are exercising the exact cognitive process that speaking requires.

This is why the exercises that feel harder tend to be the ones that work. Translating a word you half-remember is more productive than reading a sentence in which it appears and nodding along. Writing the word from memory consolidates it more deeply than ticking a box that says you recognised it. The additional effort is not incidental difficulty — it is the mechanism. As we explored in why words slip away, the testing effect is one of the sturdiest findings in memory research: the act of retrieval, including retrieval that requires a struggle, strengthens a memory more than passive re-exposure ever does.

Frequency matters here too. Words in your passive vocabulary were most likely learned in only one or two contexts — you heard them in a film, you saw them in an article. Words in your active vocabulary were encountered from many angles: you heard the word, you looked it up, you wrote it in a sentence, you got corrected, you wrote it again. Multiple encounters from multiple directions weave a richer web of connections around a word, and it is that web that makes retrieval fast and reliable rather than effortful and fragile.

The key distinction: consuming language grows your receptive vocabulary. Producing language — being asked to recall, to translate, to use — is what converts it to vocabulary you can actually reach for when you speak.

What this means in practice

If you want to close the comprehension-production gap, the priority is clear: spend less time encountering words passively and more time being asked to produce them. That does not mean abandoning reading and listening — those remain essential, and rich input is one of the best ways to meet new words. But input alone, even lots of it, will not automatically unlock the ability to speak. Production practice has to be deliberate.

Talk to people, even badly. Write things down. When you learn a word, try to use it in a sentence of your own before the day is out. If you are using a vocabulary app, pay attention to which exercises ask you to produce — to type the word, to fill in the blank, to recall the translation — rather than simply to recognise. Recognition exercises can tell you how much you understand; they cannot tell you how much you can say.

This is the reasoning behind how vocabla structures its sessions. Every word is practised from five angles — translation, definition, synonym, antonym, and open recall — and each angle is an act of production, not recognition. You are not asked to pick the right answer from a list; you are asked to find the word yourself, from the meaning, from context, from a related word. It is more demanding than flashcards, and it is supposed to be. The difficulty is the point: each successful retrieval deepens the pathway, and over dozens of spaced sessions, what began as a word you vaguely recognised becomes a word you reach for without thinking.

The gap between understanding and speaking is not a sign that you have been learning wrong. It is a natural feature of how language builds. The question is simply whether your practice is doing the specific work of closing it — and now you know what that work looks like.

On learning