In the Netherlands, movies and TV shows are subtitled, not dubbed. In Germany, they are dubbed, not subtitled. Both countries start from the same place — children who speak Dutch or German and know almost no English — yet the average Dutch person consistently outperforms the average German on English proficiency tests, year after year, in the EF English Proficiency Index. The Netherlands and Sweden, the two most enthusiastic subtitling countries in Europe, routinely sit at the very top of the ranking. The most common explanation? Screen time. Decades of watching English-language television with subtitles has functioned as a giant, unpaid language course for an entire population.

If an entire country can learn English from subtitled TV, what can you learn from a Netflix subscription? More than you might think — and there is now a solid decade of peer-reviewed research on exactly how much, under what conditions, and how to make it work with flashcards.


The Case for Screen Time: Incidental Learning Is Real

The idea that you can learn a language without consciously studying it — by simply understanding content — is called incidental learning, and for decades it was treated with suspicion. Classroom research kept showing that learners needed explicit instruction. But the subtitle studies told a different story.

In 1999, Dutch researchers Cees Koolstra and Tom Beentjes published a study in Educational Technology Research and Development that tested elementary school children who watched an English-language television program at home. One group watched with Dutch subtitles, one without, and a control group watched a Dutch program. The result: the children who watched with subtitles showed the highest vocabulary acquisition and English word recognition — without a single minute of instruction. The subtitle did the teaching.

Around the same time, the Belgian psychologist Géry d'Ydewalle and his colleagues at KU Leuven ran a landmark series of experiments showing that children could acquire foreign-language vocabulary incidentally from subtitled television — and, in a famous finding, that reversed subtitles (your native language on the soundtrack, the foreign language as subtitles) produced measurable word learning all by themselves. You read the new language while you hear the old one, and the pairing sticks.

"Television is considered an important source of comprehensible input for second language learners of English, and there is some evidence that L2 words can be learned incidentally by watching television." — Puimège & Peters, Language Learning, 2019

The single most important recent result arrived in 2025: a meta-analysis published in the journal Language Learning that pooled 89 effect sizes from 49 independent experiments on captioned viewing. The verdict: captions produce a medium, statistically reliable effect on vocabulary learning (g = 0.56). This is not one lucky study — it is nearly five decades of experiments agreeing that on-screen text + audio + moving images = learning.

What the Captions Research Actually Shows

First, a distinction that matters. Researchers separate:

  • Captions — same-language on-screen text (English audio, English text).
  • Subtitles — translated on-screen text (English audio, your language text).

Both work, but they work differently. Captions are the more powerful learning tool for one simple reason: they give you orthographic facilitation. You hear a stream of fast, blurred speech and see exactly where one word ends and the next begins. For learners, the hardest part of listening comprehension is not unknown words — it is segmenting the stream into words at all. Captions remove that bottleneck. The eye anchors the ear, and the brain stores a stronger memory trace because it has received the same information through two channels at once.

The 2025 meta-analysis also found that the captioning effect is moderated by learner level and video material — captions help most when the material is at the right difficulty for the learner, which is exactly what the input hypothesis would predict. Captions are a ladder, not a crutch: they accelerate learning at the level where you can almost understand, and their benefit shrinks once you have outgrown them.

There is a second, less obvious finding from the viewing research that matters for flashcard users. Elke Peters and Stuart Webb, in a 2018 study in Studies in Second Language Acquisition, had learners watch a single full-length television program and tested what they picked up. They found three word-level factors predicted learning: how often the word appeared, whether it was a cognate (similar to a word in the learner's own language), and how relevant it was to the plot. Prior vocabulary size mattered too — the more you already know, the more you pick up from the same episode. This mirrors what reading research has shown for forty years: the rich get richer, and exposure compounds.

How Many Words Do You Need Before TV Becomes Comprehensible?

This is where the research gets wonderfully concrete. In 2009, Stuart Webb and Michael Rodgers published "The Vocabulary Demands of Television Programmes" in Language Learning, analyzing 88 television programs totaling 264,384 running words. Their numbers:

  • Knowing the most frequent 3,000 word families plus proper nouns and marginal words gave 95.45% coverage of the programs.
  • Knowing 7,000 word families gave 98.27% coverage.
  • Genres varied: 95% coverage required between 2,000 and 4,000 word families depending on the show; 98% required between 5,000 and 9,000.

What do these percentages mean? Researchers generally treat 95% coverage as the threshold for adequate comprehension of spoken language — slightly lower than the 98% often cited for reading, because listening gives you more redundancy. So the practical translation is blunt and encouraging: if you know the 3,000 most common word families of your target language, you have enough vocabulary to watch most television with reasonable comprehension — and every hour you spend watching is now comprehensible input that builds more vocabulary.

If you are below that threshold, you are not doomed. You just need to choose your material carefully: children's shows, news with slow anchors, or shows you have already seen in your native language. Or you use subtitles in your own language to bootstrap comprehension while your ears catch up — the d'Ydewalle reversed-subtitle effect.

Narrow Viewing: Watch the Same World, Not Random Worlds

Here is a problem with general TV viewing: each new episode throws new low-frequency vocabulary at you. In the same 2009 study, Webb and Rodgers found that learners got relatively few encounters with low-frequency words across random programs — and research on vocabulary acquisition suggests a word needs roughly ten or more meaningful encounters before it is reliably learned.

The solution is one of the most practical findings in the entire field: narrow viewing. In a follow-up study (Rodgers & Webb, 2011, in Applied Linguistics), the researchers analyzed 288 television episodes — 1,330,268 running words and 203 hours of footage — and compared vocabulary in episodes of the same series versus episodes of random programs. Related episodes recycled vocabulary: watch one season of a single series, and the same words keep reappearing, giving you the repetition your memory needs. The same trick works with a single genre — all crime shows, all cooking shows, all documentaries.

Watching one season of a single series recycles vocabulary far more than watching the same number of random episodes. Narrow viewing is the television equivalent of reading a graded-reader series.

This is why the "one anime, one show, one genre" rule of sentence-mining communities is not just taste — it is the direct application of the narrow-viewing research. Your brain treats a recurring world (the same characters, settings, and situations) as a familiar context, which lowers the comprehension threshold and multiplies encounters.

The Real-World Proof: Extramural English

The most striking evidence that screen time teaches languages comes from Scandinavia, where children learn English outside the classroom so effectively that researchers gave it a name: extramural English (coined by Swedish researcher Pia Sundqvist). In a 2012 study, Sundqvist tested 86 Swedish children aged 11–12 and found that their English proficiency correlated strongly with how often they played digital games and consumed English-language media in their free time. Not with their grades. Not with their study habits. With screen time.

There is no genetic reason Swedish children are better at English than German children. There is a media-diet reason: Swedish children grow up watching subtitled English TV, so they hear English constantly from age five; German children grow up watching dubbed TV, so they hear German. The countries that subtitle their television consistently top the EF English Proficiency Index; the countries that dub do not. An entire population acquired a second language through the same mechanism this article describes.

The Sentence Mining Protocol: Turning Episodes into Flashcards

Watching alone is not enough — that is the honest part of the science. Incidental learning from viewing is real but slow, and it leaves gaps. The fix is to capture what the screen teaches you and lock it in with spaced repetition. Here is a protocol that combines the viewing research with the flashcard science this site is built on:

1. Pick your series by coverage, not by taste alone

Choose a show where you understand roughly 90–95% of what you hear. If you are a beginner, use a series you have already seen in your native language, or children's shows. If you are intermediate, pick one genre you love and stay in it (narrow viewing).

2. Set up dual-text viewing

Use a tool like Language Reactor, Migaku, or any player that shows captions and subtitles simultaneously. Native-language subtitles for comprehension, target-language captions for segmentation. When you can, turn the native subtitles off and keep only the target-language captions — that is when the real listening gains happen.

3. Mine one episode at a time — 10 to 15 cards

After each episode, extract 10–15 sentences you almost understood but didn't quite catch, or that contain a word you heard three times. That repetition signal is your guide: a word that recurs in the show is a word worth learning (the Peters & Webb frequency effect).

4. Make sentence cards, not word cards

The Puimège & Peters (2019) study found that viewing teaches not just single words but formulaic sequences — the multi-word chunks native speakers actually use. So build cards around the full sentence:

Front:  ¿Podrías pasarme la sal?
Back:   Could you pass me the salt?
        (pasarse = to pass; "la sal" = the salt)

5. Keep the audio on the card

Your memory of the word is anchored to how it sounded in the show. Add the sentence's audio to the card — FluentCards generates native TTS audio automatically — so the review recreates the original listening moment.

6. Let spaced repetition schedule the re-encounters

This is the crucial synergy. The narrow-viewing research gives you encounters inside the episode; the FSRS algorithm gives you encounters between episodes. Every flashcard review is another meeting with that sentence, at the moment your memory curve says you need it.

Where It Goes Wrong (and How to Avoid the Traps)

The research also tells you exactly how to waste this opportunity:

  • Passive viewing with native subtitles only. If you watch with your own language on screen and never hear the target language clearly, you are reading your native language with background noise. The d'Ydewalle effect works because the foreign text is on screen, where your eyes have to process it.
  • Watching content you can't understand at all. The 95% coverage finding cuts both ways: below roughly 90% coverage, comprehension collapses and learning drops with it. This is not a moral failure — it is a sign to switch material.
  • Binge-watching without mining. The forgetting curve applies to screen time like everything else. An episode you binge today will be 60% gone in a week unless the sentences land in your review queue.
  • Confusing comprehension with acquisition. Understanding an episode feels like learning, but comprehension is the input; acquisition is the durable change in your memory. Flashcards are what convert the first into the second.

Conclusion: Your Screen Time Is Study Time

Here is the summary of everything the research supports. TV and movies are not a guilty pleasure you squeeze around your "real" studying — they are a legitimate, scientifically validated source of comprehensible input, and the countries that use them as such have the proficiency rankings to prove it. The conditions for success are known: choose material near your level, use captions strategically, watch narrow (one series or one genre), mine the sentences your brain flagged, and put them into a spaced-repetition system.

So the next time someone tells you to stop wasting time on Netflix, you have an answer backed by 49 experiments, a 264,000-word corpus analysis, and the entire nation of Sweden. Pick your show, turn on the captions, and start mining. Your screen time is study time — it just needs a flashcard to make it stick.