July 21, 2026  ·  Blog

Can You Understand English Recorded on a Potato?

We ran a state-of-the-art speech recognizer over 33,664 real English recordings. It fully cleared 60% of them. The other 40% is the audio you skip on YouTube — and it is where listening is actually learned.

"Recorded on a potato" is the internet's verdict on bad audio: the laptop mic across the room, the phone in a windbreaker pocket, the voice memo with the fridge humming through it. Everyone knows the reflex it triggers. The video opens, the audio is potato, and you are gone in four seconds. Not because you couldn't have understood it — because you didn't want to work that hard.

Here is what that reflex is protecting you from. We recently scored our entire English corpus — 33,664 clips of real people reading real sentences on whatever microphone they owned — using Whisper, a speech recognizer trained on more spoken English than any human hears in a hundred lifetimes. It transcribed 60% of the clips perfectly. On the rest it stumbled: a word swapped here, a phrase mangled there, and on the true potatoes, whole clauses lost. A machine that has effectively heard all the English there is still can't cleanly parse four recordings in ten.

Those clips aren't defective. They are ordinary. That is what spoken language actually sounds like once it leaves the studio: compressed, clipped, half-swallowed, competing with a kettle. The clean audio in your textbook app is the exception — a laboratory condition that the real world never reproduces.

And here is the inversion worth savoring: the corpus that embarrasses the machine was recorded by volunteers — thousands of ordinary people, on whatever microphone they owned, in whatever room they were in. No studio could assemble this, and no studio would want to. That is precisely why volunteers on potato mics make the most comprehensive language course and the most honest assessment there is: they sample the actual distribution of human speech — every accent, every room tone, every cheap diaphragm — instead of the one approved voice reading at one approved speed.

Shannon knew

In 1948 Claude Shannon formalized what every listener does without knowing it: communication over a noisy channel. The signal degrades in transit; the receiver reconstructs the message anyway — not by hearing every bit, but by exploiting redundancy, by predicting what the corrupted parts must have been. Human language is drenched in redundancy precisely because it evolved for noisy channels: shouting across fields, whispering in markets, talking over wind. Your native ear is a magnificent error-correcting decoder. You don't hear every word your friends say in a loud bar. You reconstruct them, instantly, from priors so deep you never notice the reconstruction happening.

A learner who has only ever heard clean audio has trained the microphone's job, not the listener's. The listener's job is error correction — and error correction can only be learned on a channel that has errors.

This is why the potato clip that makes you bounce from a YouTube video is not an obstacle between you and the language. It is the exam. The gap between "I understand the podcast at 1x with earbuds" and "I understand my in-laws in a moving car" is exactly the gap between a clean channel and a noisy one — and no amount of clean-channel practice closes it, for the same reason that no amount of rallying returns a professional serve.

One more precision, because the word "dirty" oversells it: nothing about this audio is dirty. It is ordinary speech sampled by really cheap digital microphones — a faithful recording of how the language actually travels. And notice the asymmetry it creates: the audio native speakers hate listening to is nonetheless the best instrument for assessing a foreign learner, precisely because natives pass it annoyed but effortlessly while learners fail it honestly. A crisp podcast helps you learn; nobody disputes that. But it cannot tell you where you stand, because it removes exactly the conditions that separate the decoders from the subtitle-readers.

Training on potatoes, on purpose

The practical conclusion is uncomfortable and simple: some meaningful fraction of your listening practice should be audio you would normally skip. Recordings where you get 60%, guess 20% more from context, and honestly lose the rest — then see the transcript, find out what the noise was hiding, and let the missed ones come back around until your decoder catches up. The struggle isn't a defect of the material. The struggle is the reconstruction machinery being built.

We wrote elsewhere about why dirty audio is the signal — why verification against real recordings beats another vocabulary pass. The potato data sharpens that argument with a number: if a machine trained on everything clears only 60% cleanly, then a curriculum of studio audio is a curriculum with the hardest 40% of reality quietly removed. That 40% is not noise in your way. It is the half of the language that lives in the noise.

So: can you understand English recorded on a potato? Today, probably not all of it. That's not the bad news. That's the syllabus.

SiteDictation is dictation practice on real recordings by real speakers — potatoes included, on purpose. Listen, type what you heard, see exactly what the noise was hiding, and let spaced repetition bring your misses back. Test your decoder →