When you know the words but can't hear them
Why a sentence made of familiar words still slips past, and how I trained my ear.
⏱ About 7 min read
This was the most infuriating thing about my first two years.
I would hear a sentence and understand nothing. Then I would turn the subtitles on and find that I knew every single word in it. Not one was unfamiliar.
It took me a long time to see the problem: my vocabulary was not the issue. The issue was that I had never heard those words in the form they are actually spoken.
Spelling lies to you
Native speakers do not pronounce words one at a time. When words run together, sounds link, drop, change and weaken.
A few sentences I used to miss, written side by side:
| Written | What it actually sounds like |
|---|---|
| What are you going to do? | Whaddaya gonna do |
| I want to go. | I wanna go |
| Did you eat yet? | Dija eat yet |
| a lot of people | a lotta people |
| next day | nex day |
| in bed | im bed |
| Give me the cup of tea. | Gimme the cuppa tea |
| I should have known. | I shoulda known |
Look at the right-hand column: not one hard word in it. But at real speed, if your ear is waiting for "what are you", it will not catch "whaddaya".
The five things that deform a sentence
Knowing their names made me hear them far faster. This is one of the rare times theory paid off immediately.
1. Linking. A final consonant sticks to the following vowel. an apple becomes a napple. turn it off becomes tur-ni-toff.
2. Elision. A sound is dropped entirely — most often /t/ and /d/ trapped between two consonants. next day loses its /t/. friendship loses its /d/.
3. Assimilation. A sound shifts to be easier to say. in bed → /im bed/, because /n/ leans towards /m/ to match the /b/ right after it. did you → dija.
4. Weak forms. This is the one I underrated the longest. Function words — to, of, for, and, can, was, are, at, from — collapse their vowel to a tiny /ə/ when unstressed. for becomes /fə/. and becomes /ən/, or just /n/.
5. Stress timing. English packs time into the meaning-carrying words and squeezes everything else to fit the beat. Vietnamese gives roughly equal time to every syllable. Two completely different rhythmic systems, so a Vietnamese ear is off-beat from the start.
The pair worth drilling first
can and can't.
In a positive sentence, can nearly vanishes: /kən/. can't is clearly stressed. If you are listening for a final /t/ to tell them apart, you will mishear a lot of sentences — American speakers usually swallow that /t/.
The real cue is stress and vowel length, not the /t/.
Where a Vietnamese ear is specifically off
Vietnamese and English differ in exactly these places, and I fell into all of them.
- Final consonants. Vietnamese has only a handful of final sounds, and none are released. So I could neither produce English final consonants nor hear them. The direct consequence: I could not hear plural -s or past-tense -ed, so for years I did not know what tense anyone was speaking in.
- Consonant clusters. Vietnamese has no clusters. texts, asked, strengths are three consonants in a row with no vowel between them — I used to drop one of the middle ones. That is also the most common error in research on Vietnamese learners' cluster errors: about 60% of cases are deleting the second consonant of the cluster outright.
- Sounds Vietnamese does not have. /θ/ and /ð/ in think and this. /ʃ/ and /s/ in she and sea. Final /z/.
- Long and short vowels. ship and sheep, full and fool, live and leave.
Why pronunciation helps your listening
This sounds backwards, but it is why I recommend pronunciation practice early even if you are not speaking yet.
If your mouth has never produced a sound, your brain has no template for it. Meeting it in fast speech, your brain will map it onto the nearest sound it does know — usually a Vietnamese one.
Pronunciation practice is not about sounding good. It is about giving your ear something to match against.
The most effective exercise I have done: dictation
One minute of audio a day. Only a minute — but finish it.
1. Pick a 30–60 second clip with an accurate transcript or subtitles
A podcast with a transcript is easiest.
2. Listen and write it out word for word
Rewind as much as you like. Ten times is fine. Don't look at the text.
3. When you can't extract anything more, open the transcript
Circle everything you got wrong or left blank.
4. Sort every error into one of two columns
Didn't know the word → make an Anki card.
Knew the word but didn't hear it → this is a sound problem. Replay that exact spot five times and say it along until your mouth produces the same thing.
5. Count the errors in each column and write it down
After a few weeks you will see which column is your real bottleneck.
The sorting in step 4 is the most valuable part. Before that I assumed I was bad at vocabulary. It turned out nearly two thirds of my errors were in the second column — words I already knew.
Shadowing — doing it right
Shadowing is listening to a sentence and repeating it almost simultaneously, without waiting for the sentence to finish.
What I used to get wrong: I tried to pronounce each word correctly. That is reading aloud, not shadowing.
What I do now, in two stages — the split comes from a research review on shadowing:
- Stage 1 — track the rhythm. Ignore meaning. Just follow the rhythm, the intonation, where it rises and falls. Mumbling is allowed, as long as the beat matches.
- Stage 2 — track the meaning. Once the rhythm is locked in, shadow while picturing the content.
10–15 focused minutes beats 40 minutes of going through the motions. And record yourself once a week — uncomfortable, but without a teacher it is the only honest feedback you get.
How long until it changes
For me, about 6–8 weeks of daily dictation before the difference was obvious. Not a sudden understanding of everything. Rather, sentences that used to slide straight past became sentences where I could hear the boundaries between words — even when I couldn't work out the meaning fast enough yet.
Hearing the boundaries is step one. Understanding in time is step two, and it comes later.
Sources for this page
- Phonological processes in English connected speech: implications for L2 speech learning and communication. Cogent Education (2025), free to read. — the five processes that deform a sentence.
- Common mistakes in pronouncing English consonant clusters: A case study of Vietnamese learners. CTU Journal of Innovation and Sustainable Development (2022). — the 60% figure for deleting the second consonant in a cluster.
- A Systematic Review of Research on the use of Shadowing for Second Language Pronunciation Teaching (2025). — the two-stage split for shadowing.