← Back to the word list

Methodology and sources

Where the ranks, meanings, and pronunciations come from.

Frequency data

Ranks come from the Hermit Dave FrequencyWords Spanish list, built from the OpenSubtitles 2018 corpus. Because subtitles are dialogue, the list reflects spoken Spanish rather than written prose.

Coverage

Coverage is the share of all word occurrences in the full corpus (423,290,924 tokens) that the top ranks account for.

Top wordsCoverage
10050.67%
50068.6%
1,00074.92%
2,00080.65%
5,00087.38%

Meanings

English meanings come from Wiktionary, via the Kaikki (Wiktextract) dataset. High-frequency words are reviewed by hand, since automatic word-by-word translation produces errors on exactly the common words learners need most.

Pronunciation

IPA, syllable breaks, and stress are generated by a rule-based transcriber for Latin American Spanish (seseo, yeismo). Spanish spelling maps to sound predictably, which makes rule-based transcription reliable.

License and attribution

The frequency data is licensed CC BY-SA 4.0, inherited from the source corpus. Attribution: Hermit Dave FrequencyWords and the OpenSubtitles project. Wiktionary content is also CC BY-SA.

ShadowingKit app icon
Knowing the words is one thing. Saying them out loud is another. ShadowingKit trains your speaking reflexes with native audio, so the words you recognize become words you can actually say.
Download on the App Store

Word data licensed CC BY-SA 4.0. Source: Hermit Dave FrequencyWords (OpenSubtitles). Methodology and sources.