Methodology and sources
Where the ranks, meanings, and pronunciations come from.
Frequency data
Ranks come from the Hermit Dave FrequencyWords Spanish list, built from the OpenSubtitles 2018 corpus. Because subtitles are dialogue, the list reflects spoken Spanish rather than written prose.
Coverage
Coverage is the share of all word occurrences in the full corpus (423,290,924 tokens) that the top ranks account for.
| Top words | Coverage |
|---|---|
| 100 | 50.67% |
| 500 | 68.6% |
| 1,000 | 74.92% |
| 2,000 | 80.65% |
| 5,000 | 87.38% |
Meanings
English meanings come from Wiktionary, via the Kaikki (Wiktextract) dataset. High-frequency words are reviewed by hand, since automatic word-by-word translation produces errors on exactly the common words learners need most.
Pronunciation
IPA, syllable breaks, and stress are generated by a rule-based transcriber for Latin American Spanish (seseo, yeismo). Spanish spelling maps to sound predictably, which makes rule-based transcription reliable.
License and attribution
The frequency data is licensed CC BY-SA 4.0, inherited from the source corpus. Attribution: Hermit Dave FrequencyWords and the OpenSubtitles project. Wiktionary content is also CC BY-SA.
Word data licensed CC BY-SA 4.0. Source: Hermit Dave FrequencyWords (OpenSubtitles). Methodology and sources.