Subtitle Translation Similarity

This experiment compares how 20 languages express equivalent English words and phrases in subtitle translations. The scores reflect recurring patterns visible through English, not total linguistic distance or shared ancestry. Toggle between the heatmap and graph to explore.

Loading similarity data…

How This Works

Where the data comes from

The scores are calculated from the OpenSubtitles v2024 dataset: a large collection of subtitle translations from films and TV. 500,000 sentence pairs per language were sampled, covering 20 languages and 190 unique pairs.

How the scores are calculated

For every word in English, we look at which foreign words appear as its translation in each language. The cosine similarity between those patterns becomes the score. The closer to 1, the more similarly two languages express equivalent English vocabulary in this subtitle dataset.

How reliable are the scores?

Each score comes with a 95% bootstrap confidence interval, a margin of error calculated by re-running the measurement 1,000 times on random subsets of the data. The typical margin is ±0.015, so the broad clusters are more informative than tiny differences between scores. Hover any cell to see the exact range.

Reading the heatmap

Languages are grouped by default using hierarchical clustering, so pairs that score most alike end up next to each other. Switch to Family sort to group them by linguistic family instead. Darker purple means higher similarity; lighter means lower.

Interacting with the Tool

  • Hover a heatmap cell to see the score, confidence interval, and a short note on what the result means.
  • Search a language name to highlight its full row and column.
  • Threshold slider greys out cells below the selected similarity.
  • Switch to Graph view for a force-directed layout; click any node for top-5 neighbours.

Limitations

These scores measure recurring translation patterns through English. The cosine similarity calculation normalises for each language's overall relationship with English, so a language being "far from English" in absolute terms does not by itself push its scores down. The comparison still cannot capture every aspect of similarity between two languages. The main limitation is that English vocabulary maps unevenly across language families: Romance and Germanic languages share a lot of English roots, giving their translation vectors denser overlap. Families like Semitic and East Asian have sparser mappings, so their scores carry slightly more uncertainty rather than being systematically wrong. Subtitle language is also informal and conversational, which affects which vocabulary gets counted.

Read the full experiment

The methodology, findings and limitations are discussed in Why do some languages feel weirdly familiar?, published by The Cambridge Language Collective.

Learn languages with real content

SubSmith lets you study any of these languages using films, TV and your own media. Generate transcripts, mine sentences and build vocabulary that sticks.

Plan your roadmap