Views
No views yet
pip install gensim numpy1from huggingface_hub import snapshot_download
2from gensim.models import KeyedVectors
3
4repo = snapshot_download("adameubanks/YearlyWord2Vec")
5model_2020 = KeyedVectors.load(f"{repo}/word2vec-2020/word2vec_2020.model")1from gensim.models import KeyedVectors
2
3# Load a model for a specific year
4model_2020 = KeyedVectors.load("word2vec_2020.model")
5model_2024 = KeyedVectors.load("word2vec_2024.model")
6
7# Find similar words
8print(model_2020.most_similar("covid"))
9print(model_2024.most_similar("covid"))
10
11# Compare semantic drift
12word = "technology"
13similar_2020 = model_2020.most_similar(word)
14similar_2024 = model_2024.most_similar(word)1# Study semantic drift over time
2years = [2005, 2010, 2015, 2020, 2025]
3models = {}
4
5for year in years:
6 models[year] = KeyedVectors.load(f"word2vec_{year}.model")
7
8# Analyze how a word's meaning changed
9word = "smartphone"
10for year in years:
11 similar = models[year].most_similar(word, topn=5)
12 print(f"{year}: {[w for w, s in similar]}")| Year | Corpus Size | Articles | Vocabulary |
|---|---|---|---|
| 2005 | 2.3 GB | 689,905 | 23,344 |
| 2006 | 3.3 GB | 1,047,683 | 23,142 |
| 2007 | 4.5 GB | 1,468,094 | 22,998 |
| 2008 | 7.0 GB | 2,379,636 | 23,076 |
| 2009 | 9.3 GB | 3,251,110 | 23,031 |
| 2010 | 11.6 GB | 4,102,893 | 23,008 |
| 2011 | 12.5 GB | 4,446,823 | 23,182 |
| 2012 | 20.0 GB | 7,276,289 | 23,140 |
| 2013 | 15.7 GB | 5,626,713 | 23,195 |
| 2014 | 8.7 GB | 2,868,446 | 23,527 |
| 2015 | 8.7 GB | 2,762,626 | 23,349 |
| 2016 | 9.4 GB | 2,901,744 | 23,351 |
| 2017 | 10.1 GB | 3,085,758 | 23,440 |
| 2018 | 10.4 GB | 3,103,828 | 23,348 |
| 2019 | 10.9 GB | 3,187,052 | 23,228 |
| 2020 | 12.9 GB | 3,610,390 | 23,504 |
| 2021 | 14.3 GB | 3,903,312 | 23,296 |
| 2022 | 16.5 GB | 4,330,132 | 23,222 |
| 2023 | 21.6 GB | 5,188,559 | 23,278 |
| 2024 | 27.9 GB | 6,443,985 | 24,022 |
| 2025 | 16.6 GB | 3,625,629 | 24,919 |
1@misc{yearly_word2vec_2025,
2 title={Yearly Word2Vec Embeddings: Language Evolution from 2005-2025},
3 author={Adam Eubanks},
4 year={2025},
5 url={https://huggingface.co/adameubanks/YearlyWord2Vec},
6 note={Trained on FineWeb dataset with single-year segmentation}
7}1@inproceedings{
2 penedo2024the,
3 title={The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale},
4 author={Guilherme Penedo and Hynek Kydl{\'\i}{\v{c}}ek and Loubna Ben allal and Anton Lozhkov and Margaret Mitchell and Colin Raffel and Leandro Von Werra and Thomas Wolf},
5 booktitle={The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
6 year={2024},
7 url={https://openreview.net/forum?id=n6SCkn2QaG}
8}