CEFR Llama Hidden States Combined Dataset
Overview
This dataset contains high-fidelity 1D latent representations extracted from the frozen meta-llama/Llama-3.1-8B-Instruct transformer architecture. It maps 13,837 English sentences across the 6 discrete Common European Framework of Reference for Languages (CEFR) proficiency bands ($A1 \rightarrow C2$).
This resource was engineered explicitly to support Phase 1 of a Plug and Play Language Model (PPLM) architecture… See the full description on the dataset page: https://huggingface.co/datasets/MohammadKhosravi/cefr-llama3.1-8b-hidden-states-combined.