These are aimed at users with limited hardware — requiring less memory and storage than higher quantization formats. Other repositories on HF include the decoder layers, resulting in a much bigger file. If you're using this for generating embeddings for Chroma or FLUX models (as I am), you only need the encoder layers.