LEAF Embeddings - INT8 Quantized (FAILED v1 - DO NOT USE)
🚨 CRITICAL: This model FAILED quality evaluation - DO NOT USE for production.
⚠️ This is experiment v1 (512 tokens) - kept for research purposes only.
Status: Training completed successfully but MTEB evaluation shows critical quality loss. This serves as a baseline for comparison with the improved v2 model (2048 tokens, better architecture) currently in development.
Model Description
This model is a distilled and quantized version of google/embeddinggemma-300m trained using the LEAF (Layer-wise Early-exit Alignment Framework) methodology. It generates 768-dimensional embeddings optimized for fast CPU inference with INT8 quantization.
What is LEAF?
LEAF is a knowledge distillation framework that:
Compresses larger embedding models into smaller, faster versions
Preserves semantic quality through multi-objective training (distillation + alignment + contrastive losses)
Optimizes for CPU deployment with INT8 post-training quantization
Architecture
Property
This Model (LEAF)
Base Model (EmbeddingGemma-300m)
Dimensions
768D
768D (also 512D, 256D, 128D via Matryoshka)
Parameters
~75M (6 layers, compressed)
300M (full architecture)
Max Tokens
512
2048
Quantization
INT8 (441MB)
FP32 (~600MB)
Inference Speed
695 texts/s (CPU)
~50-100 texts/s (CPU)
Trade-offs:
✅ 6-10x faster inference on CPU
✅ Smaller model size (441MB vs ~600MB)
✅ Lower memory footprint
⚠️ Reduced context length (512 vs 2048 tokens)
⚠️ Possible quality loss from distillation (not yet benchmarked)
Performance
Inference Speed (CPU)
Throughput: 695 texts/second
Latency: ~1.4ms per text
Memory: ~500MB RAM
Hardware: Standard CPU, no GPU required
❌ ACTUAL QUALITY (MTEB Evaluation - FAILED)
Evaluation Date: 2025-10-12
Status: ❌ CRITICAL FAILURE - Model does not capture semantic relationships
Dataset
Metric
This Model (v1)
Base Model
Quality Loss
STSBenchmark
Spearman
0.223
0.81
-72% ❌
STS22 English
Spearman
0.373
0.75
-50% ❌
STS22 Average
Spearman
~0.21
0.65
-68% ❌
Cross-lingual
Spearman
-0.14 to 0.12
0.55
Complete loss ❌
Detailed STS22 Results by Language:
Language
Spearman
Status
🇨🇳 Chinese
0.499
🟡 Moderate (best)
🇸🇦 Arabic
0.469
🟡 Moderate
🇮🇹 Italian
0.435
🟡 Moderate
🇪🇸 Spanish
0.403
🟠 Poor
🇬🇧 English
0.373
🟠 Poor
🇫🇷 French
0.300
🔴 Very poor
🇷🇺 Russian
0.268
🔴 Very poor
🇹🇷 Turkish
0.247
🔴 Very poor
🇩🇪 German
0.163
❌ Critical
🇵🇱 Polish
0.132
❌ Critical
Cross-lingual pairs (translation tasks): All FAILED (scores 0.002 to -0.143)
Conclusion: This model cannot be used for semantic search, similarity tasks, or any production use. The embeddings do not preserve semantic meaning from the base model.
For questions or issues, please open an issue on the model repository.
Disclaimer: This is an experimental model for testing purposes. Performance and quality may vary. Thorough evaluation recommended before production use.