Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
tinyllama-1.1b-gguf-benchmarks – AI Model by Yash1bajpai | AlphaNeural AI
You can deploy this model and start earning money today!
Yash1bajpai
/
tinyllama-1.1b-gguf-benchmarks
like
0
gguf
mit
endpoints_compatible
us
conversational
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
TinyLlama 1.1B GGUF Quantized Models
This repository contains FP16, INT8, and INT4 quantized versions of TinyLlama 1.1B using llama.cpp.
Variants
FP16 → Best quality, highest memory
INT8 (q8_0) → Balanced performance
INT4 (q4_k_m) → Fastest, but degraded quality
Key Results
INT4: ~3x faster, ~66% less RAM, but weaker reasoning
INT8: Best trade-off between speed and quality
FP16: Most accurate but resource-heavy
Full Benchmark & Details
👉 See GitHub repo for full evaluation:
https://github.com/Yash1bajpai/edge-llm-benchmarks