Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Quantized-Llama-3.2-1B-15k – AI Model by ANISH-j | AlphaNeural AI
You can deploy this model and start earning money today!
ANISH-j
/
Quantized-Llama-3.2-1B-15k
like
0
gguf
AdamLucek/Orpo-Llama-3.2-1B-15k
quantized
mit
endpoints_compatible
us
conversational
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Orpo-Llama-3.2–1B-15k (Q4_K_M)
Quantized GGUF version of
Orpo-Llama-3.2–1B-15k
, optimized for efficient local inference using
:contentReference[oaicite:0]{index=0}
.
Model Overview
Base Model:
meta-llama/Llama-3.2–1B
Fine-tuning:
ORPO
Training Data:
15k samples from
mlabonne/orpo-dpo-mix-40k
Parameters:
~1B
Architecture:
Llama (Transformer)
Quantization
Method:
Post-Training Quantization (PTQ)
Format:
GGUF
Precision:
Q4_K_M (4-bit K-quant)
Tooling:
llama.cpp
Size Reduction:
~2.3 GB (FP16) → ~700–850 MB
Conversion Pipeline
Hugging Face
.safetensors
GGUF F16 (intermediate)
GGUF Q4_K_M (final)
Technical Specs
Layers:
16
Context Length:
131072
Embedding Size:
2048
FFN Size:
8192
Attention Heads:
32 (KV: 8)
Vocab Size:
128,258
RoPE Base:
500,000
Quantization Version:
2
Usage
Designed for fast CPU / low-VRAM inference via llama.cpp-compatible runtimes.
Resources
Colab (conversion reference):
https://colab.research.google.com/drive/1bZOnFp01XcLaYYWt2xJ5LAzOD9msvEZL?usp=sharing