Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Llama-3.3-8B-Instruct-GPTQ-FP8-Block – AI Model by AlphaOxO | AlphaNeural AI
You can deploy this model and start earning money today!
AlphaOxO
/
Llama-3.3-8B-Instruct-GPTQ-FP8-Block
like
0
safetensors
llama
text-generation
conversational
tatsu-lab/alpaca
shb777/Llama-3.3-8B-Instruct-128K
quantized
compressed-tensors
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Llama 3.3 8B Instruct GPTQ FP8 Block
Using Hardware
CPU: AMD Ryzen Threadripper PRO 7995WX
MB: GIGABYTE AI TOP TRX50
GPU: RTX 5090*1
RAM: RDIMM DDR5 5600 128GB*2
Using Software
CUDA version: 13.0
CUDA driver version: 580.95.05
pyTorch: 2.10.0+cu130
transformers: 5.3.0
llmcompressor: 0.10.0.1
vllm: 0.17.1
Backbone Model
shb777/Llama-3.3-8B-Instruct-128K
Using Dataset
tatsu-lab/alpaca
convert as chat template.
Compare
--
origin model(FP16)
Quantization(This Model, FP8 Block)
leaderboard_gpqa
0.3733
0.3616
VRAM Usage (Model loading -> max_model_len = 1024)
14.99 GiB
8.49 GiB