Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
Apertus-8B-Instruct-2509-W8A16 – AI Model by jgerster0 | AlphaNeural AI | AlphaNeural AI
You can deploy this model and start earning money today!
jgerster0
/
Apertus-8B-Instruct-2509-W8A16
like
0
safetensors
apertus
quantization
llm
swissai
compressed-tensors
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Apertus-8B-Instruct-2509-W8A16
This is an
INT8 weight-only quantized
version of
swiss-ai/Apertus-8B-Instruct-2509
using
llm-compressor
.
💡 What this means in practice
Only the
weights
are quantized to
8-bit integers (INT8)
Activations remain FP16/BF16
No FP8 is used in this configuration
Faster inference, reduced memory,
minimal accuracy loss
Quantization Details
Quantization Scheme
: W8A16 (INT8 weights, FP16 activations)
Method
: Weight-only INT8 quantization
Targets
: All
Linear
layers
Ignored Layers
:
lm_head
(kept in higher precision)
Tool
: llm-compressor