Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Llama-2-70b-chat-hf_FP8_MLPerf_V2 – AI Model by amd | AlphaNeural AI
You can deploy this model and start earning money today!
amd
/
Llama-2-70b-chat-hf_FP8_MLPerf_V2
like
0
safetensors
llama
meta-llama/Llama-2-70b-chat-hf
quantized
llama2
fp8
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Quark Team FP8 Llama-2-70b-chat-hf Model Overview
Model Information For MLPerf
Model Name
: Llama-2-70b-chat-hf
Version
: MLPerf v5.0
Commit
: Close Division Commit
Calibration Dataset
The calibration dataset consists of
1000 OpenOcra samples
provided by MLPerf
Quantized Tensors
The following tensors are quantized in each decoder:
MLP Layer inputs and weights
Linear (including QKVO linear) layer inputs and weights
KV Cache Entries
Ignored Layers
The following layers are ignored during quantization:
lm_head
Model Performance Comparison
Metric
Baseline Accuracy Target (%)
FP8 Quant Accuracy (%)
Open Orca (Chat)
- Rouge1
44.4312
44.6369
- Rouge2
22.0352
22.1798
- RougeL
28.6162
28.8249
License
Modifications copyright(c) 2024 Advanced Micro Devices,Inc. All rights reserved.