Views
No views yet
| Specification | Details |
|---|---|
| Source Model | https://huggingface.co/ReadyArt/Broken-Tutu-24B-Transgression-v2.0 |
| Quantization Method | INT4-AWQ |
| Precision | 4-bit weights |
| KV Cache | int8 |
| KV Cache Type | paged |
| KV Reuse | enabled |
| Block/Group Size | 128 |
| TensorRT-LLM Version | 1.2.0rc5 (used for quantization) |
| Max Batch Size | 64 |
| Max Input Length | 5525 |
| Max Output Length | 150 |
| SM Architecture | sm90 |
| GPU | NVIDIA H100 NVL |
| CUDA Toolkit | 13.0 |
| Generated | 2026-01-03 17:53:47 UTC |
trt-llm/
checkpoints/
*.safetensors
config.json
engines/sm90_trt-llm-1.2.0rc5_cuda13.0/
rank*.engine
config.json| Parameter | Value |
|---|---|
| Method | INT4-AWQ |
| Calibration Size | 64 samples |
| Calibration Seq Length | 5675 |
| AWQ Block Size | 128 |
| Calibration Batch Size | 16 |
1.2.0rc5engines/ subdirectories