Views
No views yet
| File | Quant | Size (MB) | Recall @ IoU 0.5 | Recall @ IoU 0.95 | Mean |Δscore| | Latency (median ms, T=8) |
|---|---|---|---|---|---|---|
rfdetr-medium-f32.gguf | F32 | 125.0 | 0.9389 | 0.8937 | 0.0065 | 139.9 |
rfdetr-medium-f16.gguf ← recommended | F16 | 67.2 | 0.9389 | 0.9032 | 0.0067 | 143.9 |
rfdetr-medium-q8_0.gguf | Q8_0 | 40.2 | 0.9389 | 0.9032 | 0.0083 | 145.8 |
rfdetr-medium-q4_K.gguf | Q4_K | 32.5 | 0.8992 | 0.7714 | 0.0199 | 155.8 |
rfdetr 1.7.0) on 7 COCO val2017 images at threshold 0.5. Latency is measured with rfdetr-cli bench (8 iters + 3 warmup) at T=8 threads on a single AMD Ryzen 9 9950X3D image (coco_kitchen.jpg, 640x427).ne[0] % 256 != 0 (the decoder's 128-dim MLP halves, 60 tensors) silently fall back to Q8_0 per ggml's quantizer logic — net compression is still ~3.8× over F32. Use only when the size budget is tight; expect a measurable Recall@0.95 drop relative to F16/Q8_0 (see file table above).1# 1. Clone + build rfdetr.cpp
2git clone https://github.com/mudler/rf-detr.cpp
3cd rt-detr.cpp
4cmake -B build -DRFDETR_BUILD_CLI=ON && cmake --build build -j
5
6# 2. Download a quant (F16 recommended)
7hf download mudler/rfdetr-cpp-medium rfdetr-medium-f16.gguf --local-dir models/
8
9# 3. Run detection
10build/bin/rfdetr-cli detect \
11 --model models/rfdetr-medium-f16.gguf \
12 --input my_image.jpg \
13 --threshold 0.5 --threads 8 \
14 --output detections.jsonbenchmarks/results/accuracy_sweep.json for the full sweep across all 32 (variant × quant) cells.