Views
No views yet
| File | Quant | Size (MB) | Recall @ IoU 0.5 | Recall @ IoU 0.95 | Mean |Δscore| | Latency (median ms, T=8) |
|---|---|---|---|---|---|---|
rfdetr-small-f32.gguf | F32 | 119.0 | 0.9762 | 0.9514 | 0.0114 | 100.1 |
rfdetr-small-f16.gguf ← recommended | F16 | 64.0 | 0.9762 | 0.9514 | 0.0113 | 94.3 |
rfdetr-small-q8_0.gguf | Q8_0 | 38.2 | 0.9762 | 0.9425 | 0.0115 | 110.2 |
rfdetr-small-q4_K.gguf | Q4_K | 31.2 | 0.9673 | 0.8482 | 0.0267 | 116.0 |
rfdetr 1.7.0) on 7 COCO val2017 images at threshold 0.5. Latency is measured with rfdetr-cli bench (8 iters + 3 warmup) at T=8 threads on a single AMD Ryzen 9 9950X3D image (coco_kitchen.jpg, 640x427).ne[0] % 256 != 0 (the decoder's 128-dim MLP halves, 60 tensors) silently fall back to Q8_0 per ggml's quantizer logic — net compression is still ~3.8× over F32. Use only when the size budget is tight; expect a measurable Recall@0.95 drop relative to F16/Q8_0 (see file table above).1# 1. Clone + build rfdetr.cpp
2git clone https://github.com/mudler/rf-detr.cpp
3cd rt-detr.cpp
4cmake -B build -DRFDETR_BUILD_CLI=ON && cmake --build build -j
5
6# 2. Download a quant (F16 recommended)
7hf download mudler/rfdetr-cpp-small rfdetr-small-f16.gguf --local-dir models/
8
9# 3. Run detection
10build/bin/rfdetr-cli detect \
11 --model models/rfdetr-small-f16.gguf \
12 --input my_image.jpg \
13 --threshold 0.5 --threads 8 \
14 --output detections.jsonbenchmarks/results/accuracy_sweep.json for the full sweep across all 32 (variant × quant) cells.