GGUF conversions of CREPE, a convolutional pitch (F0) estimator, for use
with CrispASR's ggml runtime.
CREPE runs directly on the raw waveform — no STFT, no CQT — and emits a
360-bin pitch activation per frame.
Files
file
capacity
quant
size
crepe-tiny-f16.gguf
tiny
f16
0.93 MB
crepe-tiny-q8_0.gguf
tiny
q8_0
0.50 MB
crepe-tiny-q4_k.gguf
tiny
q4_k
0.27 MB
crepe-full-f16.gguf
full
f16
42.4 MB
crepe-full-q8_0.gguf
full
q8_0
22.6 MB
crepe-full-q4_k.gguf
full
q4_k
12.0 MB
tiny is the recommended default; full is ~38× more compute per frame for a
modest accuracy gain and is the right choice for offline work.
Input / output contract
Input: 16 kHz mono audio. The model consumes 1024-sample frames, each
normalized per-frame (subtract mean, divide by max(std, 1e-10)). Reference
hop is 10 ms.
Output: 360 activations per frame, sigmoid-valued. The bins are spaced
20 cents apart:
Bin 0 ≈ 32.7 Hz, bin 359 ≈ 1975.5 Hz. Decode with the original CREPE
weighted-local-average around the argmax; the activation peak value doubles as
a voicing confidence.
Quantization
Only conv*.weight and classifier.weight are quantized. The per-channel
affine parameters — conv*.bias, conv*_BN.scale, conv*_BN.offset,
classifier.bias — are kept at F32 deliberately: in CREPE the ReLU comes
before the BatchNorm, so the BN cannot be folded into the conv and ships as a
standalone per-channel affine. Rounding those would apply a multiplicative
error to an entire channel.
Note on q4_k: the conv2–conv6 kernels are 64 taps wide, and 64 is not a
multiple of Q4_K's 256-element super-block, so those five tensors fall back to
Q4_0 (32-element blocks). conv1 (512 taps) and classifier are true Q4_K.
There is no size penalty — Q4_0 and Q4_K are both 4.5 bits per weight.
Measured fidelity
Two independent measurements. Prefer f16 or q8_0.
Per-frame, against the model's own f16 (crispasr-diff crepe on 1101 frames
of real speech) — cos_min and the fraction of frames whose argmax pitch bin
is unchanged:
f16
q8_0
q4_k
tiny
0.999999 · 100%
0.999807 · 98.5%
0.961643 · 85.2%
full
1.000000 · 100%
0.999937 · 99.5%
0.992563 · 91.4%
q4_k does not meet a 0.999 cosine bar at either capacity. For tiny-q4_k,
roughly 1 frame in 7 lands on a different pitch bin than f16. Ship q4_k only
if size genuinely dominates and you post-filter by voicing confidence; it is not
a drop-in for f16/q8_0. q8_0 is effectively lossless and is the right choice
whenever f16's size is inconvenient.
The f16 files themselves score cos = 1.0 against torchcrepe (max abs error
~2e-5 tiny / ~4e-6 full, i.e. f16 weight rounding).
Accuracy on real music
Evaluated on 10 monophonic instrumental recordings (violin arco + pizzicato,
piano, glockenspiel, carillon, cello, flute, three folk melodies, brass). With no
hand-labelled F0, the proxies are tiny-vs-full octave disagreement and the
in-tessitura rate over frames with voiced_prob >= 0.5:
tiny
full
in-tessitura
89.6%
89.0%
octave disagreement tiny-vs-full
2.3%
—
tiny is not meaningfully worse than full on monophonic music, despite
being ~38x cheaper — so tiny is the recommended default. Known domain limits,
shared by both capacities: plucked/percussive attacks with fast decay (violin
pizzicato scored ~50%, most frames having no sustained pitch) and inharmonic
sources such as bells, where the model correctly abstains — a carillon clip
marked only 39/1501 frames voiced at tiny — rather than inventing pitch.
Caveat: the tessitura bounds are hand-chosen, so the absolute percentages are
soft; the tiny-vs-full comparison is the robust part, both being scored
identically. A labelled MIR dataset is still needed for an absolute note-F.
Performance
Measured on an Apple M1 (quiet box), 10 s of audio at the reference 10 ms hop:
model
Metal
CPU
tiny
RTF 0.28
RTF ~2.4
full
RTF 2.0
RTF ~40
CREPE is genuinely expensive per frame (≈7.3 GFLOP per second of audio for
tiny, ≈282 GFLOP/s for full). Neither capacity is real-time on CPU — the
GPU path is not optional here.
Provenance and license
MIT, at every step of the chain:
Original model: Jong Wook Kim, Justin Salamon, Peter Li, Juan Pablo Bello,
"CREPE: A Convolutional Representation for Pitch Estimation", ICASSP 2018.
Released under the MIT license.
(paper ·
code)
Weights: taken from torchcrepe
by Max Morrison (MIT), which is itself a port of the original CREPE Keras
weights.
This conversion: models/convert-crepe-to-gguf.py in CrispASR (MIT).
If you use CREPE, please cite the original paper:
bibtex
1@inproceedings{kim2018crepe,
2 title = {{CREPE}: A Convolutional Representation for Pitch Estimation},
3 author = {Kim, Jong Wook and Salamon, Justin and Li, Peter and Bello, Juan Pablo},
4 booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
5 year = {2018}
6}
Upstream licence:mit. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.