A language model that writes Pure Data patches and SuperCollider code, and runs
locally. Qwen2.5-Coder-7B-Instruct fine-tuned with QLoRA, exported to GGUF at
Q4_K_M.
Use bidubr/repente-v0.7-GGUF
instead. This version is published as the frozen reference on which the experiments
reported in the paper were run, so that those results stay reproducible. It is not the
model to reach for if you just want to generate patches.
Why this repository exists
Version 0.5 was published first, selected by a battery of five prompts sampled once
each. A later measurement, resampling the same five prompts 30 times on every
checkpoint, showed that the selection was a lucky draw: the result that chose v0.5 had
a 1.5% chance of occurring, and v0.7 is measurably stronger.
Model
Expected score out of 5
95% interval
Qwen2.5-Coder-7B, unmodified
0.03
[0.00, 0.10]
Repente v0.5
2.67
[2.37, 2.97]
Repente v0.7
3.83
[3.53, 4.13]
Version 0.5 remains available because the baseline comparison and the prompting study
in the paper were run against these exact weights, and withdrawing them would make
those results unverifiable. The full account, including how the wrong model came to be
published for four training cycles, is in Sections 7.6 and 11 of the paper.
What it does
Give it a description of a sound and it writes the code that produces it.
4.5 GB on disk. Fits in 8 GB of VRAM with room for a 4096-token context.
Prompting
The system prompt used throughout the reported experiments is short:
You are Repente, a musical programming expert.
Format validity improves substantially when three worked examples and a short
chain-of-thought instruction are supplied together. Measured on frozen weights, the
combination raises format validity from 75% to 100% across four prompt difficulty
levels and eliminates output truncation, at no compute cost. Few-shot exemplars
supplied alone introduce a context-leak failure mode in which the model continues the
turn structure of the examples until the token limit; the reasoning instruction
suppresses it. Details in Section 8 of the paper.
Known limitations
ELSE objects are not generated. Across 1,350 measured generations, an object from
the ELSE library appears once. Five cycles of corpus weighting, including one that
weighted ELSE explicitly, did not change this. Treat library coverage as a retrieval
problem at inference time, not something the weights will supply.
Analysis responses are short. The unmodified base model produces analyses averaging
776 tokens; this model produces 151. The compression is a self-distillation artifact of
using each version's own output to train the next, and it is documented rather than
fixed.
Patch validity is not patch quality. The measurements above score a generation as
passing when it carries a canvas header, an output object, and connections. That is
necessary for a patch to make sound and not sufficient for it to make the right sound.