Standalone demonstration jupyter notebook.
It can be launched freely on colab or locally, inside our outside its repo.
Implementation of a
SoundStream-style neural audio codec.
The model operates only on 16kHz audio. FiLM and bitrate dropout are not
used in this implementation.
The model shrinks time dimension 200 times.
RVQ uses 8 codebooks, each sized 1024=2^10. Hence it needs 80 bits
per timestep. If we compress 1s audio in 16kHz sample rate we will get
16000 / 200 = 80 timesteps per second, so the implemented codec compresses
audio to 6.4kbps.