Views
No views yet
laion/clap-htsat-unfused,
publie sous la licence de l'original (Apache-2.0). Sert a
NoiseKernel pour l'auto-tagging et la recherche semantique
de samples, en local et hors ligne.| Fichier | Taille | Role |
|---|---|---|
clap-audio.onnx | 116 Mo | Branche audio. Suffit a tagger et a « des sons comme celui-ci ». |
clap-text.onnx | 501 Mo | Branche texte. N'achete QUE la recherche en texte libre. |
vocab.json | 798 Ko | Tokenizer RoBERTa BPE, copie conforme de l'amont. |
merges.txt | 456 Ko | Idem. |
labelvec.bin | 172 Ko | Vecteurs des 86 labels de la taxonomie NoiseKernel. Specifique au projet. |
labelvec.bin est ce qui permet de tagger sans jamais charger la branche texte : les 86 vecteurs
de labels sont calcules une fois ici, au lieu d'exiger 501 Mo sur chaque appareil pour retrouver
exactement les memes.entree audio input_features float32 (batch, 1, 1001, 64)
entree texte input_ids int64 (batch, tokens)
attention_mask int64 (batch, tokens)
sortie embedding float32 (batch, 512), DEJA normalise L2n_fft 1024, hop 480, fmin 50 Hz, fmax 14 kHz, clip de 10 s
(480 000 echantillons), echelle mel Slaney (pas HTK), fenetre de Hann periodique, padding par
reflexion, plancher 1e-10, sortie en dB.is_longer, seconde entree du modele amont, est absorbee dans l'enveloppe : verifiee inerte sur
la variante unfused (ecart exactement nul). Cela ne vaut pas pour la variante fused.audio ecart_max 3.390e-07 cos_min 1.000000
texte/9 ecart_max 6.221e-07 cos_min 1.000000
texte/16 ecart_max 6.650e-07 cos_min 1.000000
texte/64 ecart_max 5.364e-07 cos_min 1.0000001pip install torch transformers onnx onnxruntime numpy
2tools/clap/export-clap.py --out build/clap # les quatre premiers fichiers
3engine/build/rm_make_labelvec build/clap # labelvec.bin (taxonomie NoiseKernel)1@inproceedings{laionclap2023,
2 title = {Large-scale Contrastive Language-Audio Pretraining with
3 Feature Fusion and Keyword-to-Caption Augmentation},
4 author = {Wu, Yusong and Chen, Ke and Zhang, Tianyu and Hui, Yuchen and
5 Berg-Kirkpatrick, Taylor and Dubnov, Shlomo},
6 booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP},
7 year = {2023}
8}