Views
No views yet

Automatic Speech Recognition (ASR),
Speech Translation (ST), Spoken Question Answering (SQA),
Spoken Dialogue Summarization (SDS), Speech Instruction (SI), and Paralinguistics (PARA).Qwen2-Audio 7B, WavLLM, SALMONN, and a cascaded model.
As is shown in the following table, MERaLiON-AudioLLM performs better in the Singapore local context,
as evidenced by evaluation results on Singapore's Multitask National Speech Corpus (MNSC) datasets.[!NOTE] MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching.
| Task | Dataset | MERaLiON | Qwen2-Audio 7B | WavLLM | SALMONN-7B | Cascaded Model |
|---|---|---|---|---|---|---|
| Automatic Speech Recognition WER (↓) | LibriSpeech-Test-Clean | 0.03 | 0.03 | 0.02 | 0.10 | 0.03 |
| LibriSpeech-Test-Other | 0.05 | 0.06 | 0.05 | 0.10 | 0.05 | |
| Common-Voice-15-En-Test | 0.10 | 0.11 | 0.15 | 0.31 | 0.11 | |
| Earnings21-Test | 0.17 | 0.19 | 0.65 | 0.26 | 0.11 | |
| Earnings22-Test | 0.20 | 0.24 | 0.67 | 0.36 | 0.14 | |
| MNSC-ASR-Part 1 | 0.05 | 0.07 | - | 0.09 | 0.07 | |
| MNSC-ASR-Part 2 | 0.05 | 0.19 | - | 0.42 | 0.33 | |
| MNSC-ASR-Part 3 | 0.28 | 0.35 | - | 0.66 | 0.30 | |
| MNSC-ASR-Part 4 | 0.40 | 0.56 | - | 0.76 | 0.48 | |
| MNSC-ASR-Part 5 | 0.21 | 0.28 | - | 0.35 | 0.23 | |
| MNSC-ASR-Part 6 | 0.15 | 0.22 | - | 0.25 | 0.18 | |
| Speech Translation BLEU (↑) | CoVoST 2 En → Id | 32.62 | 16.33 | 13.84 | 14.14 | 27.62 |
| CoVoST 2 En → Zh | 37.98 | 25.77 | 31.96 | 33.89 | 35.27 | |
| CoVoST 2 En → Ta | 8.50 | 0.03 | 0.00 | 0.00 | 8.46 | |
| CoVoST 2 Id → En | 37.07 | 6.33 | 5.93 | 26.89 | 46.80 | |
| CoVoST 2 Zh → En | 15.01 | 16.47 | 2.37 | 5.30 | 15.21 | |
| CoVoST 2 Ta → En | 3.97 | 0.04 | 0.17 | 0.36 | 2.83 | |
| Spoken Question Answering LLM-as-a-Judge (↑) | SLUE-SQA-5 | 82.94 | 80.05 | 83.92 | 83.48 | 88.58 |
| Spoken-SQuAD | 70.33 | 64.86 | 77.65 | 66.40 | 88.62 | |
| CN-College-Listen-Test | 85.03 | 74.51 | 65.43 | 50.90 | 91.85 | |
| Singapore-Public-Speech-SQA | 60.32 | 58.31 | 58.55 | 59.24 | 73.11 | |
| MNSC-SQA-Part 3 | 51.4 | 42.0 | - | 40.60 | 53.20 | |
| MNSC-SQA-Part 4 | 49.0 | 39.6 | - | 36.60 | 60.20 | |
| MNSC-SQA-Part 5 | 58.2 | 51.6 | - | 44.60 | 67.20 | |
| MNSC-SQA-Part 6 | 65.2 | 53.6 | - | 46.80 | 71.60 | |
| Spoken Dialogue Summarization LLM-as-a-Judge (↑) | MNSC-SDS-Part 3 | 46.80 | 33.80 | - | 9.0 | 45.40 |
| MNSC-SDS-Part 4 | 45.80 | 24.80 | - | 7.0 | 44.00 | |
| MNSC-SDS-Part 5 | 55.2 | 40.4 | - | 17.2 | 58.00 | |
| MNSC-SDS-Part 6 | 61.8 | 46.2 | - | 24.2 | 65.40 | |
| Speech Instruction LLM-as-a-Judge (↑) | OpenHermes-Audio | 71.4 | 44.8 | 22.40 | 15.80 | 72.20 |
| Alpaca-GPT4-Audio | 73.4 | 52.6 | 21.60 | 17.20 | 73.80 | |
| Paralinguistics LLM-as-a-Judge (↑) | VoxCeleb-Gender-Test | 99.53 | 99.12 | 69.68 | 88.81 | 35.25 |
| VoxCeleb-Accent-Test | 46.35 | 29.18 | - | 34.22 | 24.64 | |
| MELD-Sentiment-Test | 42.26 | 53.49 | 50.08 | 42.07 | 56.67 | |
| MELD-Emotion-Test | 30.15 | 40.54 | 41.07 | 30.73 | 47.39 |
[!WARNING] Out of Scope use: This model is not intended for use in tool calling, math, and coding tasks.
4.46.3pip install transformers==4.46.31import librosa
2from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
3
4repo_id = "MERaLiON/MERaLiON-AudioLLM-Whisper-SEA-LION"
5
6processor = AutoProcessor.from_pretrained(
7 repo_id,
8 trust_remote_code=True,
9 )
10model = AutoModelForSpeechSeq2Seq.from_pretrained(
11 repo_id,
12 use_safetensors=True,
13 trust_remote_code=True,
14)
15
16prompt = "Given the following audio context: <SpeechHere>\n\nText instruction: {query}"
17transcribe_query = "Please transcribe this speech."
18translate_query = "Can you please translate this speech into written Chinese?"
19
20conversation = [
21 [{"role": "user", "content": prompt.format(query=transcribe_query)}],
22 [{"role": "user", "content": prompt.format(query=translate_query)}],
23]
24
25chat_prompt = processor.tokenizer.apply_chat_template(
26 conversation=conversation,
27 tokenize=False,
28 add_generation_prompt=True
29)
30
31# Use an audio within 30 seconds, 16000hz.
32audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
33audio_array = [audio_array]*2
34inputs = processor(text=chat_prompt, audios=audio_array)
35
36outputs = model.generate(**inputs, max_new_tokens=256)
37generated_ids = outputs[:, inputs['input_ids'].size(1):]
38response = processor.batch_decode(generated_ids, skip_special_tokens=True)1import torch
2import librosa
3from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
4
5repo_id = "MERaLiON/MERaLiON-AudioLLM-Whisper-SEA-LION"
6device = "cuda"
7
8processor = AutoProcessor.from_pretrained(
9 repo_id,
10 trust_remote_code=True,
11 )
12model = AutoModelForSpeechSeq2Seq.from_pretrained(
13 repo_id,
14 use_safetensors=True,
15 trust_remote_code=True,
16 attn_implementation="flash_attention_2",
17 torch_dtype=torch.bfloat16
18).to(device)
19
20prompt = "Given the following audio context: <SpeechHere>\n\nText instruction: {query}"
21transcribe_query = "Please transcribe this speech."
22translate_query = "Can you please translate this speech into written Chinese?"
23
24conversation = [
25 [{"role": "user", "content": prompt.format(query=transcribe_query)}],
26 [{"role": "user", "content": prompt.format(query=translate_query)}],
27]
28
29chat_prompt = processor.tokenizer.apply_chat_template(
30 conversation=conversation,
31 tokenize=False,
32 add_generation_prompt=True
33)
34
35# Use an audio within 30 seconds, 16000hz.
36audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
37audio_array = [audio_array]*2
38inputs = processor(text=chat_prompt, audios=audio_array)
39
40for key, value in inputs.items():
41 if isinstance(value, torch.Tensor):
42 inputs[key] = inputs[key].to(device)
43
44 if value.dtype == torch.float32:
45 inputs[key] = inputs[key].to(torch.bfloat16)
46
47outputs = model.generate(**inputs, max_new_tokens=256)
48generated_ids = outputs[:, inputs['input_ids'].size(1):]
49response = processor.batch_decode(generated_ids, skip_special_tokens=True)@misc{he2024meralionaudiollmtechnicalreport,
title={MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models},
author={{MERaLiON Team}},
year={2024},
eprint={2412.09818},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.09818},
}