Views
No views yet
1from unsloth import FastLanguageModel
2from transformers import TextStreamer
3
4model, tokenizer = FastLanguageModel.from_pretrained(
5 model_name=f'brunoyun/Llama-3.1-Amelia-ED-8B-v1',
6 max_seq_length=2048,
7 dtype=None,
8 load_in_4bit=False,
9 gpu_memory_utilization=0.6,
10)
11
12FastLanguageModel.for_inference(model)
13
14messages = [{'role': 'system', 'content': 'You are an expert in argumentation. Your task is to determine whether the given [SENTENCE] is an Evidence or Non-evidence. Utilize the [TOPIC] and the [ARGUMENT] as context to support your decision\nYour answer must be in the following format with only Evidence or Non-evidence in the answer section:\n<|ANSWER|><answer><|ANSWER|>.'}, {'role': 'user', 'content': '[TOPIC]: We should abandon marriage\n[ARGUMENT]: Marriage preserves norms that are harmful for women\n[SENTENCE]: Roughly one in three girls in Zimbabwe is married by her eighteenth birthday. Discriminatory social norms that link a girl’s perceived “purity” to her family’s honor are among the factors that push girls into marriage. According to Human Rights Watch, young women and girls who become pregnant, stay out late, are seen in the company of suspected boyfriends, or are otherwise thought to be sexually active can be forced into marriage in order to preserve their familial honor.\n'}]
15
16txt_streamer = TextStreamer(tokenizer, skip_prompt=True)
17
18txt = tokenizer.apply_chat_template(
19 messages,
20 add_generation_prompt=True,
21 return_tensors="pt",
22).to('cuda')
23
24_ = model.generate(
25 txt,
26 streamer=txt_streamer,
27 max_new_tokens=128,
28 pad_token_id=tokenizer.eos_token_id
29)You are an expert in argumentation. Your task is to determine whether the given [SENTENCE] is an Evidence or Non-evidence. Utilize the [TOPIC] and the [ARGUMENT] as context to support your decision\nYour answer must be in the following format with only Evidence or Non-evidence in the answer section:\n<|ANSWER|><answer><|ANSWER|>.1{
2 "name": "Llama 3 V2",
3 "inference_params": {
4 "input_prefix": "<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n",
5 "input_suffix": "<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n",
6 "pre_prompt": "You are an expert in argumentation. Your task is to determine whether the given [SENTENCE] is an Evidence or Non-evidence. Utilize the [TOPIC] and the [ARGUMENT] as context to support your decision\nYour answer must be in the following format with only Evidence or Non-evidence in the answer section:\n<|ANSWER|><answer><|ANSWER|>.",
7 "pre_prompt_prefix": "<|start_header_id|>system<|end_header_id|>\n\n",
8 "pre_prompt_suffix": "",
9 "antiprompt": [
10 "<|start_header_id|>",
11 "<|eot_id|>"
12 ]
13 }
14}| Model | ACC | CD | ED | AR | ET | SD | FD_Single | FD_Multi | AQ |
|---|---|---|---|---|---|---|---|---|---|
| Llama 3.1 8B zero-shot | 73.52% | 51.50% | 17.06% | 28.32% | 37.41% | 14.10% | 44.07% | 21.77% | 15.10% |
| Llama 3.1 8B few-shot | 75.47% | 67.83% | 64.20% | 35.97% | 49.31% | 80.00% | 48.50% | 17.25% | 31.83% |
| Llama 3.1 8B fine-tuned for ACC | 89.61% | 61.35% | 68.25% | 38.51% | 41.43% | 65.82% | 38.43% | 21.58% | 33.07% |
| Llama 3.1 8B fine-tuned for CD | 50.18% | 85.16% | 68.91% | 38.29% | 33.91% | 66.97% | 38.90% | 22.67% | 31.24% |
| Llama 3.1 8B fine-tuned for ED | 63.32% | 74.94% | 78.00% | 28.60% | 38.67% | 68.42% | 39.65% | 18.47% | 29.01% |
| Llama 3.1 8B fine-tuned for AR | 50.81% | 59.98% | 67.00% | 87.20% | 35.07% | 76.00% | 35.14% | 25.86% | 27.97% |
| Llama 3.1 8B fine-tuned for ET | 56.10% | 67.08% | 61.45% | 26.88% | 75.22% | 69.82% | 46.78% | 29.68% | 29.03% |
| Llama 3.1 8B fine-tuned for SD | 50.93% | 48.88% | 57.62% | 38.26% | 39.17% | 94.63% | 43.23% | 20.99% | 20.39% |
| Llama 3.1 8B fine-tuned for FD | 66.58% | 65.13% | 64.50% | 38.64% | 46.83% | 64.32% | 82.92% | 50.77% | 41.90% |
| Llama 3.1 8B fine-tuned for AQ | 74.46% | 59.73% | 68.00% | 30.86% | 44.06% | 60.43% | 47.98% | 24.31% | 69.54% |
| GGUF_ACC | 87.73% | 63.59% | 63.75% | 36.31% | 37.98% | 64.63% | 30.19% | 29.27% | 32.94% |
| GGUF_CD | 54.10% | 81.92% | 60.70% | 36.43% | 31.99% | 63.82% | 30.00% | 31.21% | 33.20% |
| GGUF_ED | 56.20% | 63.72% | 71.62% | 34.63% | 36.22% | 61.84% | 34.10% | 34.54% | 34.77% |
| GGUF_AR | 55.19% | 60.25% | 63.70% | 84.57% | 31.71% | 76.50% | 29.94% | 34.18% | 32.15% |
| GGUF_ET | 58.23% | 64.37% | 58.59% | 29.14% | 72.47% | 68.20% | 39.05% | 32.94% | 31.48% |
| GGUF_SD | 56.70% | 50.75% | 57.75% | 38.27% | 33.67% | 93.75% | 34.66% | 30.32% | 21.43% |
| GGUF_FD | 62.20% | 59.91% | 62.88% | 35.51% | 42.52% | 64.68% | 74.08% | 62.16% | 41.69% |
| GGUF_AQ | 67.08% | 59.73% | 69.50% | 31.17% | 41.31% | 61.16% | 41.86% | 30.02% | 66.53% |
| Llama 3.1 8B fine-tuned Multi-task | 90.74% | 84.71% | 77.75% | 88.33% | 73.84% | 95.75% | 82.53% | 50.22% | 69.80% |
| Merged Model | 78.72% | 70.69% | 69.62% | 72.52% | 54.60% | 77.04% | 57.00% | 35.03% | 57.52% |
| GGUF_Merged | 65.95% | 65.83% | 62.13% | 62.93% | 49.06% | 74.38% | 50.04% | 40.75% | 44.97% |