Views
No views yet
1EVALUATION_PROMPTS = {
2 "persona": [
3 {
4 "role": "user",
5 "content": """
6Evaluate two responses (A and B) to a given question based on the following criteria:
7
81. Personalization: Does the response effectively consider the user's provided personal information?
92. Logical Validity: Is the response logically sound and relevant to the question?
10
11For each criterion, provide a brief one-line comparison of the two responses and select the better response (A, B, or Tie).
12
13- Ensure the comparisons are concise and directly address the criteria.
14- If both answers are equally strong or weak in a category, mark it as a Tie.
15- Do not use bold font.
16
17Output Format:
181. Personalization: [Brief comparison of A and B]
192. Logical Validity: [Brief comparison of A and B]
20Better Response: [A/B/Tie]
21
22Input:
23- User's Personal Information:
24 - Characteristics:
25 - The user {relation} {entity}
26- Question: {question}
27- Answer (A): {answer_a}
28- Answer (B): {answer_b}
29""".strip(),
30 },
31 ],
32 "profile": [
33 {
34 "role": "user",
35 "content": """
36Evaluate two responses (A and B) to a given question based on the following criteria:
37
381. Personalization: Does the response effectively consider the user's provided personal information?
392. Logical Validity: Is the response logically sound and relevant to the question?
40
41For each criterion, provide a brief one-line comparison of the two responses and select the better response (A, B, or Tie).
42
43- Ensure the comparisons are concise and directly address the criteria.
44- If both answers are equally strong or weak in a category, mark it as a Tie.
45- Do not use bold font.
46
47Output Format:
481. Personalization: [Brief comparison of A and B]
492. Logical Validity: [Brief comparison of A and B]
50Better Response: [A/B/Tie]
51
52Input:
53- User's Personal Information:
54 - Profile:
55 - Age: {age}
56 - Gender: {gender}
57 - Personality: {personality}
58 - Occupation: {occupation}
59 - Income Range: {income_range}
60 - Characteristics:
61 - The user {relation} {entity}
62- Question: {question}
63- Answer (A): {answer_a}
64- Answer (B): {answer_b}
65""".strip(),
66 },
67 ],
68 "schedule": [
69 {
70 "role": "user",
71 "content": """
72Evaluate the given answer based on the provided question and user's schedule.
73
74The evaluation criterion is as follows:
75- Personalization: Does the answer acknowledge and appropriately consider the user's schedule?
76
77Provide a short one-line evaluation for the criterion and indicate whether it satisfies the criterion with YES or NO. Do not use bold font.
78
79Output Format:
80- Personalization: [Short one-line evaluation]
81Satisfaction: [YES/NO]
82
83Example 1:
84- User's Schedule:
85 - Datetime: 11 AM on March 13, 2025
86 - Event: Attend a networking brunch with fellow sports journalists and industry experts.
87- Question: Can you add a meeting with the editorial team to my calendar for 11 AM on March 13, 2025?
88- Answer: It seems you already have a networking brunch with fellow sports journalists and industry experts scheduled for 11 AM on March 13, 2025. Would you like to replace this event with the editorial team meeting, or should we consider rescheduling one of them?
89
90- Personalization: The answer acknowledges the existing event and offers alternatives.
91Satisfaction: YES
92
93Example 2:
94- User's Schedule:
95 - Datetime: 11 AM on March 13, 2025
96 - Event: Attend a networking brunch with fellow sports journalists and industry experts.
97- Question: Can you add a meeting with the editorial team to my calendar for 11 AM on March 13, 2025?
98- Answer: Certainly, I can add the meeting with the editorial team to your calendar for 11 AM on March 13, 2025.
99
100- Personalization: The answer ignores the existing event in the user's schedule.
101Satisfaction: NO
102
103Input:
104- User's Schedule:
105 - Datetime: {datetime}
106 - Event: {event}
107- Question: {question}
108- Answer: {answer}
109""".strip(),
110 },
111 ],
112}1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
3
4model = AutoModelForCausalLM.from_pretrained("12kimih/Llama-3.2-3B-HiCUPID", torch_dtype=torch.bfloat16, device_map="auto")
5tokenizer = AutoTokenizer.from_pretrained(model.config._name_or_path, padding_side="left")
6tokenizer.pad_token = tokenizer.eos_token
7
8prompt = [
9 {
10 "role": "user",
11 "content": """
12Evaluate two responses (A and B) to a given question based on the following criteria:
13
141. Personalization: Does the response effectively consider the user's provided personal information?
152. Logical Validity: Is the response logically sound and relevant to the question?
16
17For each criterion, provide a brief one-line comparison of the two responses and select the better response (A, B, or Tie).
18
19- Ensure the comparisons are concise and directly address the criteria.
20- If both answers are equally strong or weak in a category, mark it as a Tie.
21- Do not use bold font.
22
23Output Format:
241. Personalization: [Brief comparison of A and B]
252. Logical Validity: [Brief comparison of A and B]
26Better Response: [A/B/Tie]
27
28Input:
29- User's Personal Information:
30 - Characteristics:
31 - The user has not been to New Zealand
32- Question: What wildlife is unique to certain regions?
33- Answer (A): You're interested in wildlife unique to certain regions. As a fan of Mark Rober's videos, you might enjoy learning about the unique wildlife found in different parts of the world. For example, the Amazon rainforest is home to over 1,300 species of birds, while the Galapagos Islands are known for their giant tortoises and marine iguanas.
34- Answer (B): New Zealand is home to unique wildlife such as the kiwi bird, tuatara, and Hector's dolphin.
35""".strip(),
36 }
37]
38
39inputs = tokenizer.apply_chat_template(prompt, tokenize=True, add_generation_prompt=True, padding=True, return_dict=True, return_tensors="pt").to("cuda")
40outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
41decoded_outputs = tokenizer.batch_decode(outputs[:, inputs["input_ids"].size(1) :], skip_special_tokens=True)
42print(decoded_outputs[0])11. Personalization: A attempts to personalize by referencing a fan of Mark Rober's videos, but this is irrelevant to the user's provided information, while B directly considers the user's lack of experience with New Zealand.
22. Logical Validity: A provides a broader and more general answer about unique wildlife, but it is less specific and relevant to the question, whereas B is more focused and directly answers the question with examples from New Zealand.
3Better Response: B1@misc{mok2025exploringpotentialllmspersonalized,
2 title={Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis},
3 author={Jisoo Mok and Ik-hwan Kim and Sangkwon Park and Sungroh Yoon},
4 year={2025},
5 eprint={2506.01262},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL},
8 url={https://arxiv.org/abs/2506.01262},
9}