Views
No views yet

googlecalendar.
| LLM | Acc % |
|---|---|
| GPT-5 Mini (medium) | 0.71 |
| Qwen3-1.7B | 0.82 |
| Funcdex-1.7B | 0.86 |
| LLM | Exact Match | String Ratio | Total Cost ($) |
|---|---|---|---|
| GPT-OSS-120B (medium) | 0.35 | 0.51 | 9.32 |
| GPT-5 Mini (medium) | 0.35 | 0.58 | 99.71 |
| GPT-5 (minimal) | 0.18 | 0.59 | 205.45 |
| Qwen3-0.6B | 0.27 | 0.59 | 2.83 |
| Qwen3-1.7B | 0.27 | 0.69 | 5.73 |
| Funcdex-0.6B | 0.39 | 0.70 | 0.19 |
| Funcdex-1.7B | 0.43 | 0.81 | 5.64 |
| Toolkit | GPT-OSS-120B (medium) | GPT-5 (minimal) | GPT-5 Mini (medium) | Qwen3-0.6B | Funcdex-0.6B | Qwen3-1.7B | Funcdex-1.7B | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EM | SR | EM | SR | EM | SR | EM | SR | EM | SR | LoRA Checkpoint | EM | SR | EM | SR | LoRA Checkpoint | |
| 0.38 | 0.47 | 0.12 | 0.68 | 0.49 | 0.71 | 0.33 | 0.63 | 0.46 | 0.69 | 🤗 | 0.30 | 0.79 | 0.52 | 0.82 | 🤗 | |
| 0.47 | 0.56 | 0.41 | 0.63 | 0.41 | 0.56 | 0.44 | 0.66 | 0.54 | 0.78 | 🤗 | 0.47 | 0.74 | 0.54 | 0.86 | ||
| 0.48 | 0.70 | 0.24 | 0.69 | 0.50 | 0.73 | 0.27 | 0.61 | 0.47 | 0.72 | 🤗 | 0.31 | 0.73 | 0.53 | 0.83 | ||
| 0.27 | 0.52 | 0.20 | 0.50 | 0.21 | 0.51 | 0.21 | 0.53 | 0.39 | 0.74 | 🤗 | 0.23 | 0.64 | 0.47 | 0.83 | ||
| 0.19 | 0.38 | 0.07 | 0.49 | 0.18 | 0.46 | 0.07 | 0.58 | 0.13 | 0.64 | 🤗 | 0.11 | 0.62 | 0.18 | 0.79 | ||
| 0.34 | 0.52 | 0.19 | 0.61 | 0.38 | 0.58 | 0.26 | 0.65 | 0.40 | 0.75 | 🤗 | 0.26 | 0.73 | 0.48 | 0.82 | ||
| 0.47 | 0.53 | 0.17 | 0.65 | 0.47 | 0.66 | 0.51 | 0.69 | 0.58 | 0.76 | 🤗 | 0.47 | 0.76 | 0.59 | 0.83 | ||
| 0.15 | 0.37 | 0.10 | 0.46 | 0.12 | 0.39 | 0.08 | 0.50 | 0.17 | 0.71 | 🤗 | 0.09 | 0.56 | 0.16 | 0.80 | ||
| 0.65 | 0.74 | 0.19 | 0.72 | 0.64 | 0.79 | 0.57 | 0.87 | 0.65 | 0.88 | 🤗 | 0.55 | 0.91 | 0.72 | 0.94 | ||
| 0.23 | 0.39 | 0.13 | 0.47 | 0.24 | 0.43 | 0.20 | 0.43 | 0.28 | 0.64 | 🤗 | 0.26 | 0.55 | 0.31 | 0.71 |
| Bundle | GPT-OSS-120B (medium) | GPT-5 (minimal) | GPT-5 Mini (medium) | Qwen3-0.6B | Funcdex-0.6B | Qwen3-1.7B | Funcdex-1.7B | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EM | SR | EM | SR | EM | SR | EM | SR | EM | SR | LoRA Checkpoint | EM | SR | EM | SR | LoRA Checkpoint | |
| 0.28 | 0.53 | 0.15 | 0.54 | 0.22 | 0.56 | 0.19 | 0.51 | 0.26 | 0.54 | 🤗 | 0.17 | 0.61 | 0.32 | 0.71 | 🤗 | |
| 0.32 | 0.45 | 0.17 | 0.52 | 0.35 | 0.47 | 0.19 | 0.49 | 0.35 | 0.60 | 🤗 | 0.15 | 0.66 | 0.40 | 0.78 | ||
| 0.28 | 0.37 | 0.12 | 0.50 | 0.33 | 0.47 | 0.18 | 0.54 | 0.34 | 0.70 | 🤗 | 0.19 | 0.68 | 0.43 | 0.76 | ||
| 0.42 | 0.60 | 0.18 | 0.66 | 0.36 | 0.66 | 0.29 | 0.61 | 0.39 | 0.71 | 🤗 | 0.28 | 0.72 | 0.44 | 0.82 | ||
| 0.32 | 0.58 | 0.19 | 0.66 | 0.35 | 0.69 | 0.26 | 0.50 | 0.41 | 0.70 | 🤗 | 0.27 | 0.68 | 0.39 | 0.77 |
(context_messages, function_calls) and use it to generate predictions. We ignore the content field and only evaluate function_calls generated by an LLM.tool_choice="auto".difflib.SequenceMatcher.ratio. The number reported is average string ratio."email_content": "This is an example." v/s "email_content": "This is an Example.", both only differ by one letter.1from transformers import AutoModelForCausalLM, AutoTokenizer
2from peft import PeftModel
3import torch
4import json
5
6# Load model and tokenizer
7base_model_name = "ojus1/Qwen3-0.6B-Instruct"
8model_name = "prem-research/Funcdex-0.6B-googlecalendar"
9
10tokenizer = AutoTokenizer.from_pretrained(model_name)
11
12base_model = AutoModelForCausalLM.from_pretrained(
13 base_model_name,
14 torch_dtype="auto",
15 device_map="auto"
16)
17
18model = PeftModel.from_pretrained(
19 base_model,
20 model_name,
21 torch_dtype="auto",
22 device_map="auto"
23)
24
25# Define tools
26tools = [
27 {
28 "type": "function",
29 "function": {
30 "name": "CREATE_EVENT",
31 "description": "Create a new calendar event",
32 "parameters": {
33 "type": "object",
34 "properties": {
35 "summary": {"type": "string", "description": "Event title"},
36 "start_time": {"type": "string", "description": "Start time in ISO format"},
37 "end_time": {"type": "string", "description": "End time in ISO format"},
38 "attendees": {"type": "array", "items": {"type": "string"}, "description": "List of attendee emails"}
39 },
40 "required": ["summary", "start_time", "end_time"]
41 }
42 }
43 },
44 {
45 "type": "function",
46 "function": {
47 "name": "LIST_EVENTS",
48 "description": "List calendar events",
49 "parameters": {
50 "type": "object",
51 "properties": {
52 "calendar_id": {"type": "string", "description": "Calendar ID"},
53 "time_min": {"type": "string", "description": "Start time filter"},
54 "max_results": {"type": "integer", "description": "Maximum number of events"}
55 }
56 }
57 }
58 }
59]
60
61# Define conversation
62messages = [
63 {"role": "system", "content": "You are a helpful assistant that can help with tasks by using tools."},
64 {"role": "user", "content": "Create a team meeting event for tomorrow at 2 PM for 1 hour. The title should be 'Q4 Planning Session'."}
65]
66
67# Apply chat template with tools
68formatted_input = tokenizer.apply_chat_template(
69 messages,
70 tools=tools,
71 tokenize=False,
72 add_generation_prompt=True
73)
74
75# Tokenize and generate
76input_tokens = tokenizer(formatted_input, return_tensors="pt").to(model.device)
77output = model.generate(**input_tokens, max_new_tokens=256, do_sample=False)
78response = tokenizer.decode(output[0][input_tokens['input_ids'].shape[1]:], skip_special_tokens=True)
79
80print("Response:", response)vllm serve ojus1/Qwen3-0.6B-Instruct --enable-lora --lora-modules prem-research/Funcdex-0.6B=prem-research/Funcdex-0.6B-googlecalendar --enable-auto-tool-choice --tool-call-parser hermes