Views
No views yet

gmail, jira, calendar, docs, etc.
0.5M and 0.1M input/output tokens, then cost is 1 * 0.5 + 0.1 * 10 = $1.5.| LLM | Acc % |
|---|---|
| GPT-5 Mini (medium) | 0.71 |
| Qwen3-1.7B | 0.82 |
| Funcdex-1.7B | 0.86 |
| LLM | Exact Match | String Ratio | Total Cost ($) |
|---|---|---|---|
| GPT-OSS-120B (medium) | 0.35 | 0.51 | 9.32 |
| GPT-5 Mini (medium) | 0.35 | 0.58 | 99.71 |
| GPT-5 (minimal) | 0.18 | 0.59 | 205.45 |
| Qwen3-0.6B | 0.27 | 0.59 | 2.83 |
| Qwen3-1.7B | 0.27 | 0.69 | 5.73 |
| Funcdex-0.6B | 0.39 | 0.70 | 0.19 |
| Funcdex-1.7B | 0.43 | 0.81 | 5.64 |
| Toolkit | GPT-OSS-120B (medium) | GPT-5 (minimal) | GPT-5 Mini (medium) | Qwen3-0.6B | Funcdex-0.6B | Qwen3-1.7B | Funcdex-1.7B | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EM | SR | EM | SR | EM | SR | EM | SR | EM | SR | LoRA Checkpoint | EM | SR | EM | SR | LoRA Checkpoint | |
| 0.38 | 0.47 | 0.12 | 0.68 | 0.49 | 0.71 | 0.33 | 0.63 | 0.46 | 0.69 | 🤗 | 0.30 | 0.79 | 0.52 | 0.82 | 🤗 | |
| 0.47 | 0.56 | 0.41 | 0.63 | 0.41 | 0.56 | 0.44 | 0.66 | 0.54 | 0.78 | 🤗 | 0.47 | 0.74 | 0.54 | 0.86 | ||
| 0.48 | 0.70 | 0.24 | 0.69 | 0.50 | 0.73 | 0.27 | 0.61 | 0.47 | 0.72 | 🤗 | 0.31 | 0.73 | 0.53 | 0.83 | ||
| 0.27 | 0.52 | 0.20 | 0.50 | 0.21 | 0.51 | 0.21 | 0.53 | 0.39 | 0.74 | 🤗 | 0.23 | 0.64 | 0.47 | 0.83 | ||
| 0.19 | 0.38 | 0.07 | 0.49 | 0.18 | 0.46 | 0.07 | 0.58 | 0.13 | 0.64 | 🤗 | 0.11 | 0.62 | 0.18 | 0.79 | ||
| 0.34 | 0.52 | 0.19 | 0.61 | 0.38 | 0.58 | 0.26 | 0.65 | 0.40 | 0.75 | 🤗 | 0.26 | 0.73 | 0.48 | 0.82 | ||
| 0.47 | 0.53 | 0.17 | 0.65 | 0.47 | 0.66 | 0.51 | 0.69 | 0.58 | 0.76 | 🤗 | 0.47 | 0.76 | 0.59 | 0.83 | ||
| 0.15 | 0.37 | 0.10 | 0.46 | 0.12 | 0.39 | 0.08 | 0.50 | 0.17 | 0.71 | 🤗 | 0.09 | 0.56 | 0.16 | 0.80 | ||
| 0.65 | 0.74 | 0.19 | 0.72 | 0.64 | 0.79 | 0.57 | 0.87 | 0.65 | 0.88 | 🤗 | 0.55 | 0.91 | 0.72 | 0.94 | ||
| 0.23 | 0.39 | 0.13 | 0.47 | 0.24 | 0.43 | 0.20 | 0.43 | 0.28 | 0.64 | 🤗 | 0.26 | 0.55 | 0.31 | 0.71 |
| Bundle | GPT-OSS-120B (medium) | GPT-5 (minimal) | GPT-5 Mini (medium) | Qwen3-0.6B | Funcdex-0.6B | Qwen3-1.7B | Funcdex-1.7B | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EM | SR | EM | SR | EM | SR | EM | SR | EM | SR | LoRA Checkpoint | EM | SR | EM | SR | LoRA Checkpoint | |
| 0.28 | 0.53 | 0.15 | 0.54 | 0.22 | 0.56 | 0.19 | 0.51 | 0.26 | 0.54 | 🤗 | 0.17 | 0.61 | 0.32 | 0.71 | 🤗 | |
| 0.32 | 0.45 | 0.17 | 0.52 | 0.35 | 0.47 | 0.19 | 0.49 | 0.35 | 0.60 | 🤗 | 0.15 | 0.66 | 0.40 | 0.78 | ||
| 0.28 | 0.37 | 0.12 | 0.50 | 0.33 | 0.47 | 0.18 | 0.54 | 0.34 | 0.70 | 🤗 | 0.19 | 0.68 | 0.43 | 0.76 | ||
| 0.42 | 0.60 | 0.18 | 0.66 | 0.36 | 0.66 | 0.29 | 0.61 | 0.39 | 0.71 | 🤗 | 0.28 | 0.72 | 0.44 | 0.82 | ||
| 0.32 | 0.58 | 0.19 | 0.66 | 0.35 | 0.69 | 0.26 | 0.50 | 0.41 | 0.70 | 🤗 | 0.27 | 0.68 | 0.39 | 0.77 |
(context_messages, function_calls) and use it to generate predictions. We ignore the content field and only evaluate function_calls generated by an LLM.tool_choice="auto".difflib.SequenceMatcher.ratio. The number reported is average string ratio."email_content": "This is an example." v/s "email_content": "This is an Example.", both only differ by one letter.vllm serve ojus1/Qwen3-1.7B-Instruct --enable-lora --lora-modules prem-research/Funcdex-1.7B=prem-research/Funcdex-1.7B --enable-auto-tool-choice --tool-call-parser hermes1from transformers import AutoModelForCausalLM, AutoTokenizer
2from peft import PeftModel
3import torch
4import json
5
6# Load model and tokenizer
7base_model_name = "ojus1/Qwen3-1.7B-Instruct"
8model_name = "prem-research/Funcdex-1.7B"
9
10tokenizer = AutoTokenizer.from_pretrained(model_name)
11
12base_model = AutoModelForCausalLM.from_pretrained(
13 base_model_name,
14 torch_dtype="auto",
15 device_map="auto"
16)
17
18model = PeftModel.from_pretrained(
19 base_model,
20 model_name,
21 torch_dtype="auto",
22 device_map="auto"
23)
24
25# Define tools (supports all toolkits)
26tools = [
27 {
28 "type": "function",
29 "function": {
30 "name": "CREATE_SHARED_DRIVE",
31 "description": "Create a new shared drive in Google Drive",
32 "parameters": {
33 "type": "object",
34 "properties": {
35 "name": {"type": "string", "description": "Name of the shared drive"},
36 "requestId": {"type": "string", "description": "Unique request ID"}
37 },
38 "required": ["name", "requestId"]
39 }
40 }
41 },
42 {
43 "type": "function",
44 "function": {
45 "name": "CREATE_A_FOLDER",
46 "description": "Create a folder in Google Drive",
47 "parameters": {
48 "type": "object",
49 "properties": {
50 "folder_name": {"type": "string", "description": "Name of the folder"},
51 "parent_id": {"type": "string", "description": "Parent drive or folder ID"}
52 },
53 "required": ["folder_name", "parent_id"]
54 }
55 }
56 }
57]
58
59# Define conversation
60messages = [
61 {"role": "system", "content": "You are a helpful assistant that can help with tasks by using tools."},
62 {"role": "user", "content": "Create a shared drive named 'Partner-Alpha-Integration' with request ID 'req-12345'."}
63]
64
65# Apply chat template with tools
66formatted_input = tokenizer.apply_chat_template(
67 messages,
68 tools=tools,
69 tokenize=False,
70 add_generation_prompt=True
71)
72
73# Tokenize and generate
74input_tokens = tokenizer(formatted_input, return_tensors="pt").to(model.device)
75output = model.generate(**input_tokens, max_new_tokens=256, do_sample=False)
76response = tokenizer.decode(output[0][input_tokens['input_ids'].shape[1]:], skip_special_tokens=True)
77
78print("Response:", response)
79# Expected output includes: <tool_call>{"name": "CREATE_SHARED_DRIVE", "arguments": {"name": "Partner-Alpha-Integration", "requestId": "req-12345"}}</tool_call>