Views
No views yet

jira and gmail.
| LLM | Acc % |
|---|---|
| GPT-5 Mini (medium) | 0.71 |
| Qwen3-1.7B | 0.82 |
| Funcdex-1.7B | 0.86 |
| LLM | Exact Match | String Ratio | Total Cost ($) |
|---|---|---|---|
| GPT-OSS-120B (medium) | 0.35 | 0.51 | 9.32 |
| GPT-5 Mini (medium) | 0.35 | 0.58 | 99.71 |
| GPT-5 (minimal) | 0.18 | 0.59 | 205.45 |
| Qwen3-0.6B | 0.27 | 0.59 | 2.83 |
| Qwen3-1.7B | 0.27 | 0.69 | 5.73 |
| Funcdex-0.6B | 0.39 | 0.70 | 0.19 |
| Funcdex-1.7B | 0.43 | 0.81 | 5.64 |
| Toolkit | GPT-OSS-120B (medium) | GPT-5 (minimal) | GPT-5 Mini (medium) | Qwen3-0.6B | Funcdex-0.6B | Qwen3-1.7B | Funcdex-1.7B | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EM | SR | EM | SR | EM | SR | EM | SR | EM | SR | LoRA Checkpoint | EM | SR | EM | SR | LoRA Checkpoint | |
| 0.38 | 0.47 | 0.12 | 0.68 | 0.49 | 0.71 | 0.33 | 0.63 | 0.46 | 0.69 | 🤗 | 0.30 | 0.79 | 0.52 | 0.82 | 🤗 | |
| 0.47 | 0.56 | 0.41 | 0.63 | 0.41 | 0.56 | 0.44 | 0.66 | 0.54 | 0.78 | 🤗 | 0.47 | 0.74 | 0.54 | 0.86 | ||
| 0.48 | 0.70 | 0.24 | 0.69 | 0.50 | 0.73 | 0.27 | 0.61 | 0.47 | 0.72 | 🤗 | 0.31 | 0.73 | 0.53 | 0.83 | ||
| 0.27 | 0.52 | 0.20 | 0.50 | 0.21 | 0.51 | 0.21 | 0.53 | 0.39 | 0.74 | 🤗 | 0.23 | 0.64 | 0.47 | 0.83 | ||
| 0.19 | 0.38 | 0.07 | 0.49 | 0.18 | 0.46 | 0.07 | 0.58 | 0.13 | 0.64 | 🤗 | 0.11 | 0.62 | 0.18 | 0.79 | ||
| 0.34 | 0.52 | 0.19 | 0.61 | 0.38 | 0.58 | 0.26 | 0.65 | 0.40 | 0.75 | 🤗 | 0.26 | 0.73 | 0.48 | 0.82 | ||
| 0.47 | 0.53 | 0.17 | 0.65 | 0.47 | 0.66 | 0.51 | 0.69 | 0.58 | 0.76 | 🤗 | 0.47 | 0.76 | 0.59 | 0.83 | ||
| 0.15 | 0.37 | 0.10 | 0.46 | 0.12 | 0.39 | 0.08 | 0.50 | 0.17 | 0.71 | 🤗 | 0.09 | 0.56 | 0.16 | 0.80 | ||
| 0.65 | 0.74 | 0.19 | 0.72 | 0.64 | 0.79 | 0.57 | 0.87 | 0.65 | 0.88 | 🤗 | 0.55 | 0.91 | 0.72 | 0.94 | ||
| 0.23 | 0.39 | 0.13 | 0.47 | 0.24 | 0.43 | 0.20 | 0.43 | 0.28 | 0.64 | 🤗 | 0.26 | 0.55 | 0.31 | 0.71 |
| Bundle | GPT-OSS-120B (medium) | GPT-5 (minimal) | GPT-5 Mini (medium) | Qwen3-0.6B | Funcdex-0.6B | Qwen3-1.7B | Funcdex-1.7B | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EM | SR | EM | SR | EM | SR | EM | SR | EM | SR | LoRA Checkpoint | EM | SR | EM | SR | LoRA Checkpoint | |
| 0.28 | 0.53 | 0.15 | 0.54 | 0.22 | 0.56 | 0.19 | 0.51 | 0.26 | 0.54 | 🤗 | 0.17 | 0.61 | 0.32 | 0.71 | 🤗 | |
| 0.32 | 0.45 | 0.17 | 0.52 | 0.35 | 0.47 | 0.19 | 0.49 | 0.35 | 0.60 | 🤗 | 0.15 | 0.66 | 0.40 | 0.78 | ||
| 0.28 | 0.37 | 0.12 | 0.50 | 0.33 | 0.47 | 0.18 | 0.54 | 0.34 | 0.70 | 🤗 | 0.19 | 0.68 | 0.43 | 0.76 | ||
| 0.42 | 0.60 | 0.18 | 0.66 | 0.36 | 0.66 | 0.29 | 0.61 | 0.39 | 0.71 | 🤗 | 0.28 | 0.72 | 0.44 | 0.82 | ||
| 0.32 | 0.58 | 0.19 | 0.66 | 0.35 | 0.69 | 0.26 | 0.50 | 0.41 | 0.70 | 🤗 | 0.27 | 0.68 | 0.39 | 0.77 |
(context_messages, function_calls) and use it to generate predictions. We ignore the content field and only evaluate function_calls generated by an LLM.tool_choice="auto".difflib.SequenceMatcher.ratio. The number reported is average string ratio."email_content": "This is an example." v/s "email_content": "This is an Example.", both only differ by one letter.1from transformers import AutoModelForCausalLM, AutoTokenizer
2from peft import PeftModel
3import torch
4import json
5
6# Load model and tokenizer
7base_model_name = "ojus1/Qwen3-0.6B-Instruct"
8model_name = "prem-research/Funcdex-0.6B-jira_gmail"
9
10tokenizer = AutoTokenizer.from_pretrained(model_name)
11
12base_model = AutoModelForCausalLM.from_pretrained(
13 base_model_name,
14 torch_dtype="auto",
15 device_map="auto"
16)
17
18model = PeftModel.from_pretrained(
19 base_model,
20 model_name,
21 torch_dtype="auto",
22 device_map="auto"
23)
24
25# Define tools (Jira + Gmail combined)
26tools = [
27 {
28 "type": "function",
29 "function": {
30 "name": "LIST_DRAFTS",
31 "description": "List Gmail draft messages",
32 "parameters": {
33 "type": "object",
34 "properties": {
35 "max_results": {"type": "integer", "description": "Maximum number of drafts"},
36 "verbose": {"type": "boolean", "description": "Verbose output"}
37 }
38 }
39 }
40 },
41 {
42 "type": "function",
43 "function": {
44 "name": "CREATE_PROJECT",
45 "description": "Create a Jira project",
46 "parameters": {
47 "type": "object",
48 "properties": {
49 "key": {"type": "string", "description": "Project key"},
50 "name": {"type": "string", "description": "Project name"},
51 "project_type_key": {"type": "string", "description": "Project type"}
52 },
53 "required": ["key", "name"]
54 }
55 }
56 },
57 {
58 "type": "function",
59 "function": {
60 "name": "GET_CONTACTS",
61 "description": "Get Gmail contacts",
62 "parameters": {
63 "type": "object",
64 "properties": {
65 "person_fields": {"type": "string", "description": "Fields to retrieve"}
66 }
67 }
68 }
69 }
70]
71
72# Define conversation
73messages = [
74 {"role": "system", "content": "You are a helpful assistant that can help with tasks by using tools."},
75 {"role": "user", "content": "List drafts from Gmail with verbose output, maximum 20 results."}
76]
77
78# Apply chat template with tools
79formatted_input = tokenizer.apply_chat_template(
80 messages,
81 tools=tools,
82 tokenize=False,
83 add_generation_prompt=True
84)
85
86# Tokenize and generate
87input_tokens = tokenizer(formatted_input, return_tensors="pt").to(model.device)
88output = model.generate(**input_tokens, max_new_tokens=256, do_sample=False)
89response = tokenizer.decode(output[0][input_tokens['input_ids'].shape[1]:], skip_special_tokens=True)
90
91print("Response:", response)vllm serve ojus1/Qwen3-0.6B-Instruct --enable-lora --lora-modules prem-research/Funcdex-0.6B=prem-research/Funcdex-0.6B-jira_gmail --enable-auto-tool-choice --tool-call-parser hermes