A Malay edition of the Berkeley Function-Calling Leaderboard (BFCL) v3 — the standard benchmark for evaluating LLM function/tool calling.
5,011 test entries across 21 categories (simple, multiple, parallel, parallel_multiple, irrelevance, java, javascript, rest, sql, live_, chatable, multi_turn_), with ground-truth answer files unchanged so the official BFCL scoring harness runs unmodified.
What is translated to Malay: