Evaluating tooling capabilities.embedding model: BAAI/bge-m3collections: chunks-v6, limit=10
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-Small-3.1-24B-Instruct-2503
Metrics: generation_time, judge_precision, nb_tool_calls, output_length
model
generation_time… See the full description on the dataset page:
https://huggingface.co/datasets/AgentPublic/evalap-mfs_tooling_v7-50.