I wanted to train a small agent to use a browser effectively, most smaller models I tried <32b struggled to call the tools correctly.
I created this dataset for two main reasons:
To help with finetuning smaller models to use the browser specific tools in playwright.
To look at the security implications of giving browser access to untrusted open-weight models, see blog post.
I am ironing out the kinks, but I will leave the older versions here in… See the full description on the dataset page:
https://huggingface.co/datasets/jdaddyalbs/playwright-mcp-toolcalling.