The training corpus for ProCreations/auto-1b,
a 1B encoder that decides whether an AI agent's next tool call is safe to run — scoring
96.40% on approve-or-deny,
ahead of DeepSeek V4 Flash and within 0.57 points of GPT-5.6-Luna.
Each row is one proposed tool call judged in context: the user's original request, the
agent's recent actions, and the call itself. The label answers "can this run without asking
the human?"
712,000 training examples (355,078 approve / 356… See the full description on the dataset page:
https://huggingface.co/datasets/ProCreations/auto-1b-data.