A curated set of 83 test scenarios for evaluating AI-powered financial agents. Covers portfolio queries, market data lookups, transaction searches, risk analysis, benchmark comparisons, safety guardrails, and multi-turn conversations.
Designed to be framework-agnostic — use it with any LLM agent, not just AgentForge.
This dataset is released under the CC-BY-4.0 license. You are free to use, share, and adapt it for any… See the full description on the dataset page:
https://huggingface.co/datasets/sripathii/agentforge-financial-eval-dataset.