The current paradigm for evaluating Large Language Models (LLMs) and AI Agents in financial analysis is constrained by its reliance on static, historical datasets. This approach primarily assesses a model's capacity to interpret past events rather than forecast future outcomes. This methodological misalignment with real-world practice fails to simulate the dynamic, looking-forward environments that analysts and economists face. To address… See the full description on the dataset page:
https://huggingface.co/datasets/OpenFinArena/FinDeepForecast.