CFIR v0.1 — Coupled Failure Intervention Reasoning Benchmark
What this repo does
CFIR v0.1 is a synthetic benchmark designed to evaluate whether models can reason about stability and intervention geometry in coupled systems.
Many real-world failures occur not because systems lack information, but because they fail to interpret interacting pressures, buffers, delays, and couplings correctly.
This dataset tests whether a model can determine when an intervention will stabilize or fail to… See the full description on the dataset page:
https://huggingface.co/datasets/ClarusC64/cfir-stability-intervention-geometry-v0.1.