What this repo is for
Detect scheduler pressure before customer-visible incidents.
Focus
• queue depth vs wait time
• utilization vs SLO breaches
• preemption vs job failure
• fragmentation where capacity exists but cannot be used
Why it matters
Operators often see queue problems too late.
This dataset spots the drift earlier and links it to real failure signals.