ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability
across three production-style task types:
json: schema-constrained structured extraction
sql: text-to-SQL validated by SQLite execution
codestub: Python function generation validated by unit tests
This dataset is designed for verifier-based evaluation: outputs must
work, not just look plausible.
reliability_v1_60.jsonl
Canonical split with 60 tasks: