A micro-benchmark for sequential boundary integrity in LLMs.
Most evals test a single prompt.BCT tests the slope of collapse — how models degrade over turns.
This dataset measures:
• boundary acknowledgement• policy memory across conversation• resistance to relational pressure• consistency under emotional or identity bait• collapse shape (early-warning → collapse_event → recovery)
Columns: