What this dataset tests
Whether a model can score the integrity of a multi-doctor diagnostic processusing dialogue structure, hypothesis competition, and objection handling.
Required outputs
process_integrity_score_0_100
primary_reasoning_strength
primary_reasoning_weakness
Strength labels
evidence_coverage
hypothesis_competition
objection_closure
cross_specialty_synthesis
counterfactual_testing
bias_resistance
uncertainty_tracking
Weakness labels