This model is a fine-tuned version of
nlpconnect/roberta-base-squad2-nq on the BioASQ 10b dataset.
Training: BioASQ 10B with SQUAD sampled evenly to match the same samples as BioASQ 10B
Eval: BioASQ 9B Eval with SQUAD Eval sampled evenly to match the same samples as BioASQ 9B Eval
Went from untrained exact match: 60.9% (f1 71.8%) to exact match: 95.2% (96.6% f1) on BioASQ 9B held out training set.
Scores on SQUAD+BioASQ remained stable at exact match: 72.5% (f1 81.4%) to 88.5% (f1 93.3%).