This model is a fine-tuned version of
distilbert-base-uncased on the squad dataset.
It achieves the following results on the evaluation set:
Most base model weights were frozen leaving only to finetune the last layer (qa outputs) and 3 last layers of the encoder.