The HealthHistoryRoBERTa-pt model was pre-trained from scratch using the RoBERTa structure and with patient data from health insurances organized in the form of historical sentences. The initial objective of the training was to predict hospitalizations, however, due to the possibility of applications in other tasks, we made these models available to the scientific community. This model was trained in Portuguese, however, there is an
English version trained from translated data. There are also other training approaches that can be seen at:
The model was pre-trained from 837,159 historical sentences from health insurance patients generated using the approach described in this paper
Predicting Hospitalization from Health Insurance Data.
For this model, our own tokenizer was generated, trained from aggregated data from health plans, see about the aggregated data in
Predicting Hospitalization from Health Insurance Data. Therefore, the tokenizer used is closely related to health insurance data.
The model was trained on a GeForce NVIDIA RTX A5000 24GB GPU from laboratories of
IT departament at UFPR (Federal University of Paraná). The model parameters were initialized randomly as this training is performed without any previous pre-trained model.
We use a batch size of 16, a maximum sequence length of 512, accumulation steps of 4, masked language model probability = 0.15, number of epochs = 2 and a learning rate of 10−4 to pre-train this model.
The training time was 3 hours 04 minutes and 30 seconds per epoch.
1from transformers import AutoTokenizer, AutoModel
2tokenizer = AutoTokenizer.from_pretrained("efbaro/HealthHistoryRoBERTa-pt")
3model = AutoModel.from_pretrained("efbaro/HealthHistoryRoBERTa-pt")