This model leverages a Stanford Alpaca style instruction tuning dataset, the format is as follows:
###Translate English Text to German:{text} ###Output: {translated_text}
The format is slightly modified to reduce the additional tokens required for the instructions as GPT2 context size is very limited.
The model is trained on small ~5k sample to showcase the impact of instruction tuning on overall alignment of the model towards requested task
Intended uses & limitations
This is only for learning purposes. The model seems to have picked up German vocabulary as well as sentence structures to a good extent but the actual translations are at time grossly incorrect.
The model also attempts at completing the news headlines given as prompt and has a high tendency to hallucinate.
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate: 5e-05
train_batch_size: 16
eval_batch_size: 16
seed: 42
optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08