## Dataset Location: ufldl-stanford/svhn cropped_digits
## Dummy Dataset: SUN397
## Dummy Dataset Location: tanganke/sun397
## Loss Term: 1e-0
## Merge Method: arithmetic
## Test-Set Accuracy: 0.9731529355049133
## Test-Set Loss: 0.11128436164976141
## Training Hyperparameters:
- Seed: 599121577
- Device: cuda
- Machine: L40S
- Number of Devices: 1
- Precision: Float32
- Gradient Norm Clip Value: 10.0
- Accumulate Gradient Batches: 1
- Batch Size: 32
- Loss: Cross Entropy
- Epochs: 4
- Optimizer: AdamW
- Learning Rate: 1e-5
- Weight Decay: 0.1
- Scheduler: CosineAnnealingLR
- Train Split Percent: 90%
- Validation Split Percent: 10%
- Test Split Percent: 100%
## Method:
1. Compute standard loss on target dataset using current model.
2. Merge current model with fine-tuned dummy dataset model through arithmetic, with a coefficient of 0.5.
3. Compute dummy loss on dummy dataset using merged encoder.
4. Multiply the dummy loss by the (1e-0) and add it to the standard loss.
5. Backpropagate based on new loss.