Views
No views yet
| Step | Checkpoint Type | Loss | Perplexity | Char Acc | Word Acc | Improvement vs Pre |
|---|---|---|---|---|---|---|
| Pre | pre_training | 2.6261 | 13.82 | 31.1% | 12.9% | +0.0% |
| 1,000 | checkpoint | 0.9395 | 2.56 | 20.1% | 4.1% | +64.2% |
| 2,000 | checkpoint | 0.8058 | 2.24 | 21.2% | 4.0% | +69.3% |
| 3,000 | checkpoint | 0.7305 | 2.08 | 23.0% | 6.1% | +72.2% |
| 4,000 | checkpoint | 0.6669 | 1.95 | 20.6% | 3.4% | +74.6% |
| 5,000 | checkpoint | 0.5341 | 1.71 | 21.4% | 3.6% | +79.7% |
| 6,000 | checkpoint | 0.4656 | 1.59 | 20.9% | 3.8% | +82.3% |
| 7,000 | checkpoint | 0.3917 | 1.48 | 21.4% | 3.5% | +85.1% |
| 8,000 | checkpoint | 0.3310 | 1.39 | 21.6% | 4.8% | +87.4% |
| 9,000 | checkpoint | 0.2892 | 1.34 | 20.7% | 4.0% | +89.0% |
| 10,000 | checkpoint | 0.2566 | 1.29 | 20.9% | 4.7% | +90.2% |
| 11,000 | checkpoint | 0.2199 | 1.25 | 20.2% | 4.9% | +91.6% |
| 12,000 | checkpoint | 0.2033 | 1.23 | 20.3% | 3.2% | +92.3% |
| 13,000 | checkpoint | 0.1891 | 1.21 | 19.4% | 3.4% | +92.8% |
| Step | Training Loss | Timestamp |
|---|---|---|
| 12,991 | 0.154684 | 2025-08-19T10:40 |
| 12,992 | 0.183019 | 2025-08-19T10:40 |
| 12,993 | 0.157314 | 2025-08-19T10:40 |
| 12,994 | 0.168899 | 2025-08-19T10:40 |
| 12,995 | 0.116096 | 2025-08-19T10:40 |
| 12,996 | 0.122316 | 2025-08-19T10:40 |
| 12,997 | 0.149480 | 2025-08-19T10:40 |
| 12,998 | 0.166267 | 2025-08-19T10:40 |
| 12,999 | 0.152927 | 2025-08-19T10:40 |
| 13,000 | 0.124526 | 2025-08-19T10:40 |


training_curves.png - 4-panel view: Training loss with eval points, Character accuracy, Word accuracy, Perplexityevaluation_comparison.png - 4-panel comparison: Loss, Character accuracy, Word accuracy, Perplexity across all checkpoints1from transformers import AutoModelForCausalLM, AutoTokenizer
2# For vision-language models, use appropriate imports
3
4model = AutoModelForCausalLM.from_pretrained("./model_step_13000")
5tokenizer = AutoTokenizer.from_pretrained("./model_step_13000")
6
7# Your inference code here1{
2 "dataset_name": "johnlockejrr/yiddish_synth_v2",
3 "model_name": "LiquidAI/LFM2-VL-450M",
4 "max_steps": 100000,
5 "eval_steps": 1000,
6 "num_accumulation_steps": 1,
7 "learning_rate": 7e-05,
8 "train_batch_size": 15,
9 "val_batch_size": 1,
10 "train_select_start": 0,
11 "train_select_end": 100000,
12 "val_select_start": 100001,
13 "val_select_end": 105000,
14 "train_field": "train",
15 "val_field": "train",
16 "image_column": "image",
17 "text_column": "text",
18 "user_text": "Please transcribe all the Yiddish text you see in this historical manuscript image. Provide only the transcribed text without any additional commentary or description.",
19 "max_image_size": 250
20}