Views
No views yet
cuda2llm. The code path uses hand-written CUDA kernels for the transformer forward pass, backward pass, AdamW update, checkpoint loading, and text generation.prepare_data, grad_check, train_gpt, and generate_gpt all run on an NVIDIA RTX A4000.1step 1000 | train loss 1.6028 | val loss 1.7558
2step 2000 | train loss 1.2770 | val loss 1.5911
3step 3000 | train loss 1.1062 | val loss 1.6381
4step 4000 | train loss 0.7658 | val loss 1.8825
5step 5000 | train loss 0.7931 | val loss 2.04831ROMEO:
2Ay not sustom's master and by earor-woo's love;
3She damned life, lest have had some removed,
4Lest in neither slander man, and might have
5Because the moon his commission.1model.bin
2vocab.bin
3train.bin
4val.bin
5full_run.log
6checkpoints/model-interim-20260706-011839.bin
7checkpoints/model-interim-20260706-020538.bin
8samples/interim_sample_20260706-011839.txt
9samples/interim_sample_20260706-020538.txt175eb8311e2bdec22b0c0e40084fb39fb1ed190662d3cd51f26e9f4b37ee06691 model.bin
2464eb51d66b5b643889064b82c75607e2d756df11ec960e5e6aaaf03c71734be vocab.bin
3a3f591a5c0dfa7f54a2ec0ab39c4561d6475f1af28264f38076932ca425d5037 train.bin
4836220b36a012c19cd285816001700117edef4f8e1098130b234d4b8fa766d86 val.bin