Views
No views yet
├── main.py
├── pyproject.toml
├── .gitignore
├── README.md
├── ROADMAP.md
├── ml/ # Classical Machine Learning (pure NumPy)
│ ├── mlp/ # MLP (MNIST, manual backprop)
│ └── basics/ # 12 standalone models (lin/log reg, SVM, K-Means, PCA, RF, GBDT, etc.)
├── cv/ # Computer Vision
│ ├── simplecnn/ # SimpleCNN (CIFAR-10, Conv×3+Pool×3+FC×2)
│ ├── resnet18/ # ResNet18 (CelebA, 15 attrs, skip connections)
│ ├── resnet34/ # ResNet34 (CelebA, 40 attrs, [3,4,6,3] blocks)
│ ├── resnet50/ # ResNet50 (Bottleneck block 1×1→3×3→1×1)
│ ├── mobilenet/ # MobileNet (depthwise separable conv, CIFAR-10)
│ ├── vit/ # Vision Transformer (patch embed + BERT encoder, CIFAR-10)
│ ├── unet/ # UNet (Oxford-IIIT Pet segmentation)
│ └── yolo/ # YOLO (Pascal VOC object detection)
├── gen/ # Generative Models
│ ├── dcgan/ # DCGAN (CelebA, transposed conv)
│ ├── vae/ # VAE (reparameterization trick, KL divergence)
│ ├── ddpm/ # DDPM (CIFAR-10, denoising diffusion)
│ └── simclr/ # SimCLR (CIFAR-10, contrastive learning)
├── graph/ # Graph Neural Networks
│ └── gcn/ # GCN (Cora, spectral graph convolution)
├── rl/ # Reinforcement Learning
│ └── dqn/ # DQN (CartPole, experience replay)
├── nlp/ # Natural Language Processing
│ ├── bert/ # BERT (MLM pretrain + classification finetune)
│ ├── gpt/ # GPT (decoder-only, causal attention, KV cache)
│ ├── lstm/ # LSTM (hand-written gates, IMDB sentiment)
│ ├── word2vec/ # Word2Vec (CBOW + Skip-gram, negative sampling)
│ ├── seq2seq/ # Seq2Seq Transformer (EN→DE translation)
│ └── lora/ # LoRA (parameter-efficient GPT fine-tuning)
├── utils/ # Shared Infrastructure
│ ├── config.py # YAML config loading/saving
│ ├── seed.py # Reproducibility seed locking
│ └── device.py # CUDA → MPS → CPU auto-detection
└── scripts/ # Notebook generation scripts
│ ├── __init__.py
│ ├── config.yaml # DCGAN hyperparameters
│ ├── model.py # Generator + Discriminator
│ ├── data.py # CelebA images (64×64, no labels)
│ ├── train.py # Adversarial training loop (G/D alternating)
│ └── generate.py # Generate sample grid from trained model
├── vit/
│ ├── __init__.py
│ ├── config.yaml # ViT hyperparameters (patch_size, d_model, n_layers, etc.)
│ ├── model.py # ViT: PatchEmbed → Transformer encoder (reused from BERT) → CLS head
│ ├── data.py # CIFAR-10 via HF datasets
│ ├── train.py # Training loop
│ └── eval.py # Per-class accuracy on test split
├── unet/
│ ├── __init__.py
│ ├── config.yaml # UNet hyperparameters
│ ├── model.py # U-Net: encoder–decoder with skip connections
│ ├── data.py # Oxford-IIIT Pet (image + mask) with augmentation
│ ├── train.py # Training loop (pixel-wise CrossEntropy)
│ └── eval.py # IoU and pixel accuracy
├── cnn/
│ ├── __init__.py
│ ├── data.py # CIFAR-10 via HF datasets (uoft-cs/cifar10)
│ ├── model.py # Plain CNN (Conv×3 + Pool×3 + FC×2)
│ ├── train.py # Training script (Adam + CosineAnnealingLR)
│ └── eval.py # Test evaluation + confusion matrix
├── mlp/
│ ├── __init__.py
│ ├── data.py # MNIST via HF datasets (ylecun/mnist)
│ ├── model.py # MLP — pure NumPy (Linear, ReLU, SoftmaxCrossEntropy, SGD)
│ ├── train.py # Training script
│ └── eval.py # Test evaluation (per-digit accuracy)
├── utils/
│ ├── __init__.py
│ ├── config.py # YAML config loading/saving (load_config / save_config)
│ └── seed.py # set_seed() — lock torch + numpy + random + cudnn
├── nlp/
│ ├── bert/
│ ├── word2vec/
│ ├── lstm/
│ ├── gpt/
│ └── seq2seq/
│ ├── __init__.py
│ ├── tokenizer.py # Word-level tokenizer (5000 vocab, from text8)
│ ├── model.py # Decoder-only Transformer (Causal Attention + KV Cache)
│ ├── train.py # Autoregressive LM on text8
│ └── generate.py # Text generation (temperature + top-k + [SEP] blocked)
│ └── seq2seq/
│ ├── __init__.py
│ ├── config.yaml # Transformer hyperparameters
│ ├── model.py # Encoder (from BERT) + Decoder (cross-attention) → Seq2Seq
│ ├── data.py # Multi30k EN→DE, word-level tokenizer
│ ├── train.py # Teacher forcing training
│ └── generate.py # Greedy decoding translation demo
├── basics/
│ ├── __init__.py
│ ├── logistic_regression.py # Single Linear layer + Softmax (92.3% on MNIST)
│ ├── linear_regression.py # California Housing (Normal Equation + GD, R²=0.583)
│ ├── k_means.py # Unsupervised clustering (pure NumPy)
│ ├── svm.py # SVM — GD (primal) + SMO (dual, Linear/RBF kernels)
│ ├── decision_tree.py # ID3/CART on Iris (ASCII tree, ~93% acc)
│ ├── random_forest.py # Bagging + random feature subsets, Iris 93.3%
│ ├── gbdt.py # Gradient boosting (MSE reg + binary logloss cls)
│ ├── naive_bayes.py # Gaussian NB on MNIST (generative classifier)
│ ├── pca.py # SVD-based dimensionality reduction (MNIST 2D visualisation)
│ ├── knn.py # k-Nearest Neighbors (instance-based, MNIST)
│ └── perceptron.py # Single neuron (Rosenblatt 1958, step activation)
├── .gitattributes # LFS: *.zip *.pt
└── uv.lock| Feature | Description |
|---|---|
| Config system | Each model directory has a config.yaml with its hyperparameters (seed, lr, batch_size, epochs, etc.). Edit the YAML to change training params without touching code. |
| TensorBoard | Every PyTorch training script logs loss/accuracy per epoch to runs/{model_name}/. Run tensorboard --logdir runs to visualize all experiments. |
| Reproducibility | utils/seed.py provides set_seed() that locks torch + numpy + random + cudnn. Called at the start of every train script. Config is saved alongside model weights (_config.yaml). |
1# View training curves (all models)
2tensorboard --logdir runs
3
4# Edit hyperparameters in YAML instead of code
5vim cv/resnet18/config.yaml
6# then train as usual:
7uv run python -m cv.resnet18.train| Item | Value |
|---|---|
| Model | ResNet18 (11.2M params) |
| Dataset | CelebA via HF datasets — 1,000 images |
| Attributes | 15 binary (Smiling, Male, Young, Eyeglasses, etc.) |
| Split | 800 train / 200 val |
| Val Accuracy | 91.2% |
| Training | MPS (Mac M4) + AMP |
| Item | Value |
|---|---|
| Model | ResNet34 (~21M params, [3,4,6,3] BasicBlock) |
| Dataset | CelebA via HF datasets — full 200K |
| Attributes | All 40 binary attributes |
| Optimizer | SGD + Momentum (0.9, weight_decay=1e-4) |
| Training | CosineAnnealingLR + Gradient Accumulation + Early Stopping + Loss Weighting |
| Item | Value |
|---|---|
| Model | ResNet50 (~23.6M params, [3,4,6,3] Bottleneck) |
| Dataset | CelebA via HF datasets — full 200K |
| Attributes | All 40 binary attributes |
| Optimizer | SGD + Momentum (0.9, weight_decay=1e-4) |
| Architecture | Bottleneck block: 1×1 → 3×3 → 1×1 (contrast with BasicBlock's two 3×3) |
| Item | Value |
|---|---|
| Model | Variational Autoencoder (2.6M params) |
| Dataset | CelebA via HF datasets — 10K images (64×64) |
| Architecture | Conv Encoder → μ,logσ² → reparameterize → Deconv Decoder → Sigmoid |
| Loss | Reconstruction (BCE) + KL divergence |
| Training | Adam(lr=2e-4), 50 epoch |
| Item | Value |
|---|---|
| Model | Encoder-Decoder Transformer (1M params) |
| Dataset | Multi30k EN→DE — 29K train / 1K test |
| Architecture | Encoder (from BERT) + Decoder (causal + cross-attention) |
| Training | Teacher forcing, weight-tying, Adam(lr=1e-4) |
| Item | Value |
|---|---|
| Model | Denoising Diffusion (16.1M params) |
| Dataset | CIFAR-10 via HF datasets — 50K images (32×32) |
| Architecture | UNet + timestep embedding + sinusoid positional encoding |
| Training | Noise prediction (MSE), T=1000, linear β schedule |
| Sampling | Reverse diffusion (x_T → x_0), 1000 steps |
| Item | Value |
|---|---|
| Model | 2-layer Graph Convolutional Network (23K params) |
| Dataset | Cora via URL — 2708 nodes, 1433 features, 7 classes |
| Architecture | GraphConv × 2: Â @ H @ W (spectral graph convolution) |
| Training | Semi-supervised (20 labels/class), CrossEntropyLoss |
| Item | Value |
|---|---|
| Model | Deep Q-Network (17K params) |
| Environment | CartPole-v1 via Gymnasium — 4-dim state, 2 actions |
| Architecture | 3-layer MLP (4→128→128→2) |
| Training | Experience replay, target network, ε-greedy decay |
| Item | Value |
|---|---|
| Model | SimCLR (11M params: ResNet18 encoder + MLP projector) |
| Dataset | CIFAR-10 via HF datasets — self-supervised (no labels) |
| Architecture | ResNet18 → Projector(512→256→128) → NT-Xent loss |
| Training | 100 epoch, temperature=0.5, dual random augmentation |
| Item | Value |
|---|---|
| Model | Simplified YOLO (59M params) |
| Item | Value |
|---|---|
| Model | Low-Rank Adaptation on GPT (32K trainable / 5.7M frozen) |
| Dataset | text8 via HF datasets — 5K chunks |
| Architecture | LoRALayer: frozen Linear + low-rank B×A |
| Key concept | Parameter-efficient fine-tuning, 0.58% trainable params |
| Comparison | Full fine-tune: 5.7M vs LoRA r=8: 32K |
| Item | Value |
|---|---|
| Model | MobileNetV1 (135K params, width=1.0) |
| Dataset | CIFAR-10 via HF datasets — 50K train / 10K test |
| Architecture | DepthwiseSeparableConv (depthwise 3×3 + pointwise 1×1) |
| Key concept | Depthwise separable convolution, ~8.4× fewer ops than standard conv |
| Comparison | SimpleCNN 620K params → MobileNet 135K (4.6× smaller) |
| Dataset | Pascal VOC via HF datasets — 20 classes |
| Architecture | CNN backbone → FC detection head → 7×7×30 output |
| Training | YOLO loss (coord + obj + noobj + class), NMS at inference |
| Item | Value |
|---|---|
| Model | Generator (3.5M params) + Discriminator (2.8M params) |
| Dataset | CelebA via HF datasets — 10K images (64×64) |
| Architecture | Transposed conv G / Conv D, BN, LeakyReLU |
| Optimizer | Adam(lr=2e-4, β₁=0.5) — separate for G and D |
| Training | BCELoss, label smoothing, fixed noise grid for monitoring |
| Item | Value |
|---|---|
| Model | Vision Transformer (807K params, 4 layers, 4 heads, 128-dim) |
| Dataset | CIFAR-10 via HF datasets — 50K train / 10K test |
| Architecture | PatchEmbed(4×4) → [CLS] → Transformer Encoder (from BERT) → CLS head |
| Key concept | Self-attention for vision, no convolutions, patch embeddings |
| Item | Value |
|---|---|
| Model | U-Net (31M params, 5 encoder/decoder stages) |
| Dataset | Oxford-IIIT Pet via HF datasets — image + segmentation mask |
| Architecture | Encoder: Conv+MaxPool × 4, Decoder: UpConv+skip × 4, output: pixel-wise logits |
| Loss | CrossEntropy (ignore_index=0 for unlabeled) |
| Metrics | Pixel accuracy, mean IoU |
| Item | Value |
|---|---|
| Model | SimpleCNN (620K params) |
| Dataset | CIFAR-10 via HF datasets — 50K images |
| Classes | 10 (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck) |
| Test Accuracy | 82.4% (30 epochs) |
| Training | Adam + CosineAnnealingLR |
| Item | Value |
|---|---|
| Model | MLP (235K params, pure NumPy) |
| Dataset | MNIST via HF datasets — 60K images |
| Classes | 10 digits (0-9) |
| Test Accuracy | 97.9% (20 epochs) |
| Framework | NumPy only (hand-written backward pass) |
| Item | Value |
|---|---|
| Model | BERT mini (834K params, 4 layers, 4 heads, 128-dim) |
| Pre-training | MLM on text8 (90M chars, HuggingFace) |
| Fine-tuning | Sentiment classification on IMDB (HuggingFace) |
| Test Accuracy | ~50% (character-level; word-level would be higher with subword tokenization) |
| Core components | Self-Attention (semantic aggregation) + MLM (entropy increase noise reduction) |
| Item | Value |
|---|---|
| Model | Word2Vec (50-dim embeddings, 97K vocab) |
| Architectures | CBOW + Skip-gram with Negative Sampling |
| Dataset | text8 via HF datasets (~90M chars) |
| Training | Adam, 5 epochs, k=5 negative samples |
| Evaluation | Cosine similarity search in embedding space |
| Key concept | Static word embeddings from distributional semantics |
| Item | Value |
|---|---|
| Model | LSTM (145K params, hand-written gates) |
| Dataset | IMDB via HuggingFace (9K train / 1K test) |
| Architecture | Embedding(128) → LSTM(128→128) → FC(128→2) |
| Test Accuracy | ~50-60% (character-level, harder than word-level) |
| Key concepts | Input/forget/output gates, cell state, gradient flow through gating |
| Item | Value |
|---|---|
| Model | Decoder-only Transformer (5.7M params, word-level) |
| Dataset | text8 via HuggingFace (15M words, 20K chunks) |
| Training | Autoregressive (predict next token), PPL 4.63 |
| Generation | Temperature + top-k sampling with KV Cache, [SEP] blocked |
| Key concepts | Causal Self-Attention, KV Cache, autoregressive generation, word-level tokenization |
| Algorithm | File | Datasets | Metric |
|---|---|---|---|
| Logistic Regression | ml/basics/logistic_regression.py | MNIST | 92.3% test accuracy |
| Linear Regression | ml/basics/linear_regression.py | California Housing | R²=0.583 |
| K-Means | ml/basics/k_means.py | MNIST | 57.8% cluster purity |
| SVM (GD + SMO) | ml/basics/svm.py | MNIST 3v5 | 93.3% (RBF kernel) |
| Decision Tree | ml/basics/decision_tree.py | Iris | 93.3% test acc |
| Random Forest | ml/basics/random_forest.py | Iris | 93.3% test acc (50 trees) |
| GBDT | ml/basics/gbdt.py | sin(x) / synthetic 2D | MSE 0.21 / 100% bin cls |
| Naive Bayes | ml/basics/naive_bayes.py | MNIST | 53.0% test acc |
| PCA | ml/basics/pca.py | MNIST | 17.3% variance in 2 components |
| k-NN | ml/basics/knn.py | MNIST | ~87% (k=5, 2000 train) |
| Perceptron | ml/basics/perceptron.py | MNIST 0v1 | 100% (linearly separable) |
| Method | Type | Kernel | Notes |
|---|---|---|---|
SVM_GD | Primal GD | Linear only | Fast, robust, ~80 lines |
SVM_SMO | Dual SMO | Linear + RBF | Platt SMO, ~150 lines, supports kernel trick |
| Module | Model | Key concepts |
|---|---|---|
basics/ | Logistic Regression | Linear decision boundary, Softmax, Cross-Entropy, closed-form vs gradient descent |
basics/ | Linear Regression | Normal Equation, MSE, R² score, feature standardisation |
basics/ | K-Means | Unsupervised learning, Euclidean distance, iterative centroid refinement, cluster purity |
basics/ | SVM (GD) | Hinge loss, max-margin classification, L2 regularisation, primal gradient descent |
basics/ | SVM (SMO) | Dual formulation, Lagrange multipliers, KKT conditions, kernel trick (RBF) |
basics/ | Decision Tree | Entropy, Information Gain, recursive partitioning, interpretable ASCII tree |
basics/ | Random Forest | Bagging, bootstrap sampling, feature randomness, ensemble diversity → variance reduction |
basics/ | GBDT | Gradient boosting, stage-wise additive model, pseudo-residuals, learning rate shrinkage |
basics/ | Naive Bayes | Bayes' theorem, generative vs discriminative models, Gaussian likelihood, log-space prediction |
basics/ | PCA | Singular Value Decomposition (SVD), eigenvalue, dimensionality reduction, variance explained |
basics/ | k-NN | Instance-based learning, distance metrics, curse of dimensionality, bias-variance tradeoff |
basics/ | Perceptron | Single neuron, step activation, online learning, Perceptron Convergence Theorem |
mlp/ | MLP (NumPy) | Manual backpropagation, chain rule, gradient descent without autograd, softmax cross-entropy |
cv/simplecnn/ | SimpleCNN | Convolution, max-pooling, BatchNorm, Dropout, CosineAnnealing LR schedule |
cv/resnet18/ | ResNet18 | Residual connections (skip connections), BatchNorm in deep networks, bottleneck design, AMP |
cv/resnet34/ | ResNet34 | SGD+Momentum, CosineAnnealingLR, gradient accumulation, early stopping, ROC AUC, F1 |
cv/resnet50/ | ResNet50 | Bottleneck block (1×1→3×3→1×1), deeper residual networks |
gen/vae/ | VAE | Reparameterization trick, KL divergence, latent space interpolation |
nlp/seq2seq/ | Seq2Seq Transformer | Encoder-decoder, cross-attention, teacher forcing, weight-tying |
gen/ddpm/ | DDPM | Denoising Diffusion, UNet + timestep embedding, noise prediction |
gen/dcgan/ | DCGAN | Transposed convolution, adversarial training, generator/discriminator dynamics |
cv/vit/ | Vision Transformer (ViT) | Patch embedding, self-attention for vision, Transformer without convolutions |
cv/unet/ | U-Net | Encoder-decoder, skip connections, pixel-wise classification, IoU metric |
nlp/bert/ | BERT mini | Self-Attention (semantic aggregation), Masked Language Model (entropy increase + denoising), LayerNorm, positional encoding |
nlp/word2vec/ | Word2Vec | Embedding lookup tables, Negative Sampling, CBOW vs Skip-gram, subsampling frequent words, cosine similarity |
nlp/lstm/ | LSTM | Input/forget/output gates, cell state, gradient flow through gating, sequential processing vs parallel attention |
graph/gcn/ | GCN | Graph convolution, message passing, semi-supervised node classification |
rl/dqn/ | DQN | Q-Learning, experience replay, target network, ε-greedy |
gen/simclr/ | SimCLR | Contrastive learning, NT-Xent loss, data augmentation |
nlp/lora/ | LoRA | Low-rank adaptation, parameter-efficient fine-tuning, GPT adapter |
cv/mobilenet/ | MobileNet | Depthwise separable convolution, efficient CNN, width multiplier |
cv/yolo/ | YOLO | Single-stage object detection, grid-based regression, NMS |
nlp/gpt/ | GPT | Causal Self-Attention, KV Cache, autoregressive generation, word-level tokenizer, temperature + top-k sampling, bad-token blocking |
uv sync1# Train / Evaluate ResNet18
2uv run python -m cv.resnet18.train
3uv run python -m cv.resnet18.eval
4
5# Train / Evaluate ResNet34
6uv run python -m cv.resnet34.train
7uv run python -m cv.resnet34.eval
8
9# Train / Evaluate ResNet50
10uv run python -m cv.resnet50.train
11uv run python -m cv.resnet50.eval
12
13# Train / Generate VAE
14uv run python -m gen.vae.train
15uv run python -m gen.vae.generate
16
17# Train / Translate Seq2Seq
18uv run python -m nlp.seq2seq.train
19uv run python -m nlp.seq2seq.generate
20
21# Train / Evaluate GCN
22uv run python -m graph.gcn.train
23uv run python -m graph.gcn.eval
24
25# Train DQN
26uv run python -m rl.dqn.train
27
28# Train SimCLR
29uv run python -m gen.simclr.train
30
31# Train / Demo YOLO
32uv run python -m cv.yolo.train
33uv run python -m cv.yolo.demo --image my_image.jpg --conf 0.3 --iou 0.5
34
35# Train / Generate LoRA (requires nlp/gpt/gpt_text8.pt)
36uv run python -m nlp.lora.train
37uv run python -m nlp.lora.generate
38
39# Train / Evaluate MobileNet
40uv run python -m cv.mobilenet.train
41uv run python -m cv.mobilenet.eval
42uv run python -m cv.yolo.train
43
44# Train / Generate DDPM
45uv run python -m gen.ddpm.train
46uv run python -m gen.ddpm.generate
47
48# Train / Generate DCGAN
49uv run python -m gen.dcgan.train
50uv run python -m gen.dcgan.generate
51
52# Train / Evaluate / Demo ViT
53uv run python -m cv.vit.train
54uv run python -m cv.vit.eval
55uv run python -m cv.vit.demo --image my_image.jpg
56
57# Train / Evaluate / Demo UNet
58uv run python -m cv.unet.train
59uv run python -m cv.unet.eval
60uv run python -m cv.unet.demo --image my_image.jpg
61
62# Train / Evaluate CNN
63uv run python -m cv.simplecnn.train
64uv run python -m cv.simplecnn.eval
65
66# Train / Evaluate MLP (pure NumPy)
67uv run python -m mlp.train
68uv run python -m mlp.eval
69
70# Basics
71uv run python -m basics.logistic_regression
72uv run python -m basics.k_means
73uv run python -m basics.linear_regression
74uv run python -m basics.svm
75uv run python -m basics.decision_tree
76uv run python -m basics.random_forest
77uv run python -m basics.gbdt
78uv run python -m basics.naive_bayes
79uv run python -m basics.pca
80uv run python -m basics.knn
81uv run python -m basics.perceptron
82
83# NLP
84uv run python -m nlp.bert.pretrain
85uv run python -m nlp.bert.finetune
86uv run python -m nlp.bert.eval
87
88# Word2Vec
89uv run python -m nlp.word2vec.train
90uv run python -m nlp.word2vec.eval
91
92# LSTM
93uv run python -m nlp.lstm.train
94uv run python -m nlp.lstm.eval
95
96# GPT
97uv run python -m nlp.gpt.train
98uv run python -m nlp.gpt.generate.gitignore'ed). Each model saves its weights
locally after training; paths are shown below for reference.| Model | Local path | Size |
|---|---|---|
| ResNet18 (15 attrs, 1K samples) | cv/resnet18/resnet18_celeba.pt | 45 MB |
| ResNet34 (40 attrs, 200K samples) | cv/resnet34/resnet34_celeba.pt | ~80 MB |
| ResNet50 (40 attrs, 200K samples) | cv/resnet50/resnet50_celeba.pt | ~90 MB |
| VAE (CelebA, 64×64) | gen/vae/vae_celeba.pt | 10 MB |
| Seq2Seq Transformer (Multi30k) | nlp/seq2seq/seq2seq_multi30k.pt | 4 MB |
| GCN (Cora) | graph/gcn/gcn_cora.pt | 0.1 MB |
| DQN (CartPole) | rl/dqn/dqn_cartpole.pt | 0.07 MB |
| SimCLR (CIFAR-10) | gen/simclr/simclr_cifar10.pt | 22 MB |
| YOLO (Pascal VOC) | cv/yolo/yolo_voc.pt | 226 MB |
| LoRA (GPT-adapted, text8) | nlp/lora/lora_gpt.pt | 0.2 MB |
| MobileNet (CIFAR-10) | cv/mobilenet/mobilenet_cifar10.pt | 0.5 MB |
| DDPM (CIFAR-10, 32×32) | gen/ddpm/ddpm_cifar10.pt | 62 MB |
| DCGAN (CelebA, 64×64) | gen/dcgan/dcgan_celeba.pt | ~23 MB (G+D) |
| ViT (CIFAR-10, 32×32) | cv/vit/vit_cifar10.pt | 3.2 MB |
| UNet (Oxford-Pet, 128×128) | cv/unet/unet_oxford_pet.pt | 119 MB |
| SimpleCNN (CIFAR-10) | cv/simplecnn/simple_cnn_cifar10.pt | 2.4 MB |
| MLP (MNIST, NumPy) | mlp/mlp_mnist.npz | 0.9 MB |
| Logistic Regression | basics/logistic_regression.npz | 63 KB |
| K-Means centers | basics/kmeans_centers.npz | 32 KB |
| Linear Regression | basics/linear_regression.npz | 2 KB |
| SVM | basics/svm.npz | 45 KB |
| Decision Tree | — | N/A (no weights) |
| Naive Bayes | — | N/A (no weights) |
| PCA | — | N/A (data-dependent) |
| k-NN | — | N/A (no training) |
| Perceptron | — | N/A (no weights) |
| BERT (MLM) | nlp/bert/bert_mlm.pt | 3.2 MB |
| BERT (finetuned) | nlp/bert/bert_finetuned.pt | 3.2 MB |
| Word2Vec (SG) | nlp/word2vec/skipgram.pt | 19 MB |
| Word2Vec (CBOW) | nlp/word2vec/cbow.pt | 19 MB |
| LSTM | nlp/lstm/lstm_sentiment.pt | 0.6 MB |
| GPT | nlp/gpt/gpt_text8.pt | 3.3 MB |