factorjepa-lpft-vjepa21-vitG-2B-poc
FactorJEPA — V-JEPA 2.1 ViT-G (2B) adapted on Indian-context urban driving / walking / monument clips.
This is the lpft arm of the iter18 FactorJEPA ablation: LP-FT baseline: linear-probe warmup then full fine-tune (Kumar et al. ICLR'22) (B2).
The study compares factor-surgery against strong fine-tuning baselines on the claim
vjepa_surgery >> vjepa_pretrain >> vjepa_frozen for motion / temporal features on Indian urban video.
Every non-pretrain arm initializes from the SAME continual-pretrain checkpoint (fair duel — identical
data, identical starting weights).
🎯 Training summary
| Field | Value |
|---|
| Base model | facebook/v-jepa-2-vitg |
| Stage | lpft |
| Architecture | V-JEPA 2.1 ViT-G (~2B params, 1664-dim, 48 layers, hierarchical concat 6656-dim) |
| Training data | Indian-context urban clips (10k POC pool, leakage-safe train/val/test split) |
| Steps | 481 |
| Batch size | 32 |
| Training stages | stage0_head_only → lpft_full |
| Factor/train clips | 7,021 |
| Train/val split | 7,021 / 451 |
| Final loss | 0.45375 |
| KEPT ckpt (selector) | step 43 · future_l1=0.5307 · top1=0.2705 · motion_cos=0.1395 |
| Early stop | not triggered |
📈 Training trajectory (initial → final, from probe_history.jsonl)
| Metric | Description | Initial | Final | Δ |
|---|
probe_top1 | motion-flow 16-class probe top-1 | 0.271 | 0.333 | +23.0% 📈 |
motion_cos | intra-vs-inter motion cosine | 0.1395 | 0.2860 | +105.0% 📈 |
val_jepa_loss | validation JEPA loss (L1) | 0.4836 | 0.4510 | -6.7% 📉 |
future_l1 | future-frame L1 (per clip) | 0.5307 | 0.5418 | +2.1% 📈 |
(5 checkpoints across training.)
🧪 Held-out test evaluation (N=1825 clips · 95% BCa bootstrap CI)
| Metric | Value | 95% CI (±) | better |
|---|
| action top-1 | 0.4904 | 0.0233 | ↑ |
| taxonomy F1 | 0.7888 | — | ↑ |
| motion-cos margin | 0.1192 | 0.0044 | ↑ |
| future-frame MSE | 0.5362 | 0.0009 | ↓ |
| rollout drift | 0.0096 | 0.0003 | ↓ |
| causal L1 | 0.5615 | 0.0014 | ↓ |
| t-dist error | 0.0080 | 0.0004 | ↓ |
| mask-ratio slope | 0.0609 | 0.0006 | ↓ |
| order sensitivity | 0.0172 | 0.0013 | · |
| teacher-free drift | 0.0477 | 0.0007 | ↓ |
Direction: ↑ higher better · ↓ lower better · · signed diagnostic. — = not computed for this arm.
⚡ Quick start — self-contained, no other code needed
This repo ships everything: the weights, the architecture (vjepa2_src/, vendored Meta V-JEPA 2 source, MIT), and a loader. Download it and run — no private package, no separate clone.
1huggingface-cli download anonymousML123/factorjepa-lpft-vjepa21-vitG-2B-poc --local-dir factorjepa-model
2cd factorjepa-model && pip install -r requirements.txt
3python load_factorjepa.py --encoder student_encoder.pt # builds 2B ViT-G, loads, forwards (no video needed)
1from load_factorjepa import load_encoder, preprocess_frames, extract_features
2encoder = load_encoder("student_encoder.pt", device="cuda") # bf16 on cuda, fp32 on cpu
3clip = preprocess_frames(frames_uint8)[None] # (T,H,W,3) uint8 -> (1, 16, 3, 384, 384)
4feats = extract_features(encoder, clip) # (1, 4608, 1664) token features
NATIVE V-JEPA 2.1 ViT weights — NOT transformers.VJEPA2Model (AutoModel.from_pretrained fails: different
keys + no 2.1 deep-supervision head). No xformers (SDPA attention). student_encoder.pt is encoder-only —
for an actual next-frame prediction heatmap also load the predictor from m09c_ckpt_best.pt (key predictor):
from load_factorjepa import load_predictor; predictor = load_predictor("m09c_ckpt_best.pt", device="cuda").
🏗️ Architecture
| |
|---|
| Encoder | V-JEPA 2.1 ViT-G — embed_dim=1664, depth=48, num_heads=26, RoPE, 2B params (1.84B exact) |
| Input | (B, 3, T=16, 384, 384) — 16 frames, 384² center-crop, ImageNet-normalized; patch 16, tubelet 2 |
| Tokens | 8 × 24 × 24 = 4608 tokens × 1664-dim (final layer); deep-supervision concat = 4608 × 6656 |
| Predictor | 2.1 predictor — predictor_embed_dim=384, depth=24, num_heads=12, dense-loss (return_all_tokens) |
| Attention | scaled_dot_product_attention (SDPA) — no xformers |
The exact constructor kwargs are in load_factorjepa.py (verified against the eval pipeline that produced these
weights). student_encoder.pt wraps the weights under the key student_state_dict — the loader unwraps it, strips
module./backbone. prefixes, and asserts ≥90% of params load (fail-loud).
📦 Files in this repo
| File | Size | Purpose |
|---|
student_encoder.pt | ~7 GB | Inference-ready ViT-G encoder weights (key student_state_dict) — load this for features |
m09c_ckpt_best.pt | ~8-14 GB | Best-selected ckpt incl. predictor (key predictor) — for next-frame / JEPA prediction |
load_factorjepa.py | ~8 KB | Self-contained loader — build model + load weights + preprocess + forward |
vjepa2_src/ | ~100 KB | Vendored V-JEPA 2 architecture (Meta, MIT) — the encoder/predictor classes |
requirements.txt | <1 KB | Pinned deps (exact versions that load these weights; no xformers) |
motion_aux_head.pt | ~2 MB | Motion auxiliary head (paired with student_encoder) |
training_summary.json | ~2 KB | Final-step metrics |
probe_history.jsonl | ~few KB/step | Per-checkpoint probe + drift metrics |
loss_log.{jsonl,csv} | ~several KB | Per-step JEPA loss trajectory |
*.png / *.pdf | ~few MB | Training trajectory plots (loss, drift, probe trio) |
🧪 Reproducibility
This checkpoint was produced by:
1CACHE_POLICY_ALL=2 ./scripts/run_train.sh lpft --FULL \
2 2>&1 | tee logs/lpft_full.log
Pipeline source: iter/iter14_surgery_on_pretrain/plan_HIGH_LEVEL.md
📜 Attribution & license — the links below are provenance/credit only, NOT a setup step
✅ 100% self-contained. Everything needed to load this model is already in THIS repo
(vjepa2_src/ + load_factorjepa.py + the weights). You do not need to visit, clone, pip install,
or download anything from the two links below — they are license/credit only. Loading touches no other repo.
- Adapted weights (
student_encoder.pt, m09*_ckpt_best.pt, motion_aux_head.pt) — Apache-2.0 (this repo).
Derived from facebook/v-jepa-2-vitg (provenance only — not needed to load).
- Vendored architecture (
vjepa2_src/) — Meta Platforms' V-JEPA 2, MIT, copied unmodified from
github.com/facebookresearch/vjepa2 @ 204698b (credit only — the code is already in vjepa2_src/; its MIT
license is at vjepa2_src/LICENSE). © Meta Platforms, Inc. and affiliates.
📝 Citation
1@misc{factorjepa2026,
2 title = {FactorJEPA: Factor-disentangled SSL for Indian-context urban video},
3 author = {Wanaskar, Kapil and others},
4 year = {2026},
5 note = {HF model card auto-generated by src/utils/hf_finetuned_push.py}
6}
Model card auto-generated by src/utils/hf_finetuned_push.py at 2026-06-12T11:11:32+00:00.