This is the 120M parameter model from the Navdyut Foundational Suite, trained by Dicom Pathak at Navdyut AI Labs.
This model was trained from scratch using Maximal Update Parametrization ($\mu$P) and Square Root Batch Sizing.
It was trained exclusively on a synthetic, highly-curated dataset of 14.4 Billion tokens containing Cosmopedia and CodeSearchNet traces.
It is optimized for deterministic structural logic and Python syntax execution.
You can view the full training logs and metrics for this model run here:
Weights & Biases Training Log