Model Card: Compressed Sarvam-30b Submission
Model Details
- Model name: Sarvam-30b Compressed Submission
- Base model: Sarvam-30b
- Track: Resilient AI Challenge – Sarvam category
- Task type: Text-to-text
- Inference backend: vLLM or llama.cpp
- Compression type: architecture-preserving inference optimization
Intended Use
This model is intended exclusively for challenge evaluation in the Sarvam-30b category. It is designed to preserve the original functional behavior of the base model while improving inference efficiency.
Training / Adaptation
No architectural redesign was introduced. The submission focuses on compression and inference-time optimization while keeping the original model structure intact.
Evaluation Considerations
The submission was prepared to balance two goals:
- maintain quality relative to the Sarvam-30b baseline;
- reduce energy consumption during inference.
Repository Contents
- model weights
- serving configuration file
- README describing the methodology
Access
This repository is shared with the organizing team for evaluation purposes.