This repository provides quantized GGUF versions of the OpenReasoning-Nemotron-7B model. These 4-bit and 5-bit quantized variants retain the original model’s strengths in multimodal medical reasoning, while reducing memory and compute requirements—ideal for efficient inference on resource-constrained devices.
This model is intended for developers and researchers who work on competitive math, code and science problems. It has been trained via only supervised fine-tuning to achieve strong scores on benchmarks.
The original Qwen2.5-7B-Instruct model is built on top of the Qwen architecture and Post-trained on OpenReasoning datasets by NVIDIA:
These quantized models are optimized for efficient inference while Maintaining Coding and mathamathics capabilities. Suggested use cases include:
-
Scientific question answering
Scientific Research and mathamatics concepts, coding lessions , etc.
-
Chatbot and assistant prototypes
Build interactive reasoning chat systems with coding capabilities.
-
Research & fine-tuning
Serve as a lightweight base for further task-specific tuning in coding.
-
Low-resource deployment
Run reasoning models on CPUs, edge devices, and lightweight GPUs.
For any inquiries or support, please contact us at
support@sandlogic.com or visit our
Website.