GGUF (quantized for llama.cpp, Ollama, llama-cpp-python)
License
Apache 2.0
Version
2.0
Overview
Neutrino-Instruct is a 7B-parameter language model fine-tuned from the Neutrino base model for conversational and instruction-following tasks. It is distributed in GGUF format for efficient local inference on consumer hardware via llama.cpp, Ollama, and llama-cpp-python.
The model is designed to hold coherent, contextual dialogue across multiple turns and to follow natural-language instructions reliably at chat scale, while remaining light enough to run on a single consumer GPU or CPU-only machine.
Local/offline research prototypes where data cannot leave the device
Educational and hobbyist LLM experimentation
Out of scope
Medical, legal, or financial advice, or any use where model error could cause real-world harm
Autonomous decision-making without human review
Generation of content intended to deceive, impersonate, or manipulate
High-stakes classification (e.g., hiring, credit, law enforcement)
Neutrino-Instruct is a general-purpose research and hobbyist model. It has not been evaluated for production or safety-critical deployment, and Fardeen NB makes no warranty as to its factual accuracy, safety, or fitness for any particular purpose.
Quickstart
llama.cpp
bash
1git clone https://github.com/ggerganov/llama.cpp
2cd llama.cpp &&make34# Single prompt5./main -m ./neutrino-instruct.gguf -p "Hello, who are you?"67# Interactive chat8./main -m ./neutrino-instruct.gguf -i -p "Let's chat."910# Control output length11./main -m ./neutrino-instruct.gguf -n 256 -p "Write a poem about stars."1213# Adjust temperature14./main -m ./neutrino-instruct.gguf --temp 0.7 -p "Explain quantum computing simply."1516# GPU offload (if built with CUDA/Metal)17./main -m ./neutrino-instruct.gguf --gpu-layers 50 -p "Summarize this article."
Ollama
ollama run fardeen0424/neutrino
Python (llama-cpp-python)
python
1from llama_cpp import Llama
23llm = Llama(model_path="./neutrino-instruct.gguf")45response = llm("Who are you?")6print(response["choices"][0]["text"])78# Streaming9for token in llm("Tell me a story about Neutrino:", stream=True):10print(token["choices"][0]["text"], end="", flush=True)
Training Data
Neutrino-Instruct's base model was pretrained on a mixture of web, encyclopedic, code, and long-document text, and instruction-tuned on conversational data. Component sources include:
Dataset
Type
Role
finepdfs
Long-form documents
Pretraining
finewiki
Encyclopedic text
Pretraining
fineweb-edu-100b-shuffle
Educational web text
Pretraining
wikipedia
Encyclopedic text
Pretraining
github-code
Source code
Pretraining
TinyStories
Short narrative text
Fine-tuning / eval
awesome-chatgpt-prompts
Instruction/persona prompts
Instruction tuning
Full dataset links are available in the metadata panel on the right side of this page.
Training procedure
Field
Value
Context length
32,768 tokens
Precision
bfloat16
Architecture
Neutrino-Instruct is built on a standard decoder-only Transformer stack. The core computations are as follows.
Scaled dot-product self-attention, applied per head:
Neutrino-Instruct uses grouped-query attention (GQA): the 32 query heads are partitioned into $g=8$ groups, each group sharing a single key/value projection ($KW^K_g, VW^V_g$) instead of every query head having its own — trading a small amount of expressivity for a much smaller KV cache at inference time.
Position encoding — rotary position embeddings (RoPE), applied to query/key vectors by rotating consecutive coordinate pairs as a function of token position $m$ and frequency $\theta_i$:
Hallucination: Like all LLMs, Neutrino-Instruct can generate plausible-sounding but incorrect or fabricated information. Outputs should be verified before use in any consequential context.
Bias: Training data drawn from web and code sources may encode social, cultural, or representational biases present in that data. The model has not undergone a formal bias audit.
Safety tuning: No dedicated safety alignment (red-teaming, RLHF/DPO refusal training, etc.) has been performed beyond the base instruction tuning. The model may comply with harmful, unsafe, or policy-violating requests.
Language coverage: English only; behavior on other languages is unevaluated and likely degraded.
Context length: Supports up to 32,768 tokens; quality may degrade toward the upper end of that window, as is typical for RoPE-based models — this has not been separately verified for Neutrino-Instruct.
No factual grounding: The model has no live access to external tools, retrieval, or the internet; all knowledge is frozen at training time.
This model has not been evaluated by a third party or independent safety team. Users deploying it in any user-facing product are responsible for their own safety evaluation, content filtering, and monitoring.
License
Released under the Apache License 2.0. See LICENSE for full terms.
Citation
bibtex
1@misc{fardeennb2025neutrino,
2 title = {Neutrino-Instruct: A 7B Instruction-Tuned Conversational Model},
3 author = {Fardeen NB},
4 year = {2025},
5 howpublished = {Hugging Face},
6 url = {https://huggingface.co/neuralcrew/neutrino-instruct}
7}
Acknowledgements
Built on top of the Neutrino base model. Thanks to the maintainers of the open datasets listed above, and to the llama.cpp, Ollama, and llama-cpp-python projects for inference tooling.