Neon Veil is a finetuned roleplay model built on Google's Gemma 4 E2B Instruct base, quantized to GGUF format for efficient local inference. It is designed to deliver immersive, multi-character roleplay with strong prose, coherent long-form storytelling, and a rich understanding of pacing, tension, and character drift.
The model is available in the following GGUF quantizations, ranging from highly efficient (Q2_K_M) to near-lossless (Q8_0).
Quant
File
Size
Recommended VRAM
2-bit
Q2_K_M
~1.4 GB
2 GB
3-bit
Q3_K_M
~1.8 GB
3 GB
4-bit
Q4_K_M
~2.4 GB
4 GB
5-bit
Q5_K_M
~2.8 GB
5 GB
8-bit
Q8_0
~4.3 GB
8 GB
Recommendation: For roleplay, Q4_K_M or Q5_K_M offer the best balance between quality and resource usage. Q8_0 is recommended if you have sufficient VRAM and prioritize output fidelity.
Download your preferred quant to text-generation-webui/models/neon-veil/
Launch the web UI
Select neon-veil from the model dropdown
Recommended settings:
Temperature: 0.8
Top P: 0.9
Top K: 40
Repetition Penalty: 1.05
Min P (Mooncake / SillyTavern): 0.03
Using LM Studio
Open LM Studio and navigate to "Download Models"
Search for titan087/neon-veil
Download your preferred quant
Load and start chatting
✍️ Recommended Sampling Settings
Neon Veil performs best with the following sampling configuration for creative roleplay:
Parameter
Value
Temperature
0.7–0.9
Top P (Nucleus)
0.85–0.9
Top K
40–50
Repetition Penalty
1.03–1.07
Min P
0.02–0.05
Frequency Penalty
0.3
Presence Penalty
0.3
Max New Tokens
512–1024
Tip: Lower temperature (0.6–0.7) for more coherent, controlled responses. Higher temperature (0.85–0.95) for creative, unpredictable storytelling.
🎭 Roleplay System Prompt Template
Neon Veil follows the Gemma 4 instruction format. Here is a recommended roleplay system prompt structure:
You are an immersive roleplay assistant. You will portray the characters and world described by the user, maintaining consistent characterization, rich environmental detail, and natural dialogue. Never break character. Never summarize or rush through scenes. Write in vivid, sensory prose with a natural pacing suitable for interactive fiction.
[Character profiles, setting, and scenario details go here]
For best results, provide clear character definitions, setting context, and scene setup in your initial prompt. The model responds well to structured formatting and explicit tone/direction cues.
🧠 Training & Fine-tuning
Neon Veil is fine-tuned from the Gemma 4 E2B Instruct checkpoint using a curated roleplay dataset. The training emphasizes:
Multi-turn conversational coherence
Character consistency across long contexts
Descriptive prose quality and pacing
Emotional nuance and tone matching
Instruction adherence and format compliance
📦 Hardware Compatibility
Hardware
Minimum Quant
Recommended Quant
Ryzen Zen 3 5000 (128GB)
Q2_K_M
Q8_0
8GB VRAM GPU
Q3_K_M
Q4_K_M
16GB+ VRAM GPU
Q4_K_M
Q8_0
32GB+ Unified Memory
Q5_K_M
Q8_0
The model is optimized to run efficiently on both CPU-only and GPU-accelerated setups. Apple Silicon (M-series) users will experience excellent performance with any quantization level.
⚠️ Limitations
The model may occasionally drift from character during very long roleplay sessions
Complex multi-character dialogue can sometimes lose speaker attribution
NSFW content generation is limited based on the base model's safety guidelines
Quality may vary for languages other than English
This is a first-iteration release; future versions will address current limitations
📜 License
This model is released under the GPL-3.0 License. The underlying Gemma 4 base model is subject to Google's licensing terms. Please review both licenses before use.