Neon Veil v2 is the second iteration of the finetuned roleplay model built on Google's Gemma 4 E2B Instruct base, quantized to GGUF format for efficient local inference. Building on the foundation of v1, this release delivers significantly improved character consistency, richer prose quality, better long-context retention, and enhanced instruction adherence for immersive, multi-character roleplay.
The model is available in the following GGUF quantizations, ranging from highly efficient (Q2_K_M) to near-lossless (Q8_0).
Quant
File
Size
Recommended VRAM
2-bit
Q2_K_M
~1.4 GB
2 GB
3-bit
Q3_K_M
~1.8 GB
3 GB
4-bit
Q4_K_M
~2.4 GB
4 GB
5-bit
Q5_K_M
~2.8 GB
5 GB
8-bit
Q8_0
~4.3 GB
8 GB
Recommendation: For roleplay, Q4_K_M or Q5_K_M offer the best balance between quality and resource usage. Q8_0 is recommended if you have sufficient VRAM and prioritize output fidelity.
🆕 What's New in v2
Compared to Neon Veil v1, v2 introduces the following improvements:
Improved Character Consistency — Characters maintain their voice, personality, and mannerisms more reliably across extended sessions
Enhanced Prose Quality — Richer descriptive language with better sensory detail and more natural pacing
Better Long-Context Retention — Reduced drift when running at or near the 65,536 token context limit
Refined Instruction Adherence — More faithful compliance with system prompts and formatting directives
Expanded Dialogue Fluency — Smoother multi-turn conversations with better turn-taking and emotional continuity
Tighter NSFW Filter Handling — More nuanced content filtering that preserves creativity while respecting safety boundaries
Download your preferred quant to text-generation-webui/models/neon-veil-v2/
Launch the web UI
Select neon-veil-v2 from the model dropdown
Recommended settings:
Temperature: 0.8
Top P: 0.9
Top K: 40
Repetition Penalty: 1.04
Min P (Mooncake / SillyTavern): 0.03
Using LM Studio
Open LM Studio and navigate to "Download Models"
Search for titan087/neon-veil
Download the v2 quant of your choice
Load and start chatting
✍️ Recommended Sampling Settings
Neon Veil v2 performs best with the following sampling configuration for creative roleplay:
Parameter
Value
Temperature
0.7–0.9
Top P (Nucleus)
0.85–0.9
Top K
40–50
Repetition Penalty
1.02–1.06
Min P
0.02–0.05
Frequency Penalty
0.2–0.3
Presence Penalty
0.2–0.3
Max New Tokens
512–1024
Tip: Lower temperature (0.6–0.7) for more coherent, controlled responses. Higher temperature (0.85–0.95) for creative, unpredictable storytelling.
🎭 Roleplay System Prompt Template
Neon Veil v2 follows the Gemma 4 instruction format. Here is a recommended roleplay system prompt structure:
You are an immersive roleplay assistant. You will portray the characters and world described by the user, maintaining consistent characterization, rich environmental detail, and natural dialogue. Never break character. Never summarize or rush through scenes. Write in vivid, sensory prose with a natural pacing suitable for interactive fiction.
[Character profiles, setting, and scenario details go here]
For best results, provide clear character definitions, setting context, and scene setup in your initial prompt. The model responds well to structured formatting and explicit tone/direction cues.
🧠 Training & Fine-tuning
Neon Veil v2 is fine-tuned from the Gemma 4 E2B Instruct checkpoint using an expanded and refined roleplay dataset. Compared to v1, the v2 training run incorporated:
A larger and more diverse roleplay corpus
Additional data focusing on character voice preservation
Enhanced long-context attention optimization
Improved multi-turn dialogue coherence
Refined NSFW handling with nuanced content filtering
Training emphasizes:
Multi-turn conversational coherence
Character consistency across long contexts
Descriptive prose quality and pacing
Emotional nuance and tone matching
Instruction adherence and format compliance
📦 Hardware Compatibility
Hardware
Minimum Quant
Recommended Quant
Ryzen Zen 3 5000 (128GB)
Q2_K_M
Q8_0
8GB VRAM GPU
Q3_K_M
Q4_K_M
16GB+ VRAM GPU
Q4_K_M
Q8_0
32GB+ Unified Memory
Q5_K_M
Q8_0
The model is optimized to run efficiently on both CPU-only and GPU-accelerated setups. Apple Silicon (M-series) users will experience excellent performance with any quantization level.
⚠️ Limitations
The model may still experience minor character drift during very extended sessions (100K+ tokens)
Complex multi-character dialogue with frequent speaker switches may occasionally lose attribution
NSFW content generation remains limited based on the base model's safety guidelines
Quality may vary for languages other than English
As with v1, this is an iterative release; future versions will continue to address remaining limitations
📜 License
This model is released under the GPL-3.0 License. The underlying Gemma 4 base model is subject to Google's licensing terms. Please review both licenses before use.