This is the merged full-weight model of L3.3-70B-PippaMaid-2.0: a DPO-aligned LoRA adapter merged into its base model L3.3-70B-PippaMaid-1.0. Ready for direct inference or GGUF quantization. No adapter loading required.
PippaMaid 2.0 uses Anthropic's Constitutional AI methodology to improve prose quality in uncensored RP models. This is not a safety alignment project. This is a writing quality project.
The Problem
RP finetunes converge on the same failure modes because they train on AI-contaminated data:
Em-dash abuse: Every sentence connected with dashes instead of periods
Italic overload: Emphasis on every other word, diluting actual emphasis
Synonym dumping: "aching, needing, craving, wanting" instead of picking one good word
No dynamic range: Every paragraph at maximum intensity, no tension/release
Abstract sensation over concrete detail: "heat pooling" instead of specific physical grounding
Cliche phrase recycling: The same AI-default phrases every time ("barely above a whisper", "swallowed thickly", "pupils blown wide", etc.)
The Method
The Constitutional AI pipeline:
Phase 1 (Generation): Generated ~172K raw outputs from Maginum-Cydoms-24B across diverse RP scenarios
Phase 2 (SL-CAI): A critic model (GLM-4.7 via Fireworks) critiqued and revised each output against a prose quality constitution. The revised outputs became SFT training data for PippaMaid 1.0
Phase 3 (RLAIF): Generated preference pairs from the Phase 2 model, scored by the critic against the same constitution
Phase 4 (DPO): Trained a LoRA adapter on the preference pairs to push the model's distribution toward constitutional prose, then merged into the base
The Constitution
The full constitution defines 7 sections of prose quality principles. Key rules:
Structure: Narration in plain text (no asterisk actions), dialogue in quotation marks, internal monologue in asterisks (sparingly). Consistent narration perspective. Respect user character autonomy.
Formatting bans: Zero em-dashes. No italic emphasis in narration. No ellipsis spam. No single-sentence paragraphs.
Anti-slop: 50+ banned AI cliche phrases. No synonym chains. No sensation stacking. No purple dialogue tags ("breathed", "husked", "purred"). Use "said"/"asked" or action beats.
Prose quality: Concrete physical detail over abstract sensation. Sentence length variation. Dynamic intensity (not every paragraph at 10/10). Spatial grounding. Show emotions through behavior.
Multi-turn: Narrative continuity across turns. Varied intensity across conversation arcs. Response length proportionate to input. Natural callbacks to earlier details.