1 "prompt": "1girl, solo, character: lucy, white background, cowboy shot, hip cutout, sad expressions, pantyhose, best quality, beautiful color, by bigroll"
2},
3{
4 "prompt": "(detailed features, flat color:1.2), (lineart, flat-pasto:0.3),The red giant panda holds its bamboo and falls asleep on the water, best quality, "
5},
6{
7 "prompt": "gufeng, 2.5d, impasto, old withered vines and ancient trees, a murder of crows descending into darkness, a small bridge over gently flowing water, quaint houses nestled by the river, best quality, masterpiece, highres, "
8},
9{
10 "prompt": "aesthetic, detailed, beautiful color, amazing quality, best quality, high quality, cinematic quality, LUT, fine texture, crisp detail, by rella, by_dkxlek, by makoto shinkai, ayanami_rei, beautiful, parted lips, detailed face, detailed eyes, Large eyes, deep crimson eyes, introspective eyes, eyes designed by Ilya Kuvshinov style, (detailed long hair), blue hair, white hair, (holographic hair), multicolored hair, glow hair, transparent hair, floating hair, wind effect, sad expressions, lonely, melancholy atmosphere, subtle grace, elegance, ethereal beauty, transparent, clarity, from side, looking at viewer, off shoulder, bare shoulders, splashing collarbone, strapless, white dress, wet dress, standing in lake, Calm Lake, water, White Lily, petals, barefoot, ripples, Misty Forest, soft shadows, (clear blue sky), cloud, cloudy sky, starry sky, particle, feather, [moon], (high-resolution sci-fi landscape:1.3), post-human landscapes, thousands of years in the future, Earth, overgrown cities, reclaimed by nature, futuristic decay, ancient modern buildings, high-resolution ruins, (rust in ruins, collapsed building), [2.5d:celluloid:0.6], Depth of Field, sharp focus, bokeh,"
11},
12{
13 "prompt": "(from side:1.0), (Red long hair:1.0), (Hair adorned with fine pearls:1.0), (sacred:1.0), (Blue light gauze dress:1.0), (🌊:1.1), (Play the transparent harp:1.2), (Fragmented light dots:1.0), (Dawn:1.0), (fluorescence:0.8), (Crystal and transparent:1.0), (Flowing light and overflowing colors:1.0), (sit:1.0), (Close range:1.0), , (by wlop:0.6), (impasto, oil painting:1.2), detailed features, pseudo-impasto, best quality, (detailed features, flat color:1.2), (lineart, flat-pasto:0.3)"
14}]```
15
16Negative prompts: (worst quality:1.3), low quality, lowres, messy, abstract, ugly, disfigured, bad anatomy, draft, deformed hands, fused fingers, signature, text, multi views
17
18Sampler: Eular a normal as default, 28+ steps recommended.
19
20One additional merit of Neta Art XL is that it supports a very wide range of CFGs (5 - 20 compared to 7 - 9 of previous models). While we empirically found higher CFG leads to more details and higher contrast, generally CFG 9 - 14 (important!) can be used for best results.
21
22# II. Highlight: Style Versatility
23
24We carefully selected 13 style keys with good orthogonality and are commonly used in many scenarios, justified by usage data from Nieta AI (30M+ generations).
25Having orthogonal styles means each style is effectively different from the others, allowing you to easily combine and create new styles without interference.
26impasto
27
28Please refer to https://civitai.com/models/124189/anime-illust-diffusion-xl for a complete list of supported artists.
29
30# III. Expression, Posing, and Camera Angles
31
32# IV. Multi-Character Scenes
33
34# V. Text & Typography
35
36# VI. Training
37- Data annotation combining multiple sources (Original prompt, CogVLM captions, WaifuTagger tags)
38- Post-processing techniques like semantic deduplication and hierarchical tag organization
39 1. Semantic Deduplication: This removed redundant tags by intelligently detecting when a higher-level tag (e.g. very long hair) semantically covered a lower-level one (e.g. long hair).
40 2. Tag Layering Algorithm: Tags were organized into hierarchical layers based on their priorities and related semantics (eg. by wlop influence the whole picture styling, while frills influence a small fraction). More dominant tags were placed in higher layers to prioritize their influence during training.
41
42- Using high-quality regularization data from AIDXL: High-quality regular datasets with "best" and "amazing" quality ratings from AIDXL. These datasets are manually selected and come with detailed annotations and natural language descriptions.
43- Finetuning on more knowledgeable base models like AAM, blending with AnimagineXL 3.1 Character Knowledge.
44
45### Challenges Faced:
46- Imbalance in learning different styles
47- Poor generalization for some styles to diverse scenes
48- Lack of details/texture in generations
49- Trigger word overlap with base model knowledge
50
51### Solutions Explored:
52- Data reweighting to balance style learning, and supplement diverse data per style.
53- Tuning sampling hyperparameters like minimum gamma and rectified flow. Rectified Flow is a training parameter that increases the sampling frequency in the middle time steps but weakens the weight of the model's learning ability for small noises in the low time steps. This technique helps to improve the model's ability to restore styles but requires the use of a knowledge-rich base model.
54- Randomizing / drop off trigger words during training.
55
56
57# VII. Evaluation
58Nine models are evaluated using 16 styles (some unique to Neta Art, thus biased :P) and 80 prompts. Each prompt generates 3 samples with different aspect ratios, resulting in an XYZ plot (Generated from https://github.com/talesofai/comfyui-browser).
59
60XYZPlot of 3,840 samples, just for one-time evaluation. Here's an online Example.
61
62Evaluation is done on a 10-point scale across four axes under objective counting rules:
631. Prompt Following: Points deducted for every 5 images unrelated to the prompt.
642. Stability: Points deducted for distortions or breakages in heads, hands, feet, or composition.
653. Diversity: Points deducted for poor style restoration within each style column.
664. Generalizability: Points deducted for poor semantic generalization within each style column.
67
68Additionally, 10 top model trainers provide subjective scores for Aesthetics, with an average calculated for each model.
69
70
71# VIII. License
72
73- Developed with ❤️ by: Neta.art Lab - https://civitai.com/user/nieta_art
74- In collaboration with:
75 - Euge: https://civitai.com/user/Euge_
76 - 汤人烂: https://space.bilibili.com/8594480
77 - Chenkin: https://civitai.com/user/Chenkin
78 - Bo Dai: https://daibo.info/
79- Thanks to:
80 - https://blog.novelai.net/introducing-novelai-diffusion-anime-v3-6d00d1c118c3
81 - https://cagliostrolab.net/posts/animagine-xl-v3-release
82 - https://civitai.com/models/269232/aam-xl-anime-mix
83 - https://civitai.com/models/124189/anime-illust-diffusion-xl
84 - https://github.com/deepghs/waifuc
85 -
86- Model type: Diffusion-based text-to-image generative model
87- License: We merged 0.05 CLIP and 0.15 UNet input layers from Animagine 3.1, thus Fair AI Public License 1.0-SD
88
89# IX. Conclusion and Future Work
90Shortcomings:
911. Some characters are underfitted.
922. Styles are not activated well with long prompts.
933. Certain styles appear grayish at low CFG and short prompts. Partly explained in https://civitai.com/articles/4969.
94Future Work:
95- Prepare larger training sets and more knowledge-based data to improve character, style, and detail handling.
96- Welcome others to join discussions, provide suggestions, and contribute to model advancement.
97
98_Neta Art XL 2.0 is on the way._
99
100Stay tuned with us, and test our product for FREE: http://neta.art/
101
102Discord: https://discord.gg/AtRtbe9W8w
103
104Twitter: https://twitter.com/netaart_ai
105
106Civitai:https://civitai.com/user/neta_art