Views
No views yet
1docker run --gpus all \
2 --shm-size 32g \
3 --ipc=host \
4 -e SGLANG_OVERLAP_PLAN_STREAM=1 \
5 -p 8000:8000 \
6 -v $HF_HOME:/root/.cache/huggingface \
7 dogacel/sglang:v0.5.14-dspark-cu130 \
8 python3 -m sglang.launch_server \
9 --model-path fal/Qwen3.6-35B-A3B-Magic-Prompt-FP8 \
10 --served-model-name qwen36-35b-magic-fp8 \
11 --trust-remote-code \
12 --speculative-algorithm DSPARK \
13 --speculative-draft-model-path fal/Qwen3.6-35B-A3B-Magic-Prompt-FP8-DSpark \
14 --speculative-dflash-block-size 8 \
15 --linear-attn-prefill-backend flashinfer \
16 --linear-attn-decode-backend flashinfer \
17 --mamba-radix-cache-strategy extra_buffer \
18 --cuda-graph-backend-prefill tc_piecewise \
19 --tp-size 1 \
20 --max-running-requests 32 \
21 --cuda-graph-max-bs-decode 32 \
22 --weight-loader-disable-mmap \
23 --host 0.0.0.0 \
24 --port 8000 \
25 --speculative-draft-attention-backend fa4 \
26 --attention-backend trtllm_mha \
27 --mamba-ssm-dtype bfloat16 \
28 --enable-flashinfer-allreduce-fusion \
29 --mem-fraction-static 0.61docker run --gpus all \
2 --shm-size 32g \
3 --ipc=host \
4 -e SGLANG_OVERLAP_PLAN_STREAM=1 \
5 -p 8000:8000 \
6 -v $HF_HOME:/root/.cache/huggingface \
7 dogacel/sglang:v0.5.14-dspark-cu130 \
8 python3 -m sglang.launch_server \
9 --model-path fal/Qwen3.6-35B-A3B-Magic-Prompt-FP8 \
10 --served-model-name qwen36-35b-magic-fp8 \
11 --trust-remote-code \
12 --speculative-algorithm DSPARK \
13 --speculative-draft-model-path fal/Qwen3.6-35B-A3B-Magic-Prompt-FP8-DSpark \
14 --speculative-dflash-block-size 8 \
15 --linear-attn-prefill-backend flashinfer \
16 --linear-attn-decode-backend flashinfer \
17 --mamba-radix-cache-strategy extra_buffer \
18 --cuda-graph-backend-prefill tc_piecewise \
19 --tp-size 1 \
20 --max-running-requests 32 \
21 --cuda-graph-max-bs-decode 32 \
22 --weight-loader-disable-mmap \
23 --host 0.0.0.0 \
24 --port 8000 \
25 --speculative-draft-attention-backend fa3 \
26 --attention-backend fa3 \
27 --mem-fraction-static 0.81import json
2from openai import OpenAI
3
4client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
5
6SYSTEM_PROMPT = """\
7You convert a natural-language user idea into a detailed structured JSON caption for an image renderer. Emit one JSON object only.
8
9OUTPUT CONTRACT
10
11The first character must be { and the last character must be }.
12Emit exactly this final stored schema:
13{"high_level_description":"...","compositional_deconstruction":{"background":"...","elements":[...]}}
14
15Never output aspect_ratio. The user provides aspect ratio only to guide composition.
16Never output style_description, bbox, markdown, comments, explanations, or any extra top-level keys.
17Never output top-level background or top-level elements; they belong only inside compositional_deconstruction.
18Use valid JSON escaping. Do not put raw newline characters inside strings. If visible text needs line breaks, use the two-character escape sequence \\n inside the text field.
19
20Each element must use one of these exact schemas:
21{"type":"obj","desc":"..."}
22{"type":"text","text":"...","desc":"..."}
23
24No other element keys are allowed. Do not use obj, object, style, role, position, font, color, children, notes, label, or bbox as keys.
25
26TARGET STYLE
27
28Match Ideogram V4 reference magic prompts: concrete, medium-aware, visually specific, and directly renderable. Do not write a short summary. Expand underspecified ideas into a plausible finished image while preserving the user's intent.
29
30Typical target density:
31- Ordinary prompt: 7-14 elements and roughly 2500-4200 JSON characters.
32- Poster, packaging, infographic, UI, menu, map, complex scene: 12-24 elements when appropriate.
33- A one-element answer is wrong unless the user explicitly asks for a lone isolated object on a blank or transparent background.
34
35Avoid generic art-direction drift. Do not automatically add cyberpunk, fantasy, surrealism, luxury advertising, golden-hour nostalgia, hyper-realism, bokeh, lens jargon, dramatic cinematic lighting, famous-artist references, or random named brands unless the user asks for them or the medium clearly requires them.
36
37Prefer neutral observational specificity like the reference captions: concrete materials, colors, placement, lighting, surfaces, local visual anchors, and typography details.
38
39FIELD RULES
40
41high_level_description:
42- One sentence, 40-65 words.
43- Name the subject, medium, composition, style or era when relevant, and the main visual hook.
44- It should read like a polished image prompt, not analysis.
45
46compositional_deconstruction.background:
47- 90-170 words.
48- Describe the scene shell and global treatment: sky, walls, ground, road, floor, surface, weather, ambient light, palette, depth, print/camera/render treatment, and broad atmosphere.
49- Important visible objects still need elements. Do not hide all content in background.
50
51elements:
52- Include every important visible component.
53- Each obj.desc should be 30-65 words.
54- Each text.desc describes visual typography and placement, not a semantic explanation.
55
56SUBJECT COHERENCE
57
58One coherent subject is one element. A person, face, animal, car, building, bridge, bottle, lantern, signboard, product, flower, or piece of furniture is one obj. Describe attached parts inside that object's desc.
59
60Do not split eyes, hair, face, hands, cap, collar, wheels, windows, headlights, cap, nozzle, atomizer, liquid, label substrate, road, floor, rain streaks, shadows, highlights, reflections, or lens flare into separate objects unless they are truly independent visible objects.
61
62Ground, floor, pavement, road, marble surface, sky, clouds, horizon, weather, distant geography, distant cityscape, distant crowd, broad walls, global lighting, surface reflections, and scene-wide shadows belong in background.
63
64CATEGORY RECIPES
65
66Typography, poster, cover, invitation, label, packaging, sign, menu:
67- 9-16 elements.
68- Use 3-7 separate text elements: headline, subheadline or location, date, tagline, footer, credits, label details when plausible.
69- Add substrate, border/frame/rules, local iconography, illustration, badge, texture, and a concrete palette.
70- For travel posters, prefer recognizable local landmarks and culturally plausible destination lettering.
71
72Infographic, diagram, UI, educational poster, map:
73- 14-24 elements.
74- Use 8-16 text elements when the topic has labeled stages.
75- Include title, stage labels, short captions or callouts, legend/source note when plausible, arrows/connectors, icons, panels, leader lines, and simplified scene pieces.
76
77Complex scene, marketplace, festival, parade, battle, newsroom, banquet:
78- 12-18 elements.
79- Include foreground anchor, central action, secondary group, props/goods, architecture, animals or vehicles when relevant, and readable signage when plausible.
80
81Architecture, interior, room, lobby, workshop, courtyard:
82- 9-14 elements.
83- Put shell architecture in background.
84- Elements are major fixtures, furniture, built-ins, decor, signage, tools, people, focal architectural features, and independent light sources.
85
86Landscape, nature, weather, large environment:
87- 6-10 elements.
88- Include foreground anchor, midground path/water/vegetation, distinctive local or seasonal detail, and atmosphere/weather.
89- Put sky, horizon, distant geography, and broad weather in background.
90
91Portrait, character, fashion:
92- 5-9 elements.
93- Main person is one object with age, skin tone, facial texture, eyes, hair, clothing, expression, pose, and worn details inside the same desc.
94- Add only independent setting anchors, props, signage/text, foreground/background objects, or important accessories as separate elements.
95- Never collapse the whole portrait into only background plus one generic sentence.
96
97Product, commercial, food, vehicle still life:
98- 6-10 elements.
99- Main product, food plate, or vehicle is one object. Put attached parts inside it.
100- Use 2-4 text elements for brand/name/product type/volume/variant when plausible.
101- Add packaging, independent props, surface, ingredient/accessory, and studio/background treatment.
102
103Photoreal street, documentary, cinematic still:
104- 6-10 elements.
105- Main person or vehicle is one object.
106- Put road/ground/rain/global lighting/lens effects/reflections in background unless they are discrete visible objects.
107- Avoid random extra text unless a sign, plate, billboard, or display is plausible and useful.
108
109Abstract, surreal, fantasy, sci-fi:
110- 9-14 elements.
111- Include central impossible subject, scale anchors, secondary motifs, environment, light/energy effects, and depth layers.
112
113TEXT HANDLING
114
115Every visible text block is a separate text element.
116Preserve user-quoted text exactly, including capitalization, punctuation, apostrophes, line breaks when useful, and non-Latin script.
117Prose stays English except literal text inside text fields.
118
119For designed artifacts and products, invent plausible secondary visible text if the format normally contains it, but keep it restrained and useful. Do not invent unrelated real brands, public figures, or copyrighted franchise names.
120
121Do not duplicate the same visible text as several elements unless there are genuinely separate visible instances.
122
123INVENTION RULES
124
125Invent enough concrete detail to make the image renderable, but every addition must serve the user prompt.
126Pick one style, one palette, one era, one material direction, and one layout.
127Avoid hedges such as could, might, various, for example, or similar, maybe, suggested, implied.
128
129When a real place or recognizable subject is named, add factual visual anchors rather than random decoration. When no specific facts are known, choose plausible invented details with restrained specificity.
130
131FINAL CHECK BEFORE OUTPUT
132
133Silently verify:
134- exactly two top-level keys: high_level_description, compositional_deconstruction;
135- compositional_deconstruction contains exactly background and elements;
136- every element has exactly the allowed keys;
137- no aspect_ratio, no style_description, no bbox;
138- no obj key inside an element;
139- no empty text fields;
140- ordinary scenes are not one-element answers;
141- coherent subjects are not split into parts;
142- category element count and text count are reasonable;
143- the output is valid JSON and nothing else."""
144
145USER_TEMPLATE = """\
146TARGET IMAGE ASPECT RATIO: {aspect_ratio} (width:height). Use this only for composition; do not output it.
147User idea: {original_prompt}"""
148
149original_prompt = "a cat on a skateboard"
150aspect_ratio = "16:9"
151
152resp = client.chat.completions.create(
153 model="qwen36-35b-magic-fp8",
154 messages=[
155 {"role": "system", "content": SYSTEM_PROMPT},
156 {"role": "user", "content": USER_TEMPLATE.format(
157 aspect_ratio=aspect_ratio, original_prompt=original_prompt)},
158 ],
159 extra_body={"chat_template_kwargs": {"enable_thinking": False}},
160)
161
162caption = json.loads(resp.choices[0].message.content)
163print(json.dumps(caption, indent=2))