Ni ne tenas aŭ aliras datumojn de iu ajn uzanto, kaj ni ne suspendas kontojn krom se laŭleĝa aŭtoritato postulas devigan agon.
Guide · Image Prompting

The AI Image Prompt Guide

Shannon's Image Model reads a prompt like a language model. That single fact changes everything about how you should write one — starting with dropping the comma-separated tags.

Updated September 5, 2026Practical GuideShannon 3 · Image Generation

TL;DR

Write flowing prose in full sentences, not keyword tags. Aim for 30–80 words. Order it the way the model reads it: subject → action → style → setting and lighting → technical, most important first. Be concrete about the subject. If you want a photograph, write it as a photograph — name a camera, a lens, and a film stock or a light. Attach exact colors to named objects as color #RRGGBB. Pick one of five aspect ratios. And drop the SDXL-era habits: no "masterpiece, best quality, 8k", no (word:1.3) weights, and no negative prompt in text-to-image — describe only what you want. Editing and combining do take a negative prompt.

Most people arrive with prompt habits learned on SDXL: a comma-separated string of nouns, a quality-token preamble, a few parenthesized weights, and a negative prompt as long as the positive one. Those habits were correct for those models and are counterproductive here. This is the short version of what actually works — and because Shannon writes the generation prompt for you and shows it before anything renders, it is equally a guide to reading and correcting what Shannon proposes.

30–80
Target words
5
Aspect ratios
0
Negative prompts in T2I
1
Confirmation card first

01Why prose beats tags: the model reads like a language model

The older diffusion models most prompt advice was written for used a text encoder that behaved, in practice, like a bag of concepts: grammar barely mattered, and the winning strategy was to pile up relevant nouns and bias them with weights. Tag soup was the correct shape of prompt for that encoder.

Shannon's Image Model does not work that way. It reads a prompt the way a language model reads text, so syntax carries meaning. "Holding an umbrella against her shoulder" is a relationship between a person, an object and a body part; "umbrella, shoulder" is two things that happen to be present. Prepositions place things. Verbs create poses. Clauses attach detail to the right noun instead of letting it float across the whole image.

Here is the same picture requested both ways.

Tag style — the old habit
woman, 30s, tokyo, rain, night, crosswalk, neon signs, umbrella,
wet asphalt, reflections, masterpiece, best quality, 8k, ultra
detailed, (bokeh:1.3), cinematic lighting, photorealistic
27 comma-separated fragments. No verb, no relationships, four dead tokens, one weight syntax that is not parsed.
Prose style — what works here
A woman in her early thirties waits at a rain-soaked Tokyo
crosswalk at night, holding a clear vinyl umbrella against her
shoulder and glancing down the street. Neon storefront signs
bleed red and cyan into the wet asphalt at her feet. Shot on
Kodak Portra 400, 85mm f/1.4, shallow depth of field, the
traffic signal behind her thrown out of focus.
61 words. Every element is placed relative to another one, and the framing is stated as a camera decision.

Look at what the tag version cannot express. It never says the umbrella is held, so it may end up lying on the ground. It never says the neon is reflected in the asphalt, only that both exist, so you get signs and a wet street with no optical relationship between them. It never says where the woman is looking, so her pose is a coin flip. "Cinematic lighting" is a mood label, not a light source. And the last five tokens are noise, read as English words about photography rather than control handles.

The prose version is not longer because it is flowery. It is longer because it answers questions the tag version left open. That is the entire trick.

02How long should an image prompt be?

Aim for 30 to 80 words. Do not go much past 100. This is the single easiest quality lever available to you, and both ends of the range matter.

Under about 30 words the composition, light and framing are unspecified, so the model picks whatever is most statistically ordinary. "A woman in a coffee shop" is a valid prompt that produces a forgettable image, because you delegated every interesting decision.

Above roughly 100 words a different failure appears: late clauses compete with early ones for the same regions of the image, and instead of accumulating detail you get dilution. The third lighting description undoes the first, and the subject you established in the first eight words gets less relative attention with every sentence you add. Long prompts read as thorough. They do not render as thorough.

The discipline: write the prompt you want, then cut it back to the sentences that would change the picture if removed.

03The ordering discipline: most important thing first

Because the model reads sequentially, position is emphasis: what comes first anchors the image, what comes last adjusts it. Put things in the order the model needs them:

  1. Subject — who or what this picture is of, concretely.
  2. Action or pose — what they are doing, how they are held.
  3. Style or mood — photograph, oil painting, technical illustration, and the register it is in.
  4. Setting and lighting — where this happens and what light falls on it.
  5. Technical — framing, lens, depth of field, color specifics.

The most common self-inflicted wound is starting with the technical clause. Opening with "Cinematic wide shot, anamorphic lens flare, teal and orange grade, of a man reading a newspaper" tells the model that the look is the subject and the man is a detail. You get a gorgeous, empty frame with a small indistinct figure in it.

Technical first — the subject arrives too late
Cinematic wide shot, anamorphic lens flare, teal and orange
color grade, shallow depth of field, golden hour, of a man
reading a newspaper on a park bench.
Subject first — same elements, correct order
A man in his sixties in a gray wool overcoat sits on a park
bench reading a folded broadsheet newspaper, one ankle crossed
over his knee. Late golden-hour sun rakes across him from the
left through bare plane trees. Photographed wide on a 35mm
anamorphic lens, shallow depth of field, warm highlights and
cool shadows.

Every element from the first version survives in the second; nothing was dropped. The model simply now knows what the picture is of before it is told how to shoot it — and that reordering alone moves the subject from incidental to central.

04Be concrete about the subject

"A person" is not a subject, it is a placeholder — and the model fills placeholders with the average of everything it has seen, which is exactly where the glossy, ageless, characterless look comes from. Concreteness costs five words and buys a specific human being:

VagueConcrete
a persona man in his forties with a salt-and-pepper beard and deep smile lines
a womana woman in her late twenties with a dark undercut and a chipped front tooth
an old buildinga four-story brick warehouse with painted-over signage and boarded upper windows
fooda bowl of ramen with a soft-boiled egg halved on top, steam rising

Note the pattern: age or era, one structural feature, one distinguishing imperfection. That third item is what makes a result feel photographed rather than generated. A chipped tooth, a worn hem, painted-over signage — small asymmetries are the strongest anti-plastic signal available, and they cost almost nothing in your word budget.

05Photorealism: write the request as a photograph

Photorealism is the default register unless you ask for art, illustration or anime. But the word "photorealistic" does little work. What produces a photograph is describing the act of photography: a real camera, a real lens, and a real film stock or light source.

The reason is straightforward. Every image captioned "Kodak Portra 400" was made by a specific process with a specific grain structure, highlight roll-off and skin rendition, and naming the stock recruits all of it. "8k ultra detailed" only ever appeared on stock-site spam, so it recruits stock-site spam. A few reliable recipes, each producing a distinctly different image of the same subject:

RecipeWhat it gives you
Kodak Portra 400, 85mm f/1.4, soft window lightWarm, forgiving skin tones; creamy separation; editorial portrait
Ilford HP5 pushed to 1600, 35mm, available lightGrainy black and white, contrasty, street-documentary energy
Digital medium format, 80mm, single softbox at 45°Clean commercial studio look, controlled falloff, product-grade
135mm f/2, backlit at golden hour, atmospheric hazeCompressed background, rim light, romantic telephoto look

Match the grounding to the recipe — pore-level skin texture means nothing in a 24mm wide where the face is eighty pixels tall — and then add it. Grounding detail is — the texture evidence that a lens actually recorded this. Visible skin pores and fine facial hair. The weave of a linen shirt. Dust motes in a shaft of light. A scuff on a leather boot. Flyaway hairs against a bright background. Two or three are enough, and they are the difference between a rendering of a person and a photograph of one.

06Layered scenes: foreground, midground, background

When a scene has depth, describe the three planes separately and in order. This is the most effective technique for images that should feel like places rather than backdrops, and almost nobody does it. Without it you get flatness: everything the prompt mentions lands at roughly the same distance, competing for the same middle band of the frame. Naming planes gives the model a depth budget to spend.

Three planes, named explicitly
In the foreground, a cracked concrete ledge with a half-empty
paper coffee cup and a scattering of cigarette ash, slightly out
of focus. In the midground, a bicycle courier in a yellow rain
jacket wheels her bike across a wet plaza, caught mid-stride. In
the background, a glass office tower dissolves into low fog.
Shot on a 50mm lens at f/2.8, flat overcast light.
70 words. The out-of-focus foreground element is what sells the depth — it tells the model where the camera is standing.

Two refinements. Put something in the foreground that is deliberately out of focus — real photographs of deep scenes almost always have one, and its absence is a common tell. And say which plane the subject occupies: if the courier is the point of the picture she belongs in the midground, otherwise the tower and the coffee cup will fight her for the frame.

07Exact colors: hex codes, attached to objects

For brand work, costume continuity, or anywhere "red" is not specific enough, write the hex code as color #RRGGBB and attach it to a named object:

Correct: hex bound to a noun
...wearing a heavy wool peacoat, color #8B0000, over a cream
turtleneck, against a studio backdrop, color #1E3A5F.
Wrong: hex floating free
...#8B0000, #1E3A5F, dark red and navy color palette, moody...
Nothing binds these codes to anything, so they act as a vague tint suggestion over the whole frame.

Three rules of thumb. Use hex for the two or three colors that genuinely matter and plain words for the rest — a prompt with six hex codes spends most of its budget on paint chips. Bind the code to an object with edges, never to "the lighting" or "the mood". And remember it specifies the object's color, not the final pixel value: a #8B0000 coat in golden-hour backlight renders warmer than the swatch, exactly as it would in life.

08Choosing the aspect ratio

Five ratios are available and the right one is usually implied by the subject, but an explicit request from you always wins — ask for a wallpaper and you get tall regardless of what the content suggests.

RatioBest forWhy
squareDefault. Single centered objects, icons, symmetrical compositions, social postsNo dominant axis, so nothing is cropped by the frame's own bias
portraitPeople, full or half body; tall subjects like towers, trees, bottlesGives a standing figure room head to foot without shrinking it
landscapeScenery, interiors, groups of people, product-in-context shotsHorizontal room for the setting to actually be a setting
wideCinematic frames, film stills, banners, header imagesThe widest option; strong negative space, anamorphic feel
tallPosters, book covers, phone wallpapers, vertical social formatsTaller than portrait; leaves deliberate space for type or sky

Getting this wrong is more damaging than it looks. A full-length standing figure in a wide frame does not give you a small figure in a big frame — it usually gives you a cropped one, because the model fills the width and the head or feet leave the frame. Match the frame to the subject's dominant axis and most composition complaints disappear.

Two behaviors worth remembering: editing keeps the source image's aspect ratio automatically, so you cannot reframe by editing, while combining lets you pick a new canvas — often the cleanest way to change the shape of something you already have.

09Anti-patterns: what to stop doing

Quality tokens: "masterpiece, best quality, 8k, ultra detailed"

These were community-discovered handles for a specific older training distribution, where certain caption phrases correlated with certain image sources. They do not transfer. Here they are read as ordinary English words, contribute nothing, and eat part of a word budget you have already been told is tight. The replacement is to say which detail you want: not "ultra detailed" but "visible weave in the coat fabric and fine stubble along the jaw" — the same intent, expressed as something the model can render.

Weight syntax: (word:1.3), [word], {{word}}

Not parsed. The parentheses, brackets, colons and numbers land as literal text and, if anything, make the surrounding clause harder to read. If you want more emphasis on something, use the two levers that do work: move it earlier in the prompt, and give it more words. "A crimson silk scarf, its fringed ends lifting in the wind" outweighs "(red scarf:1.4)" by every measure that matters.

Conflicting directions

Mixing incompatible instructions does not blend them; it produces an unstable average that usually satisfies neither. The frequent offenders:

  • Medium — "photorealistic" plus "watercolor"; "oil painting" plus "shot on 85mm". Pick one and commit.
  • Framing — "wide establishing shot" plus "extreme close-up on her eyes"; "full body" plus "detailed facial pores".
  • Lighting — "harsh midday sun" plus "soft golden hour glow"; "low key, deep shadows" plus "bright and airy".
  • Era — "1970s film aesthetic" plus "modern minimalist" plus "cyberpunk".
  • Style stacking — four artists or four genres at once. Two related references can work; four cancel out into generic.
  • Sprawl — "rustic, weathered, ancient, timeworn wooden door" is one idea written four times, and each repetition dilutes the rest of the prompt.

10There is no negative prompt in text-to-image

This is the anti-pattern that catches experienced users hardest, because on older stacks the negative prompt was half the craft.

OperationNegative promptWhat to do
Text-to-image (generate)Not supportedDescribe only what you want. Never write "no X".
Editing one imageSupportedPut unwanted artifacts there, keep the instruction positive.
Combining two or moreSupportedSame — artifacts in the negative, roles in the positive.

For generation the rule is absolute and slightly counterintuitive: writing "no text, no watermark, no extra fingers" into the prompt makes those things more likely, not less. The model reads "no text" as a sentence containing the concept text; negation is a weak signal and the noun is a strong one. You have effectively asked for text. Always state the positive form instead:

Instead ofWrite
no background cluttera clean seamless gray backdrop
not blurrytack sharp on the eyes, f/8
no other peoplealone in an empty street at dawn
no modern objectsa 1930s kitchen with enamelware and a cast-iron stove
no cartoon styleshot on Kodak Portra 400, 85mm, natural skin texture

For editing and combining the negative prompt exists and you should use it — but for artifacts, not content. Good material: watermark, text overlay, duplicated limbs, warped hands, oversharpening halos, JPEG blocking, seam lines along a composite edge. Bad material: anything about what the picture is of, which belongs in the instruction.

11Reading the prompt Shannon writes for you

Here is what makes this guide immediately usable rather than theoretical. You do not have to write the generation prompt yourself. You ask in plain language, in any language; Shannon recognizes that an image is being requested, composes the English generation prompt, and shows it to you on a confirmation card before anything renders. Nothing is generated and nothing is billed until you confirm.

So the skill that pays off most is not writing prompts from scratch — it is reading the one on the card. Run down this list:

  1. Is the subject concrete? If the card says "a person" or "a woman", you have found the single biggest available improvement. Give it an age, a build, a feature.
  2. Is the subject first? If the prompt opens with a style or a camera, ask for it to lead with the subject instead.
  3. Is it in range? Under 30 words, ask for more specificity on light and setting. Over 100, ask for it to be cut.
  4. Is there a light source? Not a mood word — an actual light: window, streetlamp, flash, overcast sky, softbox.
  5. Any conflicts? Scan for two mediums, two framings, or two lighting conditions in the same prompt.
  6. Any smuggled negatives? If the card contains "without", "no", or "avoid", that is a text-to-image prompt asking for the thing it names.
  7. Is the ratio right for the subject? Full-length figure in wide is the classic mismatch.

Corrections are conversational: "make her early thirties with a scar through one eyebrow, shoot it on Portra at 85mm, and use portrait" is a perfectly good instruction. You get a new card and confirm that one instead. There is no penalty for iterating, because nothing has rendered yet.

And Shannon does not generate uninvited. A greeting, a question, or a photo you attach for discussion does not trigger a render. Asking Shannon to look at an image — describe it, read it, analyze it — is ordinary vision work, not generation, and costs nothing extra.

12How the craft shifts for editing and combining

The rules above are written for generation. Two adjacent modes change the shape of the prompt.

Editing one image

An edit prompt is an instruction, not a description: state what changes and what stays. Naming the unchanged parts is the single biggest quality lever in the product — "Change the background to a rainy night city street with neon reflections, keep the woman exactly as she is" holds together where the same request without that clause quietly redraws her face. Be specific about the anchor: keep the face, pose, hair and clothing identical. Exactly one source image is used, and the output keeps its aspect ratio.

Combining two or more images

Whenever elements come from more than one picture — outfit transfer, an object into a scene, two people in one frame, product-on-person, a style from one image with a subject from another — the prompt must name the role of each image, in order: "The woman from the first image wearing the red knitted sweater from the second image, standing in the same park." The first image is the main subject or scene; later images supply the parts being brought in. Do not collapse a multi-image request into one source plus a text description — combining the real images is what preserves how those things actually look.

13Uncensored prompting: say what you mean

There is no content filter on the generation path and no refusal layer on the model that writes the prompt. That is the product's whole point: filtered generators routinely refuse legitimate work — art nudes, medical and forensic illustration, security-research imagery, violence in fiction, political satire, brand and likeness work by the rights-holder. Half the prompt craft on filtered platforms is euphemism engineering, and euphemism is the enemy of an accurate image.

So there is one more thing to check on the card: that it describes the actual subject in plain words rather than the hedged vocabulary people learn from filtered tools. A forensic illustration of a wound should say what the wound is. Hedged language produces a worse image everywhere, because the model renders what the words say.

Capability and judgment are separate things. Shannon Lab LLC operates this under a published Responsible Use Policy: unfiltered generation is a professional tool, and consent, rights and the law all still apply to what you make.

The one-page checklist

  • Prose in full sentences, not comma-separated tags.
  • 30–80 words. Cut anything that would not change the picture.
  • Subject → action → style → setting and lighting → technical.
  • Concrete subject: age, one structural feature, one imperfection.
  • Photos: name a camera, lens, and film stock or light source.
  • Deep scenes: foreground, midground, background, named separately.
  • Exact colors as color #RRGGBB bound to a named object.
  • Pick the ratio from the subject's dominant axis.
  • Two or three grounding textures, matched to the framing.
  • No quality tokens, no weight syntax, no conflicting directions.
  • No negatives in text-to-image. Negatives are for editing and combining, and only for artifacts.

14Where the API fits

One point of accuracy for readers arriving from an API search: image generation, editing and combining live in the Shannon chat app, not on the /v1 API. The API serves text and vision — it can read, describe and analyze images you send it across all three dialects (/v1/chat/completions, /v1/messages, /v1/responses, all streaming) — but it does not create them. If you are building a pipeline that needs to understand images, the API documentation is the place to start. If you want to make images, that is the chat app, and everything in this guide applies there. Images render on our own GPU cluster.

15Frequently asked questions

How long should an AI image prompt be?

Aim for 30 to 80 words and never go much past 100. Below 30 words you are leaving the composition, lighting and framing to chance. Past 100 words the later clauses start competing with the earlier ones and detail begins to wash out rather than accumulate.

Should I write keyword tags or full sentences?

Full sentences. Shannon's Image Model reads a prompt the way a language model reads text, so grammar carries information. "A woman in her 30s at a rain-soaked Tokyo crosswalk at night" produces a coherent scene; "woman, 30s, Tokyo, rain, night" produces a pile of loose attributes with no relationships between them.

Can I use a negative prompt for text-to-image?

No. Text-to-image on Shannon takes no negative prompt, so describe only what you want and never write "no X" in the prompt itself — naming a thing tends to summon it. Image editing and image combining do support a negative prompt, and that is the right place for unwanted artifacts such as watermarks, text overlays or duplicated limbs.

Do tokens like "masterpiece, best quality, 8k" improve the image?

No. Those were handles for an older generation of models and they carry no information here — they are simply read as words, and they spend part of your 30–80 word budget saying nothing. Weight syntax such as (word:1.3) is not parsed either; it lands as literal text. Describe the actual detail you want instead.

How do I get an exact color?

Write the hex code as color #RRGGBB and attach it to a specific named object, for example "a wool jacket, color #8B0000". A hex code floating free of an object has nothing to bind to. Use it for the two or three colors that actually matter and describe the rest in words.

Can I generate images through the Shannon API?

No. Image generation, editing and combining happen in the Shannon chat app. The /v1 API serves text and vision — it can read and analyze images you send it, but it does not create them. Everything in this guide applies to the chat app.

Try it on a real prompt

Ask in any language. Read the card. Adjust it. Then confirm.

Open Shannon Chat More Research

Nothing renders, and nothing is billed, until you confirm the prompt


Related reading: Uncensored AI image generation · Uncensored AI image editing · Responsible Use Policy · API documentation · All Shannon research. Guidance here is drawn from our own testing of Shannon's Image Model; advice written for earlier diffusion stacks generally does not transfer.

Ĉiuj esplorligiloj