Nedržíme ani nezpřístupňujeme data žádného uživatele a nepozastavujeme účty, pokud to nevyžaduje zákonný orgán.
Image Guide · Multi-Image Composition

AI Outfit Transfer, Uncensored

How to combine two or more images into one on Shannon — clothing transfer, object-in-scene, person beside person, product-on-person — and the one rule that makes all of it work.

Updated September 5, 2026Image GuideShannon 3 · chat app

TL;DR

Attach the images and describe the result while naming what each image contributes, in order. The first image is the main subject or scene; later images supply the parts being brought in. "The woman from the first image wearing the red knitted sweater from the second image, standing in the same park." Do not collapse a two-image job into one image plus a text description of the other — a described garment becomes a generic garment; the real one stays the real one. Combining supports a negative prompt and lets you pick a new canvas aspect. There is no content filter and no refusal layer — and no video generation.

Most of the image work people actually need is not "make me a picture of a thing." It is "put this jacket on this person," "drop this chair into this room," "photograph this product in that hand." All of those need material from more than one photograph, and all of them fail the same way: one image gets attached, the other gets described in words, and back comes something plausible that is not the thing you own. The fix is smaller than you would expect, and it works everywhere.

2+
Images per composition
5
Canvas aspects
Yes
Negative prompt
0
Refusal layers

01What "combining images" is, and where it happens

Shannon 3 does three distinct things with images: it generates a new image from a description, it edits one existing image, and it combines two or more images into a single new one. They have different rules, and picking the wrong one is the first place quality goes.

Use combining whenever the finished picture needs material from more than one photograph: outfit and clothing transfer, putting an object into a scene, person A beside person B, a real product held by a real model, the style of one image applied to the subject of another. If you find yourself wishing you could point at two pictures at once, that is the operation.

All of this happens in the Shannon chat app, in ordinary conversation. It is not on the /v1 API — the API serves text and vision, meaning it reads images you send rather than producing new ones. If you arrived from an API search, that is the honest answer; see the API documentation for what the endpoints do serve. Images render on our own GPU cluster.

You ask in plain language, in whatever language you speak. Shannon works out that an image is being requested, writes the generation instruction itself in English, and shows it to you on a confirmation card before anything renders. Nothing is generated and nothing is billed until you confirm — and for multi-image work that card is where you check that the roles and the order came out as you meant. Shannon never generates uninvited: a greeting, a question, or a casually attached photo does not trigger a render.

02The rule: name each image's role, in order

The instruction must say what each image contributes, and it must say it in the order the images were given. That is the whole technique; everything below is that rule applied to a specific job.

The first image is the base — the main subject, or the scene, whichever the finished picture is mostly made of. It sets identity, pose, framing, lighting and everything you did not ask to change. Later images are sources: the parts being brought in. Write the sentence so a stranger holding the images in order could carry out the job by hand.

Works
The woman from the first image wearing the red knitted sweater
from the second image, standing in the same park, same pose and
same afternoon light, her face and hair unchanged.
Fails
Combine these two images.   Merge the photos.   Put them together.

The failing versions fail not because they are short but because they never say which picture is the person and which is the sweater. Three habits make the rule reliable:

  • Use ordinal language explicitly. "From the first image", "from the second image". Four words, and the ambiguity is gone.
  • Say what stays the same. "Keep her face, hair and pose identical." Naming the unchanged parts is the biggest quality lever in every image operation Shannon does, and it matters most here, because a composition has two or more things it could wander away from.
  • Write prose, not tags. The Image Model reads instructions like a language model reads a sentence, so full sentences beat comma-separated keywords. Aim for 30 to 80 words, and skip SDXL-era junk ("masterpiece, best quality, 8k") and weight syntax like (word:1.3).

When an image contains more than one thing, name the part you want out of it. If the second image is a catalog shot of a model wearing the sweater, "the red knitted sweater worn by the woman in the second image" says you want the garment, not the model. Leave it at "the second image" and you may get her face too.

Role templates by job

JobFirst imageSecond imageInstruction shape
Outfit transferThe personThe garment"The [person] from the first image wearing the [garment] from the second image, [what stays]"
Object in sceneThe sceneThe object"The [room] from the first image with the [object] from the second image placed [where], matching [light]"
Person + personPerson A and settingPerson B"The [man] from the first image standing beside the [woman] from the second image, [positions, eyelines]"
Product on personThe model or sceneThe product"The [model] from the first image holding the [product] from the second image, [grip, label facing]"
Style transferThe subjectThe style reference"The [subject] from the first image rendered in the [style axes] of the second image, [subject unchanged]"

03The mistake that ruins most attempts: describing instead of attaching

This is the error worth making a fuss about, because it is by far the most common and because it looks like it worked. The user has two photographs — their client and their client's product, or themselves and a jacket they own. They attach one, and they type the other.

The mistake
[attaches one photo of a woman]
"Put a red knitted sweater on her, cable knit, kind of chunky,
with a rolled collar."

What comes back is a woman in a red knitted sweater. It is a good sweater. It is not the sweater. The cable pattern is invented, the dye is a different red, the ribbing depth is wrong, the collar rolls the wrong way, and the small asymmetries that make a real garment recognizable — a repair, a pull, the stretch at the elbows — are gone, because nothing in the sentence contained them.

Text is a lossy channel and a photograph is not. A careful sentence carries maybe a dozen attributes of an object; the photograph carries thousands, including every one you would never think to write down and the ones you cannot name. Describing instead of attaching throws away everything you did not think of, which is most of it. Attaching preserves the real appearance because the real appearance is present.

The test is simple. If the point of the picture is that it is that one — that specific garment, that person, that exact product, that room — attach it. If you could pick the item out of a lineup and it matters that the result matches the one you would pick, a description will not do it.

The one-line version

A described garment becomes a generic garment. The real one stays the real one. The same holds for the second face, the product, the chair and the room: attaching preserves identity, describing replaces it with an average.

There is a legitimate case for describing: when the thing does not exist yet, or when a generic version is what you want. Designing a sweater nobody has knitted, roughing out a concept — that is text-to-image work, and a description is right because there is no real object whose identity you are keeping. The moment a real object enters the brief, so does its photograph.

A related half-mistake: attaching both images and then still describing the second one in detail. The adjectives can fight the photograph — write "chunky" when the real garment is fine-gauge and you have introduced a contradiction. Describe the second image only enough to identify which part of it you want.

04Outfit and clothing transfer

The canonical two-image job. Attach the person first and the garment second. The garment image can be a flat lay, a hanger shot, a catalog photo, or someone else wearing it — all work, as long as you say which part of it you mean.

Outfit transfer
The woman from the first image wearing the red knitted sweater
from the second image. Keep her face, hair, pose and the park
background exactly as they are, with the same soft overcast light.
The sweater should hang naturally on her frame, with the cable
pattern and collar of the original garment intact.

Negative prompt: warped cable pattern, extra sleeves, distorted
hands, plastic-looking fabric, mismatched lighting

Three things do the work: the roles are named in order, the unchanged parts are listed out loud (which is what stops the face drifting), and the last clause tells the model the garment's own details are the point — without it, compositions tend to keep color and silhouette while quietly smoothing the texture.

  • Fit is a physical claim, so make it. "Hangs naturally on her frame", "fitted at the waist", "sleeves ending at the wrist" — otherwise the garment can arrive at its source proportions rather than the wearer's.
  • Occlusion is the hard case. If arms are crossed or a bag strap crosses the chest, say what happens to it. Ambiguity there produces the classic melted-sleeve artifact.
  • More garments means more images. Person first, then top, then trousers, then shoes — and refer to them as the second, third and fourth image consistently. Do not switch to "the shoe picture" halfway.

05Putting an object into a scene

Here the roles invert: the scene goes first, because the scene is what the finished picture is mostly made of, and the object is the part being brought in. Interior mockups, staging a room with furniture you have not bought, dropping a vehicle onto a road, placing a sculpture in a gallery — same shape every time.

Object in scene
The living room from the first image with the green velvet
armchair from the second image placed beside the window on the
left, angled slightly toward the camera and standing about waist
height in frame. Match the room's warm late-afternoon light, with
a soft shadow falling to the right of the chair onto the rug.
Keep the rest of the room and the camera angle unchanged.

Negative prompt: floating object, cut-out edges, duplicate
shadows, mismatched color temperature

The failure mode here is not identity, it is physics. An object that keeps its own lighting looks pasted on, and viewers spot it instantly even when they cannot say why. Spend your words on the three things a photograph of an isolated object cannot know: placement (where in the frame, relative to what), scale in terms of the scene rather than centimeters ("about waist height", "as tall as the door handle"), and light and shadow — direction, quality, and where the shadow falls. That last one is what converts a composite into a photograph.

06Person A beside person B

Two people who were never in the same room: family photographs across generations, a founder portrait when the co-founder was traveling, a poster with two performers. It works, and it needs two things the single-subject jobs do not.

Two people
The man from the first image standing beside the woman from the
second image, both on the same stone terrace from the first image.
He stands on the left, she stands on his right, shoulders almost
touching, both facing the camera at the same eye level. Keep both
faces, hairstyles and clothing exactly as in their own photographs,
and light both with the same soft evening light from the left.

Canvas: landscape
Negative prompt: mismatched skin tone, duplicated limbs, seam
line between subjects, uneven lighting

The two extras: relative geometry — who is on which side, how far apart, whose eyeline goes where, and relative height, since two differently cropped portraits give the model no way to know one person is a head taller — and a single shared light, stated explicitly, because you are merging two lighting environments and one has to win.

This is also where choosing a new canvas aspect earns its keep: two portrait-orientation sources make a bad portrait-orientation group shot, so ask for landscape or wide. And state the obvious once — putting a real person into an image they were not in is powerful and is not automatically harmless. Have the rights or the consent; our responsible use policy is the short version.

07Product-on-person, and uncensored product photography

Commercially, this is the composition that pays for itself. You have a product shot on white and you have a model, a hand, a kitchen, a workbench. Combining them gets you a lifestyle photograph without a studio day — and with your actual product in it rather than a generated approximation, which is the whole difference between usable marketing material and a mockup you have to apologize for.

Product on person
The woman from the first image holding the amber glass serum
bottle from the second image in her right hand, raised to chest
height, the label facing the camera and fully legible, her fingers
wrapped naturally around the bottle without covering the label.
Keep her face, hair and the kitchen background unchanged. Shot on
Kodak Portra 400, 85mm f/1.4, soft window light from the left.

Canvas: portrait
Negative prompt: warped label text, extra fingers, floating
bottle, duplicated reflections, plastic skin

Two additions specific to product work. The grip: hands holding objects is where composition most often goes wrong, so state which hand, at what height, and that the fingers wrap naturally without occluding the label. Photographic treatment: name a camera, a lens and a film or light source. The output should read as a photograph that was taken, not two photographs that were combined, and writing it as a photograph is what gets you there.

Uncensored matters here in a more mundane way than it sounds. Filtered generators refuse whole categories of ordinary, legal commerce: lingerie and swimwear catalogs, tobacco and vape hardware, alcohol, firearms accessories, adult wellness products, cannabis where it is legal, medical devices, tactical and hunting gear. They also refuse brand and likeness work done by the rights-holder. None of that is an edge case; it is a Tuesday at an agency. Shannon has no content filter on the generation path and no refusal layer on the model that writes the prompt, so the work goes through and the rights question stays with you, where it belongs.

08Style from one image, subject from another

Subject first, style reference second. What makes this job different is that you are borrowing a quality rather than an object, so you have to name which qualities — and say that the reference's own subject and composition are not coming along.

Style reference
The portrait of the man from the first image rendered in the
visual style of the second image: the same muted teal and ochre
palette, visible dry-brush texture, and flat even lighting. Keep
his face, expression and pose recognizable and unchanged. Do not
carry over the subject, background or composition of the second
image.

Negative prompt: second face, borrowed background, photographic
texture

Name the style axes explicitly, because "in the style of the second image" alone is under-specified: palette (or exact colors, written as color #8B0000 attached to a named object), brushwork or grain, lighting quality, line weight, level of abstraction, edge treatment. Pick the two or three that define the look; naming all six averages them. And do not mix conflicting directions — "photorealistic" plus "watercolor" produces a muddle, not a blend. If you want a photograph that merely borrows a palette, say exactly that.

09Negative prompts and choosing a new canvas aspect

Two capabilities that exist for combining and editing but not for text-to-image, and both are underused.

The negative prompt

Combining images supports a negative prompt. Unwanted artifacts go there — and it matters that they go there rather than into the instruction, because writing "no extra fingers" in the main instruction names the thing you are trying to avoid in the same breath as everything you want. Put exclusions in their own channel. The entries that consistently earn their place:

  • Anatomy — extra fingers, duplicated limbs, malformed hands, second face.
  • Compositing tells — cut-out edges, visible seam line, floating object, duplicate shadows, halo around subject.
  • Lighting mismatch — mismatched color temperature, inconsistent shadow direction, uneven exposure between subjects.
  • Texture failures — plastic skin, smoothed fabric, warped label text, illegible logo.

Keep it to the handful that are plausible for your image. A negative prompt listing thirty things is a diffuse instruction, not a strong one.

The canvas aspect

Editing a single image inherits the source image's aspect ratio automatically — there is one source, and its shape is the answer. Combining has no such default, so you can choose a new canvas aspect, and you often should, because the shape that suited one source rarely suits the composition.

AspectBest forTypical composition use
squareDefault, social feedsSingle subject with one transferred element
portraitPeople, tall subjectsOutfit transfer, product held by one model
landscapeScenery, groupsTwo people side by side, object placed in a room
wideCinematic framingGroup shots, environmental product scenes, banners
tallPosters, phone wallpapersFull-length outfit shots, poster layouts

An explicit request always wins. Say nothing and you get square — the wrong shape for most two-person and full-length outfit work, so it is worth a word.

10Which images Shannon uses: the conversation versus this message

There are two ways an image gets into a composition, and knowing which one you are relying on prevents most ordering mix-ups.

Attached to the current message. The direct route: attach everything to the message that carries the instruction, in the order you intend to refer to it, and "first image" and "second image" mean exactly what they look like. Default to this.

Already in the conversation. Pictures sent earlier in the thread stay available and can be referred to in ordinary language — "the photo I sent earlier", "the sweater picture from before", "the second version you made". Shannon resolves the reference against the thread, which is genuinely convenient across a long session: you can iterate on a result and then bring in a new element without re-uploading everything.

The safeguard for both is the same. The confirmation card shows the instruction before anything renders, including the ordering it chose. Read it. If the roles came out reversed — the sweater as the base, the person as the source — say so and it is fixed before a single pixel is generated or billed. Still, when a thread is long and full of images, re-attach the ones you need; six pictures back, "the second image" is ambiguous to any reader.

Asking Shannon to look at an image — describe it, read the label, tell you what the fabric is — is ordinary vision work, not a generation. It costs nothing extra and it is a good check before composing: if Shannon's description does not mention the detail you care about, your instruction needs to name it.

11Uncensored, and what that actually means

There is no content filter on the generation path and no refusal layer on the model that writes the prompt. Composition is where filtered tools fail most visibly, because a refusal system cannot tell the difference between combining images you own and combining images you should not have — so it blocks both.

The work those tools routinely refuse is not exotic: art nudes and figure study, medical and forensic illustration, security research imagery, violence in fiction for book covers and game art, political satire, and the legal product categories listed above. Every one is legitimate professional work that a keyword filter cannot distinguish from misuse.

Uncensored does not mean unaccountable. You are supplying photographs of real objects and often real people; the rights and the consent are yours to hold, and so is responsibility for the result. Shannon Lab LLC is a New Mexico company operating under a published responsible use policy — worth five minutes before you start work involving other people's likenesses.

12Limits worth knowing before you start

There is no video generation. Shannon does not animate images and does not produce video, in any form, from any of these operations — not from a single image, not from a composition, not as an experimental option. If your brief ends in motion, this is not the tool for the last step. That is what the product is, not a roadmap hint.

Other honest boundaries, all with workarounds:

  • Fine text is the weakest area. Small print, dense packaging copy and long serial numbers can come out imperfect. Frame the product so the important text is large, put warped label text in the negative prompt, and check at full size.
  • Extreme scale mismatch is hard. A 4000-pixel product shot combined with a small thumbnail of a room gives the model little to work with. Use the best copy of each source you have.
  • Heavy occlusion limits what can be transferred. A garment photographed folded, back never visible, cannot have its back invented faithfully. If the composition shows the back, supply the back.
  • More images means more room for confusion. Four sources is workable with disciplined ordinal naming, but it is not the place to be casual — use the number the job needs and no more.
  • Two big changes at once compete. Transfer the outfit, confirm, then change the background. Sequential edits beat one overloaded instruction almost every time, and each step is cheap to check.

13Frequently asked questions

How do I combine two images into one with AI?

Attach both images in the Shannon chat and describe the result while naming what each image contributes, in order. The first image is the main subject or scene; later images supply the parts being brought in. For example: the woman from the first image wearing the red knitted sweater from the second image, standing in the same park. Shannon writes the final English instruction itself, shows it to you on a confirmation card, and nothing renders until you confirm.

Why does the outfit look wrong when I describe the garment instead of attaching it?

Because a sentence is a lossy description of a photograph. Writing "a red knitted sweater" gets you a plausible red knitted sweater, not yours — the cable pattern, ribbing, dye, drape, seam placement and wear are all invented. Attaching the second image keeps the real garment's real appearance. If the point of the picture is that it is that specific item, person or product, attach it rather than describing it.

Does multi-image composition support a negative prompt and a new aspect ratio?

Yes to both. Combining images supports a negative prompt, so unwanted artifacts such as extra fingers, duplicated limbs, warped logo text or visible seams belong there rather than in the instruction. You can also choose a new canvas aspect — square, portrait, landscape, wide or tall. This differs from editing a single image, which inherits the source image's aspect ratio automatically.

Can Shannon use images already in the conversation, or only ones I attach now?

Both. You can attach the images to the current message, or refer to pictures already in the thread — "the photo I sent earlier", "the sweater picture from before". Shannon resolves those references and the confirmation card shows the ordering it chose, so you can correct the order before anything is rendered. Attaching everything in one message, in the order you will refer to it, is the most reliable way to work.

Can Shannon turn the combined image into a video?

No. Shannon does not generate video and does not animate images. It generates new images, edits one existing image, and combines two or more images into one. Image work happens in the Shannon chat app; the /v1 API serves text and vision, meaning it reads images rather than producing them.

Is uncensored image composition really unfiltered?

There is no content filter on the generation path and no refusal layer on the model that writes the prompt. Filtered generators routinely refuse legitimate composition work — lingerie and swimwear catalogs, art nudes, medical and forensic illustration, regulated consumer goods, security research imagery, political satire, and brand or likeness work done by the rights-holder. Shannon does not. The rights and the responsibility for the source images are yours; see our responsible use policy.

Stated plainly

No benchmark numbers appear in this guide, because there is no honest public benchmark for "did the sweater stay the same sweater." Everything above is operational guidance from real use. Where a capability is absent — video, image generation on the API — we say so rather than leaving it open.

Combine your images

Attach them, name each role in order, confirm the instruction, render. No refusal layer.

Open Shannon Chat Read the API Docs

Shannon Lab LLC · New Mexico, USA · images render on our own GPU cluster


Image generation, editing and multi-image composition are features of the Shannon chat app; the /v1/chat/completions, /v1/messages and /v1/responses endpoints serve text and vision, meaning they read images rather than generate them. There is no video generation. Aspect options, negative-prompt support and confirmation behavior are current as of September 5, 2026. Related: uncensored AI image generation · uncensored AI image editing · responsible use · API documentation · research index.

Všechny výzkumné odkazy