guideTutorials & how-to guides10 min read

How to add a person to a photo with AI: step by step

Add a person to a photo with AI: composite someone into an image on OmniArt using multi-image fusion, reference photos, region edits, and matched lighting.

OmniArt Team
How to add a person to a photo with AI: step by step

To add a person to a photo with AI, you are not really pasting a cutout — you are asking an image model to re-render one photo so that a second person belongs in it. The model has to place them somewhere plausible, size them to the scene, relight them to match, and put a shadow under their feet. Get those four things right and the result reads as a photograph. Miss one and it reads as a sticker.

On OmniArt this is an image-model job, not a layer-based editor task. You upload the base photo, upload a reference photo of the person, and describe the composite in the prompt. Seedream 5.0 Pro handles this well because it supports multi-image fusion — up to 10 reference images fused into one output — plus region and anchor language for controlling exactly where the edit lands. GPT Image 2 takes natural-language editing instructions, and Nano Banana 2 does reference-guided edits when you want a second opinion on the same pair of inputs.

This guide walks the whole workflow in order: pick the photos, pick the model, write the placement prompt, fix what's off, then lock a final render. Three copy-pasteable prompts are included.

What the model is actually doing

It helps to understand the mechanism before you write a prompt. When you feed two images and one instruction, the model isn't compositing pixels — it's generating a new image conditioned on both inputs. That has two practical consequences.

First, the base photo can change slightly even in areas you didn't mention. Backgrounds shift, a table edge moves, text on a sign scrambles. You reduce this by explicitly protecting the rest of the frame in the prompt, the same discipline that makes region edits work.

Second, identity comes entirely from the reference photo. If the reference is small, blurry, heavily filtered, or shot from an angle nowhere near the one you want in the output, the model will invent the difference — and invented faces drift. The reference does more work than the prompt here.

Step 1: choose the base photo and the person reference

Bad inputs cost more iterations than a bad prompt, so spend a minute here.

The base photo should have an obvious empty space where a person can stand or sit, a readable light direction, and visible ground or floor so a shadow has somewhere to land. Avoid photos where every plausible position is occluded by furniture, and avoid extreme wide-angle shots — the perspective distortion is hard to match.

The person reference should be one clean, well-lit photo, ideally shot from roughly the angle and distance you want in the final image. Full body if the person will be full body in the output. Neutral lighting beats dramatic lighting, because flat light is easier to relight than a hard-shadowed portrait. If the person needs to appear in a specific outfit, the reference should already show that outfit.

Tip

If you have several usable photos of the same person, upload two or three rather than one. Seedream 5.0 Pro accepts up to 10 references, and a second angle measurably improves facial consistency. Keep them consistent though — mixing a 2019 photo with a current one gives the model two identities to average.

Step 2: pick the model

All three of these are in the OmniArt image workspace, and all three take uploaded reference images alongside the prompt.

ModelStrength for this jobReach for it when
Seedream 5.0 ProMulti-image fusion, region and anchor language, exact material and color instructionsThe composite is the main job — several references, precise placement, or a specific outfit color
GPT Image 2Natural-language editing that follows conversational instructions closelyYou want to describe the change in plain sentences and iterate by talking
Nano Banana 2Reference-guided edits with fast turnaroundYou want a cheap second interpretation of the same inputs before committing

A reasonable default: start with Seedream 5.0 Pro for anything where placement and identity both matter. If the composite is simple — one person, obvious spot, forgiving lighting — GPT Image 2 often gets there in fewer words.

Step 3: write the placement prompt

A prompt that works for this task has four parts, in this order:

  1. The instruction — add the person from the reference into the base photo.
  2. Placement and scale — where they stand, what they're doing, how large relative to a known object in the frame.
  3. Light and shadow matching — the direction and quality of light in the base photo, and where the shadow falls.
  4. Protection — everything in the base photo that must stay identical.

Here is a working version you can adapt:

Add the woman from the second reference photo into the first photo. Place her standing on the left side of the wooden deck, facing the camera, weight on her right leg, with her head roughly level with the top of the deck railing behind her. Match the base photo's lighting exactly: late-afternoon sun from the upper right, warm tone, soft-edged shadows. Cast her shadow down and to the left across the deck boards. Keep the deck, railing, plants, sky, and every other element of the base photo unchanged.

Notice that "her head roughly level with the top of the deck railing" does more for scale than "medium size" ever will. Anchoring the person to something already measurable in the frame is the single most useful scale trick in these prompts.

A seated variant, for a smaller and often easier composite:

Insert the man from the reference photos into the café scene, seated in the empty chair on the right side of the table, three-quarter view turned slightly toward the window, hands resting on the table. Scale him so his shoulders sit just below the top of the chair back. Light him from the window on the left with soft diffused daylight matching the base image, with a faint contact shadow where his forearms meet the table. Do not alter the table, cups, other people, or the background.

Step 4: refine with region-targeted edits

The first render is rarely final, and the mistake most people make is re-rolling the whole prompt. Don't. Take the output that got placement right and correct one thing at a time with a region-scoped instruction. Seedream 5.0 Pro is built for this — naming a region, an anchor point, or a location the way a person would read it out loud ("bottom left," "second from the right") keeps the edit local.

Using this image, adjust only the added figure: soften the edge along her right shoulder so it blends with the background wall, warm her skin tones slightly to match the surrounding light, and extend her shadow another twenty percent to the left. Change nothing else in the frame.

Two or three of these targeted passes usually beat ten full re-rolls, and they cost less. The Seedream 5.0 Pro prompt guide covers the region, anchor, and hex-color patterns in more depth if you want to push this further.

Note

OmniArt exposes standard image generation and reference-image generation. Coordinate boxes, sketch overlays, and anchor pins are upstream API workflows, not dedicated OmniArt UI controls — inside the workspace you express them as text in the prompt. That works, it just means precise language is doing the pointing.

Step 5: lock the final render

Once placement, identity, and lighting all hold, do a final pass at the resolution you actually need rather than continuing to edit a small draft. Re-render the approved composite at the largest size the model supports — Seedream 5.0 Lite offers native 4K if output resolution is the deciding factor — and check the result at 100% before you ship it. Edge halos and mismatched grain are invisible at thumbnail size and obvious at full size.

Matching lighting, perspective, scale, and shadow

These four failure modes account for most composites that look wrong without the viewer knowing why.

  • Light direction. Read the base photo's shadows first, then state the direction in the prompt ("from the upper right," "from the window on the left"). If the person's reference photo was lit from the opposite side, say so explicitly: "relight the figure from the right; the reference photo's left-side lighting should not carry over."
  • Light quality. Hard sun gives sharp-edged shadows; overcast and indoor light give soft ones. A soft-lit person in a hard-sun scene looks pasted. Name the quality, not just the direction.
  • Perspective and camera height. If the base photo was shot from waist height, a person referenced from a phone held at eye level will sit wrong in the frame. Add "match the base photo's low camera angle, viewed slightly from below" and the model will adjust the figure's foreshortening.
  • Scale. Always tie height to something in the frame — a doorway, a railing, a chair back, another person's shoulder. Absolute descriptions don't survive translation into pixels.
  • Contact shadow. The small dark area where a body meets a surface is what sells physical presence. Ask for it by name: "a soft contact shadow under both shoes where they meet the pavement."
  • Grain and color. Ask the model to match the base photo's grain, color temperature, and any lens softness. A clean figure in a grainy photo is a giveaway.

Troubleshooting

The face doesn't look like the person. Your reference is doing too little work. Upload additional angles of the same person, prefer higher resolution, and drop any heavily filtered shots. Then ask specifically: "preserve the facial features, hairline, and hair color from the reference photos without stylizing them."

The background changed. Your protection clause was too vague. List the elements explicitly instead of saying "keep the background." Then run a region-scoped correction pass rather than regenerating.

The person is floating. Missing contact shadow, or feet cropped ambiguously. Ask for the contact shadow directly and confirm the ground plane is visible where you placed them.

Wrong size. Replace any absolute size language with a comparison to an object already in the frame.

Edges look cut out. Ask for a softened boundary and matched grain along the figure's outline, and check whether the reference photo had a hard cutout background — white-background cutouts encourage cutout-looking results.

The pose is stiff. Describe weight distribution and hand position. "Weight on her right leg, one hand in her jacket pocket" produces a more natural stance than "standing naturally."

Warning

Adding a real, identifiable person to a photo they were not in creates an image that implies something untrue. Get that person's explicit permission before you generate, and before you publish. Do not composite public figures, colleagues, ex-partners, or children into scenes without consent. Do not create images implying attendance at an event, endorsement of a product, or presence in a place. Check the platform rules where you post — many require disclosure of AI-altered images of real people — and follow OmniArt's terms alongside any local law on likeness rights.

Getting started on OmniArt

Pick one base photo with an obvious empty space and one clean, well-lit reference photo of the person. Open the OmniArt image workspace, choose Seedream 5.0 Pro, upload both, and run the deck prompt from step 3 with your own details swapped in. Judge only placement and scale on the first render — identity and lighting are what the refinement passes are for.

Then work in small corrections instead of new prompts. Two region-scoped passes and one final full-resolution render is the normal path, and it keeps the base photo intact while the figure comes together.

Ready to Create?

Start generating amazing content with AI

Get started free