Neither of these setups can be filmed by one person with a phone: one needs a moon the size of a building and a second silhouette to dance with, the other needs rain, a lamp behind it and a street with nobody on it. Both were built in CapCut at 16:9 from generated stills, and both turned out to depend on the same thing, which is where the light sits relative to the people. The first setup finished as a five-second intercut. The second stopped at a still, for a reason documented at the end.
Setup one: a solo dance in a room, the same hold as two silhouettes under a red moon, and back. Shown here as frames in sequence; the full video and a GIF version are supplied alongside this document.
Three shots, two cuts, one moon
The finished piece has this shape. Each shot is listed with what the viewer is given and what the picture takes away, because the cut only works if the second column changes and the first does not.
Everything below is about producing a middle shot that survives having that much removed, and a first shot that makes the same shape as it.
A silhouette is judged with the detail removed
A silhouette shot has no faces, no fabric and no interior lines, so the pose is carrying the entire shot on its outline. The way to know whether it can is to remove the detail before the model does: take the still, turn it into pure black and white at one threshold, and look at what remains. Two versions of the duet were generated for this, one with the dancers in profile with their limbs apart and one with them facing the camera in an embrace.
The two-tone test on both stills. Limbs apart in profile survive it; an embrace facing the camera becomes one shape with two heads.
The test settles the pose before any credits go on video. In profile, with sky visible between the arms and between the two bodies, the figures stay readable as two people and the arms stay readable as arms. Facing the camera and pressed together, the same two dancers become a single column against the moon, and no amount of drama in the moon fixes that. The still that passed is the one the clip was made from.
Put the moon where the light comes from
A silhouette is what a camera sees when the only light is behind the subject. The moon in this setup is that light, so it has to sit behind the dancers and low enough to fill the frame around them, not overhead where a moon usually is. The prompt places it there, asks for no detail inside the figures, and describes the pose in terms of the gaps.
Cinematic wide shot, 16:9 landscape. Two dancers in pure black silhouette on a flat rooftop at night, and a huge deep red full moon low on the horizon directly behind them, filling most of the frame. A man and a woman in profile to the camera in a ballroom hold, arms extended away from their bodies with clear sky visible between their bodies and between every arm and leg, so each limb reads as a separate shape against the moon. No detail inside the silhouettes, no faces, no clothing detail. Thin clouds crossing the moon, a dark city skyline far below. Camera at their level, fixed.
The still was generated in Video Studio's Image mode with Seedream 4.5 at 16:9 and 2K, through the CapCut AI image generator flow, and then animated from the home composer as a video request with the still attached, through the CapCut AI video generator in Video mode. The video prompt keeps the outlines separate and pins everything that is not a dancer.
The two silhouettes dance slowly: he leads her into a slow turn under his raised arm and they come back into the hold, their outlines staying separate against the moon the whole time with sky visible between their arms and bodies. They stay pure black silhouettes with no detail inside them. The moon, the clouds and the camera do not move. No zoom, no pan, no cut.
The silhouette clip across five seconds. The outlines stay apart until the last second, when the dance closes into the pose the test rejected.
Four of the five seconds held the shape. In the last second the choreography closed into an embrace, which is the pose the two-tone test had rejected, so the finished piece uses the clip only up to 3.8 seconds. Two-tone frames from the clip itself show the same thing the stills showed: at 1.5 seconds two figures, at 4.8 seconds two heads on one body.
The same test applied to two frames of the clip. The point where the outlines merge is the point to cut before.
The room has to make the same shape
The cut from a real room to a moon only reads as one dance if the person in the room is doing what the silhouette is doing. The solo still was written to match the silhouette woman: in profile, one arm resting on a partner who is not there, the other arm out, mid-turn, with one warm lamp and a plain wall so the outline stays clean even in color.
Photoreal medium-wide shot, 16:9 landscape. A young woman dances alone in a warm living room at night, in profile to the camera, mid-turn, one arm extended as if resting on an invisible partner's shoulder and the other held out to the side. A single floor lamp with a warm fabric shade lights the room, soft shadows on a plain wall behind her, a sofa and a rug at the edges of the frame. She wears a plain dark dress with no logos. Camera at chest height, fixed. Shallow depth of field.
The room clip, a single slow turn. It starts and ends in the pose the silhouettes hold, which is what lets the cut land on either side of it.
The video prompt asked for one slow turn on the spot, hair and hem following, ending facing the same way she started, with the lamp and camera fixed. The clip that came back does that: her back is to the lamp at three seconds and she faces the starting direction again by the end, so both cuts in the finished piece land on the hold rather than in the middle of the turn.
Cut the room away where the moon should show
Intercutting two clips does not need the clips rearranged on one track. In the web editor the room clip went on the main track, and the silhouette clip was dropped onto the timeline with the playhead at 00:01:08, which placed it on the track above, starting there. Splitting that upper clip and deleting the pieces where the room should show is the whole edit: the room is visible wherever the upper track has a gap, and the moon is visible wherever it does not. Here the upper clip runs from 00:01:08 to 00:03:23, and the room shows on either side of it.
The silhouette clip on the upper track, trimmed to the window where it should show. Nothing was moved on the main track.
Two things about the timing. The editor timeline in this project ran at 30 frames per second while both clips were 24, so the cut points sit on the timeline's grid at 00:01:08 and 00:03:23. And the upper clip's own first frame is the open hold, which is why dropping it at the playhead rather than at zero matters: its content starts where the room's first pose ends.
The finished piece sampled across its three shots. The hold carries across the first cut; the turn carries across the second.
On cost: each of the two clips in this setup was requested at Video clip, Seedance 2.0 Mini, 16:9, 5s and 720p, and the chat stated 40 credits before each on 5 September 2026. The stills, at Image mode with Seedream 4.5 at 16:9 and 2K, were stated at 1 credit each on the same date. Across the whole session the header balance fell from 1,406 to 1,320 for six stills and two finished clips, and the two requests described in the last section were charged and then returned. Other durations, models and modes are priced differently, so read the line before each send rather than reusing these figures.
The panel behind the composer chip, as set for both clips.
Rain shows where the light is behind it
The second setup removes color instead of detail, and it lives or dies on whether the rain is visible. Rain is visible when it is lit from behind, because each drop then becomes a bright streak against a dark street; lit from the camera side it disappears into the dark. Two stills were requested from the same description to check this, one with a street lamp placed behind the couple and one with the light asked for from the camera side.
Black and white photograph, 16:9 landscape. Heavy rain at night on an empty city street. A man and a woman stand close together in profile to the camera, foreheads almost touching, about to kiss, both soaked, hair flat with water. A single street lamp stands behind them and a little to one side, so the rain is lit from behind and every drop shows as a bright streak against the dark street, and the edges of their faces catch the lamp. Film grain, deep blacks, plain dark coats, no logos, no text anywhere.
Lamp behind them on the left, light requested from the camera side on the right, with the sky above the heads enlarged. The rain is in the first and mostly gone from the second.
The request for light from the camera side did not come back front-lit. What came back had no lamp and a low glow between the two bodies, so the light was still behind them, only lower and weaker, and the rain above their heads all but vanished. Measured on the band of sky above the heads, about five percent of that area is bright in the lamp version and under one percent in the other. Read the pair as one session's evidence, but the direction of the difference is the point: the lamp behind them is what makes the rain a subject rather than a texture.
The rain kiss as a still, at 2560 by 1440. This is where setup two stops.
The kiss that would not animate
That still went back into the composer twice as a video request, at the same settings as the duet. The first prompt asked the couple to close the distance and kiss while the rain kept streaking through the lamplight. The chat priced it at 40 credits, started, and then returned Couldn't generate. Credits returned. with a message that the prompt did not align with the safety and community guidelines, and a suggested rewrite without physical contact.
The first refusal. The credits shown as consumed before the request came back to the balance.
The second prompt removed the kiss and every contact verb: they stay forehead to forehead, eyes closed, and do not move apart, while the rain and the lamp do the work. It was refused in the same way, with the same credit return, and the chat offered three alternatives in its place: a romantic rainy scene, a couple standing in the rain, and a cinematic rain scene in black and white. The still had been generated without objection minutes earlier from a description of the same two people in the same position.
So the honest shape of setup two, in this session, is a finished still and no clip. If you take it further, the suggestions the tool made are the place to start, and they point the same way the rain measurement did: describe the street, the lamp and the rain, keep the two people where they are, and ask for the weather to move rather than the couple.
Written 5 September 2026. Every person in this article is AI-generated in CapCut and depicts no real individual; if you rebuild either setup with photographs of real people, use only your own and your partner's, with their agreement. Neither the rain nor the dance was filmed. Interface labels, timings, credit figures and the refusal messages reflect a single session on that date and the configurations named above, and may change. Music, beat matching and color grading are not covered here.