How to Make a Hand-Drawn Style Animated Scene

A step-by-step guide to creating a hand-drawn style animated scene in CapCut, using layered prompts to avoid a rendered look.

*No credit card required
AI image generator interface with a prompt box and style/category buttons
CapCut
CapCut
Sep 7, 2026

The usual complaint about a hand-drawn style animated scene from a generator is that it comes back rendered: soft volumetric clouds, a photographic sky, a glow on every edge. The prompt said "hand-drawn" and the tool heard "pretty". This article builds one such scene in CapCut at 16:9, a girl running across a sea of clouds at sunset, the way a background painter would build it: sky first, then clouds, then the figure, then the motion. Each layer gets the words that made it read as drawn in this run, next to the same layer written the other way, so you can see which words did the work.

The finished example appears in the last section, after an earlier attempt that preserved the style but changed the running motion. Both clips started from a still generated in CapCut during a session on 5 September 2026. A separate image used 3D-render wording for comparison. The prompts describe a character with short dark hair, a white shirt, and a navy skirt without naming an animation studio, film, or series.

Paint the sky first, and paint it in bands

For this scene, the sunset is described as horizontal bands of color with soft edges. Naming the colors from the horizon upward gives the generator a more specific direction than "hand-drawn sunset" alone. The following prompt was entered in Video Studio's Image mode; its sky sentence sets out that color sequence.

Hand-drawn animation background painting, 16:9 landscape. A sea of clouds at sunset seen from slightly above. The sky is painted in flat horizontal color bands: pale yellow at the horizon, then apricot, then rose pink, then lavender, then deep blue at the top of the frame, with soft edges between the bands and no photographic haze. The clouds are stacked rounded shapes with clear silhouettes, shaded in two tones only: warm cream on the sun side and lilac-gray on the shadow side, no volumetric glow, no fog. A girl with short dark hair, a white shirt and a knee-length navy skirt runs across the top of the nearest cloud from left to right. She is drawn small, about one fifth of the frame height, placed in the lower-left part of the frame, arms swinging, one foot off the cloud. Flat color fills, thin clean outlines on the character, visible brush texture in the sky, matte surfaces, no lens flare, no depth-of-field blur, no 3D render look. No text, no logos.

For the control image the scene, the girl and her placement stayed word for word and only the environment sentences were swapped: a "3D animated film render" with a "realistic sunset atmosphere with warm haze, the sky a smooth photographic gradient from yellow at the horizon to deep blue at the top", "volumetric clouds with soft light scattering through their edges, glowing rims", and "glossy highlights, shallow depth of field, cinematic lens flare from the sun". Both requests ran through the CapCut AI image generator flow at Seedream 4.5, 16:9 and 2K, and both came back at 2560 by 1440.

AI image generator interface with style tabs and sample thumbnails

The home composer in Image mode. Select 16:9 explicitly when you want a landscape still, and check the returned image before continuing.

Comparison of hand-drawn and 3D sky scenes with color-step and continuous-blend gradients

Each sky reduced to the average color of its rows, top of frame on the left, horizon on the right. The hand-drawn wording came back as steps; the 3D wording came back as one blend with a sun in it.

The row-color comparison shows a more segmented palette in the hand-drawn example and a smoother blend in the 3D-style example. This is a comparison of two outputs, not proof that one phrase will always produce the same effect. Listing the colors in order gives you a concrete starting point; review the generated sky and revise the description if the bands are too sharp, too blended, or in the wrong order.

Shade the clouds with two tones and a pen line

Clouds are where the render look usually gets back in, because volume is the thing renderers do well. The drawn prompt refuses volume twice, once by naming the shading ("two tones only: warm cream on the sun side and lilac-gray on the shadow side") and once by naming what is not wanted ("no volumetric glow, no fog"). The control prompt asks for the opposite on purpose.

Side-by-side clouds comparison: hand-drawn outline style versus glowing 3D render

The same region of the frame at full pixel size. Left: two fills, a dark edge line and short hatch strokes. Right: a glowing rim, fog between the banks, and no line anywhere.

At full size, the drawn example has simplified fills, contour lines, and short texture strokes. The rendered example has bright cloud rims and softer edges that fade into haze. These are different interpretations of the same scene, guided by different environment descriptions. To explore another look, change the style wording and compare the next result; suggested follow-up actions may also appear in your workspace, but their labels and availability can vary.

The girl retained a drawn appearance in both examples, even though the control used 3D-style environment wording. That result shows why you should inspect the character and background separately: changing the overall style prompt may not change every element equally. The next comparison focuses on how the figure's size changes the composition.

Keep the figure small enough that the sky stays the subject

The brief for a scene like this is a mood, not a portrait, and the size of the figure decides which one the viewer reads. The source still asks for "about one fifth of the frame height", and the girl measured 22 percent of the frame from hair to shoe. A third image was generated from the same prompt with one clause changed, "drawn large, filling about half of the frame height, centered in the frame", and she came back at 50 percent.

Animated girl running above clouds at different sizes, with two zoomed-in head close-ups below

Left: the source still, figure at 22 percent of frame height. Right: the same prompt with the figure at half the frame height. The head from each image is enlarged below.

At the smaller size, the face has fewer visible details and the sky remains the main subject. At the larger size, facial features and hair become more prominent, shifting the emphasis toward a character illustration. A small figure can make minor facial changes less noticeable, but it does not reduce the quoted generation cost at otherwise identical settings or guarantee fewer retries.

An earlier version placed the girl "on the lower-left third line," and an unwanted thin stroke appeared in the results. A later request used "placed in the lower-left part of the frame" and returned an image without that stroke. This suggests a useful wording change to try if a composition instruction produces an unwanted line, although a single retry does not establish the cause.

Side-by-side animation frames showing a girl on clouds, with a line drawn versus no line.

An unwanted line appeared in the earlier result. The revised prompt describes a region of the frame.

Use plain descriptions of position and size when you can. If a composition term produces an unwanted object or line, simplify it and check the next image rather than assuming the same wording will fail every time.

Set the clip to 16:9 and ten seconds before you write the motion

The corrected still was attached to the home composer, and the mode was switched to Video. In the session used here, the settings panel offered Video clip and Full video; the short scene was generated with Video clip, Seedance 2.0 Mini, 720p, 16:9, and 10s. Choose the equivalent options available in your workspace and check the settings before sending the request, since model and duration choices can vary.

Video mode settings with 16:9 aspect ratio and 10s duration selected

The Video mode settings panel with 16:9 and 10s selected. The chip above it repeats the choice as Seedance 2.0 Mini, 16:9, 10s, 720p.

CapCut interface showing generate images and generate videos panels with credits cost lines highlighted

The price line the chat prints before each request. Left, the still; right, the ten-second clip.

The session recorded a quoted cost of 1 credit for a Seedream 4.5 still at 16:9 and 2K, and 80 credits for a Seedance 2.0 Mini clip at 16:9, 10s, and 720p on 5 September 2026. These were the generation prompts shown for that session, not a current price guarantee. Check the cost displayed for your own request before generating; account-balance changes alone are not a reliable way to establish a per-item price.

Describe different motion for the foreground and background

To suggest depth, the prompt asks the near clouds to drift faster than the distant ones while the sky stays still. Here, "layers" means visual regions within one uploaded image, not separate editable layers or individual speed controls. The generator interprets those motion instructions, so review the result to see whether the foreground, background, and character move as intended.

The girl keeps running from left to right along the top of the nearest cloud, arms swinging, hair and skirt trailing behind her. The nearest cloud bank drifts slowly to the left under her feet, the far cloud banks drift more slowly, and the sky bands do not move. The camera pans slowly to the right so she stays in the left third of the frame. Colors stay flat, outlines stay thin and clean, no added glow, no haze, no lens flare, no camera shake, no cut.

Animated girl running across fluffy clouds under a pastel sunset sky

The first ten-second clip from the corrected still, made through the CapCut AI video generator at Video clip, Seedance 2.0 Mini, 16:9, 10s, 720p. Shown here as frames in sequence.

Girl runs across fluffy clouds at sunset in a hand-drawn style scene

Nine frames of the first clip at equal intervals. The bands, the two-tone clouds and the pen line survive the whole ten seconds.

The output was a 720p clip. In this example, the foreground clouds appeared to move faster than the distant banks, creating a sense of depth, while the banded sky and outlined clouds remained recognizable. The sky has little horizontal texture, so its movement is difficult to judge precisely. These observations describe the result of this request; they are not independently adjustable speed settings.

The clip has continuous motion, but differences between consecutive full frames do not establish whether the character is animated on ones or twos. Camera and background movement can change every frame even when a character pose is held. The source image also carried a small "Ai" mark that appeared in the generated clip in this session. Check the actual output for labels and their placement before adding text or other elements.

Make the running motion more specific

The first clip began with a running pose but shifted toward walking later in the scene. The next request used the same still and settings, with a more detailed description of the stride. The second prompt below is a useful alternative to try, although one comparison cannot prove why the first result changed pace or guarantee that the revised wording will prevent it.

The girl runs the whole clip at a steady running pace and never slows to a walk: knees lifted, both feet leaving the cloud between strides, arms pumping, hair and skirt trailing behind her. Cloud tops keep passing under her feet from right to left. The nearest cloud bank drifts left, the far cloud banks drift more slowly, and the sky bands do not move. The camera pans right at her pace so she stays in the left third of the frame. Colors stay flat, outlines stay thin and clean, no added glow, no haze, no lens flare, no camera shake, no cut.

Girl running across fluffy clouds at sunset with a pastel sky

The second ten-second clip, same still, same settings, motion sentence rewritten. This is the finished scene. Shown here as frames in sequence.

Girl running across fluffy clouds at sunset with a pastel sky

Nine frames of the second clip at equal intervals. The stride is a stride in each one, from the first frame to the last.

The sampled frames of the second example show a more consistent running pose, with the girl staying near the left third as the clouds pass beneath her. The camera and background also appear to move faster than in the first example. This result better matched the intended motion, but the difference cannot be attributed to a single phrase with certainty; another generation may behave differently.

The revised prompt specifies lifted knees, an airborne phase between strides, pumping arms, and clouds passing beneath the feet. Those details give you more precise instructions to test than "running" alone. If your result loses the intended action, revise the motion description and review the next clip while keeping the style wording consistent. This walkthrough ends with the generated clip; it does not cover sound, scene editing, or a separate workflow for creating held-frame animation timing.

This example was created on 5 September 2026 using prompts that do not name an animation studio, film, series, or existing character. Interface labels, available models, settings, and quoted credit costs may vary by account, region, and version. The described visual results apply to these generations and are not guaranteed outcomes. Sound, music, and editing after generation are not covered here.

Hot and trending