Mood clips are the AI cat videos where nothing is supposed to happen, and that is exactly what makes them hard to prompt: a request full of words like cozy and peaceful comes back as a flat picture of a cat, with no rain you can hear and no reason to keep watching. This article builds one such clip in CapCut at 16:9, two kittens in a sleeping bag inside a tent on a rainy night, ten seconds long, with one thunderclap and the sound to go with it. It starts at the only moment in the clip where something happens, then works outward to the quiet around it, the light, the sound the clip arrived with, and the rain that was added to it in the web editor.
Seven seconds in, the tent lights up
At 7.2 seconds the tent walls go blue-white for a fraction of a second, a bolt shows through the mesh window, and the two kittens flinch and press together. That is the whole event, and it was scheduled in the video prompt with one plain sentence about time: nothing for six seconds, then one flash.
The camera does not move at all for the whole clip. Rain keeps running down the outside of the tent fabric in drops and streaks, the lantern flame flickers slightly, and the two kittens breathe slowly with their eyes half open. Nothing else happens for the first six seconds. Then one bright blue-white lightning flash lights up the tent walls from outside for a fraction of a second, both kittens flinch, and they press together side by side in the sleeping bag, ears back, and stay pressed together, still, for the rest of the clip. The warm lantern light stays warm, the outside stays cool blue-gray. No cut, no zoom, no pan.
Frame brightness and the loudness of the clip's own soundtrack, measured from the downloaded file. The flash and the thunder land within a tenth of a second of each other.
Measured from the downloaded frames, the average brightness of the picture sits at about 60 on a 0 to 255 scale for the first seven seconds, jumps to 117 at frame 172, which is 7.2 seconds at 24 frames per second, flickers twice, and falls back by 8.1 seconds. The prompt said six seconds; the clip delivered a little after seven. The reaction follows the light: by 8.8 seconds the kittens are pressed side by side, and in the last frame their eyes are closed again. One sentence with a number in it was enough to put a single event late in a ten-second clip in this run, and the number was treated as approximate rather than exact.
The six quiet seconds are the clip
Everything before the flash is what a viewer actually receives from a mood clip, so it has to move without doing anything. In this clip three things keep moving: rain runs down the outside of the fabric on the right wall, the lantern flame flickers, and the kittens breathe and blink. Their eyes are closed by 1.2 seconds, open again at about 5 seconds, and the camera does not move at all. Counted as change between consecutive frames, seconds one through six each average between 0.5 and 0.7 on a scale where the flash second scores 9, which is the number a mood clip should look like: not zero, and nowhere near an event.
The finished clip, exported from the web editor with the library rain added under the clip's own soundtrack. Turn the sound on; the page mutes it by default. Shown here as frames in sequence; the full video and a GIF version are supplied alongside this document.
Nine frames at equal intervals. The lantern, the rain on the right wall and the kittens' eyes are the only things that change until the seventh second.
The clip was made from a still rather than from text alone, so the room could be checked for one credit before ten seconds of it were paid for. The still went through the CapCut AI image generator flow in Video Studio's Image mode at Seedream 4.5, 16:9 and 2K, and came back at 2560 by 1440. The still was then attached to the composer for the CapCut AI video generator in Video mode, with the settings panel set to Video clip, Seedance 2.0 Mini, 720p, 16:9 and 10s, and the video prompt above.
The Video mode settings panel with 16:9 and 10s selected before the video prompt was sent.
The price line the chat prints before each request: 1 credit for the still, 80 credits for the ten-second clip at the settings above.
On 5 September 2026, at Image mode with Seedream 4.5, 16:9 and 2K, the chat printed "Generating 1 item will consume 1 credit" before the still; at Video clip with Seedance 2.0 Mini, 16:9, 10s and 720p it printed "Generating 1 item will consume 80 credits" before the clip. Those lines are the only credit figures used here. Other durations, models, resolutions and modes are priced differently, and the balance shown in the page header did not track these lines during the session, so read the line before you send.
Warm on the left, cold on the right
When nothing moves, the picture's contrast has to come from light, and a tent at night gives you two sources for free: a lamp inside and the night outside. The still prompt spends two sentences on them and names a side for each.
Photoreal still, 16:9 landscape, night. The inside of a small camping tent seen from the foot end at the kittens' eye level, camera low and level. Two kittens, one gray tabby and one cream-colored, lie side by side inside an open dark-green sleeping bag in the center of the frame, heads out, eyes half open, not touching yet, a hand's width apart. The only light inside is a small warm lantern on the left, orange-yellow, lighting the sleeping bag and the kittens' faces from the left. The tent walls are thin fabric lit from outside by a cool blue-gray night, brighter on the right wall, with rain running down the outside of the fabric in visible drops and streaks. A mesh window at the back shows blurred dark pine trees and falling rain. Everything is still. No people, no text, no logos, no brand marks.
The average color inside each box. Left wall: red 68 points above blue. Right wall: blue 52 points above red. The two kittens sit on the line between them.
Averaging the pixels inside the two boxes puts a number on what the eye sees: the lantern wall is red 68 points above blue, the rain wall is blue 52 points above red, on the 0 to 255 scale. The line between them runs through the sleeping bag, which is why the kittens read as lit from one side rather than flatly. Two phrases did most of the work, "the only light inside is a small warm lantern on the left" and "lit from outside by a cool blue-gray night, brighter on the right wall". The word "only" is there to rule out a second interior light; in this run the prompt came back with one lantern and the split above, and nothing in the frame competes with it.
The AI cat video came back with its own sound
The downloaded clip carried an audio track. That is not obvious, since the canvas card plays silently by default, but the file has an AAC stereo stream alongside the video, and on this clip it is worth listening to. Measured in 50 millisecond windows, the track holds a quiet, steady bed at about minus 40 dBFS from the second second to the seventh, a rain-like hiss, then climbs to minus 10 dBFS at 7.3 seconds, a tenth of a second after the flash, and rolls off through the last three seconds like a thunderclap fading. The first second is louder than the bed, at about minus 25 dBFS, a short settling sound before the rain takes over.
Two things follow from this. The generator wrote a soundtrack that matched its own picture in this run, with the loud low sound landing where the light did, so the one event in the clip already has its sound. And the rain in that soundtrack is faint, about minus 40 dBFS, some 30 dB under its own thunder. Earlier clips in this series had audio streams too; they went unchecked because the pages that showed them were silent. Check yours with any player that shows a volume meter before deciding what to add.
Add the rain from the library, and then measure it
The web editor's Audio panel has two tabs, Music and Sound effects. Searching Sound effects for "rain" listed a page of rain beds: a 3 minute 9 second track named "rain", a one-minute "RAIN", a 17-minute recording labeled with a polycarbonate roof, and more. Searching the same tab for "thunder", "storm" and "lightning" returned "No search results yet" in this session, which is one more reason to keep the thunder the clip came with. Hovering a result shows a play button and a blue plus labeled "Add to timeline".
Audio, then Sound effects, then a search for rain. The plus on a hovered row adds the effect to the timeline at the playhead.
Three details of the add step cost time in this session and are worth knowing in advance. The effect lands at the playhead, so park the playhead at zero before pressing the plus, or the rain starts wherever the playhead happened to be. The effect keeps its full length, so a 3 minute 9 second rain makes the project 3 minutes 19 seconds long until it is trimmed: move the playhead to the end of the video, click the rain clip to select it, click the Split icon at the left end of the timeline toolbar, click the tail and press Delete. The keyboard shortcut for Split did not take on the audio clip in this session; the toolbar icon did. And clicking an audio clip selects it without moving the playhead, while clicking empty timeline moves the playhead, so the order above matters.
The trimmed rain track under the video, and the audio clip's Basic panel where Volume was typed in.
Selecting the rain clip opens a Basic panel on the right with a Volume field in dB, a Fade in and out section with fade-in and fade-out durations, a Noise reduction toggle marked with a gem, and an Enhance voice entry. The first export used Volume at minus 10 dB, which is what a background layer usually gets, and it made no audible difference: a test export with the video track muted measured the effect alone at about minus 49 dBFS, quieter than the rain the clip already had. This library file is a quiet recording. Typed as 10, the field shows 10 dB, and the export with that setting raised the rain bed between seconds two and six from minus 37 dBFS to minus 31 dBFS while the thunder window stayed at minus 14 dBFS, still the loudest thing in the clip. Export settings were left at 720p, Recommended quality, 30fps and MP4; the file came back at 1280 by 720 and 10.07 seconds.
The number to carry away is not 10 dB, which belongs to this one file, but the method: after an export, measure the bed and the event in any audio tool that reads dBFS, and move the library layer until the bed sits under the event by a margin you can hear, here about 17 dB.
Check the first second before anything else
A mood clip has one second in which the viewer decides, and it is the first, not the seventh. In this clip the first second is the busiest of the quiet ones: the frame-to-frame change is 1.1 against 0.5 to 0.7 for the seconds that follow, the kittens' eyes are open and about to close, the rain streaks are already on the right wall, and the soundtrack's first second is its loudest passage outside the thunder. None of that was scheduled; it is what this run happened to put there, and it is what makes the clip start rather than begin.
So the last check is the first second, played with sound, before the light split or the thunder timing is judged. If it is still and silent, no amount of editing later in the clip fixes it, and the right move is a new generation with a first-second motion written into the prompt, a drop sliding down the window or the lantern flaring, rather than a longer edit. What this article does not cover: camping gear, outdoor safety, music, color grading, and the Noise reduction toggle, which carries a gem badge and was left off.
Written 5 September 2026. The kittens, the tent and the scene are generated images and video from CapCut made for this article; no real animals, products or brands appear. Interface labels, panel contents, search results, credit lines and levels reflect a single session on that date and the configurations named above, and may change. Loudness figures are RMS in dBFS measured from the downloaded files with a standard audio library; they describe these files, not the editor's meters.