How I Made the "Dumb Dumb Dumb" Stacking Reel With Claude Code, in Full 4K HDR

Published September 2026 · avivashishta.com

Overview

You've probably seen this trend. The audio goes "don't be dumb, dumb, dumb, dumb", the video cuts every two seconds, and on every "dumb" a copy of something in the shot lands on top of the last one. A laptop becomes a tower of laptops. A keyboard fans out like a deck of cards. By the last shot, there are four of you.

Most people make it in CapCut with cut-outs they trace by hand. I gave Claude Code my raw iPhone clips and one prompt, and it did the whole edit: measured the beat timings from the reference Reel, cut out the objects, stacked the copies and encoded a 13-second master at the same 4K, 60fps HDR the phone shot. I added the official audio myself in Instagram's Edits app.

This post has the prompt I used (copy it below), what Claude did for each shot, and the rules I had to add after the first attempts looked wrong.

What it is: 7 shots, 780 frames, exactly 13.0 seconds at 60fps. In each shot one object (or me) gets 3 copies that snap in on the beat. The video stays live the whole time. No freeze frames.

Stack: Claude Code, ffmpeg, OpenCV, rembg (SAM ViT-B and u2net human segmentation as ONNX, on the CPU with no PyTorch) and x265. Output is 3840×2160, 10-bit HEVC with HLG HDR, about 65 MB, and has no audio.

The prompt

Paste in the link to the Reel you're copying, point Claude at a folder of 5 to 7 raw clips, and send this. The clips can be long and unedited. Claude picks the windows it needs.

I want to recreate the viral "stacking" Reel trend (like the "DON'T BE DUMB – dumb, dumb, dumb" audio) using my own clips. Reference reel: <paste Instagram link> My raw clips: in the folder I've connected (5–7 clips, each can be long/unedited). Audio: I'll add the official Instagram audio myself in Edits – give me a silent master, plus a preview with my audio file if I attach one. How the trend works: the video is ~13s, cut into ~2s shots on the beat. In each shot one object (or me) is cut out and, on every "dumb" hit (~every 0.5s), a new copy appears on top of the last. Copies either stack straight up/sideways or fan/curve around a pivot. The original stays where it is. Rules: - Keep everything as LIVE video. No freeze frames, and the copies of anything that moves must move too. - ONLY the stacking effect: no zoom, shadows, stabilization crop, filters or animated pop-ins. Copies snap in exactly on the beat. - Keep my original quality (4K 60fps HDR if that's what I shot). - Copies must be solid, never see-through. Cut-outs must be clean: no background or desk bits, and don't drop parts like my watch, headphones or a shiny/mirrored object. - When I'm in the shot, the real me (and anything I'm holding or sitting in) always stays in front of the copies. - If a clip has a moment where the desk/objects are empty, use that frame for the object copies and keep me live on top. - First watch the reference and measure its cut and beat timings, then match them exactly. - Show me which object and stacking direction you plan for each clip before you build. - Check every shot frame by frame yourself before sending it, and send shots one by one as they're done, then the full edit.

Every rule in that list is there because an earlier version broke it. The section below on what went wrong explains each one.

Step 1: measuring the reference

Claude opened the reference Reel in my browser and pulled the video file out of the page. It then did two things with it:

The result: about 120 BPM, a hit every 0.5 seconds, and cuts at 0, 1.55, 3.58, 5.58, 7.58, 9.54, 11.04 and 13.0 seconds. At 60fps that's shots of 93, 122, 120, 120, 117, 90 and 118 frames, and I checked that the finished shots have exactly those lengths. Inside each shot the copies land at roughly frames 30, 60 and 88, which is 0.5, 1.0 and 1.5 seconds in.

If you attach your own copy of the song, Claude lines up its loudness curve with the reference's to find where the song should start. That gives you the offset to use when you add the official audio in Edits.

Step 2: the plan for each shot

Before building anything, Claude made a contact sheet of every raw clip and proposed an object and a direction for each slot, with a few options where it wasn't sure. One thing I learned: a monitor on its own reads as vague, so it works better with something happening in front of it. This is what shipped:

#FramesWhat gets copiedDirectionHow it was cut out
193 The open laptop on its stand Straight up into a tower Rigid-object tracking
2122 A potted plant, held in my hand Up and to the side, curving Hand, arm and pot tracked together
3120 A paper lantern, as I reach past it Straight up Rigid-object tracking, my arm kept on top
4120 The keyboard and the laptop, while I solve a mirror cube Diagonally away from the centre Clean plate, my hands and the cube live on top
5117 Keyboard and trackpad, top-down A fan around the keyboard's left end Rigid-object tracking
690 The monitor, while I put my headphones on Straight up, behind me Clean plate, me live on top
7118 Me, leaning back with my hands behind my head Stacked up the frame Human segmentation, live copies of me

Step 3: cutting things out

Everything ran on my laptop's CPU. There's no green screen and no manual rotoscoping. Claude used three approaches, depending on the shot.

Rigid objects: segment a few frames, track the rest

For the laptop, lantern and keyboard, it ran SAM (Meta's Segment Anything model) on a keyframe every 3 or 4 frames and used optical flow to carry the mask across the frames in between. Each new SAM result was checked against the tracked mask. If the two overlapped by less than 75%, SAM had probably wandered onto the desk, so it kept the tracked mask instead.

Hands holding things: track the points, not the mask

For the plant, the mask kept creeping into the wood of the desk. What worked was tracking the click points SAM was given (points on my hand, my watch and the pot, plus "not this" points on the desk) and running SAM fresh from those points every other frame. The held object gets its own box. Without a point on the watch strap, SAM would cut the watch off my wrist.

Clean plate: the most reliable trick

If a clip has a moment where the desk is clear of hands, that frame is a "clean plate". Claude cut the objects out of that one frame and, for every other frame, worked out how the camera had moved since (matching features between the two frames) and warped the cut-out to fit. The copies come from the plate, so hands passing in front never damage them.

On top of that goes the live foreground: me, found by a person segmentation model plus anything that differs from the plate. If the foreground mask is slightly too big, nothing goes wrong, because it just shows more of the original frame. That's what finally made the cube shot and the headphones shot look right. Both copied objects are static, so a copy from the plate looks exactly like a live one.

Step 4: compositing without losing quality

Every frame is decoded to 16-bit RGB at full 4K and composited while still in HLG, so the HDR is never squashed. Copies are drawn from the newest to the oldest, each one moved by its own shift or rotation. Then the parts that must stay in front, meaning the original object and, in the shots I'm in, me and whatever I'm holding or sitting on, are painted back over the top. Each shot is encoded with x265 at 10-bit with the BT.2020 and HLG tags the iPhone uses. Because every shot uses identical settings, the final join just copies the streams and never re-encodes them.

Claude also made a 1080p SDR preview with my audio file lined up, so I could check the timing on my phone. The silent 4K file is the one I uploaded.

What went wrong first

The first versions didn't look like the trend. These are the problems and the rules they led to:

That's why the prompt makes Claude check four frames of every shot itself before sending it, and send the shots one at a time. It's faster to reject one shot than to find the problem in the full edit.

What to know before you try it

Recap

More camera and video experiments: scrubbing video with a pinch, rain that splashes off your body, and turning yourself into Spider-Man in React.

← Portfolio · Blog