How I Made the "Dumb Dumb Dumb" Stacking Reel With Claude Code, in Full 4K HDR
Overview
You've probably seen this trend. The audio goes "don't be dumb, dumb, dumb, dumb", the video cuts every two seconds, and on every "dumb" a copy of something in the shot lands on top of the last one. A laptop becomes a tower of laptops. A keyboard fans out like a deck of cards. By the last shot, there are four of you.
Most people make it in CapCut with cut-outs they trace by hand. I gave Claude Code my raw iPhone clips and one prompt, and it did the whole edit: measured the beat timings from the reference Reel, cut out the objects, stacked the copies and encoded a 13-second master at the same 4K, 60fps HDR the phone shot. I added the official audio myself in Instagram's Edits app.
This post has the prompt I used (copy it below), what Claude did for each shot, and the rules I had to add after the first attempts looked wrong.
What it is: 7 shots, 780 frames, exactly 13.0 seconds at 60fps. In each shot one object (or me) gets 3 copies that snap in on the beat. The video stays live the whole time. No freeze frames.
Stack: Claude Code, ffmpeg, OpenCV, rembg (SAM ViT-B and u2net human segmentation as ONNX, on the CPU with no PyTorch) and x265. Output is 3840×2160, 10-bit HEVC with HLG HDR, about 65 MB, and has no audio.
The prompt
Paste in the link to the Reel you're copying, point Claude at a folder of 5 to 7 raw clips, and send this. The clips can be long and unedited. Claude picks the windows it needs.
Every rule in that list is there because an earlier version broke it. The section below on what went wrong explains each one.
Step 1: measuring the reference
Claude opened the reference Reel in my browser and pulled the video file out of the page. It then did two things with it:
- Found the cuts by comparing each frame with the previous one at thumbnail size. A big jump is a cut. A small jump is a copy landing.
- Found the beats by decoding the audio and picking out the sharp onsets, which are the "dumb" hits.
The result: about 120 BPM, a hit every 0.5 seconds, and cuts at 0, 1.55, 3.58, 5.58, 7.58, 9.54, 11.04 and 13.0 seconds. At 60fps that's shots of 93, 122, 120, 120, 117, 90 and 118 frames, and I checked that the finished shots have exactly those lengths. Inside each shot the copies land at roughly frames 30, 60 and 88, which is 0.5, 1.0 and 1.5 seconds in.
If you attach your own copy of the song, Claude lines up its loudness curve with the reference's to find where the song should start. That gives you the offset to use when you add the official audio in Edits.
Step 2: the plan for each shot
Before building anything, Claude made a contact sheet of every raw clip and proposed an object and a direction for each slot, with a few options where it wasn't sure. One thing I learned: a monitor on its own reads as vague, so it works better with something happening in front of it. This is what shipped:
| # | Frames | What gets copied | Direction | How it was cut out |
|---|---|---|---|---|
| 1 | 93 | The open laptop on its stand | Straight up into a tower | Rigid-object tracking |
| 2 | 122 | A potted plant, held in my hand | Up and to the side, curving | Hand, arm and pot tracked together |
| 3 | 120 | A paper lantern, as I reach past it | Straight up | Rigid-object tracking, my arm kept on top |
| 4 | 120 | The keyboard and the laptop, while I solve a mirror cube | Diagonally away from the centre | Clean plate, my hands and the cube live on top |
| 5 | 117 | Keyboard and trackpad, top-down | A fan around the keyboard's left end | Rigid-object tracking |
| 6 | 90 | The monitor, while I put my headphones on | Straight up, behind me | Clean plate, me live on top |
| 7 | 118 | Me, leaning back with my hands behind my head | Stacked up the frame | Human segmentation, live copies of me |
Step 3: cutting things out
Everything ran on my laptop's CPU. There's no green screen and no manual rotoscoping. Claude used three approaches, depending on the shot.
Rigid objects: segment a few frames, track the rest
For the laptop, lantern and keyboard, it ran SAM (Meta's Segment Anything model) on a keyframe every 3 or 4 frames and used optical flow to carry the mask across the frames in between. Each new SAM result was checked against the tracked mask. If the two overlapped by less than 75%, SAM had probably wandered onto the desk, so it kept the tracked mask instead.
Hands holding things: track the points, not the mask
For the plant, the mask kept creeping into the wood of the desk. What worked was tracking the click points SAM was given (points on my hand, my watch and the pot, plus "not this" points on the desk) and running SAM fresh from those points every other frame. The held object gets its own box. Without a point on the watch strap, SAM would cut the watch off my wrist.
Clean plate: the most reliable trick
If a clip has a moment where the desk is clear of hands, that frame is a "clean plate". Claude cut the objects out of that one frame and, for every other frame, worked out how the camera had moved since (matching features between the two frames) and warped the cut-out to fit. The copies come from the plate, so hands passing in front never damage them.
On top of that goes the live foreground: me, found by a person segmentation model plus anything that differs from the plate. If the foreground mask is slightly too big, nothing goes wrong, because it just shows more of the original frame. That's what finally made the cube shot and the headphones shot look right. Both copied objects are static, so a copy from the plate looks exactly like a live one.
Step 4: compositing without losing quality
Every frame is decoded to 16-bit RGB at full 4K and composited while still in HLG, so the HDR is never squashed. Copies are drawn from the newest to the oldest, each one moved by its own shift or rotation. Then the parts that must stay in front, meaning the original object and, in the shots I'm in, me and whatever I'm holding or sitting on, are painted back over the top. Each shot is encoded with x265 at 10-bit with the BT.2020 and HLG tags the iPhone uses. Because every shot uses identical settings, the final join just copies the streams and never re-encodes them.
Claude also made a 1080p SDR preview with my audio file lined up, so I could check the timing on my phone. The silent 4K file is the one I uploaded.
What went wrong first
The first versions didn't look like the trend. These are the problems and the rules they led to:
- Frozen copies. Copying one frame is the easy way to do it, but a still copy of a moving hand looks dead next to the live original. Now copies of anything that moves come from the live frame.
- Extra effects. Early renders added a slow zoom, a drop shadow and a pop-in animation. The trend has none of those. Copies just appear on the beat.
- See-through copies. Soft mask edges made the copies look like ghosts. Masks are now made fully solid first, and the edges are refined after that.
- The mirror cube vanished. A mirrored Rubik's cube reflects the desk, so the model decided half of it was background. Filling the cube's outline solid fixed it.
- Missing watch and headphones. The person model doesn't count things you wear or hold as part of you. Comparing each frame with the clean plate picks them back up.
- Copies over my face. In the shots I'm in, the real me has to be in front of every copy, and so does anything touching me, like the chair or whatever I'm holding. Copies stacked upward go behind the original.
- Flicker in the person shot. Rotating around my body's centre moved the pivot a little every frame, which made the copies shake. A fixed pivot point stopped it.
That's why the prompt makes Claude check four frames of every shot itself before sending it, and send the shots one at a time. It's faster to reject one shot than to find the problem in the full edit.
What to know before you try it
- It's slow on a laptop. SAM takes around 8 seconds per frame on the CPU (on a 1080p working copy) and needs about 3.6 GB of RAM, so only one copy of it can run at a time. Budget a long session, not a quick edit.
- Shoot for it. Give each clip a second where the object is fully in frame, with nothing covering it, and ideally a moment of empty desk to use as a clean plate. Plain backgrounds behind the object help a lot.
- Mixed frame rates. Some of my clips were 60fps and some were 59.94. The output is a steady 60, which is fine for a Reel.
- The audio is on you. Upload the silent master and add the trending sound inside Instagram or Edits, so the Reel is linked to the original audio.
Recap
- The trend is 7 cuts in 13 seconds with a copy on every beat. Claude measured the exact timings from the reference: shots of 93 to 122 frames, copies at 0.5, 1.0 and 1.5 seconds into each.
- Objects are cut out with SAM plus tracking, or from a clean plate when the clip has an empty-desk moment. People are cut out with a person segmentation model plus a difference against the plate.
- Everything stays live 4K60 HDR. No freeze frames, zooms, shadows or filters, and the real me always stays in front of the copies.
- Copy the prompt above, add your reference link and your clips, and ask Claude to show you the plan before it builds.
More camera and video experiments: scrubbing video with a pinch, rain that splashes off your body, and turning yourself into Spider-Man in React.