← Learn · Build · 28 August 2026 · 9 min

Pinch Time Travel: Scrubbing Video With Your Fingers in 300 Lines of Python

Spread your fingers and a timelapse runs forward; squeeze and it rewinds. How Avi Vashishta built a pinch-to-scrub video player with MediaPipe hand tracking and OpenCV — including the four bugs that only show up in front of a real camera.

TL;DR: Spread your thumb and index finger apart and a video runs forward. Squeeze them together and it rewinds. Point it at a timelapse — a sunflower opening over 83 days, a mushroom pushing out of the soil, a mushroom cloud blooming — and it stops feeling like operating a scrubber and starts feeling like pulling time itself back and forth. Built with MediaPipe hand tracking, OpenCV and NumPy in about 300 lines of Python. This post is the four places where the obvious implementation is quietly wrong.
Python MediaPipe 0.10.21 OpenCV NumPy Computer Vision yt-dlp

The Core Trick: Pinch Distance Is Not the Distance Between Your Fingers

The naive version measures the pixel distance between the thumb tip and the index tip and maps it onto the timeline. It feels broken immediately, and the reason is worth understanding: moving your hand toward the camera makes your fingers farther apart in pixels. So leaning forward scrubs the video. Every unrelated hand motion becomes input. The gesture never stabilizes.

The fix is to divide by something that scales the same way. MediaPipe gives 21 hand landmarks, so we use the hand itself as the ruler:

thumb   = lm[4]   # thumb tip
index   = lm[8]   # index tip
wrist   = lm[0]
mid_mcp = lm[9]   # base knuckle of the middle finger

hand_size = ||mid_mcp - wrist||
pinch     = ||thumb - index|| / hand_size

wrist → middle_MCP is the palm's spine. It is rigid — it does not change when you open your fingers — and it grows and shrinks with distance from the camera at exactly the same rate the pinch gap does. Dividing one by the other cancels the depth term out. The measurement now means "how open are your fingers, as a fraction of your palm", which is what we actually wanted, and you can wave your whole arm around without touching the timeline.

This is a general pattern for gesture work: never use a raw distance, always use a ratio against a rigid part of the same body.

Calibration: The Constants You Hardcode Will Be Wrong

The first version mapped the normalized pinch from 0.12 (closed) to 0.85 (fully spread) onto the timeline. Reasonable-looking numbers. Then I measured an actual hand in front of an actual webcam for 90 frames:

hand detected in 80/90 frames
raw pinch: min 0.12  max 0.43  mean 0.35
maps to timeline pos: 0.00 .. 0.43

A fully spread hand reached 0.43, not 0.85. Which means the back 57% of every video was physically unreachable — you could open your fingers as wide as they go and the sunflower would never finish blooming. The gesture was not broken; the range was a guess, and the guess was off by a factor of two.

Guessing better numbers just moves the problem to the next hand and the next camera angle. So the range calibrates itself:

class PinchRange:
    def update(self, raw, dt):
        self.lo = min(self.lo, raw)          # widen instantly for any new extreme
        self.hi = max(self.hi, raw)
        pull = RELAX * dt * (self.hi - self.lo)
        self.lo, self.hi = self.lo + pull, self.hi - pull   # relax slowly back in
        return clip((raw - self.lo) / (self.hi - self.lo), 0, 1)

Widen instantly, relax slowly. Any new extreme is admitted on the frame it happens, so the first time you fully spread your fingers the far end of the timeline becomes reachable. Meanwhile the range creeps inward at 2% per second, so if you settle into a smaller comfortable motion the mapping re-tightens around it and you get full travel without full extension. A RANGE_MIN floor stops it from collapsing to a point when your hand is still.

The practical upshot: wiggle open-and-closed once at launch and it is tuned to you. Press r to reset if you hand it to someone else.

Two Ways to Map a Gesture Onto a Timeline

These feel completely different and both are worth having.

Absolute — your hand is the scrubber

Absolute mapping puts pinch distance straight onto the timeline. Fingers touching is frame one, fully spread is the last frame. It is a direct positional grip on time. This is the magic one, and it is best on short clips, because the entire duration has to fit inside a few centimetres of finger travel.

target = pinch * (n - 1)
pos += (target - pos) * 0.5     # ease toward it, don't snap

Velocity — pinch controls speed, not position

Velocity mapping has a neutral pose with a deadzone around it. Spread past it and you fast-forward, close past it and you rewind, and the further you go the faster it runs.

delta = pinch - neutral
if abs(delta) > DEADZONE:
    mag = (abs(delta) - DEADZONE) / (1 - DEADZONE)
    pos += sign(delta) * mag**1.6 * MAX_FPS_SCRUB * dt

The **1.6 exponent matters more than it looks. A linear ramp makes slow, precise scrubbing nearly impossible, because the region right outside the deadzone is already moving fast. The exponent flattens the response near neutral — fine control where you need it — and lets it climb steeply at the extremes for fast travel. The deadzone itself exists because a hand is never perfectly still, and without it the video drifts constantly.

Absolute for short clips and demos, velocity for anything over a minute. m toggles between them; c recalibrates neutral to wherever your hand is now.

Why the Whole Video Lives in RAM

Scrubbing backward through an H.264 file is miserable. Inter-frame compression means frame N is stored as a diff against earlier frames, so seeking backward one frame makes the decoder jump to the previous keyframe and decode forward again. Do that every display frame and you get a stuttering mess exactly when the gesture should feel most fluid.

So the clip is decoded once, up front, into a list of NumPy arrays. After that, "go to frame N" is a list index — and crucially, backward costs exactly what forward costs. That symmetry is the entire feel of the thing.

The price is memory, and it is steep: one 640×360 BGR frame is 691 KB, so 900 frames is 622 MB. MAX_FRAMES caps it, and longer clips get evenly subsampled (keep every k-th frame) rather than truncated — you would rather have the whole video slightly coarser than the first third at full rate.

There is a subtle trap here I fell into. The loader originally normalized every frame to a 720px height — which upscaled the 360p test clips, turning a 622 MB buffer into a 2.5 GB one to store interpolated pixels containing no additional information. Storage resolution and display resolution are different concerns:

if h > target_h:            # only ever downscale; upscaling just wastes RAM

Frames are stored at native resolution and upscaled per-frame at display time, which is one cheap resize of one image.

Four Bugs, and What They Have in Common

Every one of these was invisible in the code and obvious the moment something real ran.

1. The webcam's first frame is always empty

On macOS, the first cap.read() after opening a capture device returns False. The main loop said if not ok: break — the textbook idiom — so the app exited instantly on launch, every time, before drawing a single frame. It now warms the camera up with a few discarded reads and tolerates transient drops mid-session, giving up only after ~2 seconds of silence.

2. The overlays were sized in pixels

The webcam inset was hardcoded to 180px tall. The camera ignores the 640×480 request and hands back 16:9, so the inset computed to 320×180 — on a 320×240 video. It covered the entire width and 75% of the height; the HUD ate another 92px of the remaining 60. The video was completely buried, and the symptom reported was "I just see myself and not the video." Nothing was wrong with the video pipeline at all. Everything is now a fraction of the canvas: the inset is 22% of canvas width, and the HUD's band, font, padding and timeline all scale off canvas height.

3. A negative slice index that didn't crash

The same inset math computed its x-origin as canvas_width - pip_width - 16 = 320 - 320 - 16 = -16. NumPy happily interprets that as an offset from the right edge and produces a silently wrong (often empty) slice instead of raising. draw_pip now bails out if the inset cannot fit.

4. The unreachable back half of the timeline

Described above — the hardcoded 0.85 ceiling against a real maximum of 0.43.

The common thread: none of these are logic errors you would catch by rereading the file. They are all collisions between the code and the messy specifics of a real camera, a real screen, and a real hand. The headless smoke test — stubbing imshow/waitKey and driving the actual loop against the actual webcam — caught three of them in about a minute, and is by far the highest-leverage thing in the repo.

Getting Test Footage

Timelapses are the ideal material, and yt-dlp is the tool. As of now the defaults fail with HTTP Error 403 — YouTube requires solving a JS challenge, and yt-dlp needs both a JS runtime and a separately-distributed solver script:

yt-dlp --js-runtimes node --remote-components ejs:github \
       --extractor-args "youtube:player_client=web_safari,tv,mweb" \
       -f "bv*[height<=720]+ba/b[height<=720]" --merge-output-format mp4 \
       -o "%(title).60s.%(ext)s" --restrict-filenames "<url>"

Forcing player_client matters: the default android_vr client's URLs get rejected outright. Even then the high-res streams demand a PO token and it falls back to format 18 (360p combined). For this project that is not a real loss — 360p keeps the frame buffer small, and the footage gets upscaled for display anyway.

Running It

python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
./run.sh                                    # first clip in ./videos
./run.sh videos/explosion.mp4
./run.sh videos/mushroom.mp4 --mode velocity
./run.sh videos/rocket.mp4 --max-frames 1800 --display-height 1000
KeyAction
q / ESCQuit
mSwitch absolute ↔ velocity
cRecalibrate neutral (velocity mode)
rReset the learned pinch range
hShow/hide the HUD scrubber (hidden by default)
fFlip the webcam horizontally
SPACEFreeze the timeline

MediaPipe must be pinned to 0.10.21. Version 1.x removed the entire mp.solutions.* namespace, so mp.solutions.hands no longer exists. On macOS your terminal also needs Camera permission under System Settings → Privacy & Security.

Tuning Dials

All at the top of pinch_player.py:

ConstantDefaultWhat it does
SMOOTH0.35Response speed. Lower is smoother and laggier
MAX_FRAMES900RAM budget vs. scrub resolution
DEADZONE0.06Velocity mode: how still "still" has to be
MAX_FPS_SCRUB90.0Velocity mode: top scrub speed
RELAX0.02How fast the learned range re-tightens
PIP_FRAC0.22Webcam inset width, as a fraction of the canvas

SMOOTH is the one to reach for first. Everything else is usually fine.

Why This Kind of Project Is Worth Building

Three hundred lines of Python, one model, one webcam — and almost all of the difficulty lives in the gap between what the code says and what the physical world does. The scale-invariant ratio, the self-widening range, the exponent on the velocity curve, the decision to burn 600 MB of RAM so that rewinding costs what fast-forwarding costs: none of those are visible in a feature list, and all of them are the difference between a gesture demo and a gesture that feels like a grip on time.

If you liked this, the same MediaPipe stack drives the gesture-controlled version of my portfolio and the 3D painting orb you rotate with your hands.