The Core Trick: Pinch Distance Is Not the Distance Between Your Fingers
The naive version measures the pixel distance between the thumb tip and the index tip and maps it onto the timeline. It feels broken immediately, and the reason is worth understanding: moving your hand toward the camera makes your fingers farther apart in pixels. So leaning forward scrubs the video. Every unrelated hand motion becomes input. The gesture never stabilizes.
The fix is to divide by something that scales the same way. MediaPipe gives 21 hand landmarks, so we use the hand itself as the ruler:
thumb = lm[4] # thumb tip
index = lm[8] # index tip
wrist = lm[0]
mid_mcp = lm[9] # base knuckle of the middle finger
hand_size = ||mid_mcp - wrist||
pinch = ||thumb - index|| / hand_size
wrist → middle_MCP is the palm's spine. It is rigid — it does not change when you open your fingers — and it grows and shrinks with distance from the camera at exactly the same rate the pinch gap does. Dividing one by the other cancels the depth term out. The measurement now means "how open are your fingers, as a fraction of your palm", which is what we actually wanted, and you can wave your whole arm around without touching the timeline.
This is a general pattern for gesture work: never use a raw distance, always use a ratio against a rigid part of the same body.
Calibration: The Constants You Hardcode Will Be Wrong
The first version mapped the normalized pinch from 0.12 (closed) to 0.85 (fully spread) onto the timeline. Reasonable-looking numbers. Then I measured an actual hand in front of an actual webcam for 90 frames:
hand detected in 80/90 frames
raw pinch: min 0.12 max 0.43 mean 0.35
maps to timeline pos: 0.00 .. 0.43
A fully spread hand reached 0.43, not 0.85. Which means the back 57% of every video was physically unreachable — you could open your fingers as wide as they go and the sunflower would never finish blooming. The gesture was not broken; the range was a guess, and the guess was off by a factor of two.
Guessing better numbers just moves the problem to the next hand and the next camera angle. So the range calibrates itself:
class PinchRange:
def update(self, raw, dt):
self.lo = min(self.lo, raw) # widen instantly for any new extreme
self.hi = max(self.hi, raw)
pull = RELAX * dt * (self.hi - self.lo)
self.lo, self.hi = self.lo + pull, self.hi - pull # relax slowly back in
return clip((raw - self.lo) / (self.hi - self.lo), 0, 1)
Widen instantly, relax slowly. Any new extreme is admitted on the frame it happens, so the first time you fully spread your fingers the far end of the timeline becomes reachable. Meanwhile the range creeps inward at 2% per second, so if you settle into a smaller comfortable motion the mapping re-tightens around it and you get full travel without full extension. A RANGE_MIN floor stops it from collapsing to a point when your hand is still.
The practical upshot: wiggle open-and-closed once at launch and it is tuned to you. Press r to reset if you hand it to someone else.
Two Ways to Map a Gesture Onto a Timeline
These feel completely different and both are worth having.
Absolute — your hand is the scrubber
Absolute mapping puts pinch distance straight onto the timeline. Fingers touching is frame one, fully spread is the last frame. It is a direct positional grip on time. This is the magic one, and it is best on short clips, because the entire duration has to fit inside a few centimetres of finger travel.
target = pinch * (n - 1)
pos += (target - pos) * 0.5 # ease toward it, don't snap
Velocity — pinch controls speed, not position
Velocity mapping has a neutral pose with a deadzone around it. Spread past it and you fast-forward, close past it and you rewind, and the further you go the faster it runs.
delta = pinch - neutral
if abs(delta) > DEADZONE:
mag = (abs(delta) - DEADZONE) / (1 - DEADZONE)
pos += sign(delta) * mag**1.6 * MAX_FPS_SCRUB * dt
The **1.6 exponent matters more than it looks. A linear ramp makes slow, precise scrubbing nearly impossible, because the region right outside the deadzone is already moving fast. The exponent flattens the response near neutral — fine control where you need it — and lets it climb steeply at the extremes for fast travel. The deadzone itself exists because a hand is never perfectly still, and without it the video drifts constantly.
Absolute for short clips and demos, velocity for anything over a minute. m toggles between them; c recalibrates neutral to wherever your hand is now.
Why the Whole Video Lives in RAM
Scrubbing backward through an H.264 file is miserable. Inter-frame compression means frame N is stored as a diff against earlier frames, so seeking backward one frame makes the decoder jump to the previous keyframe and decode forward again. Do that every display frame and you get a stuttering mess exactly when the gesture should feel most fluid.
So the clip is decoded once, up front, into a list of NumPy arrays. After that, "go to frame N" is a list index — and crucially, backward costs exactly what forward costs. That symmetry is the entire feel of the thing.
The price is memory, and it is steep: one 640×360 BGR frame is 691 KB, so 900 frames is 622 MB. MAX_FRAMES caps it, and longer clips get evenly subsampled (keep every k-th frame) rather than truncated — you would rather have the whole video slightly coarser than the first third at full rate.
There is a subtle trap here I fell into. The loader originally normalized every frame to a 720px height — which upscaled the 360p test clips, turning a 622 MB buffer into a 2.5 GB one to store interpolated pixels containing no additional information. Storage resolution and display resolution are different concerns:
if h > target_h: # only ever downscale; upscaling just wastes RAM
Frames are stored at native resolution and upscaled per-frame at display time, which is one cheap resize of one image.
Four Bugs, and What They Have in Common
Every one of these was invisible in the code and obvious the moment something real ran.
1. The webcam's first frame is always empty
On macOS, the first cap.read() after opening a capture device returns False. The main loop said if not ok: break — the textbook idiom — so the app exited instantly on launch, every time, before drawing a single frame. It now warms the camera up with a few discarded reads and tolerates transient drops mid-session, giving up only after ~2 seconds of silence.
2. The overlays were sized in pixels
The webcam inset was hardcoded to 180px tall. The camera ignores the 640×480 request and hands back 16:9, so the inset computed to 320×180 — on a 320×240 video. It covered the entire width and 75% of the height; the HUD ate another 92px of the remaining 60. The video was completely buried, and the symptom reported was "I just see myself and not the video." Nothing was wrong with the video pipeline at all. Everything is now a fraction of the canvas: the inset is 22% of canvas width, and the HUD's band, font, padding and timeline all scale off canvas height.
3. A negative slice index that didn't crash
The same inset math computed its x-origin as canvas_width - pip_width - 16 = 320 - 320 - 16 = -16. NumPy happily interprets that as an offset from the right edge and produces a silently wrong (often empty) slice instead of raising. draw_pip now bails out if the inset cannot fit.
4. The unreachable back half of the timeline
Described above — the hardcoded 0.85 ceiling against a real maximum of 0.43.
The common thread: none of these are logic errors you would catch by rereading the file. They are all collisions between the code and the messy specifics of a real camera, a real screen, and a real hand. The headless smoke test — stubbing imshow/waitKey and driving the actual loop against the actual webcam — caught three of them in about a minute, and is by far the highest-leverage thing in the repo.
Getting Test Footage
Timelapses are the ideal material, and yt-dlp is the tool. As of now the defaults fail with HTTP Error 403 — YouTube requires solving a JS challenge, and yt-dlp needs both a JS runtime and a separately-distributed solver script:
yt-dlp --js-runtimes node --remote-components ejs:github \
--extractor-args "youtube:player_client=web_safari,tv,mweb" \
-f "bv*[height<=720]+ba/b[height<=720]" --merge-output-format mp4 \
-o "%(title).60s.%(ext)s" --restrict-filenames "<url>"
Forcing player_client matters: the default android_vr client's URLs get rejected outright. Even then the high-res streams demand a PO token and it falls back to format 18 (360p combined). For this project that is not a real loss — 360p keeps the frame buffer small, and the footage gets upscaled for display anyway.
Running It
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
./run.sh # first clip in ./videos
./run.sh videos/explosion.mp4
./run.sh videos/mushroom.mp4 --mode velocity
./run.sh videos/rocket.mp4 --max-frames 1800 --display-height 1000
| Key | Action |
|---|---|
q / ESC | Quit |
m | Switch absolute ↔ velocity |
c | Recalibrate neutral (velocity mode) |
r | Reset the learned pinch range |
h | Show/hide the HUD scrubber (hidden by default) |
f | Flip the webcam horizontally |
SPACE | Freeze the timeline |
MediaPipe must be pinned to 0.10.21. Version 1.x removed the entire mp.solutions.* namespace, so mp.solutions.hands no longer exists. On macOS your terminal also needs Camera permission under System Settings → Privacy & Security.
Tuning Dials
All at the top of pinch_player.py:
| Constant | Default | What it does |
|---|---|---|
SMOOTH | 0.35 | Response speed. Lower is smoother and laggier |
MAX_FRAMES | 900 | RAM budget vs. scrub resolution |
DEADZONE | 0.06 | Velocity mode: how still "still" has to be |
MAX_FPS_SCRUB | 90.0 | Velocity mode: top scrub speed |
RELAX | 0.02 | How fast the learned range re-tightens |
PIP_FRAC | 0.22 | Webcam inset width, as a fraction of the canvas |
SMOOTH is the one to reach for first. Everything else is usually fine.
Why This Kind of Project Is Worth Building
Three hundred lines of Python, one model, one webcam — and almost all of the difficulty lives in the gap between what the code says and what the physical world does. The scale-invariant ratio, the self-widening range, the exponent on the velocity curve, the decision to burn 600 MB of RAM so that rewinding costs what fast-forwarding costs: none of those are visible in a feature list, and all of them are the difference between a gesture demo and a gesture that feels like a grip on time.
If you liked this, the same MediaPipe stack drives the gesture-controlled version of my portfolio and the 3D painting orb you rotate with your hands.