← Learn · Build · 25 August 2026 · 8 min

Motion Gallery: A 3D Painting Orb You Rotate With Your Hands (MediaPipe + React Three Fiber)

Thirteen famous paintings on a Fibonacci sphere you spin by pinching the air. The object-fit: cover maths for Three.js textures, pinch-and-drag and two-hand-zoom gesture handling with MediaPipe, an alpha WebGL canvas over a live webcam, and the CSS fallback for when there is no WebGL context.

Overview

Motion Gallery is a browser-based art experience: thirteen famous paintings arranged around a slowly rotating sphere, which you spin by dragging with a mouse — or by pinching the air in front of your webcam. Nothing about the camera feed leaves the machine; hand detection runs locally in the browser through MediaPipe.

Stack: React + TypeScript on Vite, three with @react-three/fiber and @react-three/drei for the scene, @mediapipe/tasks-vision for hand tracking, Tailwind for layout, lucide-react for the one icon, and hand-written CSS for the no-WebGL fallback.

The interface is one button

The finished screen is a full-bleed 3D orb and a single circular camera toggle in the top-right corner. No wordmark, no heading, no instructions, no bottom control bar, no tracking dots, no gesture labels, no status readout. When the camera is off, the orb sits on the dark gallery background. When it's on, the mirrored webcam feed fills the screen behind the paintings.

That restraint was the last and largest edit, not the first draft — more on the subtraction pass below.

The orb: a Fibonacci sphere, not a ring

The obvious way to place paintings in 3D is a circle at eye level, and it looks obvious. Instead the positions come from a Fibonacci-style spherical distribution: for each painting, derive a vertical spherical angle from its index, add a horizontal golden-angle rotation, and convert to X/Y/Z. Points land near-evenly over the whole sphere with no clustering at the poles and no visible seam. Each painting is then oriented to face outward from the centre.

Every artwork is the same physical object:

  • a 3 × 2 plane geometry
  • a Three.js physical material, rendered double-sided
  • a dark plane just behind it, acting as a frame

Lighting is a warm orange directional light, a cool cyan directional light from the opposite side, ambient fill, and a city-style environment map for reflections. The two-tone rig is what keeps a sphere of flat rectangles from reading as flat.

Eased state instead of direct state

Rotation and zoom are each stored twice — a target value and a current value. Input writes only to the target; every frame, useFrame interpolates the current value toward it. Dragging and gestures both go through the same two variables, which is why a jittery hand landmark doesn't produce a jittery orb: the easing is a low-pass filter that the gesture code gets for free. When nothing is driving the target, the orb keeps rotating slowly on its own.

The bit that actually took thought: not stretching the paintings

Every frame in the scene is 3:2. Real paintings are not. The Great Wave off Kanagawa is wide, The Birth of Venus is wider, The Creation of Adam is a different shape again — and mapping a texture straight onto the plane stretches all of them. On a Botticelli, stretching is not a rendering artefact, it's vandalism.

In CSS this is one line: object-fit: cover. In Three.js you implement it yourself, per texture, from four numbers:

  1. the source image's aspect ratio
  2. the destination frame's aspect ratio (always 1.5 here)
  3. how much of the texture should stay visible — horizontally if the source is wider than the frame, vertically if it's taller
  4. the repeat and offset needed to keep the visible window centred

Set texture.repeat to that fraction and texture.offset to half the remainder, and you have cover behaviour: original proportions intact, frame completely filled, the excess cropped evenly off both edges, nothing squashed. The CSS fallback gallery gets the same result the easy way, with the actual object-fit: cover.

Thirteen paintings, stored locally

The gallery started life with eleven photographs of cars. Swapping the subject for public-domain art is what turned a tech demo into something worth looking at:

  • The Starry Night
  • The Last Supper
  • The School of Athens
  • The Birth of Venus
  • Liberty Leading the People
  • The Third of May 1808
  • The Fighting Temeraire
  • The Hay Wain
  • Impression, Sunrise
  • A Sunday Afternoon on the Island of La Grande Jatte
  • The Great Wave off Kanagawa
  • The Garden of Earthly Delights
  • The Creation of Adam

All thirteen files live in the repository under attached_assets/paintings/. Nothing is fetched from a remote image host at runtime, so the gallery can't be broken by someone else's hotlink policy.

Mouse controls, so a webcam is optional

Gesture control is the headline, not the requirement. Dragging anywhere on the scene maps pointer movement on X to horizontal orb rotation and movement on Y to vertical rotation. The scroll wheel moves the camera's Z position, clamped to a range that stops you from passing through the orb, drifting infinitely far away, or landing anywhere the camera becomes unstable.

Camera plumbing

The webcam is a full-screen HTML <video> element, mounted at all times so the tracking hook always holds a valid reference — mounting it on demand is how you end up with a hook racing a ref that is still null. It sits behind the Three.js canvas, uses object-fit: cover, is mirrored with scaleX(-1) so movement matches the viewer's intuition, and is muted with playsInline. Invisible while disabled, fully visible while enabled.

The canvas above it requests an alpha-enabled WebGL context. That single flag is what lets the video show through between the paintings.

Pressing the toggle:

  1. requests webcam permission
  2. attaches the stream to the full-screen video
  3. initialises hand recognition
  4. reveals the video behind the orb

Pressing it again stops the stream and the tracking loop. The button swaps between a camera and a camera-off icon and carries an accessible label, but shows no visible text.

Gestures

MediaPipe Tasks Vision loads only after the user enables the camera, and runs on the CPU delegate deliberately — the GPU is already busy rendering the orb, and competing for it costs more frames than the faster inference wins back.

Pinch and drag

Each frame, measure the distance between the relevant fingertip landmarks. When it falls below the grab threshold, store the hand position; on subsequent frames, feed the delta from the stored position into horizontal and vertical rotation, exactly as a mouse drag would. Releasing the pinch clears the stored position. Because it reuses the drag path rather than paralleling it, the orb behaves like a physical object you've grabbed.

Two-hand zoom

With both hands visible, measure the distance between the two hand centres and compare it to the previous frame. Hands apart zooms in, hands together zooms out, clamped to the same safe range as the scroll wheel — one clamp, two inputs.

All of this is invisible. There are no tracking circles, hand counts, gesture names, or telemetry anywhere on screen, and no backend endpoint receives a single frame.

Two passes of taking things away

The version before this one had a Spiral mode, a dark translucent scrim over the camera with a saturation filter and a vignette, tracking crosshairs, hand-position markers, live gesture names, a Sensor Status panel with Standby and Tracking Active states, corner decoration lines, Motion Gallery branding, two rows of instructions, and a hardware-acceleration warning.

The scrim existed to keep overlaid text readable. Once the text was gone, so was the reason for the scrim — and the camera now shows at full brightness with the painting orb rendering straight over it. The dark vignette stays for the camera-off state, where it's doing real work as a gallery wall rather than fighting a video.

Spiral mode came out completely: its state, its button, its geometry, its scroll maths, its animation loop, its hand and mouse handlers, its fallback rendering, its CSS, and its keyframes. Half-removing a mode is how you get dead branches that break the next refactor. One experience, one code path.

Framer Motion went with the animated panels it was animating. The remaining transitions are Tailwind opacity classes.

When there is no WebGL context

Some preview environments and older devices can't reliably create one. Rather than mounting the canvas and hoping, the app tests for WebGL first and, if it's missing, renders a CSS orb instead: the same thirteen paintings, correct proportions via object-fit: cover, rotating on a CSS animation, with the camera still able to show through behind it. A React error boundary around the Three.js scene catches anything that fails after mount and switches to the same fallback.

Notably, the fallback displays no warning text. A visitor on an old laptop gets a slightly simpler gallery, not an apology.

Where things live

  • src/pages/gallery-page.tsx — page composition: the full-screen video, the gallery, the toggle, and the two visual states
  • src/components/gallery-3d.tsx — WebGL detection, the canvas, orb geometry and distribution, texture cropping, lighting, mouse and hand controls, smoothing, the static fallback, and the error boundary
  • src/hooks/use-hand-tracking.ts — stream acquisition, MediaPipe setup, per-frame recognition, pinch distance, per-hand positions, start/stop, cleanup
  • src/components/camera-status.tsx — now just the toggle button. The filename is a fossil of the telemetry panel it used to be
  • src/index.css — theme, dark gallery background, fallback orb, painting cropping, rotation keyframes
  • attached_assets/paintings/ — the thirteen local images

What I verified, and what I couldn't

TypeScript passes, Vite starts, the workflow runs, and the page loads clean. All thirteen paintings appear, mouse drag rotates, wheel zoom responds, the crops hold their proportions, the video element covers, no Spiral code path survives anywhere in the tree, and the visible UI is the camera button and nothing else. The final layout was inspected at 1440 × 1000.

The honest gap: the automated browser I tested in has no physical camera, so no real webcam stream was ever displayed during verification. The stream acquisition, full-screen video layer, start/stop lifecycle, and MediaPipe activation path are all implemented and exercised up to the point where a device is required — but the gesture loop itself has only been verified by hand, on a laptop with a webcam.

Takeaways

  • Ease everything through a target value. Noisy gesture input needs no dedicated smoothing if the renderer is already interpolating toward a target.
  • Route new inputs into existing ones. Pinch-drag writes to the same rotation targets as the mouse and shares the mouse's zoom clamp — two inputs, one set of edge cases.
  • Give the CPU delegate to MediaPipe when WebGL owns the GPU. Faster inference isn't a win if you pay for it in frames.
  • Re-derive the texture crop yourself. object-fit: cover has no Three.js equivalent; it's four numbers, and skipping it stretches the art.
  • Remove features, don't disable them. A mode with its state and handlers left in place is a trap for the next change.

Try it: gesture-controlled-3-d-gallery.replit.app. Enable the camera, pinch, and pull the sphere around.