Turning Yourself Into Spider-Man in React (MediaPipe Pose → a Mixamo Rig in React Three Fiber)

Published August 2026 · avivashishta.com

Overview

Step out of frame and press a key. Three-second countdown, and the app keeps a still photo of your empty room. Now walk back in — and you are not there. Spider-Man is standing exactly where you were, at your height, turning when you turn, tilting his head when you tilt yours, curling his fingers when you curl yours. Move fast and the suit moves fast. Throw a punch at the lens and the arm foreshortens at you.

No green screen. No depth camera. No suit with markers on it. One webcam, two MediaPipe models, and a rigged GLB, all running in a React 19 tab at 60 fps.

What it does: tracks 33 body landmarks and 21 per hand from the webcam, retargets them onto a Mixamo-rigged Spider-Man mesh, solves the character's on-screen size and position against your own body every frame, and composites him over a captured background plate so he replaces you rather than floating in front of you.

Stack: @mediapipe/tasks-vision (PoseLandmarker + HandLandmarker), three via @react-three/fiber, React 19, TypeScript, Vite. Zero backend — no frame ever leaves the machine.

The first wall: you cannot animate a Sketchfab embed

The character came from a Sketchfab model, and my first instinct was to drive the embed directly. That does not work, and it is worth knowing why before you waste an afternoon on it.

The Sketchfab Viewer API gives you a scene graph, node transforms (translate, rotate, setMatrix) and playback control over animation tracks that were baked into the file. It gives you no access whatsoever to skeleton joints, and the docs are explicit that only MatrixTransform nodes can be moved. You can play an animation somebody else authored. You cannot pose a rig.

So the model has to come down as a GLB and get loaded yourself. Which introduces the second wall.

368 MB of Spider-Man

The good version of the model was 368 MB. Not geometry — 632,794 triangles is about 40 MB of buffers. The other 328 MB was thirty PNG textures at 4096×4096, several of them over 20 MB each.

Two passes fixed it.

Textures. Decode every image out of the GLB's binary chunk, resize to 1024, re-encode as JPEG at quality 88 (WebP only where an alpha channel is actually in use), then rebuild the buffer with recomputed byteOffsets. 328 MB of PNG becomes about 5 MB. For a character that occupies maybe 600 px of screen height, 4K albedo maps are pure waste.

Geometry. glTF-Transform does the rest:

gltf-transform dedup in.glb t1.glb gltf-transform weld t1.glb t2.glb gltf-transform meshopt t2.glb out.glb --level medium

368 MB → 14 MB, same rig, same triangle count.

One thing to watch: meshopt re-partitioned the single shared skin into one skin per mesh (1 → 16). That is harmless — every skin still references the same bone nodes, so posing a bone moves all of them — but it looks alarming in an inspector. Loading it back needs the decoder wired up:

import { useLoader } from '@react-three/fiber' import { GLTFLoader } from 'three/examples/jsm/loaders/GLTFLoader.js' import { MeshoptDecoder } from 'three/examples/jsm/libs/meshopt_decoder.module.js' const gltf = useLoader(GLTFLoader, '/models/spiderman.glb', (loader) => { loader.setMeshoptDecoder(MeshoptDecoder) })

Reading the body: two landmarkers, one video element

@mediapipe/tasks-vision is already a dependency of this portfolio for its gesture controls, so the setup is familiar. Both models run in VIDEO mode against the same <video> element, on the same monotonically increasing timestamp:

// src/hooks/useBodyTracking.ts import { FilesetResolver, PoseLandmarker, HandLandmarker, } from '@mediapipe/tasks-vision' const WASM = 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@0.10.32/wasm' const MODELS = 'https://storage.googleapis.com/mediapipe-models' export async function createTrackers(delegate: 'GPU' | 'CPU' = 'GPU') { const files = await FilesetResolver.forVisionTasks(WASM) const pose = await PoseLandmarker.createFromOptions(files, { baseOptions: { modelAssetPath: `${MODELS}/pose_landmarker/pose_landmarker_full/float16/1/pose_landmarker_full.task`, delegate, }, runningMode: 'VIDEO', numPoses: 1, }) const hands = await HandLandmarker.createFromOptions(files, { baseOptions: { modelAssetPath: `${MODELS}/hand_landmarker/hand_landmarker/float16/1/hand_landmarker.task`, delegate, }, runningMode: 'VIDEO', numHands: 2, }) return { pose, hands } }

Always wrap that in a try and retry with delegate: 'CPU'. The GPU delegate fails on more machines than you would like, and the failure is a rejected promise, not a degraded result.

Pose gives you two arrays that matter. landmarks are normalised image coordinates — where the joint is on screen. Pose also provides worldLandmarks: metres, origin between the hips, axes parallel to the image. Those two get used for completely different jobs, and mixing them up is the root of a lot of bad mocap:

MediaPipe's axes are x-right, y-down, z-negative-toward-camera, so everything gets mapped once on the way in and never thought about again:

// MediaPipe → three.js. det = +1, so handedness is preserved. v.set(lm.x, -lm.y, -lm.z)

The retargeting core: aim, don't copy

The temptation is to place bones at landmark positions. Do not. Your arm and Spider-Man's arm are different lengths, and a skinned mesh with joints yanked to arbitrary positions tears. Bone rotations are the only thing you should ever write.

So the primitive is: rotate this bone so the direction toward its child points along the direction I measured on the human. Everything below is one function, applied down the hierarchy, parents before children.

// Precomputed once, in the rest pose, with the rig at identity: // restWorld: bone -> world quaternion // restChildDir: bone -> normalised world direction to its child // restLocal: bone -> local quaternion (to relax back to) function aimBone( bone: THREE.Bone, target: THREE.Vector3, // world-space direction, normalised weight = 1, ) { const parentNow = worldOf(bone.parent!) // resolved this frame const parentRest = restWorld.get(bone.parent!)! // how much the parent chain has already rotated since rest carry.copy(parentNow).multiply(q.copy(parentRest).invert()) // where the child direction ended up after that, then the correction dir.copy(restChildDir.get(bone)!).applyQuaternion(carry) const delta = new THREE.Quaternion().setFromUnitVectors(dir, target) const world = delta.multiply(carry).multiply(restWorld.get(bone)!) const local = q.copy(parentNow).invert().multiply(world) if (weight < 1) local.slerp(restLocal.get(bone)!, 1 - weight) bone.quaternion.slerp(local, responsiveness) }

Two details in there earn their keep. The weight lets a limb fade back to its rest pose when tracking confidence drops, instead of snapping or flailing. And after applying, you must record the bone's actual resulting world rotation — not the target you asked for — because smoothing means the two differ, and the children solve against where the parent really ended up.

The whole body is then a table:

const AIM = [ { bone: 'LeftArm', child: 'LeftForeArm', from: L_SHO, to: L_ELB }, { bone: 'LeftForeArm', child: 'LeftHand', from: L_ELB, to: L_WRI }, { bone: 'LeftUpLeg', child: 'LeftLeg', from: L_HIP, to: L_KNE }, { bone: 'LeftLeg', child: 'LeftFoot', from: L_KNE, to: L_ANK }, { bone: 'LeftFoot', child: 'LeftToeBase', from: L_ANK, to: L_FOOT }, // …and the right side ]

Turning around needs frames, not directions

A direction has no twist. Aim the spine bones from hips to shoulders and the torso will bend correctly and never turn — the whole upper body just inherits its facing from the pelvis. Which is why the first version stared straight at the camera no matter which way I sat.

The torso needs full three-axis frames. Build one from a lateral axis and an up axis, compare it against the same frame in the rig's rest pose, and you have an absolute orientation:

/** Orthonormal frame from a lateral axis and an up axis (columns: left, up, fwd). */ function basisFrom(left: THREE.Vector3, up: THREE.Vector3) { const u = up.clone().normalize() const f = left.clone().cross(u).normalize() // forward = left × up const l = u.clone().cross(f).normalize() // re-orthogonalised return new THREE.Matrix4().makeBasis(l, u, f) } const up = shoulderCentre.clone().sub(hipCentre) const shoulderLeft = lm[L_SHO].clone().sub(lm[R_SHO]) // delta = current frame × rest frame⁻¹ (rest frames are precomputed transposes) const dPelvis = new THREE.Quaternion().setFromRotationMatrix( m.copy(basisFrom(pelvisLeft, up)).multiply(restPelvisInv), ) const dChest = new THREE.Quaternion().setFromRotationMatrix( m.copy(basisFrom(shoulderLeft, up)).multiply(restChestInv), )

Then spread the twist between them along the spine instead of snapping it at one joint. Because both are corrections applied to a rest orientation, you can interpolate the deltas directly:

applyWorld(hips, dPelvis.clone().multiply(restWorld.get(hips)!)) for (const [bone, t] of [[spine, 0.34], [spine1, 0.67]] as const) applyWorld(bone, dPelvis.clone().slerp(dChest, t).multiply(restWorld.get(bone)!)) applyWorld(chest, dChest.clone().multiply(restWorld.get(chest)!))

The part nobody warns you about: if you sit at a desk, your hips are behind it. MediaPipe still reports hip landmarks, but they're guesses, and a pelvis frame built on guesses does not rotate. So the pelvis borrows the shoulder line as its lateral axis in proportion to how little the hips can be trusted:

const wHip = Math.min(confidence(L_HIP), confidence(R_HIP)) pelvisLeft .copy(lm[L_HIP]).sub(lm[R_HIP]).normalize().multiplyScalar(wHip) .addScaledVector(shoulderLeft.clone().normalize(), 1 - wHip)

The head gets the same treatment, with its lateral axis from your ears and its up axis from the cross product of ear-line and nose direction. As you turn far enough that your face leaves view, the head's frame fades back to simply following the chest — otherwise it snaps the moment the face landmarks go.

Tested against a synthetic body rotated about its axis: at 30°, 60° and −45°, pelvis, chest, head and arms all follow within 0.1°. Repeated with hip confidence dropped to 0.1 to simulate the desk — still a full 60° turn.

Matching your size: least squares, not height

Getting him to be your size is the part that looks trivial and is not. My first attempt measured your on-screen height head-to-heel, scaled the model to match, and pinned his hips to your hips.

It was visibly wrong, and it jumped.

Wrong, because the proportions of a heroically-built character are not yours — match the total height and everything in between lands slightly off. Jumpy, because when your legs leave the frame MediaPipe keeps reporting hips, knees and ankles as extrapolated guesses somewhere below the bottom edge, and the fit stretches him to reach them. As those confidences flicker, the size pops.

The right tool is a weighted least-squares similarity fit: given the already-posed rig, find the single uniform scale s and 2-D offset t that best land his joints on yours. Closed form, no solver:

// p = rig joints in unscaled model space, q = your joints in screen space // s = Σ w (p − p̄)·(q − q̄) / Σ w |p − p̄|² // t = q̄ − s · p̄ function solve(rows: Row[]) { let sw = 0, pbx = 0, pby = 0, qbx = 0, qby = 0 for (const [w, px, py, qx, qy] of rows) { sw += w; pbx += w * px; pby += w * py; qbx += w * qx; qby += w * qy } pbx /= sw; pby /= sw; qbx /= sw; qby /= sw let num = 0, den = 0 for (const [w, px, py, qx, qy] of rows) { const dx = px - pbx, dy = py - pby num += w * (dx * (qx - qbx) + dy * (qy - qby)) den += w * (dx * dx + dy * dy) } return { s: num / den, pbx, pby, qbx, qby } // t is derived from the s you apply }

Four things make it behave:

  1. Reject off-frame joints outright. A landmark outside the viewport is a guess. It contributes to neither the fit nor the pose.
  2. Weight by trust, not equally. Shoulders and hips are rigid and well tracked (1.0), elbows and knees less (0.6), wrists and ankles least (0.35).
  3. One robust pass. Compute residuals, drop anything beyond ~2.5× the median, re-solve. One mis-tracked limb should not be allowed to resize a whole character.
  4. Rate-limit the scale, not the position. Scale is clamped to 3% change per frame with a slow filter on top; translation stays fully responsive. Size should drift, position should snap.

Ship a lock button too. Auto-fit is the right default, but the moment it looks right the honest feature is a key that freezes the scale and lets position keep tracking.

Fingers: solve in the hand's own frame

This is the bug I would most like to save you. Fingers were tracking, and every hand came out as the same claw regardless of what my hands did.

The cause: I assumed HandLandmarker's worldLandmarks live in the same camera-aligned space as the pose model's, and aimed finger bones straight at those vectors. They do not. The hand model reports its 21 points in its own frame, so every finger was being aimed at a fixed direction that had nothing to do with my hand.

The fix is to stop caring what frame the tracker used. Build a palm frame from the tracked hand itself, express each finger direction in that frame, then push it back out through the wrist's actual posed orientation. Only the shape of the hand survives the round trip:

// palm frame of the tracked hand — index MCP → pinky MCP across, wrist → middle MCP up const Blm = basisFrom(h[5].clone().sub(h[17]), h[9].clone().sub(h[0])) const toPalm = new THREE.Matrix4().copy(Blm).transpose() // how the wrist bone is rotated right now, relative to its rest pose const R = worldOf(handBone).clone().multiply(q.copy(restWorld.get(handBone)!).invert()) for (const f of fingerAims) { d.copy(h[f.tip]).sub(h[f.base]).normalize() .applyMatrix4(toPalm) // → palm coordinates .applyMatrix4(restPalm) // → rig rest world .applyQuaternion(R) // → follow the posed wrist aimBone(f.bone, d.normalize()) }

The wrist itself is oriented from the body model's wrist, index and pinky landmarks, which are camera-aligned and trustworthy. Body model for the wrist, hand model for the knuckles.

This is testable, and worth testing, because "the fingers look a bit off" is not a signal you can iterate on. Re-express the hand landmarks in a completely different frame — 63° about an arbitrary axis — and re-run:

frame rotated 63° → solved finger bones move 0.10° (must be ~0) hand actually curls → solved finger bones move 27.25° (must be large)

Frame-independent, shape-sensitive. Before the fix, the frame rotation moved them as much as the shape did.

Punches: perspective and speed-adaptive smoothing

Two things stop a punch from landing. The first is projection: under an orthographic camera an arm thrown at the lens does not get bigger, it just slides. Switching to perspective, matched so the character's plane spans the same range the ortho frustum did, keeps all the placement maths identical while limbs aimed at the camera foreshorten properly:

const FOV = 45 const D = 1 / Math.tan(THREE.MathUtils.degToRad(FOV / 2)) // half-height = 1 at z = 0 <PerspectiveCamera makeDefault fov={FOV} position={[0, 0, D]} near={0.02} />

The second is latency. A fixed smoothing factor is a bad trade: enough to kill jitter when you hold still is far too much when you move fast. So smooth adaptively — a cheap One-Euro-style filter that opens up as a joint accelerates:

// heavy filtering at rest, backs off automatically as the joint speeds up const speed = smoothed[i].distanceTo(incoming) / dt const alpha = base * (1 - Math.min(1, speed / 2.5)) // 2.5 m/s ≈ a fast move smoothed[i].lerp(incoming, 1 - alpha)

Jitter is a low-speed problem and lag is a high-speed problem, so tie the filter to speed and both go away.

Compositing: the plate is what erases you

Rendering the character is only half of it. If you draw him over the live feed, you get Spider-Man standing in front of a man in a t-shirt.

The trick is a background plate: a still frame of the empty room, captured before you walk in. Three layers on a 2-D canvas, every frame:

ctx.setTransform(1, 0, 0, 1, 0, 0) ctx.clearRect(0, 0, W, H) ctx.save() if (mirror) ctx.setTransform(-1, 0, 0, 1, W, 0) // ONE mirror for every layer ctx.drawImage(video, 0, 0, W, H) // 1. you, live if (suitOn) { if (plate) ctx.drawImage(plate, 0, 0, W, H) // 2. the empty room, over you if (hasPose) ctx.drawImage(gl.domElement, 0, 0, W, H) // 3. him, where you were } ctx.restore()

That single setTransform is load-bearing. Mirroring a 3-D scene is not a rotation, so do not try to mirror the character in world space — you will flip its handedness and spend an hour wondering why one hand is inside out. Do all the maths in raw camera coordinates and apply one reflection to the finished composite.

In React Three Fiber, the WebGL canvas needs gl: { alpha: true } and a transparent clear colour so it can be drawn over the plate, and you read gl.domElement straight into the 2-D context.

The five bugs, each of which looked like something else

Symptom Actual cause
Every bone lookup returned undefined three's GLTFLoader sanitises node names — mixamorig:Hips_01 arrives as mixamorigHips_01. The colon is gone. Never match on the raw Mixamo name.
Shoulders looked hunched and broken I was driving the clavicles. MediaPipe has no clavicle landmark, so I was aiming them along the shoulder line, which forces them straight out sideways and collapses the deltoid. Leave clavicles at their sculpted rest pose.
Arms wrong, but only once the torso turned Consequence of the fix above. Once a bone is not driven, its children resolve against a stale parent rotation unless you compose world rotations up the chain through undriven joints.
Size popping between frames Extrapolated off-frame leg landmarks entering the fit as if they were real.
Permanent claw hands Hand landmarks are not in the pose model's frame. Solve fingers in the palm frame.

And a sixth that was purely my own fault: in the headless test rig I built to check the maths, the synthetic face pointed away from the camera — I had the sign of the nose's z wrong. That flipped the head bone 150° in every test run while real webcam data was perfectly fine. Half an hour of debugging code that was already correct. If you build a mock, mock the coordinate conventions too.

Testing computer vision without a camera

Worth its own note, because it is what made the rest of this tractable. I never once debugged the retargeting maths by standing in front of a webcam and squinting.

Instead: a synthetic landmark generator behind a ?mock flag that produces a parametric body — arms raising, fingers curling, the whole thing rotatable by a ?yaw=60 parameter, with a ?frz=1.7 flag that freezes the clock so two runs are comparable. Then headless Chromium reads bone quaternions straight out of the scene and asserts on them.

That turns vague into numeric:

Overlaying the tracked skeleton on the posed rig in a screenshot is the other half. If the green dots sit on the joints, the fit is right; if they float, it isn't. You can read that in a second, and you cannot read it from a number.

Performance

Two MediaPipe models plus a 632k-triangle skinned mesh sounds like it should crawl. On an M-series MacBook it ran at 60–90 fps. What matters:

What it still can't do

Being honest about the edges, because they are structural:

Why bother

Because it is the same stack as the gesture controls on this portfolio, pushed one step further: from "read a hand and move a cursor" to "read a whole body and drive a character". Everything hard about it — retargeting onto a skeleton you did not author, sizing one body against another, rejecting data a model reports confidently but cannot actually see — shows up again in anything that maps the physical world onto a rig.

And it takes one webcam, one browser tab, and no server. That was not true three years ago.

If you want more builds in this shape, there is the gesture-controlled portfolio itself, a drivable 3D car game in React Three Fiber, and a punch-through ice wall with Voronoi shatter.

← Portfolio · Blog