MotionMesh
Browser-based hand gesture interaction system that uses real-time computer vision to control a 3D holographic object through natural hand movements.
- Category
- Computer Vision
- Status
- Personal
- Stack
- React, TypeScript, Vite +7
Problem
Traditional 3D interfaces often rely on keyboards, mice, or touch controls, limiting natural interaction. MotionMesh explores a more intuitive interface by combining real-time hand tracking and gesture recognition with interactive 3D graphics directly in the browser.
Architecture
The application captures live webcam input through the browser's MediaDevices API and processes each frame locally using MediaPipe Tasks Vision for real-time hand landmark detection. The detected 21-point hand landmarks are smoothed using an Exponential Moving Average and classified using geometric gesture heuristics. These gestures are then mapped to transformations in a React Three Fiber / Three.js WebGL scene, enabling real-time hologram movement, scaling, rotation, expansion, and collapse.
Key features
- Real-time tracking of up to two hands
- 21-landmark hand tracking using MediaPipe
- Gesture recognition for pinch, two-hand pinch, open palm, fist, and point
- Gesture-controlled 3D scaling and rotation
- Interactive WebGL hologram with visual effects
- Real-time fingertip connections and motion trails
- Gesture-responsive particle effects
- Camera permission and device switching support
- Landmark smoothing for stable gesture detection
- Fully client-side camera processing with no video or image uploads
My role
Designed and developed the complete browser-based application, including the hand-tracking pipeline, gesture classification logic, coordinate transformations, WebGL interaction system, camera management, and responsive user interface using React and TypeScript.
Impact
Demonstrates practical implementation of real-time computer vision, browser-based machine perception, gesture-based human-computer interaction, and interactive 3D graphics. The project combines MediaPipe hand tracking with WebGL rendering to create a privacy-conscious, camera-driven interface without requiring a backend for video processing.
Stack
Retrieval-Augmented persona simulation powered by LLMs with long-term conversational memory.