Flower arranging: real human to robot
Input: A 27.5-second real human flower-arranging video.
Output: A robot-rendered flower-arranging video aligned beside the source.
Turn first-person human video into embodied training data at scale
PhiAgent is LivSyn Robotics' embodied data engine. It transforms large-scale first-person human manipulation videos into robot motion, simulation trajectories and rendered video, generating scalable synthetic data across embodiments and actions for embodied model training.
Business Inquiry →Support PhiAgent with a GitHub Star. It is the project's public, verifiable like.
Opens the matching GitHub repository in a new tab.
Input: A 27.5-second real human flower-arranging video.
Output: A robot-rendered flower-arranging video aligned beside the source.
Input: A continuous 35.9-second first-person, two-hand tabletop demonstration.
Output: A polished silver five-finger robot-arm replacement selected from four full-stream JoyAI candidates. Only the requested 19-25-second interval receives localized SAM2 source-background refinement, with 18.5-25.5-second transition context; all other frames retain the selected direct stream. The result panel is labeled PhiAgent. This remains a PARTIAL visual result, not robot execution.
Input: A continuous 40.2-second first-person demonstration with a head-mounted device, straps, phone, cables, and connectors.
Output: The selected uninterrupted seed-73 JoyAI RV2V stream with silver five-finger robot arms, with the result panel labeled PhiAgent. Thin cable/finger occlusions remain the highest-risk regions; this is a PARTIAL visual result, not robot execution.
Input: One quilt, two full-body humanoid robots, and a frozen eight-phase collaboration plan.
Output: A photorealistic generated proposal shown for diagnosis. Strict acceptance is 0/4; it is not recorded hardware execution.
Harder input: Two pillow obstacles and one diagonal, overhanging quilt.
Output: Two robots clear the pillows, square the quilt, fold it twice, and transport the bundle. Strict acceptance is 0/4 because every candidate violates the frozen terminal-motion gate; this is a generated diagnostic, not hardware execution.
Input: The same T-shirt scene and dual-arm robot setup.
Output: Three synchronized folds that place the finished T-shirt in front, left, or right.
Simulation working · URDF-constrained IK/FK replay
Input: An eight-second PhiAgent dual-arm T-shirt-folding video with reviewed gripper points, visible wrist-to-tip axes, and open/close events.
Output: The synchronized motion replayed in MuJoCo by two six-axis RealMan RM65-B arms and AG2F90-C grippers. Axis-only orientation, branch-continuous IK, and a reviewed left tool-roll offset remove the wrist-flip solution and keep both gripper planes parallel to the tabletop. V26 preserves all arm joint states while correcting the final grasp sequence: both grippers remain closed at frame 171, the left releases by frame 191, and the right remains closed. All 192 states are finite; current mean left/right EEF FK residual is 0.72/0.00 mm.
Input: An Ego video of a human cutting cabbage.
Output: A frame-aligned robot-hand replacement video.
Input: The same Ego kitchen scene with three action instructions.
Output: Ten-second bottle-handover, cap-unscrewing, and bottle-rinsing futures.
Input: The same real first frame with two action conditions.
Output: A lift-and-carry-right future and a lift-up future.
Input: A 621-frame continuous human-hand gesture.
Output: A 24-DOF Shadow Hand and forearm retargeting.
Native HaWoR, native Dyn-HaMR, and PhiAgent 21-keypoint overlays on the same human input.
Matched MANO mesh overlays for native HaWoR, native Dyn-HaMR, and PhiAgent.
Left: Native HaWoR.
Center: Stage-3-BIR.
Right: Stage-3-BIR with Adaptive VDA. Relative to Native HaWoR, Adaptive VDA improves Train20 W-MPJPE by 8.21%, WA-MPJPE by 7.03%, RTE by 8.84%, and Accel by 9.14%. Its additional gains over Stage-3-BIR are 4.22%, 3.64%, 6.89%, and 0.56%, respectively. This is a GT-2D upper-bound development result. The demo shows a fixed 10-second segment; aggregate metrics use all Train20 sequences.
Kinematic replay · visual composite
Left: Human RGB input. Center: Dual-X5 MuJoCo replay on a gray background. Right: Robot render composited over the original scene.
Scope: Motion-phase and screen-space alignment only; garment motion is inherited from the source video, not simulated contact or cloth dynamics.
Input: One human-hand source video.
Output: Frame-aligned human, silver robot-hand, and graphite robot-hand views.
Input: The same source motion and scene.
Output: Sharpa, Wonik Allegro, and Shadow Hand variants.
Input: One manipulation source video.
Output: A four-embodiment comparison including a two-finger gripper attempt.
Input: An 89-frame human-hand motion video.
Output: Synchronized silver and graphite robot-arm conversions.
Input: The same human motion video.
Output: Silver, graphite, and Sudo R1-style robot variants.
Image-space evaluation passed · occlusion-aware completion
Left: The real human source is visibly blocked by a near-camera laptop during the middle and later folding stages. Right: The generated dual-arm robot result preserves the garment task, completes the fold continuously, and does not copy the person or foreground occluder into the result.
Claim boundary: “Completion” means generating an unoccluded, task-equivalent, temporally continuous robot video from the visible context—not uniquely reconstructing the hidden source pixels or proving 3D cloth state, contact force, collision safety, or real-robot execution.
Physical-boundary follow-up · PARTIAL: We added nine fail-closed physical gates, proved monocular scale ambiguity, measured a 9.09× visually compatible force range, and rejected eight evidence-spoof attacks. The harness and acquisition contract are ready, but 0 / 9 candidate physical gates pass without real calibration, telemetry, sensors, simulation, robot trials, and independent human ratings.