Annotated egocentric datasets
First-person video from real construction sites, segmented, labelled, and QA'd frame by frame against your schema.
Egocentric video data · physical AI
Farasa runs head-mounted capture on live construction sites, annotates every frame for hand pose, objects, gaze, contact, and action, then delivers it to teams training embodied models.
Move your cursor over the frame to turn the head.
01The gap
Third-person capture, simulation, and staged lab demonstrations share one flaw: the camera is never where the learner's camera will be. The geometry is wrong, the occlusions are wrong, and the hands, the single most informative thing in the frame, are the first thing lost.
A tripod sees a body moving. A head-mounted camera sees the task: what the hands touch, what gets occluded, and where attention goes next.
Physics engines miss the friction, deformation, and unpredictability of real materials and real tools under real load.
Controlled environments strip out the clutter, occlusion, and improvisation that define how work actually gets done.
Short scripted clips can't teach long-horizon structure: setup, correction, interruption, recovery. That is most of what a real task is.
02The annotation stack
Nothing here is a model output we forwarded on. Every layer is produced by technical annotators against your schema and cleared through a two-pass QA gate.
03What we deliver
First-person video from real construction sites, segmented, labelled, and QA'd frame by frame against your schema.
Complete multi-step takes, setup to finish, with the mistakes and corrections left in. Structured for imitation learning pipelines.
Targeted collection campaigns on active job sites, scoped to the tasks, tools, and viewpoints your models are currently missing.
Send us your own footage. It runs through the same stack, the same annotators, and the same QA gate as everything we capture.
04Why Farasa
Standing relationships with active construction sites. Not staged reenactments, not one-off shoots with a rented crew.
Technical teams who understand both the labour being performed and the models it is going to train.
Real-world data at a cost structure that makes collection at genuine scale practical rather than aspirational.
A deliberate focus on construction and demanding physical tasks: the hardest domain in embodied AI, chosen on purpose.
05Market
Almost every dataset for embodied AI was captured from the outside — tripods, simulation, or lab benches. The geometry is wrong, the occlusions are wrong, and the hands, the single most informative thing in the frame, are the first thing lost. Farasa closes that loop.
06How it works
We define the tasks, viewpoints, sensors, and annotation schema with your team before anyone puts a camera on.
Operators wear the rig through a real shift on an active site, with full consent and safety compliance.
Technical annotators label pose, objects, gaze, contact, and task structure against your specification.
Two-pass QA, then the dataset ships in your format, ready for training and evaluation on arrival.
07Who we serve
08Get started
Tell us what your models are missing and we'll send back a sample take with the full label stack attached. We reply within two business days.
ahmed.zeeshan@vanderbilt.edu