Egocentric & POV Datasets

First-person image and video, captured to your scenario, actions, and environment. Real people in real settings - built for embodied AI, robotics, and computer vision, delivered to your schema and model-ready.

Three ways to get started​

01 - See samples first

Start with sample clips.

Sample footage from existing egocentric collections. Check capture quality, framing, resolution, and metadata before scoping.

02 - Collect to spec

New data, collected to your spec.

You define the environment, activity set, contributor profile, and device. We deploy and return footage structured to your schema.


03 - Expand coverage

Scale an existing dataset.

Extend an existing POV dataset into new settings, longer sequences, or additional markets. Tell us the gap, we scope the capture.

Datasets by environment

Robotics

Household video

First-person household video across kitchens, bedrooms, bathrooms, living areas, garages, and gardens. Real people complete cooking, cleaning, laundry, sewing, repair, tool use, gardening, and everyday object-handling tasks for embodied AI and imitation learning.


Request samples →

Robotics

Industrial video

First-person capture on factory floors, workshops, and production lines. Covers assembly, inspection, machine operation, tool handling, maintenance, repair, safety procedures, and repeatable step-based workflows for industrial robotics and computer vision.


Request samples →

Robotics

Warehouse video

First-person warehouse and logistics activity across receiving, storage, picking, packing, sorting, scanning, pallet handling, loading, and inventory checks. Built for robotic manipulation, workflow recognition, navigation, and human-object interaction models.


Request samples →

Commercial vision

Commercial video

First-person activity across offices, hospitality venues, service environments, and other commercial spaces. Covers equipment setup, cleaning, stock handling, food preparation, maintenance, customer service, and repeatable workplace procedures.


Request samples →

Commercial vision

Retail video

In-store activity from shopper and staff perspectives. Includes shelf browsing, product selection, basket handling, restocking, price checking, barcode scanning, checkout, returns, and customer interactions for retail computer vision and behavior analysis.


Request samples →

Navigation

Outdoor video

Street-level first-person movement through public and outdoor environments. Covers walking, wayfinding, road crossing, obstacle avoidance, cycling, public transport, and interaction with signs, pathways, and changing terrain.


Request samples →

Capture and delivery specs

Capture devicesSmartphones, cameras – head-mounted, chest-mounted.
Resolution1080p / 30fps standard; up to 4K / 60fps on capable devices
FormatMP4 video with metadata in JSON or CSV
AudioAmbient sound and voice captured where the task requires it
FramingHands and manipulated objects kept in frame
AnnotationsPer-clip scene, location, lighting, and motion metadata as standard. Activity and sub-step labels, on-frame object inventory, hand-object interaction events, and first-person narration on request.
VolumeScaled to your project – from pilot batch to ongoing collection
LicensingProject-based or exclusive. Rights and provenance defined upfront
DeliveryExisting datasets shared on request; new collection scoped to your project, delivered in days to weeks depending on volume and complexity.
PricingPriced per hour of footage, based on industry, scope, and technical requirements. No seat licenses, no retainers.

Your data team has better things to do

Acquirox handles data collection, dataset preparation, labeling, QA, and delivery - so your engineers stay focused on model development.

FAQ

Real. Every clip is captured by a verified contributor performing a real action in a real environment. No staged studio footage and no synthetic generation - the variability comes from real people in real settings, which is what embodied models need to generalize.
Phone-based capture - head-mounted, chest-mounted, or handheld. Device, resolution, and framing are scoped to your spec before collection begins. Because the network records on consumer phones, capture stays within what contributors can reliably produce at scale.
First-person image and video across defined environments - household, industrial, warehouse, commercial, retail, and outdoor. Scenarios are scoped to what contributors can perform reliably on a phone, covering everyday actions, manipulation, and movement. Tell us the environment and action list and we will confirm feasibility.
18M+ verified contributors across 150+ countries, concentrated in Asia, Latin America, and Africa. This gives you environmental and demographic range across settings, and coverage is scoped to your target markets before a project starts.
Human in the loop on every submission. 2-3 contributors cross-check each clip against your spec, with automated integrity checks on format, resolution, and framing running first. Senior Acquirox reviewers set the golden standard on flagged and edge-case clips, so the dataset stays consistent as it scales.
Action labels, object tagging, and event marking on the clips we capture, within the annotation levels our contributor network handles reliably. Order capture and annotation together and receive labeled footage in one delivery. Where a project needs a higher annotation level, we scope it against network capability first.
Defined upfront. Project-based or exclusive licensing, with documented provenance and contributor consent. You know what rights you hold before capture begins, and exclusive datasets are held only for your team.
MP4 / H.264 video with structured metadata in JSON or CSV. Per-clip metadata covers scene, environment, and action label as standard, so footage arrives organized and ready for your pipeline.
Priced per hour of footage. The rate depends on exclusivity, technical requirements, capture setup, and scenario complexity. No seat licenses, no retainers - we return a scope-based quote against your capture brief.
Samples in 48 hours. Full collection is scoped to your volume and scenario, delivered in days to weeks. We share sample clips from the pilot so you can validate the capture before committing to full volume.

© 2026 Acquirox. All rights reserved.