Multivision Egocentric Video Dataset (First-Person, Multi-View)
First-person (egocentric) video of everyday manual tasks, captured with synchronised multi-view rigs — for embodied AI, action recognition, hand-object interaction and imitation learning.
- Rights-cleared
- Consent documented
- 🇮🇳 India data residency
- Languages
- Narration in English & Hindi
- Scale
- 220 hrs · 4-view sync · 1,600 sessions · 1600 records
- Inter-annotator agreement
- α = 0.89
- QA pass rate
- 98.70%
Measured, audited, reproducible
- IAA / Krippendorff α
- 0.89
- QA pass rate
- 98.70%
Action segments, hand-object interactions and object tracks are double-annotated at α = 0.89 with a 98.7% QA pass rate. Every session provides frame-synchronised multi-view streams with verified temporal alignment.
Captured from consenting participants performing everyday tasks under fair-pay agreements granting commercial reuse, with on-camera bystanders consented or blurred. Each session carries consent and scene-context references; all data is stored and processed within India.
Participants wear a head-mounted camera while synchronised fixed cameras capture the same scene from multiple views. Annotators label action boundaries, hand-object interactions and object tracks; a second reviewer and a temporal-consistency audit verify alignment before each versioned release.
What's inside each record
A representative schema for Multivision Egocentric Video Dataset (First-Person, Multi-View). The full data card ships the complete field dictionary, value ranges and annotation rubric.
| Field | Type | Example |
|---|---|---|
| id | string (uuid) | "txt_a3f1…" |
| prompt | string | "When will this feature ship?" |
| response | string | "…" |
| label | enum | "preferred" | "rejected" |
| language | string (ISO 639) | "nar" |
| annotator_id | string (hashed) | "anr_7f3…" |
| qa_status | enum | "passed" |
| consent_ref | string | "cns_2024_…" |
Representative schema — exact fields and value ranges are documented in the dataset card shipped with every licence.
License tiers
Choose the tier that matches your use case. Every tier ships with the full data card, provenance log, and quality report.
Commercial Non-Exclusive
Contact us
Production licence for embodied-AI, action-recognition and imitation-learning models.
Request accessEnterprise / Sovereign Exclusive
Contact us
Exclusive licensing plus bespoke tasks, environments and sensor rigs captured to spec.
Talk to salesRelated datasets
-
Physical AIPhysical AI Video 🇮🇳 Sovereign
Off-the-Shelf Egocentric Data for Physical AI
🟢 Available now and continuously growing — first-person (egocentric) video of real commercial and residential tasks for embodied and Physical-AI training. Video only; no audio.
- Languages:
- Video only (no audio / no language track)
- Size:
- 🟢 Actively collecting · continuously growing
Rights-clearedLicence from
Contact us
-
Physical AIPhysical AI Video 🇮🇳 Sovereign
Egocentric — Agriculture, Landscaping & Grounds
🟢 Growing — first-person video of outdoor agricultural and grounds tasks. Video only; no audio.
- Languages:
- Video only (no audio / no language track)
- Size:
- 🟢 Growing · 50+ hrs available
Rights-clearedLicence from
Contact us
-
Physical AIPhysical AI Video 🇮🇳 Sovereign
Egocentric — Commercial Cleaning & Janitorial
🟢 Growing — first-person video of mopping, vacuuming, sanitising and waste collection. Video only; no audio.
- Languages:
- Video only (no audio / no language track)
- Size:
- 🟢 Growing · 5–10 hrs/task
Rights-clearedLicence from
Contact us
Explore more Physical AI data
Ready to license Multivision Egocentric Video Dataset (First-Person, Multi-View)?
A senior data PM will scope access, residency, and licensing terms and respond within one business day.