Egocentric · Cinematic · Sensor-synced · Multi-camera

The annotation layer
your training data
was missing.

Structured annotation, inter-annotator agreement measurement, and machine-readable provenance documentation — for egocentric video, cinematic media, sensor-synced captures, and multi-camera datasets.

1,500+
hours delivered
manipulation-v2.3
original taxonomy
RLHF
quality loop

Case studies

Delivered to and accepted by teams building the next generation of AI

Physical AI · Household egocentric
Leading physical AI lab
1,500+ hours of residential household egocentric dataset delivered and accepted — kitchen activities, cleaning, sorting, folding, and manipulation tasks from head-mounted and wrist-mounted cameras.
egocentric · multi-camera · sensor-synced
Creative AI · Generative models
Leading creative AI platform
Annotated cinematic datasets delivered and accepted by the platform's data science team for generative model training — scene, object, and action labels on broadcast-quality footage.
cinematic · scene labels · NL captions

Four distinct pillars.
One complete pipeline.

Annotation and provenance are separate disciplines — and separate pillars. Annotation structures what the data means. Provenance documents where it came from and proves it hasn't changed. Most vendors offer one. We deliver all four.

🗄️
Data
Acquire · Structure · Enrich
Source and structure footage — from existing archives or custom collection. Build a queryable index with pre-packaged action buckets before annotation begins.
Egocentric · Cinematic · Sensor · Multi-camera
🧠
Intelligence
Annotate · Understand
Apply manipulation-v2.3 taxonomy at five granularity levels — action, scene, object, frame, gesture. RLHF feedback builds a custom model per project. Quality improves with every batch.
Shot-level metadata: actions, objects, scenes, motions
📊
Quality
QA · Evaluate
IAA measured on every project. Gold standard clips per domain. QA evidence package ships with every dataset — independently verifiable without contacting us.
Cohen's Kappa · Fleiss' Kappa · target > 0.85
🔐
Trust
Provenance · Lineage
Separate from annotation. Documents data origin, consent basis, rights scope, and version history. SHA-256 Merkle hash seals every version. EU AI Act Article 10 fields covered.
Machine-readable manifest · C2PA in development

What each pillar
looks like in practice.

Confirmed capabilities — with honest notes on what is live and what is in development.

DATA · 01
Indexing & action taxonomy
Every clip classified against manipulation-v2.3 — built from scratch, not adapted from Ego4D or EPIC-Kitchens. Five granularity levels: full session → scene → action → frame → gesture. Pre-packaged action buckets (all chopping clips together, all washing clips together) assembled before annotation begins.
manipulation-v2.35 granularity levelsaction bucketsoriginal schema
INTELLIGENCE · 02
Annotation with RLHF quality loop
Human labellers start the process. A reward model trained on feedback guides the next cycle. Custom model built per engagement — quality improves with each batch. Shot-level metadata: actions, objects, scenes, lighting, motions, contacts. Delivered datasets accepted by physical AI and creative AI customers in production.
RLHF feedback loopcustom model per projectimproving qualityconfidence scores
QUALITY · 03
Annotation quality evals
IAA measured per project — Cohen's Kappa for pairwise, Fleiss' Kappa for multi-annotator tasks, targeting above 0.85. Gold standard clips per domain. QA evidence package in every delivery. Available as a standalone engagement on existing datasets — no re-collection needed.
Cohen's / Fleiss' Kappagold standard QAQA evidence packtarget > 0.85
TRUST · 04
Provenance & lineage documentation
This pillar is distinct from annotation. Provenance documents data origin and version history — it does not describe what is in the footage.
Machine-readable manifest on every delivery: consent basis, data origin, rights scope, full version history — sealed with SHA-256 Merkle hash. EU AI Act Article 10 fields covered. C2PA-standard alignment in active development.
consent documentationversioned manifestSHA-256 MerkleEU AI Act Art.10C2PA · in development

Built for the hardest annotation
problems in physical and creative AI.

The main public egocentric datasets — Ego4D (3,670 hrs), EPIC-Kitchens (100 hrs), EgoDex — were built for research. Production AI training requires domain depth, provenance, and quality systems public datasets cannot provide.

EU AI Act Article 10 enforcement deadline: December 2, 2027. The Digital Omnibus (Regulation EU 2026/1744), which entered into force on 27 July 2026, deferred high-risk AI system obligations from August 2026. The requirement itself has not changed — only the enforcement date. Data provenance documentation must cover consent basis, data origin, pre-processing steps, and quality assurance. Providers preparing to launch after December 2027 need compliant training data documentation in place before that date. The time to build governance into your pipeline is before your model ships, not after.
🏠
Residential household egocentric
Kitchen activities, cleaning, sorting, folding, vacuuming — annotated at action and gesture level from head-mounted and wrist-mounted cameras. The domain where public datasets have least coverage and physical AI teams have greatest need.
Primary domain — deepest taxonomy coverage
🤖
Manipulation policy training
Fine-grained annotation for hand-object interaction: grasp type, contact events, hand engagement patterns, task completion states. Requires a taxonomy built for manipulation — not adapted from general activity recognition benchmarks.
NVIDIA EgoScale: egocentric pretraining (20k+ hrs) lifts task success 54% vs. no pretraining
📡
Sensor-synced & multi-camera
Head and wrist-mounted rigs synchronised to sub-0.5 second accuracy. Sensor-paired egocentric with IMU accelerometer and gyroscope streams. 2 to 6 camera configurations. Cross-camera object identity consistent across all views.
Sub-0.5s cross-device sync confirmed in delivered datasets
🎬
Cinematic & broadcast annotation
Broadcast-quality cooking and lifestyle footage annotated for generative AI training — scene classification, object labels, NL captions, temporal action labels. Same manipulation-v2.3 taxonomy depth. Delivered to and accepted by a leading generative AI platform.
Active paid customer in this domain
📋
EU AI Act provenance compliance
Machine-readable provenance manifest with consent basis, data origin, rights scope, and SHA-256 version hash — covering EU AI Act Article 10 data governance fields. Article 10 enforcement for high-risk AI systems is now December 2, 2027, deferred by the Digital Omnibus. The requirement is real; the deadline moved.
Deadline: December 2, 2027 · Digital Omnibus (EU 2026/1744)
📊
Annotation quality evals on existing data
Have annotated data and need to verify quality before training? We run IAA measurement, gold standard benchmarking, and confidence score analysis — and return a QA evidence package you can act on. No re-collection required.
Available as a standalone engagement

Three domains.
One platform.

The same annotation, quality, and provenance stack serves all three. Domain-agnostic infrastructure — domain-specific expertise.

Primary · Physical AI
Residential household egocentric
Cooking, cleaning, sorting, folding, vacuuming — from head and wrist-mounted cameras with optional IMU and audio. Activities confirmed in delivered datasets: pouring/liquid, chopping, vacuuming, washing dishes, tying shoes, folding laundry. Custom collection to specification.
Deepest coverage · 1,500+ hrs delivered
Active · Creative AI
Cinematic & broadcast
Licensed broadcast cooking shows and professional kitchen footage annotated for generative AI training — scene, object, NL captions, and temporal action labels. Same manipulation-v2.3 taxonomy as egocentric. Delivered to and accepted by a leading creative AI platform.
Active paid customer
Available · Multimodal
Spatial & multimodal
Multi-camera rigs with sub-0.5s synchronisation. Sensor-paired egocentric with IMU. LiDAR and 3D spatial capture via COLMAP structure-from-motion. Point cloud output. Custom modality sets to buyer specification.
Custom collection available

From raw footage to
delivered intelligence.

A managed four-step pipeline. Your team focuses on training models, not annotation workflows.

01
Share footage or define your collection brief
Upload existing footage via secure transfer, or specify a custom collection: activity types, modality set, participant diversity (age range, ethnicity, language), geographic range, annotation depth. We return a scoped delivery plan within five business days.
02
Indexing and taxonomy alignment
Every clip classified against manipulation-v2.3. Action boundaries detected, scene types assigned, objects identified. Pre-packaged action buckets assembled — all chopping shots together, all washing shots together — so your pipeline starts with structure, not a flat archive.
03
Expert annotation with RLHF quality loop
Domain-trained annotators label at the granularity you specify. A custom reward model trained on your domain guides each labelling cycle. Annotation quality improves with every batch — not a static ceiling. Confidence scores recorded per annotation.
04
IAA measurement, QA review, and provenance-sealed delivery
IAA measured before every delivery (target: Cohen's Kappa above 0.85). Gold standard clips reviewed. Dataset delivered with machine-readable provenance manifest, IAA scores, and QA evidence package — everything verifiable without contacting us.

What makes the output
different.

Not positioning claims. What we have actually delivered and what customers have accepted.

Quality
Annotation compounds per batch
RLHF-driven custom model building means each delivery batch is better than the last — not a static manual output. Confirmed by acceptance from physical AI and creative AI customers in production.
Schema
Original taxonomy, built from scratch
manipulation-v2.3 covers 16 parameters across action, gesture, scene, sensor, and provenance layers. Not adapted from Ego4D, EPIC-Kitchens, or any public benchmark. Validated against delivered datasets.
Domain
1,500+ hours accepted in production
Residential household egocentric data delivered, indexed, and accepted by a leading physical AI team. The domain public datasets have least coverage and physical AI teams need most.
Provenance
Manifest on every delivery, not on request
Consent basis, data origin, rights scope, SHA-256 version hash — a structured file with every dataset. C2PA-standard alignment in active development. EU AI Act Article 10 fields covered today.
Evals
QA evidence buyers can verify independently
IAA scores and gold standard QA evidence ship with every dataset. Buyers do not have to trust quality claims — the evidence package lets them verify independently.
Breadth
Physical AI and creative AI — one platform
Same platform annotates residential household egocentric for physical AI and cinematic footage for generative AI training. Sensor-synced, multi-camera, spatial — all covered. Different domains. Same rigour.

Common questions from
teams evaluating us.

Questions we hear from physical AI and creative AI teams before they engage.

How is VeriPhiLabs.AI different from Scale AI or Encord for video annotation?
Scale AI is built for high-volume managed manual annotation. Encord is a platform tool for teams running their own annotation. VeriPhiLabs.AI is a domain specialist: manipulation-v2.3 is built from scratch for household and manipulation footage, our RLHF loop improves annotation quality per batch rather than plateauing, and we ship a machine-readable provenance manifest with every dataset as standard — not as an add-on.
What is the difference between annotation and provenance? Do you do both?
Annotation describes what is in the footage: actions, objects, scenes, motions, contacts, lighting. Provenance documents where the data came from: consent basis, data origin, rights scope, version history, and a cryptographic hash that proves the data has not been altered. They are separate disciplines and separate pillars in our pipeline. Most vendors offer annotation only. We deliver both as standard on every engagement.
What is inter-annotator agreement and what score should I require?
IAA measures how consistently different annotators label the same data. Cohen's Kappa above 0.80 is strong agreement; below 0.60 is unreliable for production use. We target above 0.85 and include the IAA results in the QA evidence package with every delivery. When evaluating other vendors, ask for IAA scores per project — not a general quality guarantee statement.
Do you cover sensor data and multi-camera datasets, not just single-camera egocentric?
Yes. We handle egocentric video (head and wrist-mounted), sensor-paired egocentric (video + IMU accelerometer + gyroscope), multi-camera rigs (2 to 6 cameras synchronised to sub-0.5 second accuracy), cinematic RGB video, and 3D spatial capture via COLMAP structure-from-motion. We do not currently offer text NLP annotation, medical imaging, or satellite imagery.
What does EU AI Act Article 10 require for training data?
Article 10 requires documentation of data practices for high-risk AI systems — data sources, collection processes, consent basis, pre-processing operations, and known limitations. Article 10 enforcement for high-risk AI systems is now December 2, 2027, deferred from August 2026 by the Digital Omnibus (Regulation EU 2026/1744). The requirement has not changed — only the enforcement date. Our provenance manifest covers the required fields today. C2PA-standard alignment is in active development for future releases.
Can you annotate footage we already have, or only footage you collect?
Both. We annotate existing footage against manipulation-v2.3 or a custom schema. We also coordinate custom collection to your specification: activity types, modality set, participant diversity, and geographic range. Provenance documentation and IAA measurement are standard on both paths. Annotation quality evals are also available as a standalone engagement on your existing annotated data.

Scope a project
with our team.

Whether you need egocentric video annotation for physical AI training, cinematic dataset annotation for generative models, a quality eval on existing annotations, or a provenance-documented dataset for EU AI Act compliance — share your requirements and we will return a specific proposal within five business days.

We respond within 48 hours. Your message goes directly to our team. See live platform demos →

No spam. We do not use this form for marketing.