Evaluation NotesRESEARCH / ENGINEERINGVOL. 01 / SEPTEMBER 2026
EXPERIMENT RECORD / UPDATED 15 SEP 2026

Measured progress.
Honest results.

I am building a multimodal evaluation pipeline for egocentric data. These are the results recorded so far, with their test conditions, limitations and next steps.

Explore the findings
69/69

Focused evaluator tests passed

Separate audit · 15 September
32,179

Hand-pose baseline images

AssemblyHands · component evaluation
2,388

Depth frame/model evaluations

796 associations × 3 models
37

Historical metric-parity tests passed

Arithmetic campaign · 8 September

01 / THE RESULTS

Where we stand.

Coverage, error and completed runs measure different things. There is no invented overall readiness percentage.

01

Hand pose

ASSEMBLYHANDS / DETECTOR-CONDITIONED HAWOR

QUALITY TARGETS OPEN
32.49%

Left joint coverage

31.18%

Right joint coverage

About 67–69% of reference joints lack corresponding predictions.

Camera-space error · L / R
35.22 / 34.14 mm
Root-relative MPJPE · L / R
32.36 / 25.79 mm
Aligned MPJPE · L / R
12.96 / 14.06 mm
Missing-aware PCK @ 20 mm · L / R
8.90 / 12.60%

Errors are conditional on valid predicted joints. Aligned error does not measure absolute placement.

Recorded improvement & evidence

Saved-detection recovery raises box recall at IoU 0.5 from 30.42% to 33.34% left, and 29.51% to 33.31% right. This is not recovered 3D accuracy.

Source: assemblyhands_hawor_detector_val_v3 / evaluation_report.json; HAND_POSE_STATUS.md. Baseline: 32,179 images.

02

Segmentation

VISOR / REFERENCE-BOX SMOKE TEST

BOUNDED VALIDATION
95.84%

Official per-mask mean J/F

16 images. 17 visible-reference masks. Ground-truth boxes and absence information supplied.

Region overlap · J
96.26%
Boundary score · F
95.43%
Smoke-test image count / converted split
16 / 6,430
Historical H9 mask coverage · L / R
90.10 / 85.09%
H9 longest missing run · L / R
62 / 57 frames
Scope & evidence

The smoke test covers about 0.25% of the converted image count. It is not a full automatic video benchmark. Historical H9 mask availability is not ground-truth mask accuracy.

Source: metric_parity_v1 / segmentation.json; HAND_SEGMENTATION_BUILD_STATUS.md.

03

Metric depth

TUM FR1_XYZ / RAW PREDICTIONS

ALL THREE GATES FAILED

Same 796 RGB/depth associations per model. No reference-fitted scale. Bars show means, not the median-based gate.

Completed associations · each model
796 / 796
Depth Pro · mean RMSE
0.2206 m
VDA streaming · mean RMSE
0.3302 m
VDA offline · mean RMSE
0.3337 m

100% planned run completion. Not 100% accuracy. One scene does not establish cross-domain reliability.

Quality gates & evidence

Median AbsRel: 15.73%, 20.35% and 20.46%, against a maximum 10% median target. Boundary and scale-drift gates also fail. All three reports record release_eligible: false.

A separate 32-frame focal ablation improved mean AbsRel from 15.722% to 13.202%. That exploratory result is separate from the full baselines.

Sources: tum_depth_pro_fr1_xyz_v2, tum_vda_metric_streaming_fr1_xyz_v2, tum_vda_metric_offline_fr1_xyz_v1 / evaluation_report.json.

04

Mesh & SLAM

KNOWN-POSE RECONSTRUCTION / H9 DIAGNOSTICS

GEOMETRY NOT ACCEPTED
8 submaps

Known-pose diagnostic completed over 796 frames.

Ground-truth camera poses were inputs. This does not validate estimated SLAM.

Separate H9 run · accepted submaps
0 / 12
H9 internal median AbsRel
15.1–18.1%
H9 median raycast coverage
93.43%
H9 missing alignment edges
11
Scope & evidence

Raycast coverage and same-model depth agreement are diagnostic measures, not physical surface accuracy.

Source: MESH_STATUS.md; MESH_EVALUATION_BUILD_STATUS.md. Known-pose and H9 results are separate runs.

05

Evaluation core

SCORING ARITHMETIC / SOFTWARE VERIFICATION

VERIFIED WITHIN SCOPE
39/39

Scoring-source files unchanged during the parity campaign.

The resolved arithmetic verdict passes. Full official challenge protocol parity is not claimed.

Historical parity tests
37 passed
Separate 15 September audit
69 passed
Official evo comparisons · controlled cases
26 / 6 cases
Mesh reports replayed
8
What passed & what did not

The historical parity suite has no failures or skips; one archive-producing CLI test was deselected. Test counts describe different snapshots and must not be added together. Controlled trajectory fixtures verify the scorer, not SLAM performance.

Source: metric_parity_v1 / PARITY_VERDICT.json; separate audit tests.xml and reproduced_findings.json.

06

Ego/exo timing

COMIND / JIKU AUDIO EXPERIMENTS

REVIEW REQUIRED
83.51%

CoMind frame-mapping coverage · 3,008 / 3,602 frames

Mean internal fit residual: 0.680 ms. This is not independently measured timing accuracy.

Jiku · mean clock-curve deviation
3.453 ms
Jiku · p95 clock-curve deviation
4.143 ms
Jiku · supported clip duration
34.95%

One audio pair, limited support. Audio agreement does not certify simultaneous video frames.

Scope & evidence

Jiku values compare RMS matching with a published manual-audio clock curve. The spectral matcher was rejected on that pair. Curve samples are correlated. Process labels are imported; automatic assembly recognition is not verified.

Source: timestamp / BENCHMARK_RUNS.md, updated 9 September 2026.

02 / LOCAL RECORDINGS

From the trim runs.

H9 Trim and hvideo (1) Trim, preserved from the attachment run. These measurements describe accepted complete outputs, not independent 3D accuracy.

INTERNAL CONSISTENCY

Accepted means the source valid flag is true and all 21 joints are finite. No independent 3D reference was supplied. These are different recordings and denominators from the AssemblyHands benchmark and the H9 segmentation results.

Run identity, definitions & source

Attachment run: hand_actions_trim_v3_attachment_20260831_001, completed 31 August 2026 on an RTX 5090. Both recordings completed decoding. The report's attachment percentages use eligible-hand frames, while the coverage above uses every frame in each recording.

Source: hand_pose_trim_20260902 / hand_pose_evaluation.json. Images: h9_frame_128.jpg and hvideo_frame_500.jpg. Recorded claim: internal_consistency_not_accuracy.

The run identifies HaWoR, WiLoR and MANO as research-only in this stack. These outputs are research diagnostics, not commercially cleared delivery artifacts.

05 / READING THE NUMBERS

The scope matters.

These results document progress without treating a completed experiment as a production approval.

What is ready today?

Controlled internal evaluation and diagnosis. Scoring arithmetic has substantial verification, while model quality and production release requirements remain open.

What about face privacy?

Final-render leakage metrics and gates are implemented. No accepted dense independent benchmark establishes production recall or leakage rates. Human-reviewed face tracks and final-blur validation remain required.

How should I interpret the metrics?

Coverage is how much expected evidence has predictions. MPJPE and RMSE are errors, so lower is better. J and F measure region and boundary agreement, so higher is better. A pass on one metric does not imply that the entire pipeline passed.

Where do the results come from?

Saved local evaluation reports and status records from September 2026, reviewed on 15 September. Source report names appear under each result. No new inference was run for this page. Evaluator implementation and raw prediction files are not distributed here.

Visual source datasets: EPIC-KITCHENS VISOR and TUM RGB-D. Images are selected experimental review artifacts, not newly collected footage.

THE NEXT MILESTONE

Better coverage.
Independent ground truth.
Repeatable acceptance.

Back to the results ↑