Views
No views yet


1@misc{he2026ropedia_xperience10m_task_suite,
2 author = {Chaoyue He},
3 title = {Ropedia Xperience-10M Task Suite},
4 year = {2026},
5 url = {https://github.com/ChaoYue0307/ropedia-xperience-10m-task-suite},
6 note = {Public task and evaluation layer for Ropedia Xperience-10M: two evidence lines, 20 embodied-AI tasks, 180 scored method-task records, and selected-128 model diagnostics}
7}| Signal | Current public state |
|---|---|
Project identity![]() | The same project logo mark is used across the GitHub README, GitHub Pages dashboard, Hugging Face Space, artifact dataset, model mirrors, favicon, and social preview. Ropedia is credited as the Xperience-10M data provider and releaser; this repository is the task-suite and evaluation layer built on that dataset. Reusable assets: logo mark and social card. |
| Two-line contract | Line 1: 1 sample episode for task construction and reproducibility. Line 2: 128 selected episodes for same-split metadata/raw baselines, Qwen3-Omni v6, and Cosmos3 diagnostics. |
| Data explorer analysis | A generated analysis layer separates the public sample, selected-128 feature exports, and authenticated Hugging Face gated full-dataset metadata with scope stats, split counts, modality breakdowns, and chart assets. |
| 180 method-task records | 9 methods x 20 tasks = 180/180 scored records. The ledger separates 174 direct scores from 6 compact-proxy scores. |
| 20 task contracts | Action, procedure, transition, trajectory, contact, objects, language, retrieval, reconstruction, order, sync, long-horizon forecasting, interaction text, action-object binding, sensor bridging, camera sync, and transition timing. |
| 4 research directions | Human Modeling & Motion Understanding; 3D/4D Reconstruction & Neural Rendering; Egocentric Vision & Interaction; Scene Reconstruction & World Modeling. These are analysis groups over the same 20 tasks, not separate benchmark tiers. |
| Line 1 methods | Minimal and Neural MLP baselines cover all 20 tasks on the one public sample episode: 40/40 direct scores. |
| Line 2 methods | Metadata simple/NN, raw-feature simple/NN, Qwen3-Omni v6 LoRA, Cosmos3-Super Reasoner, and Cosmos3-Nano Future Window cover all 20 selected-128 task axes: 140/140 scores. |
| 3 foundation pipelines | Spatial intelligence, human-video world modeling, and vision-language-action pipelines are documented as training recipes with task mappings, input-output contracts, and model-evidence requirements. |
| 1 unified target | The long-term embodied foundation-model target connects perception, 3D memory, language-grounded reasoning, action, and planning without adding a new score axis. |
| Public mirrors | GitHub, GitHub Pages, HF Space, HF artifact dataset, HF baseline model repo, Qwen3-Omni and Cosmos3 model repos, and HF collection. |

| Layer | Count | Role | Exact public labels |
|---|---|---|---|
| Task contracts | 20 | Score axes used by the matrix, radars, task cards, and method rows. | Action Recognition; Procedure Step Recognition; Action Boundary Detection; Next-Action Prediction; Hand Trajectory Forecasting; Contact State Prediction; Object Relevance Prediction; Language Grounding; Cross-Modal Retrieval; Cross-Modal Reconstruction; Temporal Order Verification; Multimodal Synchronization Detection; Long-Horizon Next-Action Forecasting; Long-Horizon Next-Subtask Forecasting; Interaction Text Prediction; Action-Object Relation Prediction; Future Object-Set Forecasting; IMU-to-Hand Pose Reconstruction; Camera-View Synchronization Retrieval; Time-to-Next-Transition Regression. |
| Research directions | 4 | Ways to interpret what the 20 tasks study; not separate benchmark tiers. | Human Modeling & Motion Understanding; 3D/4D Reconstruction & Neural Rendering; Egocentric Vision & Interaction; Scene Reconstruction & World Modeling. |
| Foundation pipelines | 3 | Larger-model training tracks with separate input-output recipes and result gates. | Spatial intelligence models; Human-video world models; Vision-language-action models. |
| Unified embodied model target | 1 | Long-term integration target, not a task/method row in the 180-result matrix. | Perception; 3D memory; language-grounded reasoning; action; planning. |
| Scope | Question | Current public analysis |
|---|---|---|
| Public sample | What files and signals are directly inspectable? | 1 episode, 5,821 frames, 1,161 aligned 20-frame windows, 8,546 feature dimensions, raw-file browser, modality breakdowns, action-window distribution. |
| Selected 128 | What selected-episode surface supports model comparison? | 96/16/16 split, 34,269 Qwen3-Omni v6 multiscale rows, 106,095 dense compact rows, selected episode links, public-safe matrices. |
| Full HF dataset | How large is the official upstream dataset? | Authenticated Hub file metadata: 804 sessions, 12,103 episode-like folders, 85,257 files, 24.63 TiB training-byte view, without redistributing raw gated data. |
| Line | Data unit | Score statement | Best use | Read separately from |
|---|---|---|---|---|
| 1 sample episode | One public Xperience-10M sample episode: 5,821 frames, 1,161 aligned 20-frame windows, 8,546 feature dimensions. | 40/40 direct scores from Minimal and Neural MLP heads. | Inspect the raw sample, understand file organization, reproduce the 20 task targets, and compare Minimal vs Neural MLP behavior inside one episode. | The selected-128 comparison rows and any broader held-out model behavior. |
| 128 selected episodes | Selected held-out 96/16/16 split: 34,269 exported windows with public-safe processed features linked to official gated episode paths. The Hugging Face artifact dataset exposes these rows separately as selected_128_windows/selected_128; it is not mixed with the one-sample episode_sample/public_sample viewer. | 140/140 selected-128 scores: 134 direct + 6 compact-proxy. | Compare same-split metadata/raw baselines, Qwen3-Omni v6, Cosmos3-Super, and Cosmos3-Nano while keeping the 6 compact-proxy cells visible. | Direct raw-target measurements for the proxy-marked cells. |
| Line | Methods | Tasks | Scored records | Direct scores | Proxy scores |
|---|---|---|---|---|---|
| 1 sample episode | 2 | 20 | 40/40 | 40 | 0 |
| 128 selected episodes | 7 | 20 | 140/140 | 134 | 6 compact-proxy scores, each source-linked and reasoned. |
| Total public matrix | 9 | 20 | 180/180 | 174 | 6 |
| Evidence line | Method block | Methods | Score statement | Read as |
|---|---|---|---|---|
| 1 sample episode | Task-head baselines | Minimal; Neural MLP | 40/40 direct scores. | Task-lab reproducibility and simple-vs-neural behavior. |
| 128 selected episodes | Aligned baseline heads | Metadata simple/NN; raw-feature simple/NN | 80/80 scores: 74 direct + 6 compact-proxy. | Same-split metadata/raw-feature baseline comparison. |
| 128 selected episodes | Qwen3-Omni series | Qwen3-Omni v6 LoRA | 20/20 direct scores from verified selected-128 Qwen3-Omni LoRA and task-specific probes. | Trainable Qwen3-Omni diagnostic baseline on the selected-128 surface. |
| 128 selected episodes | Cosmos3 series | Cosmos3-Super Reasoner; Cosmos3-Nano Future Window | 40/40 direct scores from verified public-safe reasoner and future-window artifacts. | Cosmos3 reasoner and future-window diagnostics on the selected-128 surface. |
| Run | Purpose | Main change | Eval signal | Use now |
|---|---|---|---|---|
| v1 | Prove the selected-128 LoRA/eval/package loop. | First verified 96/16/16 selected-episode Qwen3-Omni LoRA run. | 448 eval; JSON 0.8750; contact 0.6451. | Lineage only. |
| v2 | Make answers schema-checked. | Structured-JSON contract with full-8-GPU LoRA on the same split. | 448 eval; JSON 0.9978; contact 0.7188. | Structured-output ablation. |
| v3 | Separate prompt/eval effects from training. | Strict-label prompt/eval over the v2 adapter; no new adapter training. | 448 eval; JSON 1.0000; contact 0.7210. | Prompt/eval ablation. |
| v4 | Test longer structured-JSON LoRA training. | New four-epoch full-8-GPU adapter on the same selected split. | 448 eval; JSON 1.0000; contact 0.7299. | Overfit/metric-tradeoff evidence. |
| v5 | Move to denser multiscale evaluation. | Multiscale cap96 export with 4,032 held-out predictions. | 4,032 eval; JSON 1.0000; contact 0.7865. | Pinned prior release; stronger on several non-contact metrics. |
| v6 | Publish the current Qwen 20-task row. | Rank64/lr5e-5 multiscale LoRA plus verified task-specific probes. | 4,032 eval; JSON 0.9990; contact 0.8177. | Current public 20-task Qwen3-Omni row. |
QWEN3_OMNI_RUN_LINEAGE.md and
qwen3_omni_run_lineage.json.TWO_EVIDENCE_LINES.md,
two_evidence_lines.json,
TWO_EVIDENCE_LINE_RESULT_SUMMARY.md,
two_evidence_line_result_summary.json,
QWEN3_OMNI_RUN_LINEAGE.md,
qwen3_omni_run_lineage.json,
single_episode_task_model_radar.json,
episode128_task_model_radar.json,
task_method_20_result_matrix.json, and
xperience10m_128_episode_feature_index.json.| Goal | Start here | Then inspect |
|---|---|---|
| Understand quickly | Project brief Project status | Dashboard |
| Choose the published mirror | Public evidence map | evidence-map data |
| Decode project terms | Glossary | glossary data |
| Inspect the 20 tasks | 20-task guide | task contract data task walkthroughs |
| Explore data scales | Data explorer analysis | website analysis section structured analysis record |
| Compare results | Research takeaways | two-line result summary 180-record result table radar data score/proxy audit |
| Understand one sample | Single-episode explorer | sample-file map feature manifest |
| Read foundation directions | Three foundation pipelines | pipeline contract data foundation model plan |
| Reproduce or audit | Reproducibility Evidence contract | quality gates publication audit mirror parity |
| Capability | What this project shows |
|---|---|
| Multimodal data understanding | Parses the public sample into synchronized windows across video, audio, depth, pose/SLAM, mocap, IMU, calibration, and language-derived signals. |
| Task design | Defines 20 human-readable tasks in one unified public-sample suite, plus four direction-extension probes with inputs, outputs, process modules, metrics, and case-study walkthroughs. |
| Model and evaluation discipline | Runs minimal and compact neural baselines, records predictions/metrics, keeps chronological split boundaries explicit, and separates the sample readout from held-out comparison rows. |
| Scale-up planning | Connects the public-sample pipeline to 32/128-episode held-out pilots, Qwen3-Omni LoRA, Cosmos-style world-model tracks, policy/VLA tracks, and the future Xperience-native foundation-model pretraining goal. |
| Goal | Best entry point |
|---|---|
| Choose the right published mirror | Public evidence map evidence-map data |
| Resolve confusing terms and abbreviations | Glossary glossary data |
| Understand the whole project quickly | Project brief |
| See the visual research dashboard | GitHub Pages dashboard |
| Navigate the unified 20 tasks, four tracks, and scale-up plan | Interactive research roadmap 20-task suite note task contract data interactive roadmap data |
| Compare current task metrics | Research takeaways summary metrics |
| Compare possible foundation backbones | Foundation-model plan foundation-model data |
| Understand the future native pretraining goal | Native pretraining plan |
| See additional concrete project directions | Additional development directions direction data |
| Understand one model input | feature manifest window table |
| Check multi-episode data status | multi-episode data status |
| Surface | What it is for |
|---|---|
| GitHub repo | Source of truth for docs, scripts, generated data, validators, and commit history. |
| GitHub Pages dashboard | Best visual overview of the sample, 20 tasks, radar results, foundation directions, and resources. |
| Hugging Face Space | Hub-hosted copy of the dashboard and static app assets. |
| HF artifact dataset | Public-safe metrics, reports, website data, result packages, and derived evidence files. |
| HF baseline model repo | Minimal/neural baseline weights, figures, metrics, and mirrored task artifacts. |
| Qwen3-Omni and Cosmos3 model repos | Adapter-specific public weights or package cards when Qwen3-Omni v6, Cosmos3-Super, or Cosmos3-Nano runs are verified and publishable. |
| Theme | Current implementation |
|---|---|
| Dataset slice | One public Xperience-10M sample episode, 5,821 frames, 1,161 windows, and an 8,546-dimensional representation. |
| Modalities | Video, audio, depth, camera pose/SLAM, hand/body mocap, IMU, calibration, and language annotations. |
| Task suite | 20 human-readable tasks form one embodied-AI public-sample suite with shared windowing, split discipline, leakage controls, and minimal/neural head pattern. |
| Baselines | Minimal linear/ridge/logistic heads plus compact PyTorch MLP task heads over the same chronological split; companion simple/NN metadata baselines are also aligned to the selected 128-episode 96/16/16 split. |
| Research directions | Task mapping and extension probes for human modeling, 3D/4D reconstruction, egocentric interaction, and world modeling. |
| Scale-up path |
|
| Published mirrors | GitHub repo, GitHub Pages dashboard, GHCR static-site package, HF Space, HF artifact dataset, HF baseline-model repo, and HF collection. |
| Layer | Current scope | Where to start |
|---|---|---|
| Data understanding | One public Xperience-10M sample episode is converted into 5,821 frames, 1,161 aligned windows, and an 8,546-dimensional multimodal representation. | PROJECT_BRIEF.md PROJECT_STATUS.md |
| Task suite |
Twenty human-readable tasks cover recognition, prediction, retrieval, reconstruction, synchronization, long-horizon forecasting, interaction text, action-object binding, sensor bridging, camera sync, and transition timing.
Historical tier2_task_suite artifact paths are kept for link stability, but they are provenance paths inside the same suite.
|
20-task suite note task contract data research takeaways summary report historical provenance baselines |
| Baselines | Minimal heads and compact PyTorch MLP heads provide a controlled single-episode comparison on the same chronological split. The selected 128-episode setup adds same-split metadata simple/NN baselines for JSON-supported tasks and raw-feature simple/NN baselines on all 20 task axes. Tasks 15 and 19 are explicitly marked as compact-proxy completions. |
neural MLP outputs baseline alignment report raw20 run summary |
| Diagnostics | Audio contribution, modality ablations, timeline overlays, object labels, and alignment stress tests show which signals are useful and which tasks remain hard. | audio ablation summary single-episode explorer |
| Scale-up |
|
research roadmap foundation-model plan selected-128 feature index note selected-128 feature-index data selected-128 enhancement note selected-128 enhancement data model comparison data verified Omni result data Qwen v5/v6 comparison data Qwen v5/v6 comparison note Omni model comparison note verified public package selected-128 enhancement run |
ropedia-ai/xperience-10m, not a local raw-data
mirror. The official gated ropedia-ai/xperience-10m card reports 31.9 TB
on the live Hugging Face dataset card and an about-1PB full-scale storage statement; the committed
API-listing snapshot records 12,103 episode folders as upstream metadata only,
not a local raw-data inventory. In other words, those episode folders are
upstream listing metadata only for this project. The public sample remains
ropedia-ai/xperience-10m-sample under cc-by-nc-4.0, with the HOMIE Toolkit
and Rerun 0.29.0 noted as source tooling. The official responsible-use note
that the data is limited in diversity is preserved.| Area | Current decision |
|---|---|
| Public-sample pipeline | Verified on one public sample episode: 5,821 frames, 1,161 windows, 8,546 dimensions. |
| 20-task suite | Verified minimal baselines with committed metrics, predictions, and manifests. |
| Neural heads | Verified compact PyTorch MLP heads over the same task contracts and chronological splits. |
| Dataset context | Official Xperience-10M links, sample-vs-gated-data boundary, modality coverage, and redistribution policy are documented. |
| Evaluation protocol | Verified generated protocol for windowing, split policy, leakage controls, and per-task metrics. |
| Website and Hub pages | Public dashboard, Hugging Face Space, artifact dataset, baseline model repo, and collection use the same project framing and links. |
| Qwen3-Omni multi-episode pilot | Final verified diagnostic result package exists for the selected 96/16/16 episode split; JSON validity meets the target, while action/subtask metrics remain weak. |
| Raw data / full Qwen weights | Raw Xperience-10M data and full Qwen weights are not redistributed. |
| Step | Question | Primary artifacts | What should be true |
|---|---|---|---|
| 1 | What is this project? | Project brief Project status Dashboard | A public-sample Xperience-10M research project with 20 tasks, baselines, and a scale-up plan. |
| 2 | What data is used? | Dataset-card alignment Official HF dataset Sample HF dataset | The implemented suite uses one public sample episode; the gated dataset is reserved for selected multi-episode training. |
| 3 | What does one model input contain? | window table feature manifest available-modality data | Each window is an aligned multimodal unit with video, audio, depth, pose/SLAM, mocap, IMU, calibration, and language-derived signals. |
| 4 | What are the 20 tasks? | 20-task suite note task contract data task walkthroughs walkthrough data | Every task has a human-readable name, input, output, metric, baseline scores, and an explicit artifact path. |
| 5 | How are tasks evaluated? | evaluation protocol note evaluation-protocol data | The window unit, chronological split, leakage controls, task metrics, and current limitations are explicit. |
| 6 | What do current results mean? | research takeaways takeaway data summary metrics | Current metrics describe sample-level task behavior and identify which signals need larger held-out experiments. |
| 7 | Which models are implemented? | summary report neural MLP outputs HF baseline repo | Each task has minimal and neural-head evidence over the same feature windows. |
| 8 | What research directions does this support? | research roadmap direction data extension-probe data task contract data | The unified tasks are mapped to human modeling, 3D/4D reconstruction, egocentric interaction, and world modeling. |
| 9 | Which foundation model comes next? | foundation-model plan foundation-model data Native pretraining plan | Qwen3-Omni is the first held-out LoRA baseline; Cosmos 3 has Nano compatibility and Super forward-dynamics LoRA; policy models wait for robot-compatible action targets. |
| 10 | How can the 128-episode suite be pushed without more data? | selected-128 enhancement note selected-128 enhancement data | The enhancement pack proposes dense windows, hierarchical action/subtask labels, raw-feature shard priorities, and multiscale_20s10_40s20_80s40 as the next export target. |
| 11 | How do I reproduce it? | reproducibility guide reproduction audit | Public commands and expected outputs are documented for the sample-episode task suite. |
| 12 | What is still pending? | verified Omni result data data access status multi-episode access status | The final held-out diagnostic Qwen pass is verified and JSON-validity target is met; strong action/subtask model quality remains pending. |
ropedia-ai/xperience-10m
dataset is a gated large-scale egocentric multimodal dataset for embodied AI,
robotics, spatial intelligence, and world modeling. The public
ropedia-ai/xperience-10m-sample
repo provides the sample episode used for the implemented task suite here.annotation.hdf5 carrying depth, SLAM/camera pose, hand/body mocap, IMU,
language/caption annotations, calibration, metadata, and timing records,| View | What to inspect | Why it matters |
|---|---|---|
| Project status | project status project-status data | Summarizes the current project state in one table. |
| Data contract | window table feature manifest modality manifests | Confirms what each sample window contains before modeling. |
| Dataset context | dataset-card alignment official dataset links | Explains the official dataset, public sample, modalities, access boundary, and what this repo uses. |
| Visual assets | figure index site assets | Shows the task-suite graphic, modality thumbnails, pipeline diagrams, charts, and logo assets. |
| Evaluation protocol | evaluation protocol note evaluation-protocol data | Defines the task unit, split, metrics, leakage controls, and current limitations. |
| Research roadmap | research roadmap roadmap data | Shows the path from sample-level task development to multi-episode work, larger model tracks, and the future native-pretraining goal. |
| Additional development directions | additional development directions direction data | Records concrete non-backbone tracks: taxonomy, benchmark protocol, representation learning, skill graphs, affordances, 3D/4D memory, QA, and policy transfer. |
| Xperience Embodied Foundation Model plan | native pretraining plan | Describes the long-term full-corpus pretraining goal, target modules, objectives, staged scale-up, hardware ranges, and evaluation protocol. |
| Minimal heads | softmax ridge projection/regression multi-label logistic heads | Keeps every input/output contract visible and inspectable. |
| Neural heads | PyTorch MLP classifiers/regressors under neural_mlp/ | Checks whether nonlinear heads improve each task without changing features. |
| Evidence | metrics predictions confusion matrices diagrams dashboard | Makes the single-episode task development inspectable without rerunning first. |
| Artifact guide | artifact guide | Groups the public evidence after the project overview. |
| Reproducibility contract | reproducibility guide reproducibility matrix | States public commands, expected outputs, exact-match reproduction evidence, and non-reproducible boundaries. |
| Citation metadata | citation metadata software metadata license | Makes the repo easier to cite, index, and reuse without confusing code license and dataset terms. |
CITATION.cff,
codemeta.json, and docs/data/project_manifest.json
for machine-readable metadata.1@misc{he2026ropedia_xperience10m_task_suite,
2 author = {Chaoyue He},
3 title = {Ropedia Xperience-10M Task Suite},
4 year = {2026},
5 url = {https://github.com/ChaoYue0307/ropedia-xperience-10m-task-suite},
6 note = {Public task and evaluation layer for Ropedia Xperience-10M: two evidence lines, 20 embodied-AI tasks, 180 scored method-task records, and selected-128 model diagnostics}
7}LICENSE and DATA_NOTICE.md.
results/episode_task_suite/summary_report.json
with scripts/render_task_suite_infographic.py,
so the published PNG is a presentation graphic with verified labels and metrics,
not a hallucinated metric sheet.TASK_SUITE_20.md
and docs/data/task_suite_20.json. Historical
tier2_task_suite paths remain only as stable provenance links inside the same
suite.sqrt(normalized_score) only for visual radius so small but real differences
are readable; raw metrics and exact linear normalized scores remain in JSON and
the table.
Cosmos3-Super forward-dynamics LoRA
remains a separate artifact card because its camera-pose proxy MSE is not one of
the 20 task metrics.
The machine-readable copies are
docs/data/unified_task_model_radar.json
and
docs/data/task_method_20_result_matrix.json;
the explicit score/proxy ledger is
docs/data/task_method_20_gap_audit.json
and TASK_METHOD_20_GAP_AUDIT.md;
the public matrix is
TASK_METHOD_20_RESULT_MATRIX.md.
The website Results section also renders the same 180 cells as a wide,
source-linked table with raw values, normalized radar values, metric keys, and
direct/proxy badges.docs/data/modality_atlas.json and
docs/assets/modalities/. Those assets are small
derived thumbnails from the public sample, not raw Xperience-10M files.


scripts/render_overview_figures.py
overlays exact labels, dimensions, and metrics from the committed result files.1scripts/
2 train_min_action_model.py # motion/IMU baseline
3 train_all_modalities_model.py # current all-feature lightweight baseline
4 episode_task_suite.py # public-sample task definitions
5 neural_task_models.py # optional PyTorch MLP heads for task contracts
6 research_direction_taxonomy.py # maps walkthrough-backed tasks to the four research tracks
7 research_direction_extension_tasks.py # one extra data-backed probe per track
8 tier2_task_suite.py # historical-name provenance builder for unified task rows
9 build_unified_task_suite.py # builds TASK_SUITE_20.md and task_suite_20.json
10 build_unified_task_model_radar.py # builds grouped 20-axis model comparison radars
11 build_task_method_20_gap_audit.py # builds the explicit 180/180 scored-cell ledger
12 task_walkthroughs.py # human-readable task-card and walkthrough-storyboard metadata
13 generate_visualizations.py # refreshes SVG charts + summary JSON
14 render_task_suite_infographic.py # renders the task-suite presentation PNG
15 export_modality_atlas_assets.py # exports responsive modality-card assets
16 render_overview_figures.py # renders polished pipeline/architecture PNGs
17 build_brand_assets.py # derives logo sizes, favicon, social card
18 build_artifact_index.py # builds the compact artifact guide data
19 build_quality_gates.py # builds release checks
20 validate_mirror_parity.py # checks prepared GitHub/HF mirror file parity
21 validate_scope_claims.py # separates setup artifacts from completed model metrics
22 validate_task_surface.py # checks readable task cards and interactive storyboard wiring
23 validate_website_integrity.py # checks local site links, anchors, and images
24 validate_publication_package.py # checks public repo + HF bundle contents
25 publish_hf_bundles.py # uploads prepared HF Space/artifact/model bundles
26 omni/
27 download_sample_modelscope.py # ModelScope sample download helper
28 build_episode_manifest.py # metadata-only multi-episode scanner
29 plan_finetune_sample_budget.py # storage/sample-count planner
30 qwen3_omni_adapter_smoke.py # real-data Qwen3-Omni adapter setup check
31 score_existing_model_output_task_probes.py # scores task targets already present in verified model outputs
32 collect_qwen3_v4_release_artifacts.py # pulls verified v4 results after remote eval
33
34results/
35 min_action_model/ # motion-only action baseline artifacts
36 min_subtask_model/ # motion-only subtask baseline artifacts
37 min_all_modalities_action_model/ # current all-feature action artifacts
38 min_all_modalities_subtask_model/ # current all-feature subtask artifacts
39 episode_task_suite/ # task-suite metrics and predictions
40 neural_mlp/ # optional neural baseline artifacts per task
41 research_directions/ # four-track taxonomy, CSV, and summary
42 research_direction_extensions/ # four extra direction probes + predictions
43 tier2_task_suite/ # provenance baseline tasks + predictions; historical path
44 task_walkthroughs/ # case-study walkthroughs for walkthrough-backed tasks
45 omni_exploration/ # ModelScope readiness-check artifacts
46 omni_finetune/model_output_task_probes_20260616/ # task-13/task-16 probes derived from verified model JSON
47
48docs/
49 index.html # GitHub Pages dashboard
50 data/additional_development_directions.json # concrete non-backbone project directions
51 data/summary_metrics.json # website-readable metrics bundle
52 data/task_suite_20.json # unified 20-task suite bundle
53 data/unified_task_model_radar.json # 20-task radar values, groups, and sources
54 data/single_episode_task_model_radar.json # 1-episode grouped radar values
55 data/episode128_task_model_radar.json # 128-episode grouped radar values
56 data/task_method_20_result_matrix.json # 9-method x 20-task result matrix
57 data/task_method_20_gap_audit.json # explicit 180/180 scored-cell ledger
58 data/task_icon_manifest.json # assigned icon asset map for all 20 tasks
59 data/evidence_contract.json # machine-readable project scope
60 data/artifact_index.json # compact project-artifact catalog
61 data/live_publication_status.json # live GitHub/HF publication verification
62 data/quality_gates.json # machine-readable release checks
63 data/task_suite_enhancement_128.json # no-new-episode 128-suite enhancement pack
64 data/task_surface_integrity.json # machine-readable task-card/storyboard integrity check
65 data/project_manifest.json # machine-readable public-surface metadata
66 data/project_packet.json # compact project path and scope summary
67 data/research_roadmap.json # multi-episode and omni-model roadmap
68 data/research_directions.json # four-track website data bundle
69 data/research_direction_extensions.json # four extra probe data bundle
70 assets/task-icons/*.svg # one crisp assigned icon per task
71 assets/task-icons/task-icon-atlas.png # generated overview atlas for the 20-task visual language
72 data/tier2_task_suite.json # provenance baseline bundle; historical path
73 data/task_walkthroughs.json # human-readable task-card and walkthrough-storyboard data
74 data/modality_atlas.json # responsive modality-card data
75 assets/brand/*.png # project logo, favicon, social card
76 assets/task_suite_infographic.png # task-suite presentation graphic
77 assets/modalities/ # public-sample derived modality thumbnails
78 assets/pipeline_diagram.png # verified episode pipeline graphic
79 assets/qwen3_omni_lora_pipeline.png # Qwen3-Omni LoRA training-flow figure
80 assets/task_architectures.png # verified task-head architecture map
81 assets/charts/unified_task_model_radar.svg # 9-method grouped small-multiple radar board
82 assets/charts/single_episode_task_model_radar.svg # 1-episode enlarged radar panel
83 assets/charts/episode128_task_model_radar.svg # 128-episode grouped radar panels
84 assets/charts/*.svg # regenerated visualizations
85
86notes/
87 min_action_model.md
88 all_modalities_model.md
89 episode_task_suite.mddocs/ site plus the main project documents; it does
not include raw Xperience-10M videos, raw annotations, gated data, or model
weights.1docker pull ghcr.io/chaoyue0307/ropedia-xperience-10m-task-suite:latest
2docker run --rm -p 8080:80 ghcr.io/chaoyue0307/ropedia-xperience-10m-task-suite:latesthttp://localhost:8080.1<workspace>/
2 HOMIE-toolkit/
3 data/sample/xperience-10m-sample/
4 annotation.hdf5
5 fisheye_cam0.mp4
6 fisheye_cam1.mp4
7 fisheye_cam2.mp4
8 fisheye_cam3.mp4
9 stereo_left.mp4
10 stereo_right.mp4fisheye_cam0.mp4, links the full
raw Hugging Face source for each MP4/HDF5/RRD file, and describes the
annotation.hdf5 group organization without copying large raw files into this
repository. The same manifest also lists what the non-playable
annotation.hdf5 and visualization.rrd files contain, how each relates to
the synchronized MP4 streams and 20-frame windows, and which external tools can
open them.ropedia-ai/xperience-10m-samplehttps://huggingface.co/datasets/ropedia-ai/xperience-10m-sample1git clone https://github.com/Ropedia/HOMIE-toolkit.git
2python3.12 -m venv .venv
3source .venv/bin/activate
4pip install -r HOMIE-toolkit/requirements.txt huggingface_hub hf_xet1hf download ropedia-ai/xperience-10m-sample \
2 --repo-type dataset \
3 --local-dir data/sample/xperience-10m-sample1python scripts/omni/download_sample_modelscope.py \
2 --output-dir data/sample/xperience-10m-sample \
3 --mode minimal--mode minimal downloads annotation.hdf5, README.md, and
fisheye_cam0.mp4. Use --mode all-training to add all six MP4 streams while
still skipping visualization.rrd.1git clone https://github.com/ChaoYue0307/ropedia-xperience-10m-task-suite.git
2cd ropedia-xperience-10m-task-suite
3python scripts/episode_task_suite.py --workspace /path/to/workspace1pip install torch
2python scripts/episode_task_suite.py \
3 --workspace /path/to/workspace \
4 --include-neural1python scripts/tier2_task_suite.py --workspace /path/to/workspace
2python scripts/build_unified_task_suite.py
3python scripts/build_evaluation_protocol.py1python scripts/train_min_action_model.py --workspace /path/to/workspace
2python scripts/train_all_modalities_model.py --workspace /path/to/workspace
scripts/omni/package_verified_omni_result.py creates a
public-safe derived-artifact package. The current verified package is listed in
docs/data/omni_finetune_verified_result.json.
The current cross-version comparison is generated at
docs/data/omni_model_comparison.json
and results/omni_finetune/OMNI_MODEL_COMPARISON.md;
it separates the single-episode task suite, 128-episode aligned simple/NN
baselines, Qwen3-Omni v6 LoRA, Cosmos3-Super Reasoner, and Cosmos3-Nano Future Window packages. The same generated
files also include model_groups: a model-first view that pairs 1-episode and
128-episode entries for the same family. Use that section when comparing task
heads against task heads, Qwen3-Omni smoke/LoRA against Qwen3-Omni LoRA, or
Cosmos3-Nano compatibility against future Cosmos weight releases. For
Qwen3-Omni specifically, read QWEN3_OMNI_RUN_LINEAGE.md: v1-v4 are
pipeline-hardening and ablation evidence, v5 is the pinned prior multiscale
release, and v6 is the current public 20-task Qwen row.docs/data/task_suite_enhancement_128.json
and TASK_SUITE_ENHANCEMENT_128.md. It keeps
the current Qwen3-Omni v6 and Cosmos3 packages as baselines, then defines dense-window
scenarios, hierarchical action/subtask targets, task bottlenecks, and experiment
cards for stronger selected-128 runs without overwriting earlier results.| Phase | Episodes/samples | Approx windows at stride 5 | Purpose |
|---|---|---|---|
| Readiness | 1-3 | 1k-3k | Verify loaders, token alignment, and task heads |
| Pilot | 16-32 | 18k-37k | First held-out-episode evaluation |
| Useful LoRA run | 64-128 | 74k-149k | Train sensor adapters plus selected Qwen3-Omni LoRA |
| Storage-heavy run | 256+ | 297k+ | Only after download layout and checkpoint size are stable |
1python scripts/omni/plan_finetune_sample_budget.py \
2 --storage-root /path/to/storage \
3 --target-free-after-download-gb 800 \
4 --all-training-per-episode-gb 2.4 \
5 --full-preview-per-episode-gb 5.11python scripts/omni/discover_xperience10m_sources.py \
2 --workspace /path/to/ropedia-xperience-10m-task-suite \
3 --data-root /path/to/xperience10m_data \
4 --output results/omni_finetune/source_discovery.jsonvisualization.rrdXPERIENCE10M_128_EPISODE_FEATURE_INDEX.md and docs/data/xperience10m_128_episode_feature_index.jsonropedia-ai/xperience-10m episode pathsresults/omni_finetune/source_discovery.jsonresults/omni_finetune/DATA_ACCESS_STATUS.mdresults/omni_finetune/MULTI_EPISODE_ACCESS_STATUS.mdvisualization.rrd, balances episode-size bands, and preserves one selected
episode per top-level session UUID.train episodes and monitoring prepared val episodes.
The final test episodes stay sealed until the end, so early development does
not contaminate held-out evaluation.1python scripts/omni/build_selection_episode_manifest.py \
2 --workspace /path/to/ropedia-xperience-10m-task-suite \
3 --data-root /path/to/xperience10m_128 \
4 --selection-json results/omni_finetune/xperience10m_128_episode_selection.json \
5 --output results/omni_finetune/trainval_progressive/episode_manifest_trainval.json \
6 --include-split train \
7 --include-split valscripts/omni/run_trainval_progressive_128.sh wraps the same guard, exports a
train/val-only Qwen3-Omni JSONL dataset, and launches LoRA training without
running final test evaluation. The exporter uses session-qualified episode IDs
and path-based split matching so repeated folder names such as ep1 cannot
collide across different sessions.scripts/omni/run_trainval_parallel_export_8gpu.sh
uses the same split guard, exports episodes in parallel CPU shards, skips and
reports episodes that contain no labeled windows under the configured label
rule, then launches Qwen3-Omni LoRA with NUM_PROCESSES=8.1RUN_ID=xperience10m_qwen3_omni_128ep_fullsplit_fast8gpu \
2DATA_ROOT=/path/to/xperience10m_128 \
3SELECTION_JSON=results/omni_finetune/xperience10m_128_episode_selection.json \
4MODEL_DIR=/path/to/Qwen__Qwen3-Omni-30B-A3B-Instruct \
5NUM_PROCESSES=8 \
6TRAIN_VAL_SPLIT=val \
7MAX_VAL_SAMPLES=512 \
8scripts/omni/run_128_fullsplit_parallel_export_8gpu.sh1python scripts/omni/monitor_omni_progress.py \
2 --run-id xperience10m_qwen3_omni_128ep_fullsplit_fast8gpuprogress.jsonl, new evaluator partial-prediction
progress, and legacy generation logs, so long held-out evals can still expose
sample-level progress even before final metrics are written.1python scripts/omni/validate_omni_finetune_run.py \
2 --run-id xperience10m_qwen3_omni_128ep_fullsplit_fast8gpu \
3 --require-stage manifest
4
5python scripts/omni/validate_omni_finetune_run.py \
6 --run-id xperience10m_qwen3_omni_128ep_fullsplit_fast8gpu \
7 --require-stage eval \
8 --min-json-validity 0.981python scripts/omni/package_verified_omni_result.py \
2 --dataset-run-id xperience10m_qwen3_omni_128ep_fullsplit_fast8gpu \
3 --train-run-id <train_run_id> \
4 --eval-run-id <eval_run_id>1python scripts/omni/watch_verified_omni_package.py \
2 --dataset-run-id xperience10m_qwen3_omni_128ep_fullsplit_fast8gpu \
3 --train-run-id <train_run_id> \
4 --eval-run-id <eval_run_id>eval_progress_observed events from
partial prediction files or legacy generation logs. This keeps the package
status file useful during long held-out evaluations.configs/omni_backbones, so Qwen3-Omni,
Cosmos-style world models, and VLA/policy tracks can share the same verified
publication gate once their model-specific evaluators exist. The package
excludes raw Xperience-10M files, base-model weights, adapter or checkpoint
weights, full checkpoints, and large archives.1CUDA_DEVICE_GROUPS="0,1 2,3 4,5 6,7" \
2SHARDS=4 \
3RUN_ID=<merged_eval_run_id> \
4scripts/omni/run_qwen3_omni_lora_eval_sharded.sh1python scripts/omni/export_model_neutral_window_index.py \
2 --dataset-jsonl results/omni_finetune/xperience10m_qwen3_omni_128ep_fullsplit_fast8gpu_dataset/dataset.jsonlwindow_index.jsonl and window_index_manifest.json so Cosmos-
style world models and VLA/policy tracks can reuse the same split-checked
windows without depending on Qwen chat-message records.cy0307/ropedia-qwen3-omni-lora-128ep, older
Qwen smoke material remains historical. Cosmos3-Nano remains an artifacts-only
compatibility result; Cosmos3-Super Forward-Dynamics now has a separate
weight-bearing model repo at
cy0307/ropedia-cosmos3-super-forward-dynamics-lora-128ep.
Metrics, predictions, audits, and reports stay in the artifact dataset.1python3 scripts/omni/upload_qwen3_omni_lora_to_hf.py \
2 --repo-id cy0307/ropedia-qwen3-omni-lora-128ep \
3 --source-dir /path/to/adapter_upload_package \
4 --message "Upload Xperience-10M Qwen3-Omni LoRA pilot"HF_TOKEN or --token.
Network availability to huggingface.co is required.| Branch | Current role | When to use it |
|---|---|---|
| Qwen3-Omni | First trainable multimodal LoRA pilot | Use for the selected 128-episode held-out baseline over video/audio/language plus sensor-bridge features. |
| Cosmos 3 | First world-model/action-generation track | Use now for future-window compatibility analysis and the verified Cosmos3-Super forward-dynamics LoRA artifact; compare its loss metrics separately from Qwen JSON-task accuracy. |
| GR00T | Humanoid/action-policy track | Use after mocap/contact retargeting creates well-defined humanoid action targets. |
| OpenVLA / openpi | Open VLA/policy baselines | Use after the project defines robot-compatible or action-token targets. |
| Gemini Robotics | External reasoning reference | Use only for qualitative comparison or annotation support unless local trainable access exists. |
| Xperience Embodied Foundation Model | Future Xperience-native pretraining goal | Use only after multi-episode pilots, full-corpus storage, distributed training infrastructure, and scaling evidence justify a from-scratch domain model. |
FOUNDATION_MODEL_PLAN.md and
docs/data/foundation_model_plan.json
for the full selection matrix, source links, and model-specific evaluation
additions. See
XPERIENCE_EMBODIED_FOUNDATION_MODEL_PRETRAINING.md
for the long-term full-corpus pretraining plan.| Pipeline track | First concrete pipeline | Current scope |
|---|---|---|
| Spatial intelligence models | Build scene/object memory targets from multiview RGB, depth, pose, calibration, object cues, and language prompts. | Ready as a geometry/reasoning pipeline; the next readout is held-out spatial QA, pose consistency, counting, and scene-memory metrics. |
| Human-video world models | Predict next action, next subtask, future object set, contact transition, and future state from observed interaction windows. | Partially evidenced by future-task probes and Cosmos-style artifacts; visual/latent future quality still needs stronger metrics. |
| Vision-language-action models | Convert egocentric video, captions, hand/body motion, contacts, and objects into action chunks or policy-compatible targets. | Feasible, but gated by action-token conversion, normalization, retargeting evidence, and held-out policy metrics. |
| Direction | One-sample input | One-sample output target |
|---|---|---|
| Spatial intelligence | 20-frame windows from windows.csv / shared_windows.npz, joined with six MP4 camera streams plus annotation.hdf5 depth, pose, SLAM/calibration, object/contact cues, and optional language questions. | Camera-view match, object relevance, object-set memory, depth/pose reconstruction proxy, caption-grounded retrieval, and spatial QA targets. |
| Human-video world model | Current observed window at time t: RGB/audio/sensor summaries, hand/body motion, camera pose, current object/contact state, and current action/subtask context only. | Shifted future targets: next action, next subtask, future object set, contact transition, time-to-transition, camera-motion delta, or latent/future feature. |
| Vision-language-action | Egocentric/fisheye video, caption/object context, hand/body mocap, contact state, and current subtask text as observation-language input. | Action-token proxies: current/next action, object-conditioned action relation, contact state, interaction-text class, subtask transition, or hand-trajectory/action-chunk proxy. |
docs/assets/foundation-pipelines. Spatial
intelligence and human-video world modeling use the clean slide PNGs supplied
for publication and are exported as 2560-pixel public images. The 2026-06-19
refresh verified that the latest uploaded Spatial and Human-video PNGs are
byte-identical to the committed clean source cache. The VLA card now uses the
clean VLA slide PNG supplied afterward and is exported through the same
2560-pixel public path. These images are
communication assets, not completed model-quality evidence; the exact task,
training, and evaluation contracts remain in the Markdown and JSON files.


configs/omni_backbones.
The extension contract is documented in
OMNI_MODEL_EXTENSION_CONTRACT.md, and the
registry can be checked with:python scripts/omni/backbone_registry.py --validate --jsonpython scripts/omni/smoke_test_backbone_packaging.py1python scripts/omni/audit_verified_omni_package.py \
2 --package-dir results/omni_finetune/verified_public/<eval_run_id>1python scripts/omni/scaffold_omni_backbone.py \
2 --template-backbone policy_vla_branch \
3 --id new_policy_branch \
4 --display-name "New Policy Branch" \
5 --model-family "Model family name" \
6 --dataset-contract xperience10m_observation_action_v1 \
7 --training-objective observation_to_action_policy \
8 --checkpoint-gate policy_checkpoint_action_space_and_normalizer \
9 --dry-run| Direction | First useful artifact | Role in the project |
|---|---|---|
| Episode taxonomy and data engine | Episode atlas, balance report, and split builder | Select representative data before training. |
| Standardized benchmark protocol | Versioned train/val/test manifests and metric scripts | Make future model results comparable. |
| Multimodal representation learning | Contrastive and masked-window encoder objectives | Learn reusable video/audio/depth/pose/mocap/IMU/language features. |
| Skill and procedure graph mining | Step graph, transitions, preconditions, and effects | Connect perception to planning and long-horizon reasoning. |
| Human-object affordance modeling | Contact, reachable-object, tool-use, and next-affordance tasks | Model what actions the scene makes possible. |
| 3D/4D scene and object memory | Persistent scene/object maps from depth, pose, multiview video, and objects | Track world state beyond single frames. |
| Data-quality and synchronization diagnostics | Per-episode QA for drift, missing streams, calibration, and corrupted files | Keep large multimodal training trustworthy. |
| Policy, retargeting, and simulation transfer | Action-token conversion and robot-compatible imitation examples | Bridge human egocentric experience to robot policy work. |
| Baseline | Role |
|---|---|
| Minimal interpretable heads | Softmax, logistic, ridge, and retrieval heads over the 8,546-dimensional multimodal representation. These expose the input/output contract cleanly. |
| Neural MLP heads | Small PyTorch MLP classifiers/regressors on the same features and splits. These check whether nonlinear heads help before moving to Qwen/Omni fine-tuning. |
| Direction | Current status | Covered task evidence | What is not solved yet |
|---|---|---|---|
| A. Human Modeling & Motion Understanding | Partially implemented | Hand Trajectory Forecasting and Contact State Prediction are direct; Action Recognition and Object Relevance Prediction are proxies. Neural MLP improves hand forecasting from 0.8647 to 0.1079 MPJPE. | No full body/shape model, SMPL/MANO target, deformation prior, or multi-episode motion-generation evaluation yet. |
| B. 3D/4D Reconstruction & Neural Rendering | Prerequisite evidence | Cross-Modal Retrieval, Cross-Modal Reconstruction, and Multimodal Synchronization Detection test alignment/reconstruction prerequisites. | No NeRF, Gaussian Splatting, TSDF, mesh, novel-view synthesis, or calibrated 4D reconstruction model yet. |
| C. Egocentric Vision & Interaction | Strongest implemented track | 6 direct tasks: action, subtask, transition, next-action, object relevance, and caption grounding, plus alignment/order diagnostics and audio ablation. | Single-episode chronological split limits generalization; stronger audio and video-language backbones still need multi-episode testing. |
| D. Scene Reconstruction & World Modeling | World-model prerequisite evidence | Procedure Step Recognition, Next-Action Prediction, Object Relevance Prediction, Cross-Modal Retrieval, Cross-Modal Reconstruction, Temporal Order Verification, and Multimodal Synchronization Detection provide state/world-model probes. | No persistent scene graph, object permanence task, long-term map, or held-out-episode world model yet. |
shared_windows.npz, windows.csv, and feature_manifest.json artifacts, so
the reported numbers are computed from sample-derived features and saved metric artifacts.research_direction_extension_results.jsonresearch_direction_extension_summary.mddocs/data/research_direction_extensions.jsonresearch_direction_extension_tasks.svg| Direction | New extension task | Input | Output | Minimal | Neural MLP | Why it matters |
|---|---|---|---|---|---|---|
| A. Human Modeling & Motion Understanding | Body and Hand Motion Intensity | non-mocap video/depth/pose/IMU/SLAM/language features | high vs low body/hand motion | 0.7827 macro-F1 | 0.7986 macro-F1 | Starts a human-motion-energy target without leaking mocap input. |
| B. 3D/4D Reconstruction & Neural Rendering | Multi-View Consistency Retrieval | fisheye camera feature query | synchronized stereo-left view rank | 0.5534 MRR | 0.3469 MRR | Tests whether multi-view features preserve synchronized 4D scene identity. |
| C. Egocentric Vision & Interaction | Action Phase Progress Estimation | non-caption multimodal window | progress inside current action segment | 0.3416 MAE | 0.3038 MAE | Adds a task-structure/intent-style target beyond class labels. |
| D. Scene Reconstruction & World Modeling | Short-Horizon Ego-Motion Forecasting | current sensors excluding camera translation and captions | future camera-translation delta | 0.1989 MAE | 0.0989 MAE | Starts a short-horizon world-model target over wearer motion. |
python scripts/research_direction_extension_tasks.pytier2_task_suite file and directory names remain only for
stable artifact links. They should be read as provenance bundles inside the
unified 20-task suite, not as a separate benchmark tier.TIER2_TASK_BASELINES.mdtier2_task_suite_results.jsondocs/data/tier2_task_suite.jsonunified_task_model_radar.svgsingle_episode_task_model_radar.svgepisode128_task_model_radar.svgtier2_task_suite.svgTASK_SUITE_20.md. Historical provenance
links remain listed above for exact source tracing, but the public task surface
should be read as one integrated 20-task suite./path/to/python-with-h5py scripts/tier2_task_suite.pyHOMIE-toolkit or an environment with h5py because
the interaction/object targets come from the raw public-sample
annotation.hdf5. The raw HDF5 and MP4 files remain excluded from the public
repo and Hugging Face mirrors.TASK_WALKTHROUGHS.mdtask_walkthroughs.jsondocs/data/task_walkthroughs.jsondocs/data/task_surface_integrity.json| Task | Case study | Input -> process -> output |
|---|---|---|
| Action Recognition | A pouring window should be named as the current action. | all-modality window -> action label builder + classifier -> action class |
| Procedure Step Recognition | A fine action is grouped into a broader drink-preparation stage. | all-modality window -> subtask label builder + classifier -> subtask label |
| Action Boundary Detection | Detect the change from preparing to pouring. | window -> boundary builder + binary classifier -> boundary/steady |
| Next-Action Prediction | A preparing window predicts what happens 20 frames later. | current window -> future-label shift + classifier -> next action |
| Hand Trajectory Forecasting | A hand moving toward a cup becomes a future 3D hand path. | current window -> future mocap target + regressor -> hand trajectory |
| Contact State Prediction | Decide whether hand/body contact is happening. | non-contact features -> contact target + binary classifier -> contact label |
| Object Relevance Prediction | Infer milk, cup, coffee, or related objects during pouring. | non-caption features -> multi-hot object target + sigmoid heads -> object set |
| Language Grounding | Query Pour milk into coffee and retrieve the matching moment. | text-like query + candidates -> projection + cosine ranker -> ranked windows |
| Cross-Modal Retrieval | Motion/IMU from pouring retrieves matching depth/video. | motion/IMU/camera -> projection + candidate index -> ranked depth/video windows |
| Cross-Modal Reconstruction | Infer depth/video features from motion, IMU, and camera pose. | source modalities -> scaler + regressor -> target modality vector |
| Temporal Order Verification | Tell whether reaching then pouring was reversed. | adjacent window pair -> pair combiner + binary classifier -> correct/reversed |
| Multimodal Synchronization Detection | Catch motion paired with visual/depth features shifted in time. | motion side + visual side -> aligned/shifted pair builder + classifier -> aligned/shifted |
1raw episode -> 20-frame windows, stride 5 -> 8,546-dimensional multimodal representation
2chronological split: first 70% train, last 30% test
3scalers are fit on train windows only| Head family | Used by | What it means |
|---|---|---|
| Linear softmax classifier | Action Recognition, Procedure Step Recognition, Action Boundary Detection, Next-Action Prediction, Contact State Prediction, Temporal Order Verification, Multimodal Synchronization Detection | z-score features, then XW+b, softmax, cross-entropy, L2 |
| Dual ridge regression/projection | Hand Trajectory Forecasting, Cross-Modal Reconstruction | z-score input/target, solve ridge regression with L2=10 |
| Ridge + cosine ranking | Language Grounding, Cross-Modal Retrieval | project one modality into another feature space, then rank candidates by cosine |
| Multi-label logistic regression | Object Relevance Prediction | z-score non-caption features, sigmoid object heads, threshold at 0.5 |
results/episode_task_suite/neural_mlp/,
and the rollup is stored in the neural_tasks section of
results/episode_task_suite/summary_report.json.| Task | Input | Minimal head | Output |
|---|---|---|---|
| Action Recognition | all featurized modalities | linear softmax | current action class |
| Procedure Step Recognition | all featurized modalities | linear softmax | current subtask class |
| Action Boundary Detection | all featurized modalities | linear softmax | steady vs action boundary |
| Next-Action Prediction | all featurized modalities at t | linear softmax | action at t+20 frames |
| Hand Trajectory Forecasting | all featurized modalities at t | ridge regression | future 10-frame left/right hand joints |
| Contact State Prediction | non-contact and non-caption signals | linear softmax | any body contact |
| Object Relevance Prediction | non-caption signals | multi-label logistic | relevant object set |
| Language Grounding | sensor windows projected to text space | ridge projection + cosine ranking | matching time window for text query |
| Cross-Modal Retrieval | motion/IMU/camera projected to visual space | ridge projection + cosine ranking | matching depth/video window |
| Cross-Modal Reconstruction | motion/IMU/camera | ridge regression | compressed depth/video target |
| Temporal Order Verification | [x_t, x_t+1, x_t+1-x_t] | binary linear softmax | correct vs reversed order |
| Multimodal Synchronization Detection | motion plus visual pair | binary linear softmax | aligned vs shifted by 8 windows |
| Experiment | Main score | Accuracy | Notes |
|---|---|---|---|
| Motion-only action | 0.9688 macro-F1 | 0.9828 | Uses motion/IMU features only |
| Current all-feature action | 0.9829 macro-F1 | 0.9863 | 8,546-dimensional multimodal representation |
| Motion-only subtask | 0.9528 macro-F1 | 0.9759 | Strong within-episode subtask signal |
| Current all-feature subtask | 0.9173 macro-F1 | 0.9828 | High accuracy, lower class-balanced score |
| Cross-modal retrieval | 0.3678 top-5 | n/a | Motion/IMU/camera/audio retrieves matching depth/video |
| Transition detection | 0.6118 macro-F1 | 0.9080 | Boundary F1 is 0.1250 |
| Hand trajectory forecast | 0.8647 MPJPE | n/a | Predicts future hand-joint trajectory |
| Neural MLP hand forecast | 0.1079 MPJPE | n/a | Same features/split, nonlinear regression head |
| Neural MLP temporal order | 0.8520 F1 | 0.8578 | Strong improvement on adjacent-window ordering |
| Neural MLP misalignment | 0.7153 F1 | 0.7009 | Detects shifted motion/visual/audio pairs better than the linear head |
| Audio ablation | +0.0418 mean delta | n/a | Current audio variant improves the primary metric on 6 walkthrough-backed task contracts |
| Alternate audio representation | +0.0936 mean delta | n/a | Alternate audio-window representation improves over the baseline audio variant on 6 walkthrough-backed task contracts |
scripts/audio_ablation_and_raw_upgrade.py
reuses the real task-suite windows and evaluates six variants for
every task: current inputs, no audio, audio-only, alternate audio-only, audio
representation replacement, and all inputs plus the alternate audio representation.| Readout | Value |
|---|---|
| Tasks where current audio improves the primary metric | 6 / 12 original contracts |
| Mean current-audio delta | +0.0418 |
| Tasks where alternate audio representation improves over baseline audio | 6 / 12 original contracts |
| Mean alternate-representation delta vs baseline audio | +0.0936 |
results/audio_ablation/AUDIO_ABLATION_SUMMARY.mdresults/audio_ablation/audio_ablation_metrics.csvresults/audio_ablation/audio_delta_summary.csvdocs/data/audio_ablation_summary.jsondocs/assets/charts/audio_ablation_delta.svg--include-neural for the original core task contracts
using 80 epochs, hidden size 128, batch size 128, and CPU execution. It is not a
foundation model result; it is a controlled nonlinear-head comparison over the
same 8,546-dimensional multimodal representation.| Task | Neural metric | Minimal metric | Readout |
|---|---|---|---|
| Action Recognition | 0.0148 macro-F1 | 0.0500 macro-F1 | Still blocked by unseen future classes |
| Procedure Step Recognition | 0.0281 macro-F1 | 0.0506 macro-F1 | Same single-episode split limitation |
| Action Boundary Detection | 0.5862 macro-F1 | 0.6118 macro-F1 | Similar to the linear baseline |
| Next-Action Prediction | 0.0419 macro-F1 | 0.0593 macro-F1 | Same unseen-label issue |
| Hand Trajectory Forecasting | 0.1079 MPJPE | 0.8647 MPJPE | Neural regression improves this target |
| Contact State Prediction | 1.0000 macro-F1 | 1.0000 macro-F1 | Degenerate one-class sample |
| Object Relevance Prediction | 0.1679 micro-F1 | 0.1803 micro-F1 | Similar weak object signal |
| Language Grounding | 0.0168 MRR | 0.0160 MRR | Similar ranking behavior |
| Cross-Modal Retrieval | 0.1300 MRR | 0.2693 MRR | Linear ridge remains stronger here |
| Cross-Modal Reconstruction | -0.0102 R2 | -0.0153 R2 | Small improvement but still weak |
| Temporal Order Verification | 0.8520 F1 | 0.5400 F1 | Neural head captures local temporal structure |
| Multimodal Synchronization Detection | 0.7153 F1 | 0.5052 F1 | Neural head improves alignment detection |
results/single_episode_diagnostics/object_labels/window_object_labels.csv
exports 1,161 real window-level object-label sets from annotation.hdf5.results/single_episode_diagnostics/modality_ablation/ablation_metrics.csv
recomputes all 96 task/modality cells, including object relevance.results/single_episode_diagnostics/timeline_overlay/timeline_overlay.csv
aligns 2,079 existing prediction rows back to the episode timeline.results/single_episode_diagnostics/alignment_stress/alignment_shift_metrics.csv
evaluates cross-modal retrieval under explicit time shifts.docs/single_episode_explorer.html is a static interactive page for
inspecting window labels, objects, predictions, modality statistics, and
diagnostic scores.notes/reproducibility_audit.md for the
commands and verification evidence.1first 70% of the episode -> train
2last 30% of the episode -> testresults/episode_task_suite/feature_manifest.json.