This is the full-staging build of MM-ArkBench: a repository-level multimodal ArkTS / ArkUI dataset for screenshot-to-code retrieval.
Current contents:
3,977 repository records from the 4,015 retained ArkTS candidates that could be acquired/parsed.
421,107 ArkTS corpus file units.
3,274,181 extracted symbols.
8,686 image queries.
Query types: repo_screenshot, runtime_screenshot, and doc_or_promo. The current build includes 122 runtime_screenshot rows… See the full description on the dataset page:
https://huggingface.co/datasets/hreyulog/Arkts-mm-ui-full-staging.