Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
DEJIMA-dataset – Dataset by MIL-UT | AlphaNeural AI
You can deploy this model and start earning money today!
MIL-UT
/
DEJIMA-dataset
like
0
image-to-text
visual-question-answering
image-captioning
visual-question-answering
monolingual
ja
apache-2.0
10M<n<100M
json
image
text
datasets
pandas
polars
mlcroissant
2512.00773
us
Views
No views yet
Model card
Files and Versions
Community
API
DEJIMA Dataset Overview
DEJIMA is a large-scale Japanese multimodal (image + text) dataset constructed through a scalable and fully reproducible pipeline combining:
Web-scale image collection
Strict filtering and deduplication
Detection-driven evidence extraction
LLM-based caption/VQA generation under grounding constraints
DEJIMA contains:
3.88M image–caption pairs (DEJIMA-Cap) 3.88M image–question–answer pairs (DEJIMA-VQA)
All annotations are in… See the full description on the dataset page:
https://huggingface.co/datasets/MIL-UT/DEJIMA-dataset
.