{
"image": <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=336x372>,
"lang": "en",
"instruction": "Summarize what is happening in the scene in a concise manner",
"response": "The image appears to be a vintage photograph...",
}
aihub-table (Only Ko)
{
"file_id": 291040,
"images": <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=1223x680>,
"type": "기본표",
"task": "markdown"… See the full description on the dataset page: https://huggingface.co/datasets/MLP-VLM/bllossom-vision.