[π Released Code]
[π€ Datasets] [π€ Checkpoints]
[π Tech Report] [π€ Paper]
Figure A. PaDT pipeline.
We are pleased to introduce Patch-as-Decodable Token (PaDT), a unified paradigm that enables multimodal large language models (MLLMs) to directly generate both textual and visual outputs.At the core of PaDT are Visual Reference Tokens (VRTs). Unlike conventional MLLMs that represent⦠See the full description on the dataset page:
https://huggingface.co/datasets/PaDT-MLLM/RefCOCO.