This is an FP8 dynamically quantized (W8A8) version of OpenGVLab/InternVL3_5-38Boptimized for high-performance inference.
The quantization process uses a specialized recipe that preserves the model's core visual understanding capabilities while reducing the memory footprint by nearly 40%.
Notes
32k max context length
reasoning parser ready to go, requires system prompt to run in thinking mode