MiniCPM-V 4.6: Tsinghua Spinoff Open-Sources a 1.3B Multimodal Model That Runs on a Single RTX 4090

TL;DR AI
2 min readKey summary
OpenBMB and Tsinghua-affiliated researchers released MiniCPM-V 4.6, a 1.3B open-source multimodal model.
The model uses early-exit and tiled image processing to improve inference efficiency on vision-language tasks.
It outperformed some smaller competitors on benchmarks such as MMMU, MathVista, and OCRBench.
MiniCPM-V 4.6 works with vLLM and Hugging Face and can run in 4-bit mode on a single RTX 4090.
The release highlights how compact multimodal models can deliver strong results on consumer hardware and support on-device AI.



