Masked Next-Scale Prediction for Self-supervised Scene Text Recognition

TL;DR AI
2 min readKey summary
Researchers introduced Masked Next-Scale Prediction, a self-supervised framework for scene text recognition.
The method combines cross-scale feature prediction, masked reconstruction, and multi-scale linguistic alignment to capture the hierarchical structure of text.
It delivers state-of-the-art results on Union14M and standard benchmarks while reducing dependence on labeled data.
The approach improves robustness to varying text scales, layouts, and image structures.
