Switch language한국어
Back to the list

Masked Next-Scale Prediction for Self-supervised Scene Text Recognition

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Masked Next-Scale Prediction, a self-supervised framework for scene text recognition.

  2. The method combines cross-scale feature prediction, masked reconstruction, and multi-scale linguistic alignment to capture the hierarchical structure of text.

  3. It delivers state-of-the-art results on Union14M and standard benchmarks while reducing dependence on labeled data.

  4. The approach improves robustness to varying text scales, layouts, and image structures.

Read the original