Image Thresholding: Understanding Bias of Evaluation Metrics towards Specific Evaluation Functions

TL;DR AI
2 min readKey summary
A study on multilevel image thresholding finds that common evaluation metrics like SSIM and PSNR are biased toward Otsu’s objective over Kapur’s entropy.
The authors tested all possible thresholds on BSDS500 images and compared thresholding objective functions against image-quality scores.
Otsu’s criterion showed a consistently stronger alignment with SSIM and PSNR than Kapur’s entropy.
The results suggest that widely used segmentation benchmarks may favor certain methods and distort fair comparisons.
