Mobile-Aptus: Confidence-Driven Proactive and Robust Interaction in MLLM-based Mobile-Using Agents

TL;DR AI
2 min readKey summary
Researchers introduced Mobile-Aptus, a two-stage framework for confidence-driven mobile agents built on multimodal large language models.
It trains agents to output both actions and confidence scores, then corrects confidence bias with semantic similarity retrieval and direct preference optimization.
The goal is to reduce two common failures: over-execution and excessive requests for human help.
Mobile-Aptus outperformed baselines on four mobile-agent benchmarks and in dynamic real-world tests.
The work improves when agents should act on their own versus ask for help, making mobile app use more reliable.
