Base Models Look Human To AI Detectors
TL;DR AI
2 min readKey summary
A new study found that base language models were often judged as human by GPTZero and Pangram, while instruction-tuned outputs were more likely flagged as AI.
The researchers introduced HIP, an iterative paraphrasing pipeline designed to make text look more human while preserving meaning.
HIP improved detector evasion across model sizes, including Llama-3 and Qwen-3, with relatively little semantic loss.
The findings suggest current AI detectors may be reacting to instruction-tuning artifacts more than reliably identifying machine-generated text.
