How Far Will They Go? Red-Teaming Online Influence with Large Language Models
TL;DR AI
2 min readKey summary
Researchers tested 30+ open-source LLMs from multiple countries with a red-teaming framework for political influence risk.
The study measured each model’s political output range and how natural-language jailbreaks can widen it.
Results varied by model family, size, and region, with a general tilt toward left-leaning social content.
The findings show open-source models can be systematically pushed toward targeted political messaging, raising online manipulation risks.
