Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training

TL;DR AI
2 min readKey summary
Researchers introduced CoPES, a cooperative parameter-subspace evolution strategy for post-training tool-using LLM agents with much lower memory use.
On Qwen3.5-4B math and QA benchmarks, CoPES recovered most of full-parameter GRPO’s validation gains under the same GPU-hour budget.
CoPES also beat standard evolution strategies and LoRA-based GRPO on pass@k, improving efficiency without sacrificing much quality.
The method offers a practical path for training agentic LLMs when compute and memory are limited.
