NVIDIA and University of Maryland Researchers Release Audio Flamingo Next (AF-Next): A Super Powerful Open Large Audio-Language Model

TL;DR AI
2 min readKey summary
NVIDIA and University of Maryland researchers released Audio Flamingo Next (AF-Next), an open large audio-language model for speech, environmental sounds, and music understanding.
AF-Next includes variants for instruction following, multi-step reasoning, and audio captioning, with a long-context architecture and timestamp-aware temporal reasoning.
Trained on large-scale audio data, it improves question answering and long-form reasoning over audio, advancing open-source multimodal AI for the audio domain.
