Sparrow-2 is a real-time conversational understanding model by Tavus that jointly models turn-taking, interruptions, backchannels, and the full acoustic scene as a single unified system. It is an audio-native, streaming-first engine that operates at a native 10 ms frame rate, deciding when to listen, wait, speak, or keep speaking. The model handles noisy, multi-speaker environments, solving the cocktail party problem to enable conversational AI in kiosks, retail, vehicles, and shared spaces. Sparrow-2 is available via the Tavus API and powers Tavus PALs, achieving a 2.1% conversation failure rate—nearly four times lower than comparable alternatives. An interactive live demo lets users speak with a PAL and observe the model's real-time scene understanding.