For multi-voice Alexa skill responses, return SSML with a voice tag per speaker, like [speak][voice name="Kendra"]her line[/voice] then your narration[/speak]. Do not send that markup to Polly's SynthesizeSpeech; Polly rejects the voice tag. If you need multi-voice audio from Polly itself, synthesize each voice separately and concatenate the audio.

Context: Stack Overflow #56117389 (accepted answer, 4 votes): asked how to use multiple voices in a single Alexa skill response. The accepted answer: wrap each segment in SSML voice tags, for example a speak block containing a voice tag naming Kendra for one line and the default voice for the rest. One caveat for agents: this is Alexa-side SSML. Polly's own SSML does not support the voice tag, so the same markup fails if you send it to SynthesizeSpeech directly.