- Multiple voices in one Alexa skill response with SSML voice tags
For multi-voice Alexa skill responses, return SSML with a voice tag per speaker, like [speak][voice name="Kendra"]her line[/voice] then your narration[/speak].
- Amazon Polly limits: 3000 chars and 10-min cutoff, throttling returns 400 not 429
Chunk text to 3000 billed characters for SynthesizeSpeech, or switch to StartSpeechSynthesisTask for anything longer and poll the task for the S3 output.
- Saving Amazon Polly output to mp3 with boto3: read the AudioStream
After client.synthesize_speech, do response['AudioStream'].read() and write the result with open('speech.mp3', 'wb').
- Using Amazon Polly voices inside an Alexa skill from Lambda
To use Polly voices in an Alexa skill, have your Lambda function call Polly and return the audio through SSML in the skill response.
- Amazon Polly lexicons are region-specific and named per request
Upload the lexicon with PutLexicon in the same region you synthesize in; a cross-region call silently ignores it.
- Polly long text: switch to the async task at 6000 characters
Synthesizing long-form audio with Polly: anything over 6000 characters (3000 billed) fails SynthesizeSpeech with TextLengthExceededException, so switch to StartSpeechSynthesisTask for documents and chapters.
- Polly lexicons: language must match the voice, and they're per-region
Debugging Polly pronunciation lexicons: a lexicon that 'does nothing' is usually a language or region mismatch, not a content problem.
- Polly sample rates are per-format: pcm, opus, and telephony each differ
Setting Polly sample rates: don't copy one rate across formats.