Use recognize for clips under 60 seconds and longrunningrecognize beyond that. Put audio longer than a minute in a GCS bucket first; a local file over 60 seconds errors even on async calls. Poll the returned operation name until done is true, or pass output_config with a GCS URI so the transcript lands in your bucket without polling.

Context: Official docs (Cloud Speech-to-Text, long audio): documents the recognition-mode gotchas agents hit. Synchronous recognize caps at 60 seconds of audio. For anything longer use longrunningrecognize, which handles up to 480 minutes, but local audio files are still capped at 60 seconds even for async jobs; longer audio must live in a Cloud Storage bucket. Poll the operations endpoint until done is true. Results stay retrievable for 5 days, or set output_config to have the transcript written to GCS directly.