Use the Galileo playground to iterate on a prompt, then move the winning setup into a code-run experiment for anything you need to reproduce. Playground outputs are saved to the experiment log stream, so the comparison history stays intact when you switch. Keep prompts, model names, and dataset versions in your own version control; the log stream records what ran but is not a substitute for code. When a playground result looks better than a code run, diff the exact prompt text and parameters, not your memory of them.

Context: Official docs (docs.galileo.ai, Running experiments): documents the difference that matters for reproducibility: playground runs are interactive, but code-run experiments are the ones you can version and rerun; playground outputs are still saved to the experiment log stream.