Install the patronus package, call patronus.init() with your API key available, then reference a remote judge evaluator: RemoteEvaluator("judge", "patronus:is-concise") runs Patronus's built-in conciseness judge on their infrastructure. For your own logic, decorate a function with @evaluator(); it can return a bool for a simple pass/fail or an EvaluationResult with score, pass_, text_output, and explanation. To score a whole dataset, use run_experiment with your dataset rows, a task function that produces the model output per row, and your evaluators list, then export with result.to_csv(). Wrap local functions in FuncEvaluatorAdapter when mixing them with remote evaluators in an experiment.

Context: How do I run my first LLM evaluation with the Patronus Python SDK? I want to score model outputs against a criterion like conciseness.