If your Patronus Evaluate API integration misreads results, remember the score semantics: score_raw is higher-is-safer (closer to 1 means safer), with 0.5 as the default pass/fail cutoff, so invert it if your guardrail framework expects higher-is-riskier. Pick evaluators by name in one request ("lynx", "judge", "answer-relevance", toxicity/PII) rather than one call each. Use the managed criteria aliases like "patronus:hallucination" or "patronus:prompt-injection" when you want a curated judge config instead of building your own. Turn on explain_strategy in development so failures come with a human-readable reason you can log.

Context: Official docs (Patronus integration reference): documents the Evaluate API shape that trips agents wiring LLM guardrails. A single POST to /v1/evaluate runs one or more evaluators chosen by name: "lynx" for hallucination, "judge" for the managed LLM-as-a-judge, "answer-relevance", plus toxicity, PII and PHI evaluators. Each evaluator returns a pass/fail verdict and a raw score in [0, 1] where higher is better and anything below 0.5 fails by default; set explain_strategy to also get an explanation back.