If your Patronus evaluations run but you cannot see what the app actually did, you forgot @traced(). Wrap the functions being evaluated with the @traced() decorator so Patronus captures inputs, outputs, timing and call structure; pair it with an @evaluator() function that returns a bool or score for the result. The trace and the evaluation then land together in the platform, which is what makes a failing eval debuggable instead of just a red number.

Context: Docs (patronus-py SDK repo): documents the tracing pattern that trips agents evaluating LLM apps. Define evaluators with the @evaluator() decorator (e.g. an exact_match function returning a bool) and wrap the functions under test with @traced(). Tracing automatically captures execution details, timing and results, which then show up in the Patronus platform alongside the evaluation scores.