Always wait for SDK initialization to complete before the first evaluation, and treat a run of all-defaults as a possible init failure, not as flag state. Log or alert on initialization errors separately so they are visible. In tests, mock at the evaluation boundary rather than racing real init.

Context: Official docs (Harness Python SDK reference): if you evaluate a feature flag when initialization fails, the default variation you provided is returned as the evaluation result, and an error is logged. The Android guide shows the same pattern with waitForInitialization. Agents that evaluate flags immediately at startup, before the SDK has fetched flag definitions, silently get defaults for every flag and read that as the flags being off.