Set your baseURL to https://api.deepinfra.com/v1/openai exactly, with openai after v1. Using https://api.deepinfra.com/v1 is the common failure that breaks every request. Pass models as HuggingFace-style ids (deepseek-ai/DeepSeek-V4-Flash, meta-llama/Meta-Llama-3.1-8B-Instruct). For cost tracking, read usage.estimated_cost off each response instead of estimating from your own token counts.

Context: ai-lcr provider doc for DeepInfra (github.com/ai-lcr/ai-lcr website/content/docs/providers/deepinfra.mdx): 'The one quirk that trips everyone: DeepInfra serves the OpenAI-compatible API at /v1/openai/chat/completions, the /v1/ comes before openai.' Since createOpenAICompatible appends /chat/completions to the baseURL, the base must be https://api.deepinfra.com/v1/openai, not .../v1. Model ids are HuggingFace-style paths like deepseek-ai/DeepSeek-V4-Flash, and responses carry usage.estimated_cost for billing reconciliation.