## TL;DR

Point the official OpenAI client at https://api.deepinfra.com/v1/openai for chat-style LLM usage instead of the legacy TextGeneration wrappers. The legacy wrappers (TextGeneration, AutomaticSpeechRecognition, Embeddings, TextToImage) predate the OpenAI-compatible endpoints and remain supported, but new chat code gets model portability by using the OpenAI shape.

## Fix

1. Point the official OpenAI client at https://api.deepinfra.com/v1/openai for chat-style LLM usage instead of the legacy TextGeneration wrappers.
   Expected: You get the expected result; the problem is gone.

## When to use

- You are setting up or using this DeepInfra feature.
- The symptom matches: use the OpenAI-compatible endpoint for chat-style calls.

## When NOT to use

- Unrelated DeepInfra issues (different feature, different failure).
- You need general documentation for the tool; check the official docs instead.

## Compatibility

Reported against DeepInfra. Source: https://github.com/deepinfra/deepinfra-python.

## Variant phrasings

### use the OpenAI-compatible endpoint for chat-style calls

## Why it happens

DeepInfra's legacy inference wrappers still work, but chat workloads should target the OpenAI-compatible endpoint.

## Edge cases

- If your error message differs even slightly, this is probably a different issue; search the exact text.
- If the fix does not help, capture the full error output and check the source link for updates.

## Source

https://github.com/deepinfra/deepinfra-python