product · inferred from evidence
meta/meta-llama-3-70b-instruct
A specific large language model version referenced in the issue.
- Llama 3 streaming repeats the previous request's first token
If your streamed Llama 3 responses start by echoing the first tokens of previous turns, you are on a client version with the streaming bug, not a model problem. The maintainer reproduced it in the client alone after ruling out stop sequence