**TL;DR:** Sparse indexes use dotproduct on high-dimensional sparse vectors for keyword-style retrieval. Generate sparse vectors consistently at ingest and query time. Create the index with metric="dotproduct" and sparse vector support. Generate sparse vectors with a sparse embedding model (e.g. pinecone-sparse-english-v0) or BM25 term weights.

## The fix

1. **Create the index** with `metric="dotproduct"` and sparse vector support. Sparse vectors are high-dimensional with mostly zero values; the cap is 2048 non-zero entries per vector.

2. **Generate sparse vectors** with a sparse embedding model (e.g. pinecone-sparse-english-v0) or BM25 term weights. The critical rule: ingest-time and query-time generation must use the same tokenizer and weighting, or nothing matches.

3. **Upsert** records with `sparse_values` (indices and values arrays).

4. **Query** with a sparse query vector built the same way.

5. **Evaluate on keyword queries.** Sparse retrieval should win on exact-term queries and lose on paraphrase queries versus dense. If it loses on exact terms, the tokenization mismatches.

6. Different tokenizers at ingest and query: the classic silent failure. Pin one pipeline and share the code between both paths.

7. Treating sparse as a drop-in for dense: paraphrases and synonyms will not match. Sparse is keyword retrieval; pair it with dense (hybrid) for general search.

8. Exceeding 2048 non-zero values: trim or re-weight; over-long sparse vectors get rejected.

## When to use this

- This covers exactly what the title says: Pinecone workflow.
- You are setting this up for the first time, or auditing an existing setup.
- You want the key gotchas in one place before you start.

## When not to use this

- You are doing a different workflow with Pinecone; these steps are specific to the title above.
- You need the full reference docs; this is the short path, not the manual.

## Compatibility

- Not pinned to a specific version; follows current Pinecone behavior.