Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

pinecone-sparse-english-v0

Use the pinecone-sparse-english-v0 embedding or reranking model with Pinecone: specs and index setup. Built on the innovations of the DeepImpact.

  `pinecone-sparse-english-v0` is the right choice for workflows that need a learned sparse-vector representation — for example, when your application already produces sparse vectors upstream of Pinecone, or when you're pairing it with a dense encoder in a [single-index hybrid](/guides/index-data-search-hybrid-search) workflow on the Vectors API.

  You can call the [`embed` operation](/guides/inference-2026-07-generate-vectors) through Pinecone Inference to turn text into vectors without writing to an index. That differs from [`upsert_records`](/guides/database-2026-07-data-plane-upsert-records) on an index with integrated embedding, where each request embeds and stores records in one step. To see how embedding consumption appears in billing and usage reports, see [Embedding tokens](/guides/manage-cost-monitor-usage-and-costs#embedding-tokens).


Built on the innovations of the [DeepImpact architecture](https://arxiv.org/pdf/2104.12016), the model directly estimates the lexical importance of tokens by using their context, unlike traditional retrieval models like BM25, which rely solely on term frequency. The model outperforms BM25 by up to 44% (average 23%) NDCG\@10 on Text Retrieval Conference (TREC) Deep Learning Tracks and up to 24% (8% on average) on BEIR. For more information see our blog post on [cascading retrieval](https://www.pinecone.io/blog/cascading-retrieval/)

When using the model to [generate embeddings](/guides/inference-2026-07-generate-vectors) directly, you must specify the `input_type` as either `query` or `passage`. When [creating an index with integrated embedding](/guides/database-2026-07-control-plane-create-for-model), `input_type` defaults to `query` for reads and `passage` for writes. Optionally, you can:

* Return the string tokens using `"return_tokens": true`.
* Raise the max input tokens limit from the default of `512` to the maximum of `2048` using `"max_tokens_per_sequence": 2048`.
* Return an error when the input exceeds `max_tokens_per_sequence` using `"truncate": "NONE"`.

### Installation
Python
pip install --upgrade pinecone
JavaScript
npm install @pinecone-database/pinecone@latest
### Create index
Python
from pinecone import Pinecone

pc = Pinecone(api_key="YOUR_API_KEY")

# Create an index for sparse vectors with integrated embedding
index_name = "pinecone-sparse-english-v0"

pc.create_index_for_model(
    name=index_name,
    cloud="aws",
    region="us-east-1",
    embed={
        "model": "pinecone-sparse-english-v0",
        "field_map": {
            "text": "text" # Map the record field to be embedded
        },
        "read_parameters": {
            "max_tokens_per_sequence": 2048 # Max input tokens for queries
        },
        "write_parameters": {
            "max_tokens_per_sequence": 2048 # Max input tokens for upserts and updates
        }
    }
)

index = pc.Index(index_name)
JavaScript
import { Pinecone } from '@pinecone-database/pinecone'

const pc = new Pinecone({ apiKey: 'YOUR_API_KEY' });

// Create an index for sparse vectors with integrated inference
const indexName = "pinecone-sparse-english-v0"

await pc.createIndexForModel({
  name: indexName,
  cloud: 'aws',
  region: 'us-east-1',
  embed: {
    model: 'pinecone-sparse-english-v0',
    fieldMap: { text: 'text' }, // Map the record field to be embedded
  },
  waitUntilReady: true,
});

const index = pc.index(indexName);
### Embed & upsert
Python
data = [
    {"id": "vec1", "text": "Apple is a popular fruit known for its sweetness and crisp texture."},
    {"id": "vec2", "text": "The tech company Apple is known for its innovative products like the iPhone."},
    {"id": "vec3", "text": "Many people enjoy eating apples as a healthy snack."},
    {"id": "vec4", "text": "Apple Inc. has revolutionized the tech industry with its sleek designs and user-friendly interfaces."},
    {"id": "vec5", "text": "An apple a day keeps the doctor away, as the saying goes."},
    {"id": "vec6", "text": "Apple Computer Company was founded on April 1, 1976, by Steve Jobs, Steve Wozniak, and Ronald Wayne as a partnership."}
]

index.upsert_records(
    namespace="example-namespace",
    records=data
)
JavaScript
const data = [
  { id: 'vec1', text: 'Apple is a popular fruit known for its sweetness and crisp texture.' },
  { id: 'vec2', text: 'The tech company Apple is known for its innovative products like the iPhone.' },
  { id: 'vec3', text: 'Many people enjoy eating apples as a healthy snack.' },
  { id: 'vec4', text: 'Apple Inc. has revolutionized the tech industry with its sleek designs and user-friendly interfaces.' },
  { id: 'vec5', text: 'An apple a day keeps the doctor away, as the saying goes.' },
  { id: 'vec6', text: 'Apple Computer Company was founded on April 1, 1976, by Steve Jobs, Steve Wozniak, and Ronald Wayne as a partnership.' }
];

await index.namespace('example-namespace').upsert(data);
### Query
Python
query_payload = {
    "inputs": {
        "text": "Tell me about the tech company known as Apple."
    },
    "top_k": 3
}

results = index.search(
    namespace="example-namespace",
    query=query_payload
)

print(results)
JavaScript

const response = await namespace.searchRecords({
  query: {
    topK: 2,
    inputs: { text: 'Tell me about the tech company known as Apple.' },
  }
});

console.log(response);

Lorem Ipsum

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu