Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Try full-text search

Create an index, load documents, and run your first full-text (keyword) search in about five minutes.

Full-text search matches exact terms: IDs, SKUs, error codes, and phrases. Create an index, load a few documents, and run a keyword (BM25) search in about five minutes.

  • A Pinecone account and API key (get one).
  • Python 3.10+.
  • The Pinecone Python SDK: pip install --upgrade pinecone.
  1. Set your API key

    Set your API key as an environment variable so the SDK can authenticate:

    Bash
    export PINECONE_API_KEY="YOUR_API_KEY"
  2. Create an index

    Define a schema with a full-text (BM25) field, then create the index. Only search fields belong in the schema; filterable metadata like category goes in the documents and is indexed automatically.

    Python
    import os, time
    from pinecone import Pinecone, SchemaBuilder
    
    pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])
    
    schema = (
        SchemaBuilder()
          .add_string_field(name="text", full_text_search={"language": "en"})
          .build()
    )
    if not pc.indexes.exists(name="quickstart"):
        pc.indexes.create(name="quickstart", schema=schema)
    
    while not pc.indexes.describe(name="quickstart").status.ready:
        time.sleep(2)
    
    index = pc.Index(name="quickstart")
  3. Load sample documents

    Each document is a JSON object with a required _id, the text field from your schema (indexed for full-text search), and any metadata you want to attach. Here, category is metadata: it's not declared in the schema, but Pinecone indexes it automatically so you can filter on it.

    These five short strings are just samples. Real documents can carry many metadata fields (up to 40 KB per document), plus dense- or sparse-vector fields if your schema declares them, and you can upsert up to 1,000 per request.

    Python
    index.documents.upsert(
        namespace="__default__",
        documents=[
            {"_id": "1", "text": "SKU KB-1024: wireless mechanical keyboard, USB-C, brown switches", "category": "product"},
            {"_id": "2", "text": "SKU MON-4200: 27-inch 4K monitor, HDMI and USB-C", "category": "product"},
            {"_id": "3", "text": "Order 88231 shipped on 2026-08-20 via express courier", "category": "order"},
            {"_id": "4", "text": "Error E1042: connection timeout after 30 seconds", "category": "log"},
            {"_id": "5", "text": "Ticket 5567 resolved: replaced faulty power adapter", "category": "support"},
        ],
    )
    
    time.sleep(5)  # documents are indexed asynchronously, so wait a moment

Swap the sample documents for your own: keep the same documents.upsert() call and replace the text and metadata with your content. Each document needs a unique _id and the text field from your schema, plus any metadata fields you want to filter on.

Python
index.documents.upsert(
    namespace="__default__",
    documents=[
        {"_id": "doc-1", "text": "Your first document...", "category": "your-category"},
        # ...your documents
    ],
)

For large sets, use Bulk import instead of upserting one request at a time.

For the full reference (query syntax, filters, analyzers, and bulk import), see the Full-text search guide.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu