Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Bring your own vectors

Create an index with a dense-vector field, upsert embeddings you already have, and run semantic search.

If you already generate embeddings, store them directly. Create an index with a dense-vector field, upsert your vectors, and rank by similarity. Pinecone does no embedding on your behalf here.

  • A Pinecone account and API key (get one).
  • Python 3.10+.
  • The Pinecone Python SDK: pip install --upgrade pinecone.
  1. Set your API key

    Set your API key as an environment variable so the SDK can authenticate:

    Bash
    export PINECONE_API_KEY="YOUR_API_KEY"
  2. Create an index with a dense-vector field

    Set dimension to match your embedding model's output, and pick a distance metric (cosine, dotproduct, or euclidean). Only the vector field goes in the schema; other fields like text are stored on the documents (non-schema fields are stored as metadata, capped at 40 KB per document).

    Python
    import os, time
    from pinecone import Pinecone, SchemaBuilder
    
    pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])
    
    schema = (
        SchemaBuilder()
          .add_dense_vector_field(name="embedding", dimension=1024, metric="cosine")
          .build()
    )
    if not pc.indexes.exists(name="byo-vectors"):
        pc.indexes.create(name="byo-vectors", schema=schema)
    
    while not pc.indexes.describe(name="byo-vectors").status.ready:
        time.sleep(2)
    
    index = pc.Index(name="byo-vectors")
  3. Upsert your vectors

    Each document carries its embedding (a list of floats matching the schema's dimension) plus any fields you want to store. Replace the truncated vectors below with your real embeddings.

    Python
    index.documents.upsert(
        namespace="__default__",
        documents=[
            # each `embedding` is a full list of 1024 floats (shown truncated)
            {"_id": "1", "embedding": [0.12, 0.04, ...], "text": "Refund requests must be submitted within 30 days."},
            {"_id": "2", "embedding": [0.08, 0.21, ...], "text": "Enterprise support responds within 4 hours."},
        ],
    )
    
    time.sleep(5)  # documents are indexed asynchronously, so wait a moment
  4. Search by vector similarity

    Embed your query with the same model you used for the documents, then rank by the dense-vector field (embedding).

    Python
    query_embedding = [0.10, 0.05, ...]  # embed your query text with the same model
    
    resp = index.documents.search(
        namespace="__default__",
        top_k=3,
        score_by=[{"type": "dense_vector", "fields": ["embedding"], "values": query_embedding}],
        include_fields=["*"],
    )
    
    for m in resp.matches:
        print(m._id, m._score, getattr(m, "text", ""))
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu