Bring your own vectors
Create an index with a dense-vector field, upsert embeddings you already have, and run semantic search.
If you already generate embeddings, store them directly. Create an index with a dense-vector field, upsert your vectors, and rank by similarity. Pinecone does no embedding on your behalf here.
Prerequisites
Section titled “Prerequisites”- A Pinecone account and API key (get one).
- Python 3.10+.
- The Pinecone Python SDK:
pip install --upgrade pinecone.
Upsert your vectors and search
Section titled “Upsert your vectors and search”Set your API key
Set your API key as an environment variable so the SDK can authenticate:
Bash export PINECONE_API_KEY="YOUR_API_KEY"Create an index with a dense-vector field
Set
dimensionto match your embedding model's output, and pick a distancemetric(cosine,dotproduct, oreuclidean). Only the vector field goes in the schema; other fields liketextare stored on the documents (non-schema fields are stored as metadata, capped at 40 KB per document).Python import os, time from pinecone import Pinecone, SchemaBuilder pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"]) schema = ( SchemaBuilder() .add_dense_vector_field(name="embedding", dimension=1024, metric="cosine") .build() ) if not pc.indexes.exists(name="byo-vectors"): pc.indexes.create(name="byo-vectors", schema=schema) while not pc.indexes.describe(name="byo-vectors").status.ready: time.sleep(2) index = pc.Index(name="byo-vectors")Upsert your vectors
Each document carries its
embedding(a list of floats matching the schema'sdimension) plus any fields you want to store. Replace the truncated vectors below with your real embeddings.Python index.documents.upsert( namespace="__default__", documents=[ # each `embedding` is a full list of 1024 floats (shown truncated) {"_id": "1", "embedding": [0.12, 0.04, ...], "text": "Refund requests must be submitted within 30 days."}, {"_id": "2", "embedding": [0.08, 0.21, ...], "text": "Enterprise support responds within 4 hours."}, ], ) time.sleep(5) # documents are indexed asynchronously, so wait a momentSearch by vector similarity
Embed your query with the same model you used for the documents, then rank by the dense-vector field (
embedding).Python query_embedding = [0.10, 0.05, ...] # embed your query text with the same model resp = index.documents.search( namespace="__default__", top_k=3, score_by=[{"type": "dense_vector", "fields": ["embedding"], "values": query_embedding}], include_fields=["*"], ) for m in resp.matches: print(m._id, m._score, getattr(m, "text", ""))