Ingest your own files
Turn a folder of documents into a searchable index: extract text, chunk it, embed the chunks, and upsert.
Make a folder of raw documents searchable. The flow is: extract text, chunk it, embed the chunks, upsert, then search. Your coding agent can run this end to end (see the Quickstart hub), or follow the steps.
Prerequisites
Section titled “Prerequisites”- A Pinecone account and API key (get one).
- Python 3.10+.
- The Pinecone Python SDK:
pip install --upgrade pinecone.
Ingest your own files and search
Section titled “Ingest your own files and search”Set your API key
Set your API key as an environment variable so the SDK can authenticate:
Bash export PINECONE_API_KEY="YOUR_API_KEY"Extract text from your files
Convert each file (PDF, DOCX, HTML, and so on) to plain text using a parser of your choice. This step happens outside Pinecone. The result should be a list of records, each with an
idandtext, like thedocslist below. It runs as-is, so you can complete the quickstart first and swap in your own extracted text after.Python docs = [ {"id": "handbook-1", "text": "Refund requests must be submitted within 30 days of purchase."}, {"id": "handbook-2", "text": "Enterprise customers get support with a 4-hour response time."}, # ...extracted from your files ]Chunk the text
Split your text into smaller pieces so each fits your embedding model's input limit. The function below is a simple length-based split you can run as-is. For smarter approaches (by sentence, token, or document structure), see chunking strategies.
Python def chunk(text, size=500): return [text[i:i + size] for i in range(0, len(text), size)] chunks = [ {"id": f"{d['id']}#{i}", "text": part} for d in docs for i, part in enumerate(chunk(d["text"])) ]Create an index with a dense-vector field
Set
dimensionto match your embedding model.llama-text-embed-v2outputs 1024 dimensions. Only the vector field goes in the schema; other fields are stored on the documents (non-schema fields are stored as metadata, capped at 40 KB per document).Python import os, time from pinecone import Pinecone, SchemaBuilder pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"]) schema = ( SchemaBuilder() .add_dense_vector_field(name="embedding", dimension=1024, metric="cosine") .build() ) if not pc.indexes.exists(name="my-files"): pc.indexes.create(name="my-files", schema=schema) while not pc.indexes.describe(name="my-files").status.ready: time.sleep(2) index = pc.Index(name="my-files")Embed the chunks and upsert
Use Pinecone Inference to embed the chunks (in batches of 96, this model's per-call limit), then
upsertthe vectors alongside the text. For a large dataset, usebatch_upsertor Import instead.Python # llama-text-embed-v2 accepts up to 96 inputs per call, so embed in batches. embeddings = [] for i in range(0, len(chunks), 96): resp = pc.inference.embed( model="llama-text-embed-v2", inputs=[c["text"] for c in chunks[i:i + 96]], parameters={"input_type": "passage"}, ) embeddings.extend(resp) index.documents.upsert( namespace="__default__", documents=[ {"_id": c["id"], "embedding": e['values'], "text": c["text"]} for c, e in zip(chunks, embeddings) ], ) time.sleep(5) # documents are indexed asynchronously, so wait a momentSearch your documents
To search, embed the query with the same model you used for the documents, then rank documents by vector similarity.
Python q = pc.inference.embed( model="llama-text-embed-v2", inputs=["what is the refund policy?"], parameters={"input_type": "query"}, ) resp = index.documents.search( namespace="__default__", top_k=3, score_by=[{"type": "dense_vector", "fields": ["embedding"], "values": q[0]['values']}], include_fields=["*"], ) for m in resp.matches: print(m._id, m._score, getattr(m, "text", ""))