Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Ingest your own files

Turn a folder of documents into a searchable index: extract text, chunk it, embed the chunks, and upsert.

Make a folder of raw documents searchable. The flow is: extract text, chunk it, embed the chunks, upsert, then search. Your coding agent can run this end to end (see the Quickstart hub), or follow the steps.

  • A Pinecone account and API key (get one).
  • Python 3.10+.
  • The Pinecone Python SDK: pip install --upgrade pinecone.
  1. Set your API key

    Set your API key as an environment variable so the SDK can authenticate:

    Bash
    export PINECONE_API_KEY="YOUR_API_KEY"
  2. Extract text from your files

    Convert each file (PDF, DOCX, HTML, and so on) to plain text using a parser of your choice. This step happens outside Pinecone. The result should be a list of records, each with an id and text, like the docs list below. It runs as-is, so you can complete the quickstart first and swap in your own extracted text after.

    Python
    docs = [
        {"id": "handbook-1", "text": "Refund requests must be submitted within 30 days of purchase."},
        {"id": "handbook-2", "text": "Enterprise customers get support with a 4-hour response time."},
        # ...extracted from your files
    ]
  3. Chunk the text

    Split your text into smaller pieces so each fits your embedding model's input limit. The function below is a simple length-based split you can run as-is. For smarter approaches (by sentence, token, or document structure), see chunking strategies.

    Python
    def chunk(text, size=500):
        return [text[i:i + size] for i in range(0, len(text), size)]
    
    chunks = [
        {"id": f"{d['id']}#{i}", "text": part}
        for d in docs
        for i, part in enumerate(chunk(d["text"]))
    ]
  4. Create an index with a dense-vector field

    Set dimension to match your embedding model. llama-text-embed-v2 outputs 1024 dimensions. Only the vector field goes in the schema; other fields are stored on the documents (non-schema fields are stored as metadata, capped at 40 KB per document).

    Python
    import os, time
    from pinecone import Pinecone, SchemaBuilder
    
    pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])
    
    schema = (
        SchemaBuilder()
          .add_dense_vector_field(name="embedding", dimension=1024, metric="cosine")
          .build()
    )
    if not pc.indexes.exists(name="my-files"):
        pc.indexes.create(name="my-files", schema=schema)
    
    while not pc.indexes.describe(name="my-files").status.ready:
        time.sleep(2)
    
    index = pc.Index(name="my-files")
  5. Embed the chunks and upsert

    Use Pinecone Inference to embed the chunks (in batches of 96, this model's per-call limit), then upsert the vectors alongside the text. For a large dataset, use batch_upsert or Import instead.

    Python
    # llama-text-embed-v2 accepts up to 96 inputs per call, so embed in batches.
    embeddings = []
    for i in range(0, len(chunks), 96):
        resp = pc.inference.embed(
            model="llama-text-embed-v2",
            inputs=[c["text"] for c in chunks[i:i + 96]],
            parameters={"input_type": "passage"},
        )
        embeddings.extend(resp)
    
    index.documents.upsert(
        namespace="__default__",
        documents=[
            {"_id": c["id"], "embedding": e['values'], "text": c["text"]}
            for c, e in zip(chunks, embeddings)
        ],
    )
    
    time.sleep(5)  # documents are indexed asynchronously, so wait a moment
  6. Search your documents

    To search, embed the query with the same model you used for the documents, then rank documents by vector similarity.

    Python
    q = pc.inference.embed(
        model="llama-text-embed-v2",
        inputs=["what is the refund policy?"],
        parameters={"input_type": "query"},
    )
    
    resp = index.documents.search(
        namespace="__default__",
        top_k=3,
        score_by=[{"type": "dense_vector", "fields": ["embedding"], "values": q[0]['values']}],
        include_fields=["*"],
    )
    
    for m in resp.matches:
        print(m._id, m._score, getattr(m, "text", ""))
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu