Data modeling
Model your data in Pinecone using documents with dense_vector, sparse_vector, full-text string, and metadata fields for efficient retrieval.
Pinecone has two ways to model your data, and the choice is made when you create the index: an index created with a document schema holds documents, while an index created with a dense or sparse vector type holds records. Both hold a unique _id and optional metadata and live in namespaces; they differ in how many ranking signals a single item can carry and which API you call. Both are fully supported.
- Documents are the unit of data in an index with a document schema, the model behind full-text search. A document is a JSON object that can carry text fields (ranked with BM25 and Lucene queries), a dense vector, a sparse vector, and metadata in a single index, and you pick the ranking signal per query with
score_by. Use documents for text-first search or any workload that needs more than one ranking signal in one index. - Records are the unit of data for indexes with dense vectors and indexes with sparse vectors, the Vectors API. A record carries one vector (or raw text for Pinecone to embed) plus metadata. Use records for vector-only workloads.
Documents
Section titled “Documents”A document is the unit of data in an index with a document schema, a JSON object with a required _id field, the ranking fields declared in the index's schema, and any number of metadata fields. Documents support multiple field types in a single record: a dense_vector field (for semantic search), a sparse_vector field (for sparse-vector search), one or more string fields with full_text_search enabled (for full-text search with BM25 and Lucene queries), plus any metadata you upsert alongside them.
The schema, declared at index creation, tells Pinecone how to rank each ranking field. Schema field types:
dense_vector, indexed for ANN similarity search.sparse_vector, indexed for sparse-vector search.stringwith a nestedfull_text_searchconfig object ({}enables with all defaults; optional sub-fields:language,stemming,stop_words), indexed for BM25 ranking and Lucene queries. Lowercasing and the token length cap are server-applied and can't be overridden.
Metadata fields aren't declared in the schema. Any field you upsert that's not declared in the schema is stored on the document, returned via include_fields, and automatically indexed for filtering. Pinecone infers the metadata field type from the values you upsert: strings, numbers (floating point), booleans, and arrays of strings are all supported.
Document fields can hold structured values: a metadata string_list field holds an array of strings; a dense_vector field holds an array of floats; a sparse_vector field is an object with two parallel arrays, indices (token positions) and values (token weights).
A schema can declare up to 100 string fields with full_text_search enabled, but at most one dense_vector field and at most one sparse_vector field per index.
Example document for an index with title, body, embedding, and category fields:
{
"_id": "document1#chunk1",
"title": "Introduction to Vector Databases",
"body": "First chunk of the document content...",
"embedding": [0.0236, -0.0329, ..., -0.0104, 0.0086],
"category": "tutorial"
}Field-name rules:
- Must be unique, non-empty strings.
- Must not start with
_(reserved for system-managed fields like_idand_score) or$(reserved for filter operators). - Limited to 64 bytes.
For the full schema reference (language and analyzer options, multi-field schemas, scoring methods), see Full-text search.
Schema validation
Section titled “Schema validation”Each document in an upsert request's documents array is validated against the schema. If any document fails validation, the entire upsert fails and nothing is written.
| Scenario | Result |
|---|---|
| Field value doesn't match declared type (for schema-declared fields) | Error, request fails |
| Document or request exceeds a size or count limit | Error, request fails |
| Field not in schema | Stored on the document and auto-indexed for filtering as metadata |
Field name starts with _ or $ |
Error, request fails |
| Schema field missing from a document | OK, schema fields are optional, as long as the document carries at least one |
Document with only _id and metadata (no schema fields) |
Error, request fails |
Document missing _id |
Error, request fails |
Schema patterns
Section titled “Schema patterns”The same document model supports several common schema shapes. Pick the one that matches the signal you want to rank by, and plan your fields up front: schema migration isn't supported after index creation. Filters are deterministic per document and apply before scoring; choose your hard yes/no constraints (including text-match operators on FTS-enabled string fields) first, then pick a score_by method to rank whatever remains. See Filters vs. scoring.
Single text field — keyword search only (FTS)
Use when you want BM25 keyword ranking on one piece of text per document (a review body, a support ticket, a product description) and you don't have embeddings to manage.
from pinecone import SchemaBuilder
schema = (
SchemaBuilder()
.add_string_field("review_text", full_text_search={"language": "en"})
.build()
)
pc.indexes.create(name="book-reviews", schema=schema)A document upserted into this index looks like:
{
"_id": "review-1234",
"review_text": "Beautifully written exploration of contact, communication, and civilization across cosmic distances. The pacing is uneven but the central premise carries you through."
}Search with a single text clause (the score_by type, not a field type — this clause runs BM25 ranking on the named string field):
index.documents.search(
namespace="reviews",
top_k=10,
score_by=[{"type": "text", "fields": ["review_text"], "query": "civilization"}],
)See Full-text search.
Multi-field FTS — score across two text fields (e.g. body + summary)
Use when a document has more than one piece of text that should both contribute to ranking, for example, a long review_text plus a short review_summary. Pinecone combines the per-field BM25 scores into one ranking per document.
schema = (
SchemaBuilder()
.add_string_field("review_text", full_text_search={"language": "en"})
.add_string_field("review_summary", full_text_search={"language": "en"})
.build()
)
pc.indexes.create(name="book-reviews-multi", schema=schema)A document upserted into this index looks like:
{
"_id": "review-1234",
"review_text": "Beautifully written exploration of contact, communication, and civilization across cosmic distances. The pacing is uneven but the central premise carries you through.",
"review_summary": "Monumental science fiction with uneven pacing",
"category": "science-fiction",
"rating": 4.5
}category and rating aren't declared in the schema. They're upserted as metadata, automatically indexed for filtering, and usable in filter expressions.
Pass two text clauses in score_by; the server combines them into one ranking, with each contributing field weighted equally in 2026-07.
index = pc.Index(name="book-reviews-multi")
index.documents.search(
namespace="reviews",
top_k=5,
score_by=[
{"type": "text", "fields": ["review_text"], "query": "disappointing"},
{"type": "text", "fields": ["review_summary"], "query": "Disappointing"},
],
include_fields=["*"],
)Dense + FTS — semantic and keyword in one index
Most workloads that combine semantic ranking with keyword matching reach for this pattern: rank by dense (or sparse) similarity, restricted to documents that contain a specific term or phrase. Common examples include semantic search over patents, regulatory filings, internal knowledge bases, or other technical literature where the right answer must contain a specific term. A single schema can include one dense_vector field plus any number of FTS-enabled string fields:
schema = (
SchemaBuilder()
.add_string_field("book_title", full_text_search={"language": "en"})
.add_string_field("review_text", full_text_search={"language": "en"})
.add_dense_vector_field("review_embedding", dimension=1024, metric="cosine")
.build()
)
pc.indexes.create(name="book-reviews-dense", schema=schema)A document upserted into this index looks like:
{
"_id": "review-1234",
"book_title": "The Three-Body Problem",
"review_text": "Beautifully written exploration of contact, communication, and civilization across cosmic distances.",
"review_embedding": [0.012, -0.087, 0.153, ...]
}review_embedding is a 1024-dim list of floats produced by your dense embedding model. Use the same model at query time so the query vector lives in the same space.
A single search request ranks by one scoring type. With this schema you have two query options:
Option A, dense ranking restricted by a text-match filter (the most common hybrid pattern):
index = pc.Index(name="book-reviews-dense")
# query_embedding is a 1024-dim list of floats from your embedding model.
query_embedding = embed("beautifully written, hard sci-fi")
index.documents.search(
namespace="reviews",
top_k=5,
score_by=[
{"type": "dense_vector", "fields": ["review_embedding"], "values": query_embedding},
],
filter={"review_text": {"$match_phrase": "beautifully written"}},
)Option B, run BM25 and dense searches separately and merge client-side (when you want both signals to contribute to ranking, for example via reciprocal rank fusion):
dense_hits = index.documents.search(
namespace="reviews", top_k=50,
score_by=[{"type": "dense_vector", "fields": ["review_embedding"], "values": query_embedding}],
)
bm25_hits = index.documents.search(
namespace="reviews", top_k=50,
score_by=[{"type": "text", "fields": ["review_text"], "query": "beautifully written"}],
)
# Merge dense_hits + bm25_hits in your application (e.g. RRF) to produce final ranking.See Hybrid search for a fuller discussion.
Multi-signal index — dense + sparse + FTS in one schema
Use when a single document is best described by more than one ranking signal, for example, a video catalog where each item has frame embeddings (dense), auto-generated captions you've encoded as sparse vectors (sparse), and a transcript text field (BM25/Lucene). One schema declares all three; you pick the ranking signal per query with score_by. You don't manage a second index or any cross-index linkage.
schema = (
SchemaBuilder()
.add_dense_vector_field("frame_embedding", dimension=1024, metric="cosine")
.add_sparse_vector_field("caption_sparse")
.add_string_field("transcript", full_text_search={"language": "en"})
.build()
)
pc.indexes.create(name="video-catalog", schema=schema)A language field upserted alongside these ranking fields is treated as metadata: stored on the document, returned via include_fields, and auto-indexed for filtering.
A document upserted into this index looks like:
{
"_id": "video-7890#scene-3",
"frame_embedding": [0.012, -0.087, 0.153, ...],
"caption_sparse": {
"indices": [42, 1077, 9821],
"values": [0.41, 0.33, 0.18]
},
"transcript": "I think we should go now before it gets dark.",
"language": "en"
}frame_embedding is a 1024-dim list of floats from your dense vision model. caption_sparse is the output of your sparse encoder, an object with parallel indices (token IDs) and values (token weights) arrays.
The same index supports three different query shapes. All three assume:
index = pc.Index(name="video-catalog")
# Replace with the outputs of your encoders.
query_embedding = embed_image(query_image) # 1024-dim list of floats
query_sparse = sparse_encode("scene with a lighthouse") # {"indices": [...], "values": [...]}Semantic frame search, ranking by visual similarity:
index.documents.search(
namespace="videos",
top_k=10,
score_by=[{"type": "dense_vector", "fields": ["frame_embedding"], "values": query_embedding}],
)Caption search, ranking by sparse-vector similarity over your encoded captions:
index.documents.search(
namespace="videos",
top_k=10,
score_by=[{"type": "sparse_vector", "fields": ["caption_sparse"], "sparse_values": query_sparse}],
)Semantic search restricted to a spoken phrase, narrowing semantic frame ranking to clips where the transcript contains a specific phrase:
index.documents.search(
namespace="videos",
top_k=10,
score_by=[{"type": "dense_vector", "fields": ["frame_embedding"], "values": query_embedding}],
filter={"transcript": {"$match_phrase": "I love you"}},
)score_by selects one ranking signal per request, but every signal stays addressable on the same documents.
Sparse + dense hybrid — single-vector index (Vectors API)
Use when you're modeling data with the Vectors API (not the Documents API) and want to combine a sparse and dense vector in one record on a single index. For new document-centric projects with text data, prefer the document-index Dense + FTS pattern above.
{
"id": "doc1#chunk1",
"values": [0.0236, -0.0329, ..., -0.0104, 0.0086],
"sparse_values": {
"indices": [822745112, 1009084850, ...],
"values": [1.7958984, 0.41577148, ...]
},
"metadata": { "document_id": "doc1", "chunk_number": 1 }
}See Hybrid search.
Records
Section titled “Records”Records are how you model data for indexes with dense vectors and indexes with sparse vectors. Each record carries one vector (dense, sparse, or both for single-index hybrid) plus optional metadata, and you can upsert raw text in place of a vector when the index is integrated with an embedding model.
When you upsert pre-generated vectors, each record consists of the following:
- ID: A unique string identifier for the record.
- Vector: A dense vector for semantic search, a sparse vector for sparse-vector search, or both for single-index hybrid search (Vectors API).
- Metadata (optional): A flat JSON document containing key-value pairs with additional information (nested objects aren't supported). You can filter by metadata when searching or deleting records.
Example:
{
"id": "document1#chunk1",
"values": [0.0236663818359375, -0.032989501953125, ..., -0.01041412353515625, 0.0086669921875],
"metadata": {
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 1,
"chunk_text": "First chunk of the document content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
}
}{
"id": "document1#chunk1",
"sparse_values": {
"values": [1.7958984, 0.41577148, ..., 4.4414062, 3.3554688],
"indices": [822745112, 1009084850, ..., 3517203014, 3590924191]
},
"metadata": {
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 1,
"chunk_text": "First chunk of the document content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
}
}{
"id": "document1#chunk1",
"values": [0.0236663818359375, -0.032989501953125, ..., -0.01041412353515625, 0.0086669921875],
"sparse_values": {
"values": [1.7958984, 0.41577148, ..., 4.4414062, 3.3554688],
"indices": [822745112, 1009084850, ..., 3517203014, 3590924191]
},
"metadata": {
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 1,
"chunk_text": "First chunk of the document content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
}
}When you upsert raw text for Pinecone to convert to vectors automatically, each record consists of the following:
- ID: A unique string identifier for the record.
- Text: The raw text for Pinecone to convert to a dense vector for semantic search or a sparse vector for sparse-vector search, depending on the embedding model integrated with the index. This field name must match the
embed.field_mapdefined in the index. - Metadata (optional): All additional fields are stored as record metadata. You can filter by metadata when searching or deleting records.
Example:
{
"_id": "document1#chunk1",
"chunk_text": "First chunk of the document content...", // Text to convert to a vector.
"document_id": "document1", // This and subsequent fields stored as metadata.
"document_title": "Introduction to Vector Databases",
"chunk_number": 1,
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
}Use structured IDs
Section titled “Use structured IDs”Use a structured, human-readable format for record IDs, including ID prefixes that reflect the type of data you're storing, for example:
- Document chunks:
document_id#chunk_number - User data:
user_id#data_type#item_id - Multi-tenant data:
tenant_id#document_id#chunk_id
Choose a delimiter for your ID prefixes that won't appear elsewhere in your IDs. Common patterns include:
document1#chunk1- Using hash delimiterdocument1_chunk1- Using underscore delimiterdocument1:chunk1- Using colon delimiter
Structuring IDs in this way provides several advantages:
- Efficiency: Applications can quickly identify which record it should operate on.
- Clarity: Developers can easily understand what they're looking at when examining records.
- Flexibility: ID prefixes enable list operations for fetching and updating records.
Include metadata
Section titled “Include metadata”Include metadata key-value pairs that support your application's key operations, for example:
- Enable query-time filtering: Add fields for time ranges, categories, or other criteria for filtering searches for increased accuracy and relevance.
- Link related chunks: Use fields like
document_idandchunk_numberto keep track of related records and enable efficient chunk deletion and document updates. - Link back to original data: Include
chunk_textordocument_urlfor traceability and user display.
Metadata keys must be strings, and metadata values must be one of the following data types:
- String
- Number (stored as a 64-bit floating point)
- Boolean (true, false)
- List of strings
Example
Section titled “Example”This example demonstrates how to manage document chunks in Pinecone using structured IDs and comprehensive metadata. It covers the complete lifecycle of chunked documents: upserting, searching, fetching, updating, and deleting chunks, and updating an entire document.
Upsert chunks
Section titled “Upsert chunks”When upserting documents that have been split into chunks, combine structured IDs with comprehensive metadata:
from pinecone.grpc import PineconeGRPC as Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# To get the unique host for an index,
# see https://docs.pinecone.io/guides/manage-data/target-an-index
index = pc.Index(host="INDEX_HOST")
index.upsert_records(
"example-namespace",
[
{
"_id": "document1#chunk1",
"chunk_text": "First chunk of the document content...",
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 1,
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
},
{
"_id": "document1#chunk2",
"chunk_text": "Second chunk of the document content...",
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 2,
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
},
{
"_id": "document1#chunk3",
"chunk_text": "Third chunk of the document content...",
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 3,
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
},
]
)from pinecone.grpc import PineconeGRPC as Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# To get the unique host for an index,
# see https://docs.pinecone.io/guides/manage-data/target-an-index
index = pc.Index(host="INDEX_HOST")
index.upsert(
namespace="example-namespace",
vectors=[
{
"id": "document1#chunk1",
"values": [0.0236663818359375, -0.032989501953125, ..., -0.01041412353515625, 0.0086669921875],
"metadata": {
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 1,
"chunk_text": "First chunk of the document content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
}
},
{
"id": "document1#chunk2",
"values": [-0.0412445068359375, 0.028839111328125, ..., 0.01953125, -0.0174560546875],
"metadata": {
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 2,
"chunk_text": "Second chunk of the document content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
}
},
{
"id": "document1#chunk3",
"values": [0.0512237548828125, 0.041656494140625, ..., 0.02130126953125, -0.0394287109375],
"metadata": {
"document_id": "document1",
"document_title": "Introduction to Vector Databases",
"chunk_number": 3,
"chunk_text": "Third chunk of the document content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"document_type": "tutorial"
}
}
]
)Search chunks
Section titled “Search chunks”To search the chunks of a document, use a metadata filter expression that limits the search appropriately:
from pinecone import Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# To get the unique host for an index,
# see https://docs.pinecone.io/guides/manage-data/target-an-index
index = pc.Index(host="INDEX_HOST")
filtered_results = index.search(
namespace="example-namespace",
query={
"inputs": {"text": "What is a vector database?"},
"top_k": 3,
"filter": {"document_id": "document1"}
},
fields=["chunk_text"]
)
print(filtered_results)from pinecone.grpc import PineconeGRPC as Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# To get the unique host for an index,
# see https://docs.pinecone.io/guides/manage-data/target-an-index
index = pc.Index(host="INDEX_HOST")
filtered_results = index.query(
namespace="example-namespace",
vector=[0.0236663818359375,-0.032989501953125, ..., -0.01041412353515625,0.0086669921875],
top_k=3,
filter={
"document_id": {"$eq": "document1"}
},
include_metadata=True,
include_values=False
)
print(filtered_results)Fetch chunks
Section titled “Fetch chunks”To retrieve all chunks for a specific document, first list the record IDs using the document prefix, and then fetch the complete records:
from pinecone.grpc import PineconeGRPC as Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# To get the unique host for an index,
# see https://docs.pinecone.io/guides/manage-data/target-an-index
index = pc.Index(host="INDEX_HOST")
# List all chunks for document1 using ID prefix
chunk_ids = []
for record_id in index.list(prefix='document1#', namespace='example-namespace'):
chunk_ids.append(record_id)
print(f"Found {len(chunk_ids)} chunks for document1")
# Fetch the complete records by ID
if chunk_ids:
records = index.fetch(ids=chunk_ids, namespace='example-namespace')
for record_id, record_data in records['vectors'].items():
print(f"Chunk ID: {record_id}")
print(f"Chunk text: {record_data['metadata']['chunk_text']}")
# Process the vector values and metadata as neededUpdate chunks
Section titled “Update chunks”To update specific chunks within a document, first list the chunk IDs, and then update individual records:
from pinecone.grpc import PineconeGRPC as Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# To get the unique host for an index,
# see https://docs.pinecone.io/guides/manage-data/target-an-index
index = pc.Index(host="INDEX_HOST")
# List all chunks for document1
chunk_ids = []
for record_id in index.list(prefix='document1#', namespace='example-namespace'):
chunk_ids.append(record_id)
# Update specific chunks (e.g., update chunk 2)
if 'document1#chunk2' in chunk_ids:
new_vector = ... # from your embedding model
index.update(
id='document1#chunk2',
values=new_vector,
set_metadata={
"document_id": "document1",
"document_title": "Introduction to Vector Databases - Revised",
"chunk_number": 2,
"chunk_text": "Updated second chunk content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-01-15",
"updated_at": "2024-02-15",
"document_type": "tutorial"
},
namespace='example-namespace'
)
print("Updated chunk 2 successfully")Delete chunks
Section titled “Delete chunks”To delete chunks of a document, use a metadata filter expression that limits the deletion appropriately:
from pinecone.grpc import PineconeGRPC as Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# To get the unique host for an index,
# see https://docs.pinecone.io/guides/manage-data/target-an-index
index = pc.Index(host="INDEX_HOST")
# Delete chunks 1 and 3
index.delete(
namespace="example-namespace",
filter={
"document_id": {"$eq": "document1"},
"chunk_number": {"$in": [1, 3]}
}
)
# Delete all chunks for a document
index.delete(
namespace="example-namespace",
filter={
"document_id": {"$eq": "document1"}
}
)Update an entire document
Section titled “Update an entire document”When the amount of chunks or ordering of chunks for a document changes, the recommended approach is to first delete all chunks using a metadata filter, and then upsert the new chunks:
from pinecone.grpc import PineconeGRPC as Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
# To get the unique host for an index,
# see https://docs.pinecone.io/guides/manage-data/target-an-index
index = pc.Index(host="INDEX_HOST")
# Step 1: Delete all existing chunks for the document
index.delete(
namespace="example-namespace",
filter={
"document_id": {"$eq": "document1"}
}
)
print("Deleted existing chunks for document1")
# Step 2: Upsert the updated document chunks
chunk1_vector = ... # from your embedding model
chunk2_vector = ...
index.upsert(
namespace="example-namespace",
vectors=[
{
"id": "document1#chunk1",
"values": chunk1_vector,
"metadata": {
"document_id": "document1",
"document_title": "Introduction to Vector Databases - Updated Edition",
"chunk_number": 1,
"chunk_text": "Updated first chunk with new content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-02-15",
"document_type": "tutorial",
"version": "2.0"
}
},
{
"id": "document1#chunk2",
"values": chunk2_vector,
"metadata": {
"document_id": "document1",
"document_title": "Introduction to Vector Databases - Updated Edition",
"chunk_number": 2,
"chunk_text": "Updated second chunk with new content...",
"document_url": "https://example.com/docs/document1",
"created_at": "2024-02-15",
"document_type": "tutorial",
"version": "2.0"
}
}
# Add more chunks as needed for the updated document
]
)
print("Successfully updated document1 with new chunks")Data freshness
Section titled “Data freshness”Pinecone is eventually consistent, so it's possible that a write (upsert, update, or delete) followed immediately by a read (query, list, or fetch) may not return the latest version of the data. If your use case requires retrieving data immediately, consider implementing a small delay or retry logic after writes.