Scale a dedicated read nodes index
Scale a Pinecone dedicated read nodes index by adding replicas for query throughput and shards for storage capacity.
When to scale
Section titled “When to scale”Dedicated read nodes don't yet scale automatically. You decide when to scale, guided by two signals:
| Scenario | What to check | What to do |
|---|---|---|
| Query load | CPU usage | Add replicas |
| Data volume | Index fullness | Add shards |
Diagnose which one you're facing before you scale. Adding shards won't relieve query latency that's driven by CPU saturation, and adding replicas won't create room on a shard that's running out of space.
Add or remove replicas
Section titled “Add or remove replicas”CPU usage reflects query pressure. As query load outgrows what your replicas can serve, query latency climbs and throughput drops.
Pinecone exposes CPU usage as the pinecone_db_drn_cpu_usage_percent metric, reported per shard. To collect it, monitor your index with Prometheus.
Alert on the highest value across your shards rather than the index-wide average, which can look healthy while a single shard is already saturated:
max(avg_over_time(pinecone_db_drn_cpu_usage_percent{index_name="docs-example"}[5m]))Add a replica when this exceeds 80%. Averaging over a window keeps a brief spike from triggering a scale-up, so shorten or lengthen the window to suit how quickly your traffic shifts.
Throughput scales approximately linearly with replicas. For high availability, allocate n+1 replicas, where n is the minimum number of replicas required to serve your expected throughput at your target latency. See Number of replicas.
To add or remove replicas, call Configure an index. This operation doesn't require downtime, but can take up to 30 minutes to complete. In the request body, set the following fields:
| Field | Value | Notes |
|---|---|---|
spec.serverless.read_capacity.mode |
Dedicated |
|
spec.serverless.read_capacity.dedicated.scaling |
Manual |
|
spec.serverless.read_capacity.dedicated.manual.replicas |
Desired number of replicas | Add replicas to increase query throughput |
Example
Section titled “Example”PINECONE_API_KEY="YOUR_API_KEY"
INDEX_NAME="YOUR_INDEX_NAME"
curl -X PATCH "https://api.pinecone.io/indexes/$INDEX_NAME" \
-H "Accept: application/json" \
-H "Content-Type: application/json" \
-H "Api-Key: $PINECONE_API_KEY" \
-H "X-Pinecone-Api-Version: 2025-10" \
-d '{
"spec": {
"serverless": {
"read_capacity": {
"mode": "Dedicated",
"dedicated": {
"scaling": "Manual",
"manual": {
"replicas": 2
}
}
}
}
}
}'{
"name": "example-dedicated-index",
"vector_type": "dense",
"metric": "cosine",
"dimension": 1024,
"status": {
"ready": true,
"state": "Ready"
},
"host": "example-dedicated-index-1c6ab6aa.svc.aped-4627-b74a.pinecone.io",
"spec": {
"serverless": {
"region": "us-east-1",
"cloud": "aws",
"read_capacity": {
"mode": "Dedicated",
"dedicated": {
"node_type": "b1",
"scaling": "Manual",
"manual": {
"shards": 1,
"replicas": 2 // <---- desired state
}
},
"status": {
"state": "Scaling",
"current_shards": 1,
"current_replicas": 1 // <---- current state
}
}
}
},
"deletion_protection": "disabled",
"tags": null,
"embed": {
"model": "llama-text-embed-v2",
"field_map": {
"text": "text"
},
"dimension": 1024,
"metric": "cosine",
"write_parameters": {
"dimension": 1024,
"input_type": "passage",
"truncate": "END"
},
"read_parameters": {
"dimension": 1024,
"input_type": "query",
"truncate": "END"
},
"vector_type": "dense"
}
}Add or remove shards
Section titled “Add or remove shards”Index fullness reflects data volume. Writes are blocked once the index reaches capacity, while reads continue normally.
To check the current value, see Monitor index fullness.
To add or remove shards, call Configure an index. This operation doesn't require downtime, but can take up to 30 minutes to complete. In the request body, set the following fields:
| Field | Value | Notes |
|---|---|---|
spec.serverless.read_capacity.mode |
Dedicated |
|
spec.serverless.read_capacity.dedicated.scaling |
Manual |
|
spec.serverless.read_capacity.dedicated.manual.shards |
Desired number of shards | Each shard provides 250 GB of storage |
Example
Section titled “Example”PINECONE_API_KEY="YOUR_API_KEY"
INDEX_NAME="YOUR_INDEX_NAME"
curl -X PATCH "https://api.pinecone.io/indexes/$INDEX_NAME" \
-H "Accept: application/json" \
-H "Content-Type: application/json" \
-H "Api-Key: $PINECONE_API_KEY" \
-H "X-Pinecone-Api-Version: 2025-10" \
-d '{
"spec": {
"serverless": {
"read_capacity": {
"mode": "Dedicated",
"dedicated": {
"scaling": "Manual",
"manual": {
"shards": 3
}
}
}
}
}
}'{
"name": "example-dedicated-index",
"vector_type": "dense",
"metric": "cosine",
"dimension": 1024,
"status": {
"ready": true,
"state": "Ready"
},
"host": "example-dedicated-index-1c6ab6aa.svc.aped-4627-b74a.pinecone.io",
"spec": {
"serverless": {
"region": "us-east-1",
"cloud": "aws",
"read_capacity": {
"mode": "Dedicated",
"dedicated": {
"node_type": "b1",
"scaling": "Manual",
"manual": {
"shards": 3, // <---- desired state
"replicas": 1
}
},
"status": {
"state": "Scaling",
"current_shards": 2, // <---- current state
"current_replicas": 1
}
}
}
},
"deletion_protection": "disabled",
"tags": null,
"embed": {
"model": "llama-text-embed-v2",
"field_map": {
"text": "text"
},
"dimension": 1024,
"metric": "cosine",
"write_parameters": {
"dimension": 1024,
"input_type": "passage",
"truncate": "END"
},
"read_parameters": {
"dimension": 1024,
"input_type": "query",
"truncate": "END"
},
"vector_type": "dense"
}
}