Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Scale a dedicated read nodes index

Scale a Pinecone dedicated read nodes index by adding replicas for query throughput and shards for storage capacity.

Dedicated read nodes don't yet scale automatically. You decide when to scale, guided by two signals:

Scenario What to check What to do
Query load CPU usage Add replicas
Data volume Index fullness Add shards

Diagnose which one you're facing before you scale. Adding shards won't relieve query latency that's driven by CPU saturation, and adding replicas won't create room on a shard that's running out of space.

CPU usage reflects query pressure. As query load outgrows what your replicas can serve, query latency climbs and throughput drops.

Pinecone exposes CPU usage as the pinecone_db_drn_cpu_usage_percent metric, reported per shard. To collect it, monitor your index with Prometheus.

Alert on the highest value across your shards rather than the index-wide average, which can look healthy while a single shard is already saturated:

Shell
max(avg_over_time(pinecone_db_drn_cpu_usage_percent{index_name="docs-example"}[5m]))

Add a replica when this exceeds 80%. Averaging over a window keeps a brief spike from triggering a scale-up, so shorten or lengthen the window to suit how quickly your traffic shifts.

Throughput scales approximately linearly with replicas. For high availability, allocate n+1 replicas, where n is the minimum number of replicas required to serve your expected throughput at your target latency. See Number of replicas.

To add or remove replicas, call Configure an index. This operation doesn't require downtime, but can take up to 30 minutes to complete. In the request body, set the following fields:

Field Value Notes
spec.serverless.read_capacity.mode Dedicated
spec.serverless.read_capacity.dedicated.scaling Manual
spec.serverless.read_capacity.dedicated.manual.replicas Desired number of replicas Add replicas to increase query throughput
Request
PINECONE_API_KEY="YOUR_API_KEY"
INDEX_NAME="YOUR_INDEX_NAME"

curl -X PATCH "https://api.pinecone.io/indexes/$INDEX_NAME" \
     -H "Accept: application/json" \
     -H "Content-Type: application/json" \
     -H "Api-Key: $PINECONE_API_KEY" \
     -H "X-Pinecone-Api-Version: 2025-10" \
     -d '{
           "spec": {
             "serverless": {
               "read_capacity": {
                 "mode": "Dedicated",
                 "dedicated": {
                   "scaling": "Manual",
                   "manual": {
                     "replicas": 2
                   }
                 }
               }
             }
           }
         }'
Response
{
  "name": "example-dedicated-index",
  "vector_type": "dense",
  "metric": "cosine",
  "dimension": 1024,
  "status": {
    "ready": true,
    "state": "Ready"
  },
  "host": "example-dedicated-index-1c6ab6aa.svc.aped-4627-b74a.pinecone.io",
  "spec": {
    "serverless": {
      "region": "us-east-1",
      "cloud": "aws",
      "read_capacity": {
        "mode": "Dedicated",
        "dedicated": {
          "node_type": "b1",
          "scaling": "Manual",
          "manual": {
            "shards": 1,
            "replicas": 2 // <---- desired state
          }
        },
        "status": {
          "state": "Scaling",
          "current_shards": 1,
          "current_replicas": 1 // <---- current state
        }
      }
    }
  },
  "deletion_protection": "disabled",
  "tags": null,
  "embed": {
    "model": "llama-text-embed-v2",
    "field_map": {
      "text": "text"
    },
    "dimension": 1024,
    "metric": "cosine",
    "write_parameters": {
      "dimension": 1024,
      "input_type": "passage",
      "truncate": "END"
    },
    "read_parameters": {
      "dimension": 1024,
      "input_type": "query",
      "truncate": "END"
    },
    "vector_type": "dense"
  }
}

Index fullness reflects data volume. Writes are blocked once the index reaches capacity, while reads continue normally.

To check the current value, see Monitor index fullness.

To add or remove shards, call Configure an index. This operation doesn't require downtime, but can take up to 30 minutes to complete. In the request body, set the following fields:

Field Value Notes
spec.serverless.read_capacity.mode Dedicated
spec.serverless.read_capacity.dedicated.scaling Manual
spec.serverless.read_capacity.dedicated.manual.shards Desired number of shards Each shard provides 250 GB of storage
Request
PINECONE_API_KEY="YOUR_API_KEY"
INDEX_NAME="YOUR_INDEX_NAME"

curl -X PATCH "https://api.pinecone.io/indexes/$INDEX_NAME" \
     -H "Accept: application/json" \
     -H "Content-Type: application/json" \
     -H "Api-Key: $PINECONE_API_KEY" \
     -H "X-Pinecone-Api-Version: 2025-10" \
     -d '{
           "spec": {
             "serverless": {
               "read_capacity": {
                 "mode": "Dedicated",
                 "dedicated": {
                   "scaling": "Manual",
                   "manual": {
                     "shards": 3
                   }
                 }
               }
             }
           }
         }'
Response
{
  "name": "example-dedicated-index",
  "vector_type": "dense",
  "metric": "cosine",
  "dimension": 1024,
  "status": {
    "ready": true,
    "state": "Ready"
  },
  "host": "example-dedicated-index-1c6ab6aa.svc.aped-4627-b74a.pinecone.io",
  "spec": {
    "serverless": {
      "region": "us-east-1",
      "cloud": "aws",
      "read_capacity": {
        "mode": "Dedicated",
        "dedicated": {
          "node_type": "b1",
          "scaling": "Manual",
          "manual": {
            "shards": 3, // <---- desired state
            "replicas": 1
          }
        },
        "status": {
          "state": "Scaling",
          "current_shards": 2, // <---- current state
          "current_replicas": 1
        }
      }
    }
  },
  "deletion_protection": "disabled",
  "tags": null,
  "embed": {
    "model": "llama-text-embed-v2",
    "field_map": {
      "text": "text"
    },
    "dimension": 1024,
    "metric": "cosine",
    "write_parameters": {
      "dimension": 1024,
      "input_type": "passage",
      "truncate": "END"
    },
    "read_parameters": {
      "dimension": 1024,
      "input_type": "query",
      "truncate": "END"
    },
    "vector_type": "dense"
  }
}
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu