Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Size and test a dedicated read nodes index

Calculate how many shards and replicas a Pinecone dedicated read nodes index needs, and load-test your workload to validate the configuration.

To determine how many shards your index requires, calculate your index size and then apply the number of shards formula.

A record can include a dense vector, a sparse vector, or both. Use the formula that matches your data to calculate total size:

An index of dense vectors contains records with one dense vector each.

Calculate size (assuming no sparse vectors)

Index size = Number of records × (
               ID size + 
               Metadata size +
               Dense vector dimensions × 4 bytes
             )

Where:

  • ID size and Metadata size are measured in bytes, averaged across all records.
  • Each Dense vector dimension uses 4 bytes.

Example calculations

These examples assume 8-byte IDs:

RecordsDense vector dimensionsAvg metadata sizeIndex size
500,000768500 bytes1.79 GB
1,000,00015361,000 bytes7.15 GB
5,000,000102415,000 bytes95.5 GB
10,000,00015361,000 bytes71.5 GB

An index of sparse vectors contains records with one sparse vector each.

Calculate size

Index size = Number of records × (
               ID size + 
               Metadata size +
               Number of non-zero sparse values × 8 bytes
             )

Where:

  • ID size and Metadata size are measured in bytes, averaged across all records.
  • Number of non-zero sparse values: Average number across all records. To find the count for a single record, check the length of the sparse vector's indices or values array. Each non-zero value uses 8 bytes.

Example calculations

These examples assume 8-byte IDs:

RecordsAvg number of non-zero sparse valuesAvg metadata sizeIndex size
500,00010500 bytes0.29 GB
1,000,000501,000 bytes1.41 GB
5,000,00010015,000 bytes79.0 GB
10,000,000501,000 bytes14.1 GB

An index with both dense and sparse vectors contains records that each have one dense vector and an optional sparse vector.

Calculate size

Index size = Number of records × (
               ID size + 
               Metadata size +
               Dense vector dimensions × 4 bytes + 
               Number of non-zero sparse values × 8 bytes
             )

Where:

  • ID size and Metadata size are measured in bytes, averaged across all records.
  • Each Dense vector dimension uses 4 bytes.
  • Number of non-zero sparse values: Average number across all records, including those without sparse vectors. To find the count for a single record, check the length of the sparse vector's indices or values array. Each non-zero value uses 8 bytes.

Example calculations

These examples assume 8-byte IDs:

RecordsDense vector dimensionsAvg number of non-zero sparse valuesAvg metadata sizeIndex size
500,00076810500 bytes1.83 GB
1,000,0001536501,000 bytes7.54 GB
5,000,000102410015,000 bytes99.5 GB
10,000,0001536501,000 bytes75.4 GB

To calculate the number of shards your index requires, divide the size of your index by 250 GB and round up:

Minimum shards = (Index size) / (250 GB per shard)

To maintain optimal performance, provision additional shards to keep your index at 70-80% capacity. For example, a 500 GB index should have three shards (750 GB capacity = 67% full), not two shards (500 GB capacity = 100% full).

Index size Minimum shards Recommended shards
~71 GB 1 (250 GB; 28% full) 1 (250 GB; 28% full)
~300 GB 2 (500 GB; 60% full) 2 (500 GB; 60% full)
~400 GB 2 (500 GB; 80% full) 3 (750 GB; 53% full)

To calculate the number of replicas your index requires, first test your workload to find the QPS a single replica can handle at your target latency. Then, use this formula, and round up:

Minimum replicas = (Required QPS) / (QPS per replica)

For example, if one replica handles 50 QPS at your target latency and you need 150 QPS, you need three replicas.

For how throughput scales with replicas and how to size for high availability, see Replicas.

To choose between on-demand and dedicated read nodes, or to optimize your dedicated read nodes configuration, test with your actual workload. Performance varies based on factors such as the size of your index, vector dimensionality, metadata characteristics, and query patterns.

  1. Calculate the size of your index

    Determine how many shards your index requires. See Calculate the size of your index.

  2. Create and populate a test index

    Populate a dedicated read nodes index with data representative of your workload.

  3. Migrate your test index to dedicated read nodes (if necessary)

    If your test index is on-demand, migrate it with a single b1 replica to start.

  4. Run a load test

    Send realistic query patterns against your test index, gradually increasing QPS. For example, start at 10 QPS for about 30 minutes, then step up in 10-QPS increments while monitoring latency. Note the QPS where latency crosses your target threshold.

  5. Calculate replicas

    From the QPS a single replica sustained, determine how many replicas you need for your target throughput.

  6. Adjust and re-test

    If you haven't hit your performance and cost goals, change the configuration and test again:

    Continue iterating until you meet your requirements with room for growth.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu