Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Decrease latency

Reduce query and upsert latency in Pinecone using namespaces, metadata filters, targeting indexes by host, and regional colocation strategies.

When you divide records into namespaces in a logical way, you speed up queries by ensuring only relevant records are scanned. The same applies to fetching records, listing record IDs, and other data operations.

In addition to increasing search accuracy and relevance, searching with metadata filters can also help decrease latency by retrieving only records that match the filter.

When you target an index by name for data operations such as upsert and query, the SDK gets the unique DNS host for the index using the describe_index operation. This is convenient for testing but should be avoided in production because describe_index uses a different API than data operations and therefore adds an additional network call and point of failure. Instead, you should get an index host once and cache it for reuse or specify the host directly.

You can get index hosts in the Pinecone console or using the describe_index operation.

The following example shows how to target an index by host directly:

Python
from pinecone.grpc import PineconeGRPC as Pineconepc = Pinecone(api_key="YOUR_API_KEY")index = pc.Index(host="INDEX_HOST")
JavaScript
import { Pinecone } from '@pinecone-database/pinecone';const pc = new Pinecone({ apiKey: 'YOUR_API_KEY' });// For the Node.js SDK, you must specify both the index host and name.const index = pc.index("INDEX_NAME", "INDEX_HOST");
Java
import io.pinecone.clients.Index;import io.pinecone.configs.PineconeConfig;import io.pinecone.configs.PineconeConnection;public class TargetIndexByHostExample {    public static void main(String[] args) {        PineconeConfig config = new PineconeConfig("YOUR_API_KEY");        config.setHost("INDEX_HOST");        PineconeConnection connection = new PineconeConnection(config);        // For the Java SDK, you must specify both the index host and name.        Index index = new Index(connection, "INDEX_NAME");    }}
Go
package mainimport (    "context"    "fmt"    "log"    "github.com/pinecone-io/go-pinecone/v4/pinecone")func main() {    ctx := context.Background()    pc, err := pinecone.NewClient(pinecone.NewClientParams{        ApiKey: "YOUR_API_KEY",    })    if err != nil {        log.Fatalf("Failed to create Client: %v", err)    }    idxConnection, err := pc.Index(pinecone.NewIndexConnParams{Host: "INDEX_HOST", Namespace: "example-namespace"})    if err != nil {        log.Fatalf("Failed to create IndexConnection for Host %v: %v", idx.Host, err)    }}

When you target an index for upserting or querying, the client establishes a TCP connection, which is a three-step process. To avoid going through this process on every request, and reduce average request latency, cache and reuse the index connection object whenever possible.

If you experience slow uploads or high query latencies, it might be because you are accessing Pinecone from your home network. To decrease latency, access Pinecone/deploy your application from a cloud environment instead, ideally from the same cloud and region as your index.

If you're batching queries, try reducing the number of queries per call to a single query vector. You can run these queries in parallel and expect roughly the same performance as with batching.

Avoid including vector values when not needed

Section titled “Avoid including vector values when not needed”

Including vector values increases response size, especially at higher top_k values, and larger responses can raise round-trip latency. If you don't need vector values in your response, leave include_values at its default of false on query. fetch always returns values, so use query when you're searching and only need IDs or metadata.

Pinecone has rate limits to protect your applications and maintain infrastructure health. Rate limits vary based on pricing plan and apply to serverless indexes only.

To handle rate limits effectively:

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu