Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Save on Pinecone costs

Save on Pinecone costs by using bulk import over upsert, namespaces for multitenancy, and query patterns that reduce read unit consumption.

In addition to the workload optimizations below, you can lower your effective rate:

  • Prepaid credits and annual commitments earn discounted usage rates. See Prepaid credits.
  • Dedicated read nodes can lower cost for sustained, high read throughput when you fully use the provisioned capacity (see Choose the right index capacity mode).
  • Volume discounts. Standard and Enterprise customers can contact Support to discuss cost optimization and discounts.

Prefer bulk import over upsert for large loads

Section titled “Prefer bulk import over upsert for large loads”

When you need to populate a new namespace or load a large dataset (for example, millions of records or hundreds of GB), importing from object storage is usually the most efficient and cost-effective path compared to streaming upserts.

  • Import is optimized for one-time or bulk loads from Parquet in your object store and is priced based on data read during the job. See Import cost.
  • Upsert is priced in write units based on request size; many small requests can cost more than fewer large ones for the same total data. See Write unit pricing.

Use upsert (including batch upsert) for ongoing, incremental ingestion after your initial load. For how import and upsert compare, see the data ingestion overview.

Partitioning tenants with namespaces instead of many separate indexes often lowers storage overhead and query cost, because cost depends in part on how much data each query scans. For patterns and rationale, see Manage cost.

  • Avoid returning vector values from a query when you don't need them (include_values=false, the default), especially at high top_k. Values are typically the largest part of a response, so leaving them out lowers egress and helps you stay within your plan's egress allowance.
  • Use metadata filters so queries scan fewer records where your workload allows.

For sustained, high read throughput, dedicated read nodes can be more cost-effective than on-demand when you fully use provisioned read capacity. For spiky or low-QPS workloads, on-demand may be cheaper. See When to use dedicated read nodes and Understanding cost.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu