Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Create an index

Create a Pinecone serverless index for full-text (BM25), semantic (dense vector), sparse-vector, or hybrid search with a document schema.

A Pinecone index can hold any combination of the following:

  • Documents are the unit of data in an index with a document schema — JSON records whose ranking fields are indexed according to a schema you declare at index creation. An index with a document schema can mix dense_vector, sparse_vector, and FTS-enabled string ranking fields in the same record, alongside any number of metadata fields (auto-indexed at upsert time). Use documents for full-text search (BM25 ranking on string fields with full_text_search enabled), and to combine multiple scoring methods on the same data via score_by.

  • Dense vectors are numerical representations of the meaning and relationships of text, images, or other data. Indexes of dense vectors are used for semantic search, or together with sparse vectors for hybrid search.

  • Sparse vectors are high-dimensional vectors with mostly zero values, produced by a sparse embedding model such as pinecone-sparse-english-v0. Indexes of sparse vectors are used for sparse-vector search, or together with dense vectors for hybrid search.

An index with a document schema stores typed JSON documents. The schema declares how each ranking field is indexed: as a string field with full_text_search enabled for BM25 ranking, a dense_vector for ANN similarity, or a sparse_vector. A single index can mix all three ranking field types; at query time, pick the ranking signal with score_by. Metadata fields (anything else you upsert) aren't declared in the schema — they're auto-indexed for filtering at upsert time.

Indexes with document schemas use API version 2026-07, supported through REST and the Python SDK. For other languages, call the REST endpoint directly.

The example below creates an articles index whose body field is indexed for BM25 ranking. Other fields included at upsert time are stored on each document and auto-indexed for filtering as metadata.

Python
import os
from pinecone import Pinecone, SchemaBuilder

pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])

schema = (
    SchemaBuilder()
      .add_string_field(name="body", full_text_search={})
      .build()
)

pc.indexes.create(name="articles", schema=schema)
curl
PINECONE_API_KEY="YOUR_API_KEY"

curl -X POST "https://api.pinecone.io/indexes" \
  -H "Api-Key: $PINECONE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Pinecone-Api-Version: 2026-07" \
  -d '{
    "name": "articles",
    "deployment": {
      "deployment_type": "managed",
      "cloud": "aws",
      "region": "us-east-1"
    },
    "schema": {
      "fields": {
        "body": {
          "type": "string",
          "full_text_search": {}
        }
      }
    }
  }'

A single index with a document schema can hold FTS-enabled string and dense_vector ranking fields together (the same schema can also include a sparse_vector field). A single search request ranks by one scoring type — multi-field BM25 is supported (multiple text clauses on different fields, or one query_string clause spanning fields), and any scoring method can be combined with metadata filters, including text-match filters ($match_phrase, $match_all, $match_any) on FTS-enabled string fields.

Python
import os
from pinecone import Pinecone, SchemaBuilder

pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])

schema = (
    SchemaBuilder()
      .add_string_field(name="title", full_text_search={})
      .add_string_field(name="body", full_text_search={})
      .add_dense_vector_field(name="embedding", dimension=1536, metric="cosine")
      .build()
)

pc.indexes.create(name="articles-multi", schema=schema)
curl
PINECONE_API_KEY="YOUR_API_KEY"

curl -X POST "https://api.pinecone.io/indexes" \
  -H "Api-Key: $PINECONE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Pinecone-Api-Version: 2026-07" \
  -d '{
    "name": "articles-multi",
    "deployment": {
      "deployment_type": "managed",
      "cloud": "aws",
      "region": "us-east-1"
    },
    "schema": {
      "fields": {
        "title":    { "type": "string", "full_text_search": {} },
        "body":     { "type": "string", "full_text_search": {} },
        "embedding":{ "type": "dense_vector", "dimension": 1536, "metric": "cosine" }
      }
    }
  }'

You can include additional fields (for example, category or year) at upsert time. All metadata fields are automatically indexed for filtering — they don't need to be declared in the schema. The schema is for ranking fields only; declaring a metadata-only field (string without full_text_search, string_list, float, or boolean) is rejected at index creation. To index only the metadata fields you filter on, see Configure metadata indexing.

For the full schema reference (all field types, language and analyzer options, dedicated read capacity, and Python SDK examples), see Full-text search.

You can create an index that stores dense vectors with integrated vector embedding, or one that stores vectors generated with an external embedding model.

If you want to upsert and search with source text and have Pinecone convert it to dense vectors automatically, create an index with integrated embedding as follows:

  • Provide a name for the index.
  • Set cloud and region to the cloud and region where the index should be deployed.
  • Set embed.model to one of Pinecone's hosted embedding models.
  • Set embed.field_map to the name of the field in your source document that contains the data for embedding.

Other parameters are optional. See the API reference for details.

Python
from pinecone import Pinecone

pc = Pinecone(api_key="YOUR_API_KEY")

index_name = "integrated-dense-py"

if not pc.has_index(index_name):
    pc.create_index_for_model(
        name=index_name,
        cloud="aws",
        region="us-east-1",
        embed={
            "model":"llama-text-embed-v2",
            "field_map":{"text": "chunk_text"}
        }
    )
JavaScript
import { Pinecone } from '@pinecone-database/pinecone'

const pc = new Pinecone({ apiKey: 'YOUR_API_KEY' });

await pc.createIndexForModel({
  name: 'integrated-dense-js',
  cloud: 'aws',
  region: 'us-east-1',
  embed: {
    model: 'llama-text-embed-v2',
    fieldMap: { text: 'chunk_text' },
  },
  waitUntilReady: true,
});
Java
import io.pinecone.clients.Pinecone;
import org.openapitools.db_control.client.ApiException;
import org.openapitools.db_control.client.model.CreateIndexForModelRequest;
import org.openapitools.db_control.client.model.CreateIndexForModelRequestEmbed;
import org.openapitools.db_control.client.model.DeletionProtection;
import java.util.HashMap;
import java.util.Map;

public class CreateIntegratedIndex {
    public static void main(String[] args) throws ApiException {
        Pinecone pc = new Pinecone.Builder("YOUR_API_KEY").build();
        String indexName = "integrated-dense-java";
        String region = "us-east-1";
        HashMap<String, String> fieldMap = new HashMap<>();
        fieldMap.put("text", "chunk_text");
        CreateIndexForModelRequestEmbed embed = new CreateIndexForModelRequestEmbed()
                .model("llama-text-embed-v2")
                .fieldMap(fieldMap);
        Map<String, String> tags = new HashMap<>();
        tags.put("environment", "development");
        pc.createIndexForModel(
                indexName,
                CreateIndexForModelRequest.CloudEnum.AWS,
                region,
                embed,
                DeletionProtection.DISABLED,
                tags
        );
    }
}
Go
package main

import (
    "context"
    "fmt"
    "log"

    "github.com/pinecone-io/go-pinecone/v4/pinecone"
)

func main() {
    ctx := context.Background()

    pc, err := pinecone.NewClient(pinecone.NewClientParams{
        ApiKey: "YOUR_API_KEY",
    })
    if err != nil {
        log.Fatalf("Failed to create Client: %v", err)
    }

  	indexName := "integrated-dense-go"
   	deletionProtection := pinecone.DeletionProtectionDisabled

    idx, err := pc.CreateIndexForModel(ctx, &pinecone.CreateIndexForModelRequest{
		Name:   indexName,
		Cloud:  pinecone.Aws,
		Region: "us-east-1",
		Embed: pinecone.CreateIndexForModelEmbed{
			Model:    "llama-text-embed-v2",
			FieldMap: map[string]interface{}{"text": "chunk_text"},
		},
        DeletionProtection: &deletionProtection,
        Tags:   &pinecone.IndexTags{ "environment": "development" },
	})
    if err != nil {
        log.Fatalf("Failed to create serverless integrated index: %v", idx.Name)
    } else {
        fmt.Printf("Successfully created serverless integrated index: %v", idx.Name)
    }
}
curl
PINECONE_API_KEY="YOUR_API_KEY"

curl -X POST "https://api.pinecone.io/indexes/create-for-model" \
     -H "Content-Type: application/json" \
     -H "Api-Key: $PINECONE_API_KEY" \
     -H "X-Pinecone-Api-Version: 2025-10" \
     -d '{
           "name": "integrated-dense-curl",
           "cloud": "aws",
           "region": "us-east-1",
           "embed": {
             "model": "llama-text-embed-v2",
             "field_map": {
               "text": "chunk_text"
             }
           }
         }'
CLI
# Target the project where you want to create the index.
pc target -o "example-org" -p "example-project"
# Create the index.
pc index create \
  --name "integrated-dense-cli" \
  --metric "cosine" \
  --cloud "aws" \
  --region "us-east-1" \
  --model "llama-text-embed-v2" \
  --field_map "text=chunk_text" \
  --tags "environment=development"

If you use an external embedding model to convert your data to dense vectors, create an index as follows:

  • Provide a name for the index.
  • Set the vector_type to dense.
  • Specify the dimension and similarity metric of the vectors you'll store in the index. This should match the dimension and metric supported by your embedding model.
  • Set spec.cloud and spec.region to the cloud and region where the index should be deployed. For Python, you also need to import the ServerlessSpec class.

Other parameters are optional. See the API reference for details.

Python
from pinecone.grpc import PineconeGRPC as Pinecone
from pinecone import ServerlessSpec

pc = Pinecone(api_key="YOUR_API_KEY")

index_name = "standard-dense-py"

if not pc.has_index(index_name):
    pc.create_index(
        name=index_name,
        vector_type="dense",
        dimension=1536,
        metric="cosine",
        spec=ServerlessSpec(
            cloud="aws",
            region="us-east-1"
        ),
        deletion_protection="disabled",
        tags={
            "environment": "development"
        }
    )
JavaScript
import { Pinecone } from '@pinecone-database/pinecone'

const pc = new Pinecone({ apiKey: 'YOUR_API_KEY' });

await pc.createIndex({
  name: 'standard-dense-js',
  vectorType: 'dense',
  dimension: 1536,
  metric: 'cosine',
  spec: {
    serverless: {
      cloud: 'aws',
      region: 'us-east-1'
    }
  },
  deletionProtection: 'disabled',
  tags: { environment: 'development' }, 
});
Java
import io.pinecone.clients.Pinecone;
import org.openapitools.db_control.client.model.IndexModel;
import org.openapitools.db_control.client.model.DeletionProtection;
import java.util.HashMap;

public class CreateServerlessIndexExample {
    public static void main(String[] args) {
        Pinecone pc = new Pinecone.Builder("YOUR_API_KEY").build();
        String indexName = "standard-dense-java";
        String cloud = "aws";
        String region = "us-east-1";
        String vectorType = "dense";
        Map<String, String> tags = new HashMap<>();
        tags.put("environment", "development");
        pc.createServerlessIndex(
            indexName,
            "cosine", 
            1536, 
            cloud,
            region,
            DeletionProtection.DISABLED, 
            tags, 
            vectorType
        );
    }
}
Go
package main

import (
    "context"
    "fmt"
    "log"

    "github.com/pinecone-io/go-pinecone/v4/pinecone"
)

func main() {
    ctx := context.Background()

    pc, err := pinecone.NewClient(pinecone.NewClientParams{
        ApiKey: "YOUR_API_KEY",
    })
    if err != nil {
        log.Fatalf("Failed to create Client: %v", err)
    }

    // Serverless index
  	indexName := "standard-dense-go"
	vectorType := "dense"
    dimension := int32(1536)
    metric := pinecone.Cosine
	deletionProtection := pinecone.DeletionProtectionDisabled

    idx, err := pc.CreateServerlessIndex(ctx, &pinecone.CreateServerlessIndexRequest{
        Name:               indexName,
        VectorType:         &vectorType,
        Dimension:          &dimension,
        Metric:             &metric,
        Cloud:              pinecone.Aws,
        Region:             "us-east-1",
        DeletionProtection: &deletionProtection,
        Tags:               &pinecone.IndexTags{ "environment": "development" },
    })
    if err != nil {
        log.Fatalf("Failed to create serverless index: %v", err)
    } else {
        fmt.Printf("Successfully created serverless index: %v", idx.Name)
    }
}
curl
PINECONE_API_KEY="YOUR_API_KEY"

curl -X POST "https://api.pinecone.io/indexes" \
     -H "Accept: application/json" \
     -H "Content-Type: application/json" \
     -H "Api-Key: $PINECONE_API_KEY" \
     -H "X-Pinecone-Api-Version: 2025-10" \
     -d '{
           "name": "standard-dense-curl",
           "vector_type": "dense",
           "dimension": 1536,
           "metric": "cosine",
           "spec": {
             "serverless": {
               "cloud": "aws",
               "region": "us-east-1"
              }
            },
            "tags": {
              "environment": "development"
            },
            "deletion_protection": "disabled"
         }'
CLI
# Target the project where you want to create the index.
pc target -o "example-org" -p "example-project"
# Create the index.
pc index create \
  --name "standard-dense-cli" \
  --vector_type "dense" \
  --dimension 1536 \
  --metric "cosine" \
  --cloud "aws" \
  --region "us-east-1" \
  --tags "environment=development" \
  --deletion_protection "disabled"

You can create an index that stores sparse vectors with integrated vector embedding, or one that stores vectors generated with an external embedding model.

If you want to upsert and search with source text and have Pinecone convert it to sparse vectors automatically, create an index with integrated embedding as follows:

  • Provide a name for the index.
  • Set cloud and region to the cloud and region where the index should be deployed.
  • Set embed.model to one of Pinecone's hosted sparse embedding models.
  • Set embed.field_map to the name of the field in your source document that contains the text for embedding.
  • If needed, embed.read_parameters and embed.write_parameters can be used to override the default model embedding behavior.

Other parameters are optional. See the API reference for details.

Python
from pinecone import Pinecone

pc = Pinecone(api_key="YOUR_API_KEY")

index_name = "integrated-sparse-py"

if not pc.has_index(index_name):
    pc.create_index_for_model(
        name=index_name,
        cloud="aws",
        region="us-east-1",
        embed={
            "model":"pinecone-sparse-english-v0",
            "field_map":{"text": "chunk_text"}
        }
    )
JavaScript
import { Pinecone } from '@pinecone-database/pinecone'

const pc = new Pinecone({ apiKey: 'YOUR_API_KEY' });

await pc.createIndexForModel({
  name: 'integrated-sparse-js',
  cloud: 'aws',
  region: 'us-east-1',
  embed: {
    model: 'pinecone-sparse-english-v0',
    fieldMap: { text: 'chunk_text' },
  },
  waitUntilReady: true,
});
Java
import io.pinecone.clients.Pinecone;
import org.openapitools.db_control.client.ApiException;
import org.openapitools.db_control.client.model.CreateIndexForModelRequest;
import org.openapitools.db_control.client.model.CreateIndexForModelRequestEmbed;
import org.openapitools.db_control.client.model.DeletionProtection;
import java.util.HashMap;
import java.util.Map;

public class CreateIntegratedIndex {
    public static void main(String[] args) throws ApiException {
        Pinecone pc = new Pinecone.Builder("YOUR_API_KEY").build();
        String indexName = "integrated-sparse-java";
        String region = "us-east-1";
        HashMap<String, String> fieldMap = new HashMap<>();
        fieldMap.put("text", "chunk_text");
        CreateIndexForModelRequestEmbed embed = new CreateIndexForModelRequestEmbed()
                .model("pinecone-sparse-english-v0")
                .fieldMap(fieldMap);
        Map<String, String> tags = new HashMap<>();
        tags.put("environment", "development");
        pc.createIndexForModel(
                indexName,
                CreateIndexForModelRequest.CloudEnum.AWS,
                region,
                embed,
                DeletionProtection.DISABLED,
                tags
        );
    }
}
Go
package main

import (
    "context"
    "fmt"
    "log"

    "github.com/pinecone-io/go-pinecone/v4/pinecone"
)

func main() {
    ctx := context.Background()

    pc, err := pinecone.NewClient(pinecone.NewClientParams{
        ApiKey: "YOUR_API_KEY",
    })
    if err != nil {
        log.Fatalf("Failed to create Client: %v", err)
    }

  	indexName := "integrated-sparse-go"
	deletionProtection := pinecone.DeletionProtectionDisabled

    idx, err := pc.CreateIndexForModel(ctx, &pinecone.CreateIndexForModelRequest{
		Name:   indexName,
		Cloud:  pinecone.Aws,
		Region: "us-east-1",
		Embed: pinecone.CreateIndexForModelEmbed{
			Model:    "pinecone-sparse-english-v0",
			FieldMap: map[string]interface{}{"text": "chunk_text"},
		},
        DeletionProtection: &deletionProtection,
        Tags:   &pinecone.IndexTags{ "environment": "development" },

	})
    if err != nil {
        log.Fatalf("Failed to create serverless integrated index: %v", idx.Name)
    } else {
        fmt.Printf("Successfully created serverless integrated index: %v", idx.Name)
    }
}
curl
PINECONE_API_KEY="YOUR_API_KEY"

curl -X POST "https://api.pinecone.io/indexes/create-for-model" \
     -H "Content-Type: application/json" \
     -H "Api-Key: $PINECONE_API_KEY" \
     -H "X-Pinecone-Api-Version: 2025-10" \
     -d '{
           "name": "integrated-sparse-curl",
           "cloud": "aws",
           "region": "us-east-1",
           "embed": {
             "model": "pinecone-sparse-english-v0",
             "field_map": {
               "text": "chunk_text"
             }
           }
         }'
CLI
# Target the project where you want to create the index.
pc target -o "example-org" -p "example-project"
# Create the index.
pc index create \
  --name "integrated-sparse-cli" \
  --cloud "aws" \
  --region "us-east-1" \
  --model "pinecone-sparse-english-v0" \
  --field_map "text=chunk_text" \
  --tags "environment=development"

If you use an external embedding model to convert your data to sparse vectors, create an index as follows:

  • Provide a name for the index.
  • Set the vector_type to sparse.
  • Set the distance metric to dotproduct. Indexes that store sparse vectors don't support other distance metrics.
  • Set spec.cloud and spec.region to the cloud and region where the index should be deployed.

Other parameters are optional. See the API reference for details.

Python
from pinecone import Pinecone, ServerlessSpec

pc = Pinecone(api_key="YOUR_API_KEY")

index_name = "standard-sparse-py"

if not pc.has_index(index_name):
    pc.create_index(
        name=index_name,
        vector_type="sparse",
        metric="dotproduct",
        spec=ServerlessSpec(cloud="aws", region="us-east-1")
    )
JavaScript
import { Pinecone } from '@pinecone-database/pinecone'

const pc = new Pinecone({ apiKey: 'YOUR_API_KEY' });

await pc.createIndex({
  name: 'standard-sparse-js',
  vectorType: 'sparse',
  metric: 'dotproduct',
  spec: {
    serverless: {
      cloud: 'aws',
      region: 'us-east-1'
    },
  },
});
Java
import io.pinecone.clients.Pinecone;
import org.openapitools.db_control.client.model.DeletionProtection;

import java.util.*;

public class SparseIndex {
    public static void main(String[] args) throws InterruptedException {
        // Instantiate Pinecone class
        Pinecone pinecone = new Pinecone.Builder("YOUR_API_KEY").build();

        // Create the index
        String indexName = "standard-sparse-java";
        String cloud = "aws";
        String region = "us-east-1";
        String vectorType = "sparse";
        Map<String, String> tags = new HashMap<>();
        tags.put("env", "test");
        pinecone.createSparseServelessIndex(indexName,
                cloud,
                region,
                DeletionProtection.DISABLED,
                tags,
                vectorType);
    }
}
Go
package main

import (
	"context"
	"fmt"
	"log"

	"github.com/pinecone-io/go-pinecone/v4/pinecone"
)

func main() {
	ctx := context.Background()

	pc, err := pinecone.NewClient(pinecone.NewClientParams{
		ApiKey: "YOUR_API_KEY",
	})
	if err != nil {
		log.Fatalf("Failed to create Client: %v", err)
	}

	indexName := "standard-sparse-go"
	vectorType := "sparse"
	metric := pinecone.Dotproduct
	deletionProtection := pinecone.DeletionProtectionDisabled

	idx, err := pc.CreateServerlessIndex(ctx, &pinecone.CreateServerlessIndexRequest{
		Name:               indexName,
		Metric:             &metric,
		VectorType:         &vectorType,
		Cloud:              pinecone.Aws,
		Region:             "us-east-1",
		DeletionProtection: &deletionProtection,
	})
	if err != nil {
		log.Fatalf("Failed to create serverless index: %v", err)
	} else {
		fmt.Printf("Successfully created serverless index: %v", idx.Name)
	}
}
curl
PINECONE_API_KEY="YOUR_API_KEY"

curl -X POST "https://api.pinecone.io/indexes" \
     -H "Accept: application/json" \
     -H "Content-Type: application/json" \
     -H "Api-Key: $PINECONE_API_KEY" \
     -H "X-Pinecone-Api-Version: 2025-10" \
     -d '{
            "name": "standard-sparse-curl",
            "vector_type": "sparse",
            "metric": "dotproduct",
            "spec": {
               "serverless": {
                  "cloud": "aws",
                  "region": "us-east-1"
               }
            }
         }'
CLI
# Target the project where you want to create the index.
pc target -o "example-org" -p "example-project"
# Create the index.
pc index create \
  --name "standard-sparse-cli" \
  --vector_type "sparse" \
  --metric "dotproduct" \
  --cloud "aws" \
  --region "us-east-1" \
  --tags "environment=development"

You can restore an index from a backup, regardless of whether it stores dense or sparse vectors. For more details, see Restore an index.

To index only the metadata fields you filter on, rather than all of them, see Configure metadata indexing.

When creating an index, you must choose the cloud and region where you want the index to be hosted. The following table lists the available public clouds and regions and the plans that support them:

Cloud Region Supported plans Availability phase
aws us-east-1 (Virginia) Starter, Builder, Standard, Enterprise General availability
aws us-west-2 (Oregon) Builder, Standard, Enterprise General availability
aws eu-west-1 (Ireland) Builder, Standard, Enterprise General availability
aws eu-central-1 (Frankfurt) Builder, Standard, Enterprise General availability
aws ap-southeast-1 (Singapore) Builder, Standard, Enterprise General availability
gcp us-central1 (Iowa) Builder, Standard, Enterprise General availability
gcp europe-west4 (Netherlands) Builder, Standard, Enterprise General availability
azure eastus2 (Virginia) Builder, Standard, Enterprise General availability

The cloud and region can't be changed after a serverless index is created.

When creating an index that stores dense vectors, you can choose from the following similarity metrics. For the most accurate results, choose the similarity metric used to train the embedding model for your vectors. For more information, see Vector Similarity Explained.

Euclidean

Querying indexes with this metric returns a similarity score equal to the squared Euclidean distance between the result and query vectors.

This metric calculates the square of the distance between two data points in a plane. It's one of the most commonly used distance metrics. For an example, see our IT threat detection example.

When you use metric='euclidean', the most similar results are those with the lowest similarity score.

Cosine

This is often used to find similarities between different documents. The advantage is that the scores are normalized to [-1,1] range. For an example, see our generative question answering example.

Dotproduct

This is used to multiply two vectors. You can use it to tell us how similar the two vectors are. The more positive the answer is, the closer the two vectors are in terms of their directions. For an example, see our semantic search example.

Dense vectors and sparse vectors are the basic units of data in Pinecone and what Pinecone was specially designed to store and work with. Dense vectors represents the semantics of data such as text, images, and audio recordings, while sparse vectors represent documents or queries in a way that captures keyword information.

To transform data into vector format, you use an embedding model. Pinecone hosts several embedding models so it's easy to manage your vector storage and search process on a single platform. You can use a hosted model to embed your data as an integrated part of upserting and querying, or you can use a hosted model to embed your data as a standalone operation.

The following embedding models are hosted by Pinecone.

multilingual-e5-large is an efficient dense embedding model trained on a mixture of multilingual datasets. It works well on messy data and short queries expected to return medium-length passages of text (1-2 paragraphs).

Details

  • Vector type: Dense
  • Modality: Text
  • Dimension: 1024
  • Recommended similarity metric: Cosine
  • Max sequence length: 507 tokens
  • Max batch size: 96 sequences

For rate limits, see Embedding tokens per minute and Embedding tokens per month.

Parameters

The multilingual-e5-large model supports the following parameters:

Parameter Type Required/Optional Description Default
input_type string Required The type of input data. Accepted values: query or passage.
truncate string Optional How to handle inputs longer than those supported by the model. Accepted values: END or NONE.

END truncates the input sequence at the input token limit. NONE returns an error when the input exceeds the input token limit.
END

llama-text-embed-v2 is a high-performance dense embedding model optimized for text retrieval and ranking tasks. It's trained on a diverse range of text corpora and provides strong performance on longer passages and structured documents.

Details

  • Vector type: Dense
  • Modality: Text
  • Dimension: 1024 (default), 2048, 768, 512, 384
  • Recommended similarity metric: Cosine
  • Max sequence length: 2048 tokens
  • Max batch size: 96 sequences

For rate limits, see Embedding tokens per minute and Embedding tokens per month.

Parameters

The llama-text-embed-v2 model supports the following parameters:

Parameter Type Required/Optional Description Default
input_type string Required The type of input data. Accepted values: query or passage.
truncate string Optional How to handle inputs longer than those supported by the model. Accepted values: END or NONE.

END truncates the input sequence at the input token limit. NONE returns an error when the input exceeds the input token limit.
END
dimension integer Optional Dimension of the vector to return. 1024

pinecone-sparse-english-v0 is a sparse embedding model for converting text to sparse vectors for sparse-vector search or hybrid search. Built on the innovations of the DeepImpact architecture, the model directly estimates the lexical importance of tokens by using their context, unlike traditional retrieval models like BM25, which rely solely on term frequency.

Details

  • Vector type: Sparse
  • Modality: Text
  • Recommended similarity metric: Dotproduct
  • Max sequence length: 512 or 2048
  • Max batch size: 96 sequences

For rate limits, see Embedding tokens per minute and Embedding tokens per month.

Parameters

The pinecone-sparse-english-v0 model supports the following parameters:

Parameter Type Required/Optional Description Default
input_type string Required The type of input data. Accepted values: query or passage.
max_tokens_per_sequence integer Optional Maximum number of tokens to embed. Accepted values: 512 or 2048. 512
truncate string Optional How to handle inputs longer than those supported by the model. Accepted values: END or NONE.

END truncates the input sequence at the the max_tokens_per_sequence limit. NONE returns an error when the input exceeds the max_tokens_per_sequence limit.
END
return_tokens boolean Optional Whether to return the string tokens. false
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu