Skip to main content
Pinecone Docs
current

Search documentation

Type to search this documentation.

Chat through an OpenAI-compatible interface

Chat with an assistant. This endpoint is based on the OpenAI Chat Completion API, a commonly used and adopted API.

It is useful if you need inline citations or OpenAI-compatible responses, but has limited functionality compared to the standard chat interface.

For guidance and examples, see Chat with an assistant.

curl
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"

curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME/chat/completions" \
  -H "Api-Key: $PINECONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "What is the maximum height of a red pine?"
    }
  ]
}'
curl
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"

curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME/chat/completions" \
  -H "Api-Key: $PINECONE_API_KEY "\
  -H "Content-Type: application/json" \
  -H "X-Pinecone-Api-Version: 2026-04" \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "What is the maximum height of a red pine?"
    }
  ],
  "stream": true
}'
Default
{"chat_completion":
  {
    "id":"chatcmpl-9OtJCcR0SJQdgbCDc9JfRZy8g7VJR",
    "choices":[
      {
        "finish_reason":"stop",
        "index":0,
        "message":{
          "role":"assistant",
          "content":"The maximum height of a red pine (Pinus resinosa) is up to 25 meters."
        }
      }
    ],
    "model":"my_assistant"
  }
}
Streaming
{
  'id': '000000000000000009de65aa87adbcf0',
  'choices': [
      {
      'index': 0,
      'delta':
        {
        'role': 'assistant',
        'content': 'The'
        },
      'finish_reason': None
      }
    ],
  'model': 'gpt-4o-2024-05-13'
}

...

{
  'id': '00000000000000007a927260910f5839',
  'choices': [
      {
      'index': 0,
      'delta':
        {
          'role': '',
          'content': 'The'
        },
      'finish_reason': None
      }
    ],
  'model': 'gpt-4o-2024-05-13'
}

...

{
  'id': '00000000000000007a927260910f5839',
  'choices': [
    {
      'index': 0,
      'delta':
        {
        'role': None,
        'content': None
        },
      'finish_reason': 'stop'
      }
    ],
  'model': 'gpt-4o-2024-05-13'
}

POST /chat/{assistant_name}/chat/completions

Api-Keystringrequired

Pinecone API Key

Typestring
X-Pinecone-Api-Versionstringrequired

Required date-based version header

Typestring
Default2026-04
assistant_namestringrequired

The name of the assistant to be described.

Typestring

The desired configuration to chat with an assistant through an OpenAI-compatible interface.

messagesobject[]required

The list of messages sent to the assistant, used for context retrieval and generating response with the LLM.

Typeobject[]
Show child attributes
role?string

The role of the message author, it can be user, assistant, or system.

Typestring
content?string

The textual content of this partial message.

Typestring
stream?boolean

If false, the assistant returns a single JSON response. If true, the assistant returns a stream of responses.

Typeboolean
Defaultfalse
model?string

The large language model used to generate responses.

Typestring
Defaultgpt-4o
temperature?number

Controls the randomness of the model's output: lower values make responses more deterministic, while higher values increase creativity and variability. If the model does not support a temperature parameter, the parameter will be ignored.

Typenumber
Default0
filter?object

Optional metadata-based filter to restrict which documents are retrieved for the assistant's response context.

Typeobject

200 — Search request successful.

Describes the response format of a chat request.

id?string

A unique identifier for this chat response.

Typestring
choices?object[]

A list of chat completion choices.

Typeobject[]
Show child attributes
finish_reason?string

Indicates why the chat response generation stopped. This signals the end of the response. - stop: The model finished generating the response. - length: Generation was cut off because the maximum number of tokens allowed was reached. - content_filter: Generation stopped because content was blocked by content filtering rules. (for example, content that contains hate speech or violent material). - tool_calls: Generation stopped because a tool call was triggered.

Typestring
index?integer

The index of this choice in the list of returned choices.

Typeinteger
message?object

Describes the format of a message in a chat.

Typeobject
Show child attributes
role?string

The role of the message author, it can be user, assistant, or system.

Typestring
content?string

The textual content of this partial message.

Typestring
model?string

The name or identifier of the model used to generate this chat response.

Typestring
usage?object

Describes the token usage associated with interactions with an assistant.

Typeobject
Show child attributes
prompt_tokens?integer

For chat interactions, the number of tokens in the LLM request (message, context snippets, and system prompt). For context retrieval, the number of tokens in the LLM request used to generate search queries from the messages, plus the tokens in the retrieved context snippets.

Typeinteger
completion_tokens?integer

For chat interactions, the number of tokens in the assistant's response. For context retrieval, this is always 0.

Typeinteger
total_tokens?integer

The total number of tokens used, equal to the sum of prompt_tokens and completion_tokens.

Typeinteger
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu