Chat through an OpenAI-compatible interface
Chat with an assistant. This endpoint is based on the OpenAI Chat Completion API, a commonly used and adopted API.
It is useful if you need inline citations or OpenAI-compatible responses, but has limited functionality compared to the standard chat interface.
For guidance and examples, see Chat with an assistant.
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"
curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME/chat/completions" \
-H "Api-Key: $PINECONE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{
"role": "user",
"content": "What is the maximum height of a red pine?"
}
]
}'PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"
curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME/chat/completions" \
-H "Api-Key: $PINECONE_API_KEY "\
-H "Content-Type: application/json" \
-H "X-Pinecone-Api-Version: 2026-04" \
-d '{
"messages": [
{
"role": "user",
"content": "What is the maximum height of a red pine?"
}
],
"stream": true
}'{"chat_completion":
{
"id":"chatcmpl-9OtJCcR0SJQdgbCDc9JfRZy8g7VJR",
"choices":[
{
"finish_reason":"stop",
"index":0,
"message":{
"role":"assistant",
"content":"The maximum height of a red pine (Pinus resinosa) is up to 25 meters."
}
}
],
"model":"my_assistant"
}
}{
'id': '000000000000000009de65aa87adbcf0',
'choices': [
{
'index': 0,
'delta':
{
'role': 'assistant',
'content': 'The'
},
'finish_reason': None
}
],
'model': 'gpt-4o-2024-05-13'
}
...
{
'id': '00000000000000007a927260910f5839',
'choices': [
{
'index': 0,
'delta':
{
'role': '',
'content': 'The'
},
'finish_reason': None
}
],
'model': 'gpt-4o-2024-05-13'
}
...
{
'id': '00000000000000007a927260910f5839',
'choices': [
{
'index': 0,
'delta':
{
'role': None,
'content': None
},
'finish_reason': 'stop'
}
],
'model': 'gpt-4o-2024-05-13'
}POST /chat/{assistant_name}/chat/completions
Authorizations
Section titled “Authorizations”Api-KeystringrequiredPinecone API Key
Headers
Section titled “Headers”X-Pinecone-Api-VersionstringrequiredRequired date-based version header
Path Parameters
Section titled “Path Parameters”assistant_namestringrequiredThe name of the assistant to be described.
The desired configuration to chat with an assistant through an OpenAI-compatible interface.
messagesobject[]requiredThe list of messages sent to the assistant, used for context retrieval and generating response with the LLM.
Show child attributes
role?stringThe role of the message author, it can be user, assistant, or system.
content?stringThe textual content of this partial message.
stream?booleanIf false, the assistant returns a single JSON response. If true, the assistant returns a stream of responses.
model?stringThe large language model used to generate responses.
temperature?numberControls the randomness of the model's output: lower values make responses more deterministic, while higher values increase creativity and variability. If the model does not support a temperature parameter, the parameter will be ignored.
filter?objectOptional metadata-based filter to restrict which documents are retrieved for the assistant's response context.
Response
Section titled “Response”200 — Search request successful.
Describes the response format of a chat request.
id?stringA unique identifier for this chat response.
choices?object[]A list of chat completion choices.
Show child attributes
finish_reason?stringIndicates why the chat response generation stopped. This signals the end of the response. - stop: The model finished generating the response. - length: Generation was cut off because the maximum number of tokens allowed was reached. - content_filter: Generation stopped because content was blocked by content filtering rules. (for example, content that contains hate speech or violent material). - tool_calls: Generation stopped because a tool call was triggered.
index?integerThe index of this choice in the list of returned choices.
message?objectDescribes the format of a message in a chat.
Show child attributes
role?stringThe role of the message author, it can be user, assistant, or system.
content?stringThe textual content of this partial message.
model?stringThe name or identifier of the model used to generate this chat response.
usage?objectDescribes the token usage associated with interactions with an assistant.
Show child attributes
prompt_tokens?integerFor chat interactions, the number of tokens in the LLM request (message, context snippets, and system prompt). For context retrieval, the number of tokens in the LLM request used to generate search queries from the messages, plus the tokens in the retrieved context snippets.
completion_tokens?integerFor chat interactions, the number of tokens in the assistant's response. For context retrieval, this is always 0.
total_tokens?integerThe total number of tokens used, equal to the sum of prompt_tokens and completion_tokens.