Chat with an assistant
This is the recommended way to chat with an assistant, as it offers more functionality and control over the assistant's responses and references than the OpenAI-compatible chat interface.
For guidance and examples, see Chat with an assistant.
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"
curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME" \
-H "Api-Key: $PINECONE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{
"role": "user",
"content": "What is the inciting incident of Pride and Prejudice?"
}
],
"stream": false,
"model": "gpt-4o"
}'PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"
curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME" \
-H "Api-Key: $PINECONE_API_KEY" \
-H "Content-Type: application/json" \
-H "X-Pinecone-Api-Version: 2026-04" \
-d '{
"messages": [
{
"role": "user",
"content": "What is the inciting incident of Pride and Prejudice?"
}
],
"stream": true,
"model": "gpt-4o"
}'{
"finish_reason": "stop",
"message": {
"role": "assistant",
"content": "The inciting incident of \"Pride and Prejudice\" occurs when Mrs. Bennet informs Mr. Bennet that Netherfield Park has been let at last, and she is eager to share the news about the new tenant, Mr. Bingley, who is wealthy and single. This sets the stage for the subsequent events of the story, including the introduction of Mr. Bingley and Mr. Darcy to the Bennet family and the ensuing romantic entanglements."
},
"id": "00000000000000004ac3add5961aa757",
"model": "gpt-4o-2024-05-13",
"usage": {
"prompt_tokens": 9736,
"completion_tokens": 105,
"total_tokens": 9841
},
"citations": [
{
"position": 406,
"references": [
{
"file": {
"status": "Available",
"id": "ae79e447-b89e-4994-994b-3232ca52a654",
"name": "Pride-and-Prejudice.pdf",
"size": 2973077,
"metadata": null,
"updated_on": "2024-06-14T15:01:57.385425746Z",
"created_on": "2024-06-14T15:01:02.910452398Z",
"signed_url": "https://storage.googleapis.com/..."
},
"pages": [
1
]
}
]
}
]
}
data:{
"type":"message_start",
"id":"0000000000000000111b35de85e8a8f9",
"model":"gpt-4o-2024-05-13",
"role":"assistant"
}
data:
{
"type":"content_chunk",
"id":"0000000000000000111b35de85e8a8f9",
"model":"gpt-4o-2024-05-13",
"delta":
{
"content":"The"
}
}
...
data:
{
"type":"citation",
"id":"0000000000000000111b35de85e8a8f9",
"model":"gpt-4o-2024-05-13",
"citation":
{
"position":406,
"references":
[
{
"file":{
"status":"Available",
"id":"ae79e447-b89e-4994-994b-3232ca52a654",
"name":"Pride-and-Prejudice.pdf",
"size":2973077,
"metadata":null,
"updated_on":"2024-06-14T15:01:57.385425746Z",
"created_on":"2024-06-14T15:01:02.910452398Z",
"signed_url":"https://storage.googleapis.com/..."
},
"pages":[1]
}
]
}
}
data:
{
"type":"message_end",
"id":"0000000000000000111b35de85e8a8f9",
"model":"gpt-4o-2024-05-13",
"finish_reason":"stop",
"usage":
{
"prompt_tokens":9736,
"completion_tokens":102,
"total_tokens":9838
}
}POST /chat/{assistant_name}
Authorizations
Section titled “Authorizations”Api-KeystringrequiredPinecone API Key
Headers
Section titled “Headers”X-Pinecone-Api-VersionstringrequiredRequired date-based version header
Path Parameters
Section titled “Path Parameters”assistant_namestringrequiredThe name of the assistant to be described.
The desired configuration to chat with an assistant.
messagesobject[]requiredThe list of messages sent to the assistant, used for context retrieval and generating response with the LLM.
Show child attributes
role?stringThe role of the message author, it can be user, assistant, or system.
content?stringThe textual content of this partial message.
stream?booleanIf false, the assistant returns a single JSON response. If true, the assistant returns a stream of responses.
model?stringThe large language model used to generate responses.
temperature?numberControls the randomness of the model's output: lower values make responses more deterministic, while higher values increase creativity and variability. If the model does not support a temperature parameter, the parameter will be ignored.
filter?objectOptional metadata-based filter to restrict which documents are retrieved for the assistant's response context.
json_response?booleanIf true, instructs the assistant to return a JSON-formatted response. Cannot be used together with streaming mode.
include_highlights?booleanIf true, instructs the assistant to include highlights from the referenced documents that support its response.
context_options?objectControls the context snippets sent to the LLM.
Show child attributes
top_k?integerThe maximum number of context snippets to use. Default is 16. Maximum is 64.
Example: 20
snippet_size?integerThe maximum context snippet size. Default is 2048 tokens. Minimum is 512 tokens. Maximum is 8192 tokens.
Example: 4096
multimodal?booleanWhether or not to send image-related context snippets to the LLM. If false, only text context snippets are sent.
include_binary_content?booleanIf image-related context snippets are sent to the LLM, this field determines whether or not they should include base64 image data. If false, only the image caption is sent. Only available when multimodal=true.
Response
Section titled “Response”200 — Search request successful.
Describes the response format of a chat request.
id?stringA unique identifier for this chat response.
finish_reason?stringIndicates why the chat response generation stopped. This signals the end of the response. - stop: The model finished generating the response. - length: Generation was cut off because the maximum number of tokens allowed was reached. - content_filter: Generation stopped because content was blocked by content filtering rules. (for example, content that contains hate speech or violent material). - tool_calls: Generation stopped because a tool call was triggered.
message?objectDescribes the format of a message in a chat.
Show child attributes
role?stringThe role of the message author, it can be user, assistant, or system.
content?stringThe textual content of this partial message.
model?stringThe name or identifier of the model used to generate this chat response.
citations?object[]Citations supporting the information in the response.
Show child attributes
position?integerThe index position of the citation in the complete text response.
references?object[]A list of file references that this citation points to.
Show child attributes
file?objectThe response format for a successful file upload request.
Show child attributes
namestringrequiredThe name of the uploaded file.
idstringrequiredThe unique identifier for the uploaded file. This may be a user-provided identifier or a system-generated ID.
size?integerThe size of the uploaded file, in bytes.
Example: 1048576
metadata?object | nullOptional metadata associated with the file. This metadata can be used to filter files when listing them or to restrict search results when querying the assistant.
created_on?stringThe timestamp when the file was uploaded, in ISO 8601 format (YYYY-MM-DDTHH:MM:SSZ).
Example: 2025-10-01T12:30:00.000Z
updated_on?stringThe timestamp of the most recent update to the file, in ISO 8601 format (YYYY-MM-DDTHH:MM:SSZ).
Example: 2025-10-01T12:45:00.000Z
status?stringThe current state of the uploaded file. Possible values: - Processing: File is being processed (parsed, chunked, embedded) - Available: Processing completed successfully; file is ready for use - Deleting: Deletion has been initiated but not yet completed - ProcessingFailed: Processing failed with an error Note: Once a file is deleted, the API returns 404 Not Found instead of a file object.
signed_url?string | nullExample: https://storage.googleapis.com/bucket/file.pdf?...
A signed URL that provides temporary, read-only access to the file. Anyone with the link can access the file, so treat it as sensitive data. Expires after a short time.
multimodal?booleanIndicates whether the file was processed as multimodal.
pages?integer[]A list of page numbers in the referenced document that contain the relevant content.
highlight?object | nullRepresents a portion of a referenced document that directly supports or is relevant to the response.
Show child attributes
typestringrequiredThe type of the highlight. Only text is supported.
contentstringrequiredThe text content of the highlighted portion from the referenced document.
usage?objectDescribes the token usage associated with interactions with an assistant.
Show child attributes
prompt_tokens?integerFor chat interactions, the number of tokens in the LLM request (message, context snippets, and system prompt). For context retrieval, the number of tokens in the LLM request used to generate search queries from the messages, plus the tokens in the retrieved context snippets.
completion_tokens?integerFor chat interactions, the number of tokens in the assistant's response. For context retrieval, this is always 0.
total_tokens?integerThe total number of tokens used, equal to the sum of prompt_tokens and completion_tokens.
context_snippet_count?integerThe number of context snippets provided to the model to generate the response. This indicates how much retrieved information was available for the generation, allowing for logic to be applied if no context was found (count is 0).
content_filter_results?objectContent filter results provided by the LLM, describing safety-related classifications applied to the content. The structure may vary depending on the model and the content being filtered. The spec field identifies the provider, and determines the structure of results.
Show child attributes
spec?stringIdentifier of the model provider.
results?anyContent filter results returned by the provider. The structure depend on the spec value.