Skip to main content
Pinecone Docs
current

Search documentation

Type to search this documentation.

Chat with an assistant

This is the recommended way to chat with an assistant, as it offers more functionality and control over the assistant's responses and references than the OpenAI-compatible chat interface.

For guidance and examples, see Chat with an assistant.

curl
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"

curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME" \
  -H "Api-Key: $PINECONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "What is the inciting incident of Pride and Prejudice?"
    }
  ],
  "stream": false,
  "model": "gpt-4o"
}'
curl
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"

curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME" \
  -H "Api-Key: $PINECONE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Pinecone-Api-Version: 2025-10" \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "What is the inciting incident of Pride and Prejudice?"
    }
  ],
  "stream": true,
  "model": "gpt-4o"
}'
Default
{
  "finish_reason": "stop",
  "message": {
    "role": "assistant",
    "content": "The inciting incident of \"Pride and Prejudice\" occurs when Mrs. Bennet informs Mr. Bennet that Netherfield Park has been let at last, and she is eager to share the news about the new tenant, Mr. Bingley, who is wealthy and single. This sets the stage for the subsequent events of the story, including the introduction of Mr. Bingley and Mr. Darcy to the Bennet family and the ensuing romantic entanglements."
  },
  "id": "00000000000000004ac3add5961aa757",
  "model": "gpt-4o-2024-05-13",
  "usage": {
    "prompt_tokens": 9736,
    "completion_tokens": 105,
    "total_tokens": 9841
  },
  "citations": [
    {
      "position": 406,
      "references": [
        {
          "file": {
            "status": "Available",
            "id": "ae79e447-b89e-4994-994b-3232ca52a654",
            "name": "Pride-and-Prejudice.pdf",
            "size": 2973077,
            "metadata": null,
            "updated_on": "2024-06-14T15:01:57.385425746Z",
            "created_on": "2024-06-14T15:01:02.910452398Z",
            "percent_done": 0,
            "signed_url": "https://storage.googleapis.com/...",
            "error_message": null
          },
          "pages": [
            1
          ]
        }
      ]
    }
  ]
}
Streaming
data:{
  "type":"message_start",
  "id":"0000000000000000111b35de85e8a8f9",
  "model":"gpt-4o-2024-05-13",
  "role":"assistant"
}

data:
{
  "type":"content_chunk",
  "id":"0000000000000000111b35de85e8a8f9",
  "model":"gpt-4o-2024-05-13",
  "delta":
  {
    "content":"The"
    }
}

...

data:
{
  "type":"citation",
  "id":"0000000000000000111b35de85e8a8f9",
  "model":"gpt-4o-2024-05-13",
  "citation":
  {
    "position":406,
    "references":
    [
      {
        "file":{
          "status":"Available",
          "id":"ae79e447-b89e-4994-994b-3232ca52a654",
          "name":"Pride-and-Prejudice.pdf",
          "size":2973077,
          "metadata":null,
          "updated_on":"2024-06-14T15:01:57.385425746Z", 
          "created_on":"2024-06-14T15:01:02.910452398Z",
          "percent_done":0.0,
          "signed_url":"https://storage.googleapis.com/...",
          "error_message":null
          }, 
      "pages":[1]
      }
    ]
  }
}

data:
{
  "type":"message_end",
  "id":"0000000000000000111b35de85e8a8f9",
  "model":"gpt-4o-2024-05-13",
  "finish_reason":"stop",
  "usage":
  {
    "prompt_tokens":9736,
    "completion_tokens":102,
    "total_tokens":9838
    }
}

POST /chat/{assistant_name}

Api-Keystringrequired

Pinecone API Key

Typestring
X-Pinecone-Api-Versionstringrequired

Required date-based version header

Typestring
Default2025-10
assistant_namestringrequired

The name of the assistant to be described.

Typestring

The desired configuration to chat an assistant.

messagesobject[]required
Show child attributes
role?string

Role of the message such as 'user' or 'assistant'

Typestring
content?string

Content of the message

Typestring
stream?boolean

If false, the assistant will return a single JSON response. If true, the assistant will return a stream of responses.

Typeboolean
Defaultfalse
model?string

The large language model to use for answer generation

Typestring
Defaultgpt-4o
temperature?number

Controls the randomness of the model's output: lower values make responses more deterministic, while higher values increase creativity and variability. If the model does not support a temperature parameter, the parameter will be ignored.

Typenumber
Default0
filter?object

Optionally filter which documents can be retrieved using the following metadata fields.

Typeobject
json_response?boolean

If true, the assistant will be instructed to return a JSON response. Cannot be used with streaming.

Typeboolean
Defaultfalse
include_highlights?boolean

If true, the assistant will be instructed to return highlights from the referenced documents that support its response.

Typeboolean
Defaultfalse
context_options?object

Controls the context snippets sent to the LLM.

Typeobject
Show child attributes
top_k?integer

The maximum number of context snippets to use. Default is 16. Maximum is 64.

Example: 20

Typeinteger
snippet_size?integer

The maximum context snippet size. Default is 2048 tokens. Minimum is 512 tokens. Maximum is 8192 tokens.

Example: 4096

Typeinteger
multimodal?boolean

Whether or not to send image-related context snippets to the LLM. If false, only text context snippets are sent.

Typeboolean
Defaulttrue
include_binary_content?boolean

If image-related context snippets are sent to the LLM, this field determines whether or not they should include base64 image data. If false, only the image caption is sent. Only available when multimodal=true.

Typeboolean
Defaulttrue

200 — Search request successful.

Describes the response format of a chat request from the citation API.

id?string
finish_reason?string
message?object

Describes the format of a message in a chat.

Typeobject
Show child attributes
role?string

Role of the message such as 'user' or 'assistant'

Typestring
content?string

Content of the message

Typestring
model?string
citations?object[]
Show child attributes
position?integer

The index position of the citation in the complete text response.

Typeinteger
references?object[]
Show child attributes
file?object

The response format for a successful file upload request.

Typeobject
Show child attributes
namestringrequired
idstringrequired
metadata?object | null
created_on?string
updated_on?string
status?string

The current state of the uploaded file. Possible values: - Processing: File is being processed (parsed, chunked, embedded) - Available: Processing completed successfully; file is ready for use - Deleting: Deletion has been initiated but not yet completed - ProcessingFailed: Processing failed with an error Note: Once a file is deleted, the API returns 404 Not Found instead of a file object.

Typestring
percent_done?number | null

The percentage of the file that has been processed

Typenumber | null
signed_url?string | null

Example: https://storage.googleapis.com/bucket/file.pdf?...

Typestring | null

A signed URL that provides temporary, read-only access to the underlying file. Anyone with the link can access the file, so treat it as sensitive data. Expires after a short time.

error_message?string | null

A message describing any error during file processing. Provided only if an error occurs.

Typestring | null
multimodal?boolean

Indicates whether the file was processed as multimodal.

Typeboolean
pages?integer[]
highlight?object | null

Represents a portion of a referenced document that directly supports or is relevant to the response.

Typeobject | null
Show child attributes
typestringrequired

The type of the highlight. Currently it is always text.

Typestring
contentstringrequired
usage?object

Describes the usage of a chat completion.

Typeobject
Show child attributes
prompt_tokens?integer
completion_tokens?integer
total_tokens?integer
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu