Skip to main content
Pinecone Docs

Search documentation

Type to search this documentation.

On this pageOverview

Chat with an assistant

This is the recommended way to chat with an assistant, as it offers more functionality and control over the assistant's responses and references than the OpenAI-compatible chat interface.

For guidance and examples, see Chat with an assistant.

curl
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"

curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME" \
  -H "Api-Key: $PINECONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "What is the inciting incident of Pride and Prejudice?"
    }
  ],
  "stream": false,
  "model": "gpt-4o"
}'
curl
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"

curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME" \
  -H "Api-Key: $PINECONE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Pinecone-Api-Version: 2026-04" \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "What is the inciting incident of Pride and Prejudice?"
    }
  ],
  "stream": true,
  "model": "gpt-4o"
}'
Default
{
  "finish_reason": "stop",
  "message": {
    "role": "assistant",
    "content": "The inciting incident of \"Pride and Prejudice\" occurs when Mrs. Bennet informs Mr. Bennet that Netherfield Park has been let at last, and she is eager to share the news about the new tenant, Mr. Bingley, who is wealthy and single. This sets the stage for the subsequent events of the story, including the introduction of Mr. Bingley and Mr. Darcy to the Bennet family and the ensuing romantic entanglements."
  },
  "id": "00000000000000004ac3add5961aa757",
  "model": "gpt-4o-2024-05-13",
  "usage": {
    "prompt_tokens": 9736,
    "completion_tokens": 105,
    "total_tokens": 9841
  },
  "citations": [
    {
      "position": 406,
      "references": [
        {
          "file": {
            "status": "Available",
            "id": "ae79e447-b89e-4994-994b-3232ca52a654",
            "name": "Pride-and-Prejudice.pdf",
            "size": 2973077,
            "metadata": null,
            "updated_on": "2024-06-14T15:01:57.385425746Z",
            "created_on": "2024-06-14T15:01:02.910452398Z",
            "signed_url": "https://storage.googleapis.com/..."
          },
          "pages": [
            1
          ]
        }
      ]
    }
  ]
}
Streaming
data:{
  "type":"message_start",
  "id":"0000000000000000111b35de85e8a8f9",
  "model":"gpt-4o-2024-05-13",
  "role":"assistant"
}

data:
{
  "type":"content_chunk",
  "id":"0000000000000000111b35de85e8a8f9",
  "model":"gpt-4o-2024-05-13",
  "delta":
  {
    "content":"The"
    }
}

...

data:
{
  "type":"citation",
  "id":"0000000000000000111b35de85e8a8f9",
  "model":"gpt-4o-2024-05-13",
  "citation":
  {
    "position":406,
    "references":
    [
      {
        "file":{
          "status":"Available",
          "id":"ae79e447-b89e-4994-994b-3232ca52a654",
          "name":"Pride-and-Prejudice.pdf",
          "size":2973077,
          "metadata":null,
          "updated_on":"2024-06-14T15:01:57.385425746Z",
          "created_on":"2024-06-14T15:01:02.910452398Z",
          "signed_url":"https://storage.googleapis.com/..."
          },
      "pages":[1]
      }
    ]
  }
}

data:
{
  "type":"message_end",
  "id":"0000000000000000111b35de85e8a8f9",
  "model":"gpt-4o-2024-05-13",
  "finish_reason":"stop",
  "usage":
  {
    "prompt_tokens":9736,
    "completion_tokens":102,
    "total_tokens":9838
    }
}

POST /chat/{assistant_name}

  • X-Pinecone-Api-Version (header, string, required) — Required date-based version header
  • assistant_name (path, string, required) — The name of the assistant to be described.
  • messages (body, object[], required) — The list of messages sent to the assistant, used for context retrieval and generating response with the LLM.
  • stream (body, boolean) — If false, the assistant returns a single JSON response. If true, the assistant returns a stream of responses.
  • model (body, string) — The large language model used to generate responses.
  • temperature (body, number) — Controls the randomness of the model's output: lower values make responses more deterministic, while higher values increase creativity and variability. If the model does not support a temperature parameter, the parameter will be ignored.
  • filter (body, object) — Optional metadata-based filter to restrict which documents are retrieved for the assistant's response context.
  • json_response (body, boolean) — If true, instructs the assistant to return a JSON-formatted response. Cannot be used together with streaming mode.
  • include_highlights (body, boolean) — If true, instructs the assistant to include highlights from the referenced documents that support its response.
  • context_options (body, object) — Controls the context snippets sent to the LLM.
  • 200 — Search request successful.
  • 400 — Bad request. The request body included invalid request parameters.
  • 401 — Unauthorized. Possible causes: Invalid API key.
  • 404 — Assistant not found.
  • 500 — Internal server error.
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu