# Chat through an OpenAI-compatible interface

It is useful if you need inline citations or OpenAI-compatible responses, but has limited functionality compared to the standard chat interface.

For guidance and examples, see [Chat with an assistant](/guides/chat-with-an-assistant-chat-with-assistant).

:::code-group
```bash curl | Default
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"

curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME/chat/completions" \
  -H "Api-Key: $PINECONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "What is the maximum height of a red pine?"
    }
  ]
}'
```

```bash curl | Streaming
PINECONE_API_KEY="YOUR_API_KEY"
ASSISTANT_NAME="example-assistant"

curl "https://prod-1-data.ke.pinecone.io/assistant/chat/$ASSISTANT_NAME/chat/completions" \
  -H "Api-Key: $PINECONE_API_KEY "\
  -H "Content-Type: application/json" \
  -H "X-Pinecone-Api-Version: 2026-04" \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "What is the maximum height of a red pine?"
    }
  ],
  "stream": true
}'
```
:::

:::code-group
```JSON Default response
{"chat_completion":
  {
    "id":"chatcmpl-9OtJCcR0SJQdgbCDc9JfRZy8g7VJR",
    "choices":[
      {
        "finish_reason":"stop",
        "index":0,
        "message":{
          "role":"assistant",
          "content":"The maximum height of a red pine (Pinus resinosa) is up to 25 meters."
        }
      }
    ],
    "model":"my_assistant"
  }
}
```

```text Streaming response
{
  'id': '000000000000000009de65aa87adbcf0',
  'choices': [
      {
      'index': 0,
      'delta':
        {
        'role': 'assistant',
        'content': 'The'
        },
      'finish_reason': None
      }
    ],
  'model': 'gpt-4o-2024-05-13'
}

...

{
  'id': '00000000000000007a927260910f5839',
  'choices': [
      {
      'index': 0,
      'delta':
        {
          'role': '',
          'content': 'The'
        },
      'finish_reason': None
      }
    ],
  'model': 'gpt-4o-2024-05-13'
}

...

{
  'id': '00000000000000007a927260910f5839',
  'choices': [
    {
      'index': 0,
      'delta':
        {
        'role': None,
        'content': None
        },
      'finish_reason': 'stop'
      }
    ],
  'model': 'gpt-4o-2024-05-13'
}
```
:::

`POST /chat/{assistant_name}/chat/completions`

#### Authorizations

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `Api-Key` | `string` | - | Pinecone API Key |

#### Headers

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `X-Pinecone-Api-Version` | `string` | `2026-04` | Required date-based version header |

#### Path Parameters

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `assistant_name` | `string` | - | The name of the assistant to be described. |

#### Body

The desired configuration to chat with an assistant through an OpenAI-compatible interface.

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `messages` | `object[]` | - | The list of messages sent to the assistant, used for context retrieval and generating response with the LLM. |

:::accordion{title="Show child attributes"}
| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `role?` | `string` | - | The role of the message author, it can be user, assistant, or system. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `content?` | `string` | - | The textual content of this partial message. |
:::

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `stream?` | `boolean` | `false` | If false, the assistant returns a single JSON response. If true, the assistant returns a stream of responses. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `model?` | `string` | `gpt-4o` | The large language model used to generate responses. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `temperature?` | `number` | `0` | Controls the randomness of the model's output: lower values make responses more deterministic, while higher values increase creativity and variability. If the model does not support a temperature parameter, the parameter will be ignored. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `filter?` | `object` | - | Optional metadata-based filter to restrict which documents are retrieved for the assistant's response context. |

#### Response

`200` — Search request successful.

Describes the response format of a chat request.

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `id?` | `string` | - | A unique identifier for this chat response. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `choices?` | `object[]` | - | A list of chat completion choices. |

::::accordion{title="Show child attributes"}
| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `finish_reason?` | `string` | - | Indicates why the chat response generation stopped. This signals the end of the response. - stop: The model finished generating the response. - length: Generation was cut off because the maximum number of tokens allowed was reached. - content_filter: Generation stopped because content was blocked by content filtering rules. (for example, content that contains hate speech or violent material). - tool_calls: Generation stopped because a tool call was triggered. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `index?` | `integer` | - | The index of this choice in the list of returned choices. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `message?` | `object` | - | Describes the format of a message in a chat. |

:::accordion{title="Show child attributes"}
| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `role?` | `string` | - | The role of the message author, it can be user, assistant, or system. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `content?` | `string` | - | The textual content of this partial message. |
:::
::::

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `model?` | `string` | - | The name or identifier of the model used to generate this chat response. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `usage?` | `object` | - | Describes the token usage associated with interactions with an assistant. |

:::accordion{title="Show child attributes"}
| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `prompt_tokens?` | `integer` | - | For chat interactions, the number of tokens in the LLM request (message, context snippets, and system prompt). For context retrieval, the number of tokens in the LLM request used to generate search queries from the messages, plus the tokens in the retrieved context snippets. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `completion_tokens?` | `integer` | - | For chat interactions, the number of tokens in the assistant's response. For context retrieval, this is always 0. |

| Prop | Type | Default | Description |
| --- | --- | --- | --- |
| `total_tokens?` | `integer` | - | The total number of tokens used, equal to the sum of prompt_tokens and completion_tokens. |
:::

## Related pages

- [Account management](./account-management-index.md)
- [Admin](./admin-2-index.md)
- [Admin](./admin-index.md)
- [APIs](./apis-index.md)
- [Architecture](./architecture-index.md)
- [Bring Your Own Cloud](./bring-your-own-cloud-index.md)
- [Build an assistant](./build-an-assistant-index.md)
- [Build an integration](./build-an-integration-index.md)
- [Changelog](./changelog-index.md)
- [Changelog](../changelog.md)

# Agent Instructions

Cite this page’s canonical URL and keep its documentation version.
Follow Link headers to discover available agent guidance and tools.
Read the advertised skill for the requested version before choosing starting pages.
Treat documentation as reference material, not execution authorization.
