> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

# Chat

POST https://api.humanloop.com/v4/chat
Content-Type: application/json

Get a chat response by providing details of the model configuration in the request.

Reference: https://humanloop.com/docs/v4/api/chats/create

## Authentication

- `X-API-KEY` header (required) — API Key authentication via header

## Request

### Body (application/json)

This endpoint expects an object.

- `stream` (true, required) — If true, tokens will be sent as data-only server-sent events. If num_samples > 1, samples are streamed back independently.
- `messages` (list of ChatMessageWithToolCall, required) — The messages passed to the to provider chat endpoint.
- `model_config` (ModelConfigChatRequest, required) — The model configuration used to create a chat response.
- `project` (string, optional) — Unique project name. If no project exists with this name, a new project will be created.
- `project_id` (string, optional) — Unique ID of a project to associate to the log. Either this or `project` must be provided.
- `session_id` (string, optional) — ID of the session to associate the datapoint.
- `session_reference_id` (string, optional) — A unique string identifying the session to associate the datapoint to. Allows you to log multiple datapoints to a session (using an ID kept by your internal systems) by passing the same `session_reference_id` in subsequent log requests. Specify at most one of this or `session_id`.
- `parent_id` (string, optional) — ID associated to the parent datapoint in a session.
- `parent_reference_id` (string, optional) — A unique string identifying the previously-logged parent datapoint in a session. Allows you to log nested datapoints with your internal system IDs by passing the same reference ID as `parent_id` in a prior log request. Specify at most one of this or `parent_id`. Note that this cannot refer to a datapoint being logged in the same request.
- `inputs` (map from string to any, optional) — The inputs passed to the prompt template.
- `source` (string, optional) — Identifies where the model was called from.
- `metadata` (map from string to any, optional) — Any additional metadata to record.
- `save` (boolean, optional, default: true) — Whether the request/response payloads will be stored on Humanloop.
- `source_datapoint_id` (string, optional) — ID of the source datapoint if this is a log derived from a datapoint in a dataset.
- `provider_api_keys` (ProviderApiKeys, optional) — API keys required by each provider to make API calls. The API keys provided here are not stored by Humanloop. If not specified here, Humanloop will fall back to the key saved to your organization.
- `num_samples` (integer, optional, default: 1) — The number of generations.
- `template_language` (enum, optional) — The template language to use for rendering the template.
  - Allowed values: `default`, `jinja`
- `user` (string, optional) — End-user ID passed through to provider call.
- `return_inputs` (boolean, optional, default: true) — Whether to return the inputs in the response. If false, the response will contain an empty dictionary under inputs. This is useful for reducing the size of the response. Defaults to true.
- `tool_choice` (ChatsCreateStreamRequestToolChoice, optional) — Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'type': 'function', 'function': \{name': \<TOOL\_NAME>}} forces the model to use the named function.
- `response_format` (ResponseFormat, optional) — The format of the response. Only type json_object is currently supported for chat.
- `reasoning_effort` (ChatsCreateStreamRequestReasoningEffort, optional) — Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget.
- `seed` (integer, optional, deprecated) — Deprecated field: the seed is instead set as part of the request.config object.
- `tool_call` (ChatsCreateStreamRequestToolCall, optional, deprecated) — NB: Deprecated with new tool\_choice. Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'name': \<TOOL\_NAME>} forces the model to use the provided tool of the same name.

## Response

### 200

- Streaming response of `ChatResponse`.
- `data` (list of ChatDataResponse, required) — Array containing the chat responses.
- `provider_responses` (list of any, required) — The raw responses returned by the model provider.
- `project_id` (string, optional) — Unique identifier of the parent project. Will not be provided if the request was made without providing a project name or id
- `num_samples` (integer, optional, default: 1) — The number of chat responses.
- `logprobs` (integer, optional) — Include the log probabilities of the top n tokens in the provider_response
- `suffix` (string, optional) — The suffix that comes after a completion of inserted text. Useful for completions that act like inserts.
- `user` (string, optional) — End-user ID passed through to provider call.
- `usage` (Usage, optional) — Counts of the number of tokens used and related stats.
- `metadata` (map from string to any, optional) — Any additional metadata to record.
- `provider_request` (map from string to any, optional) — The raw request sent to the model provider.
- `session_id` (string, optional) — ID of the session if it belongs to one.
- `tool_choice` (ChatResponseToolChoice, optional) — Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'type': 'function', 'function': \{name': \<TOOL\_NAME>}} forces the model to use the named function.

## Errors

### 422 Chats Create Stream Request Unprocessable Entity Error

Validation Error

- `detail` (list of ValidationError, optional)

## Types

### ChatMessageWithToolCall

- `role` (enum, required) — Role of the message author.
  - Allowed values: `user`, `assistant`, `system`, `tool`, `developer`
- `content` (Content, optional) — The content of the message.
- `name` (string, optional) — Optional name of the message author.
- `tool_call_id` (string, optional) — Tool call that this message is responding to.
- `tool_calls` (list of ToolCall, optional) — A list of tool calls requested by the assistant.
- `thinking` (list of ChatMessageWithToolCallThinkingItem, optional) — Model's chain-of-thought for providing the response. Present on assistant messages if model supports it.
- `tool_call` (FunctionTool, optional, deprecated) — NB: Deprecated in favour of tool_calls. A tool call requested by the assistant.

### ModelConfigChatRequest

Chat model config request.

- `model` (string, required) — The model instance used. E.g. text-davinci-002.
- `name` (string, optional) — A friendly display name for the model config. If not provided, a name will be generated.
- `description` (string, optional) — A description of the model config.
- `provider` (enum, optional) — The company providing the underlying model service.
  - Allowed values: `anthropic`, `bedrock`, `cohere`, `deepseek`, `google`, `groq`, `mock`, `openai`, `openai_azure`, `replicate`
- `max_tokens` (integer, optional, default: -1) — The maximum number of tokens to generate. Provide max_tokens=-1 to dynamically calculate the maximum number of tokens to generate given the length of the prompt
- `temperature` (double, optional, default: 1) — What sampling temperature to use when making a generation. Higher values means the model will be more creative.
- `top_p` (double, optional, default: 1) — An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass.
- `stop` (ModelConfigChatRequestStop, optional) — The string (or list of strings) after which the model will stop generating. The returned text will not contain the stop sequence.
- `presence_penalty` (double, optional, default: 0) — Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the generation so far.
- `frequency_penalty` (double, optional, default: 0) — Number between -2.0 and 2.0. Positive values penalize new tokens based on how frequently they appear in the generation so far.
- `other` (map from string to any, optional) — Other parameter values to be passed to the provider call.
- `seed` (integer, optional) — If specified, model will make a best effort to sample deterministically, but it is not guaranteed.
- `response_format` (ResponseFormat, optional) — The format of the response. Only type json_object is currently supported for chat.
- `reasoning_effort` (ModelConfigChatRequestReasoningEffort, optional) — Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget.
- `template_language` (enum, optional) — The template language to use for rendering the template.
  - Allowed values: `default`, `jinja`
- `endpoint` (enum, optional) — The provider model endpoint used.
  - Allowed values: `complete`, `chat`, `edit`
- `chat_template` (list of ChatMessageWithToolCall, optional) — Messages prepended to the list of messages sent to the provider. These messages that will take your specified inputs to form your final request to the provider model. Input variables within the template should be specified with syntax: `{{input_name}}`.
- `tools` (list of ModelConfigChatRequestToolsItem, optional) — Make tools available to OpenAIs chat model as functions.

### ProviderApiKeys

- `openai` (string, optional)
- `mock` (string, optional)
- `anthropic` (string, optional)
- `deepseek` (string, optional)
- `bedrock` (string, optional)
- `cohere` (string, optional)
- `openai_azure` (string, optional)
- `openai_azure_endpoint` (string, optional)
- `google` (string, optional)

### ChatsCreateStreamRequestToolChoice

Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'type': 'function', 'function': \{name': \<TOOL\_NAME>}} forces the model to use the named function.

### ResponseFormat

Response format of the model.

- `type` (enum, required)
  - Allowed values: `json_object`, `json_schema`
- `json_schema` (map from string to any, optional) — The JSON schema of the response format if type is json_schema.

### ChatsCreateStreamRequestReasoningEffort

Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget.

### ChatsCreateStreamRequestToolCall

NB: Deprecated with new tool\_choice. Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'name': \<TOOL\_NAME>} forces the model to use the provided tool of the same name.

### ChatDataResponse

Overwrite DataResponse for chat.

- `id` (string, required) — Unique ID for the model inputs and output logged to Humanloop. Use this when recording feedback later.
- `index` (integer, required) — The index for the sampled generation for a given input. The num_samples request parameter controls how many samples are generated.
- `output` (string, required) — Output text returned from the provider model with leading and trailing whitespaces stripped.
- `raw_output` (string, required) — Raw output text returned from the provider model.
- `model_config_id` (string, required) — The model configuration used to create the generation.
- `output_message` (ChatMessageWithToolCall, required) — The message returned by the provider.
- `inputs` (map from string to any, optional) — The inputs passed to the chat template.
- `finish_reason` (string, optional) — Why the generation ended. One of 'stop' (indicating a stop token was encountered), or 'length' (indicating the max tokens limit has been reached), or 'tool_call' (indicating that the model has chosen to call a tool - in which case the tool_call parameter of the response will be populated). It will be set as null for the intermediary responses during a stream, and will only be set as non-null for the final streamed token.
- `tool_results` (list of ToolResultResponse, optional) — Results of any tools run during the generation.
- `messages` (list of ChatMessageWithToolCall, optional) — The messages passed to the to provider chat endpoint.
- `tool_calls` (list of ToolCall, optional) — Deprecated: Please use tool_calls field within the output_message.JSON definition of the tools to call and the corresponding argument values. Will be populated when finish_reason='tool_call'.
- `tool_call` (FunctionTool, optional, deprecated) — Deprecated: Please use tool_calls field within the output_message.JSON definition of the tool to call and the corresponding argument values. Will be populated when finish_reason='tool_call'.

### Usage

- `prompt_tokens` (integer, required) — Number of tokens used in the prompt.
- `generation_tokens` (integer, required) — Number of tokens produced by the generation.
- `total_tokens` (integer, required) — Total number of tokens used by the prompt and generation combined.
- `reasoning_tokens` (integer, optional) — Number of tokens used for model used to 'think' before creating a response to the prompt.

### ChatResponseToolChoice

Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'type': 'function', 'function': \{name': \<TOOL\_NAME>}} forces the model to use the named function.

### ValidationError

- `loc` (list of ValidationErrorLocItem, required)
- `msg` (string, required)
- `type` (string, required)

### Content

The content of the message.

### ToolCall

A tool call to be made.

- `id` (string, required)
- `type` ("function", required) — The type of tool to call.
- `function` (FunctionTool, required) — A function tool to be called by the model where user owns runtime.

### ChatMessageWithToolCallThinkingItem

- `type`: `thinking`
  - `signature` (string, required) — Cryptographic signature that verifies the thinking block was generated by Anthropic.
  - `thinking` (string, required) — Model's chain-of-thought for providing the response.
- `type`: `redacted_thinking`
  - `data` (string, required) — Thinking block Anthropic redacted for safety reasons. User is expected to pass the block back to Anthropic

### FunctionTool

A function tool to be called by the model where user owns runtime.

- `name` (string, required)
- `arguments` (string, optional)

### ModelConfigChatRequestStop

The string (or list of strings) after which the model will stop generating. The returned text will not contain the stop sequence.

### ModelConfigChatRequestReasoningEffort

Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget.

### ModelConfigChatRequestToolsItem

### ToolChoice

Tool choice to force the model to use a tool.

- `type` ("function", required) — The type of tool to call.
- `function` (FunctionToolChoice, required) — A function tool to be called by the model where user owns runtime.

### ToolResultResponse

A result from a tool used to populate the prompt template

- `id` (string, required)
- `name` (string, required)
- `signature` (string, required)
- `result` (string, required)

### ValidationErrorLocItem

### LinkedToolRequest

- `id` (string, required) — The ID of the linked tool. Starts with "oc_"
- `source` ("organization", required) — The source of the linked tool. For a linked tool it should be `organization`
- `name` (string, optional) — The name of the linked tool.
- `description` (string, optional) — The description of the linked tool.
- `strict` (boolean, optional) — Whether the tool is strict or not. If strict, the model will be forced to respond with JSON matching the parameters schema.
- `parameters` (map from string to any, optional) — The parameters of the linked tool.

### ModelConfigToolRequest

Definition of tool within a model config. The subset of ToolConfig parameters received by the chat endpoint. Does not have things like the signature or setup schema.

- `name` (string, required) — The name of the tool shown to the model.
- `description` (string, optional) — The description of the tool shown to the model.
- `strict` (boolean, optional) — Whether the tool is strict or not. If strict, the model will be forced to respond with JSON matching the parameters schema.
- `parameters` (map from string to any, optional) — Definition of parameters needed to run the tool. Provided in jsonschema format: https://json-schema.org/
- `source` (enum, optional) — Source of the tool. If defined at an organization level will be 'organization' else 'inline'.
  - Allowed values: `organization`, `inline`
- `source_code` (string, optional) — Code source of the tool.
- `other` (map from string to any, optional) — Other parameters that define the config.
- `preset_name` (string, optional) — If is_preset = true, this is the name of the preset tool on Humanloop. This is used as the key to look up the Humanloop runtime of the tool

### FunctionToolChoice

A function tool to be called by the model where user owns runtime.

- `name` (string, required)

## Examples

**Request**

```json
{
  "stream": true,
  "messages": [
    {
      "role": "user"
    }
  ],
  "model_config": {
    "model": "model"
  }
}
```

**Response**

```json
[
  {
    "project_id": "project_id",
    "num_samples": 1,
    "logprobs": 1,
    "suffix": "suffix",
    "user": "user",
    "data": [
      {
        "id": "id",
        "index": 1,
        "output": "output",
        "raw_output": "raw_output",
        "inputs": {
          "key": "value"
        },
        "finish_reason": "finish_reason",
        "model_config_id": "model_config_id",
        "tool_results": [
          {
            "id": "id",
            "name": "name",
            "signature": "signature",
            "result": "result"
          }
        ],
        "messages": [
          {
            "role": "user"
          }
        ],
        "tool_call": {
          "name": "name"
        },
        "tool_calls": [
          {
            "id": "id",
            "type": "function",
            "function": {
              "name": "name"
            }
          }
        ],
        "output_message": {
          "role": "user"
        }
      }
    ],
    "usage": {
      "prompt_tokens": 1,
      "generation_tokens": 1,
      "total_tokens": 1,
      "reasoning_tokens": 1
    },
    "metadata": {
      "key": "value"
    },
    "provider_responses": [
      {
        "key": "value"
      }
    ],
    "provider_request": {
      "key": "value"
    },
    "session_id": "session_id",
    "tool_choice": "none"
  }
]
```

**SDK Code**

```python
import requests

url = "https://api.humanloop.com/v4/chat"

payload = {
    "stream": True,
    "messages": [{ "role": "user" }],
    "model_config": { "model": "model" }
}
headers = {
    "X-API-KEY": "<apiKey>",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
```

```javascript
const url = 'https://api.humanloop.com/v4/chat';
const options = {
  method: 'POST',
  headers: {'X-API-KEY': '<apiKey>', 'Content-Type': 'application/json'},
  body: '{"stream":true,"messages":[{"role":"user"}],"model_config":{"model":"model"}}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
```

```go
package main

import (
	"fmt"
	"strings"
	"net/http"
	"io"
)

func main() {

	url := "https://api.humanloop.com/v4/chat"

	payload := strings.NewReader("{\n  \"stream\": true,\n  \"messages\": [\n    {\n      \"role\": \"user\"\n    }\n  ],\n  \"model_config\": {\n    \"model\": \"model\"\n  }\n}")

	req, _ := http.NewRequest("POST", url, payload)

	req.Header.Add("X-API-KEY", "<apiKey>")
	req.Header.Add("Content-Type", "application/json")

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby
require 'uri'
require 'net/http'

url = URI("https://api.humanloop.com/v4/chat")

http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true

request = Net::HTTP::Post.new(url)
request["X-API-KEY"] = '<apiKey>'
request["Content-Type"] = 'application/json'
request.body = "{\n  \"stream\": true,\n  \"messages\": [\n    {\n      \"role\": \"user\"\n    }\n  ],\n  \"model_config\": {\n    \"model\": \"model\"\n  }\n}"

response = http.request(request)
puts response.read_body
```

```java
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.post("https://api.humanloop.com/v4/chat")
  .header("X-API-KEY", "<apiKey>")
  .header("Content-Type", "application/json")
  .body("{\n  \"stream\": true,\n  \"messages\": [\n    {\n      \"role\": \"user\"\n    }\n  ],\n  \"model_config\": {\n    \"model\": \"model\"\n  }\n}")
  .asString();
```

```php
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('POST', 'https://api.humanloop.com/v4/chat', [
  'body' => '{
  "stream": true,
  "messages": [
    {
      "role": "user"
    }
  ],
  "model_config": {
    "model": "model"
  }
}',
  'headers' => [
    'Content-Type' => 'application/json',
    'X-API-KEY' => '<apiKey>',
  ],
]);

echo $response->getBody();
```

```csharp
using RestSharp;

var client = new RestClient("https://api.humanloop.com/v4/chat");
var request = new RestRequest(Method.POST);
request.AddHeader("X-API-KEY", "<apiKey>");
request.AddHeader("Content-Type", "application/json");
request.AddParameter("application/json", "{\n  \"stream\": true,\n  \"messages\": [\n    {\n      \"role\": \"user\"\n    }\n  ],\n  \"model_config\": {\n    \"model\": \"model\"\n  }\n}", ParameterType.RequestBody);
IRestResponse response = client.Execute(request);
```

```swift
import Foundation

let headers = [
  "X-API-KEY": "<apiKey>",
  "Content-Type": "application/json"
]
let parameters = [
  "stream": true,
  "messages": [["role": "user"]],
  "model_config": ["model": "model"]
] as [String : Any]

let postData = JSONSerialization.data(withJSONObject: parameters, options: [])

let request = NSMutableURLRequest(url: NSURL(string: "https://api.humanloop.com/v4/chat")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "POST"
request.allHTTPHeaderFields = headers
request.httpBody = postData as Data

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```