> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://humanloop.com/docs/v4/api/chats/create-deployed/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/_mcp/server. # Chat Deployed POST https://api.humanloop.com/v4/chat-deployed Content-Type: application/json Get a chat response using the project's active deployment. The active deployment can be a specific model configuration. Reference: https://humanloop.com/docs/v4/api/chats/create-deployed ## Authentication - `X-API-KEY` header (required) — API Key authentication via header ## Request ### Body (application/json) This endpoint expects an object. - `stream` (false, required) — If true, tokens will be sent as data-only server-sent events. If num_samples > 1, samples are streamed back independently. - `messages` (list of ChatMessageWithToolCall, required) — The messages passed to the to provider chat endpoint. - `project` (string, optional) — Unique project name. If no project exists with this name, a new project will be created. - `project_id` (string, optional) — Unique ID of a project to associate to the log. Either this or `project` must be provided. - `session_id` (string, optional) — ID of the session to associate the datapoint. - `session_reference_id` (string, optional) — A unique string identifying the session to associate the datapoint to. Allows you to log multiple datapoints to a session (using an ID kept by your internal systems) by passing the same `session_reference_id` in subsequent log requests. Specify at most one of this or `session_id`. - `parent_id` (string, optional) — ID associated to the parent datapoint in a session. - `parent_reference_id` (string, optional) — A unique string identifying the previously-logged parent datapoint in a session. Allows you to log nested datapoints with your internal system IDs by passing the same reference ID as `parent_id` in a prior log request. Specify at most one of this or `parent_id`. Note that this cannot refer to a datapoint being logged in the same request. - `inputs` (map from string to any, optional) — The inputs passed to the prompt template. - `source` (string, optional) — Identifies where the model was called from. - `metadata` (map from string to any, optional) — Any additional metadata to record. - `save` (boolean, optional, default: true) — Whether the request/response payloads will be stored on Humanloop. - `source_datapoint_id` (string, optional) — ID of the source datapoint if this is a log derived from a datapoint in a dataset. - `provider_api_keys` (ProviderApiKeys, optional) — API keys required by each provider to make API calls. The API keys provided here are not stored by Humanloop. If not specified here, Humanloop will fall back to the key saved to your organization. - `num_samples` (integer, optional, default: 1) — The number of generations. - `template_language` (enum, optional) — The template language to use for rendering the template. - Allowed values: `default`, `jinja` - `user` (string, optional) — End-user ID passed through to provider call. - `return_inputs` (boolean, optional, default: true) — Whether to return the inputs in the response. If false, the response will contain an empty dictionary under inputs. This is useful for reducing the size of the response. Defaults to true. - `tool_choice` (ChatsCreateDeployedRequestToolChoice, optional) — Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'type': 'function', 'function': \{name': \}} forces the model to use the named function. - `response_format` (ResponseFormat, optional) — The format of the response. Only type json_object is currently supported for chat. - `reasoning_effort` (ChatsCreateDeployedRequestReasoningEffort, optional) — Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget. - `environment` (string, optional) — The environment name used to create a chat response. If not specified, the default environment will be used. - `seed` (integer, optional, deprecated) — Deprecated field: the seed is instead set as part of the request.config object. - `tool_call` (ChatsCreateDeployedRequestToolCall, optional, deprecated) — NB: Deprecated with new tool\_choice. Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'name': \} forces the model to use the provided tool of the same name. ## Response ### 200 - `data` (list of ChatDataResponse, required) — Array containing the chat responses. - `provider_responses` (list of any, required) — The raw responses returned by the model provider. - `project_id` (string, optional) — Unique identifier of the parent project. Will not be provided if the request was made without providing a project name or id - `num_samples` (integer, optional, default: 1) — The number of chat responses. - `logprobs` (integer, optional) — Include the log probabilities of the top n tokens in the provider_response - `suffix` (string, optional) — The suffix that comes after a completion of inserted text. Useful for completions that act like inserts. - `user` (string, optional) — End-user ID passed through to provider call. - `usage` (Usage, optional) — Counts of the number of tokens used and related stats. - `metadata` (map from string to any, optional) — Any additional metadata to record. - `provider_request` (map from string to any, optional) — The raw request sent to the model provider. - `session_id` (string, optional) — ID of the session if it belongs to one. - `tool_choice` (ChatResponseToolChoice, optional) — Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'type': 'function', 'function': \{name': \}} forces the model to use the named function. ## Errors ### 422 Chats Create Deployed Request Unprocessable Entity Error Validation Error - `detail` (list of ValidationError, optional) ## Types ### ChatMessageWithToolCall - `role` (enum, required) — Role of the message author. - Allowed values: `user`, `assistant`, `system`, `tool`, `developer` - `content` (Content, optional) — The content of the message. - `name` (string, optional) — Optional name of the message author. - `tool_call_id` (string, optional) — Tool call that this message is responding to. - `tool_calls` (list of ToolCall, optional) — A list of tool calls requested by the assistant. - `thinking` (list of ChatMessageWithToolCallThinkingItem, optional) — Model's chain-of-thought for providing the response. Present on assistant messages if model supports it. - `tool_call` (FunctionTool, optional, deprecated) — NB: Deprecated in favour of tool_calls. A tool call requested by the assistant. ### ProviderApiKeys - `openai` (string, optional) - `mock` (string, optional) - `anthropic` (string, optional) - `deepseek` (string, optional) - `bedrock` (string, optional) - `cohere` (string, optional) - `openai_azure` (string, optional) - `openai_azure_endpoint` (string, optional) - `google` (string, optional) ### ChatsCreateDeployedRequestToolChoice Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'type': 'function', 'function': \{name': \}} forces the model to use the named function. ### ResponseFormat Response format of the model. - `type` (enum, required) - Allowed values: `json_object`, `json_schema` - `json_schema` (map from string to any, optional) — The JSON schema of the response format if type is json_schema. ### ChatsCreateDeployedRequestReasoningEffort Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget. ### ChatsCreateDeployedRequestToolCall NB: Deprecated with new tool\_choice. Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'name': \} forces the model to use the provided tool of the same name. ### ChatDataResponse Overwrite DataResponse for chat. - `id` (string, required) — Unique ID for the model inputs and output logged to Humanloop. Use this when recording feedback later. - `index` (integer, required) — The index for the sampled generation for a given input. The num_samples request parameter controls how many samples are generated. - `output` (string, required) — Output text returned from the provider model with leading and trailing whitespaces stripped. - `raw_output` (string, required) — Raw output text returned from the provider model. - `model_config_id` (string, required) — The model configuration used to create the generation. - `output_message` (ChatMessageWithToolCall, required) — The message returned by the provider. - `inputs` (map from string to any, optional) — The inputs passed to the chat template. - `finish_reason` (string, optional) — Why the generation ended. One of 'stop' (indicating a stop token was encountered), or 'length' (indicating the max tokens limit has been reached), or 'tool_call' (indicating that the model has chosen to call a tool - in which case the tool_call parameter of the response will be populated). It will be set as null for the intermediary responses during a stream, and will only be set as non-null for the final streamed token. - `tool_results` (list of ToolResultResponse, optional) — Results of any tools run during the generation. - `messages` (list of ChatMessageWithToolCall, optional) — The messages passed to the to provider chat endpoint. - `tool_calls` (list of ToolCall, optional) — Deprecated: Please use tool_calls field within the output_message.JSON definition of the tools to call and the corresponding argument values. Will be populated when finish_reason='tool_call'. - `tool_call` (FunctionTool, optional, deprecated) — Deprecated: Please use tool_calls field within the output_message.JSON definition of the tool to call and the corresponding argument values. Will be populated when finish_reason='tool_call'. ### Usage - `prompt_tokens` (integer, required) — Number of tokens used in the prompt. - `generation_tokens` (integer, required) — Number of tokens produced by the generation. - `total_tokens` (integer, required) — Total number of tokens used by the prompt and generation combined. - `reasoning_tokens` (integer, optional) — Number of tokens used for model used to 'think' before creating a response to the prompt. ### ChatResponseToolChoice Controls how the model uses tools. The following options are supported: 'none' forces the model to not call a tool; the default when no tools are provided as part of the model config. 'auto' the model can decide to call one of the provided tools; the default when tools are provided as part of the model config. Providing \{'type': 'function', 'function': \{name': \}} forces the model to use the named function. ### ValidationError - `loc` (list of ValidationErrorLocItem, required) - `msg` (string, required) - `type` (string, required) ### Content The content of the message. ### ToolCall A tool call to be made. - `id` (string, required) - `type` ("function", required) — The type of tool to call. - `function` (FunctionTool, required) — A function tool to be called by the model where user owns runtime. ### ChatMessageWithToolCallThinkingItem - `type`: `thinking` - `signature` (string, required) — Cryptographic signature that verifies the thinking block was generated by Anthropic. - `thinking` (string, required) — Model's chain-of-thought for providing the response. - `type`: `redacted_thinking` - `data` (string, required) — Thinking block Anthropic redacted for safety reasons. User is expected to pass the block back to Anthropic ### FunctionTool A function tool to be called by the model where user owns runtime. - `name` (string, required) - `arguments` (string, optional) ### ToolChoice Tool choice to force the model to use a tool. - `type` ("function", required) — The type of tool to call. - `function` (FunctionToolChoice, required) — A function tool to be called by the model where user owns runtime. ### ToolResultResponse A result from a tool used to populate the prompt template - `id` (string, required) - `name` (string, required) - `signature` (string, required) - `result` (string, required) ### ValidationErrorLocItem ### FunctionToolChoice A function tool to be called by the model where user owns runtime. - `name` (string, required) ## Examples **Request** ```json { "stream": false, "messages": [ { "role": "user", "content": "What is the weather in SF?" } ], "project": "ai-assistant", "inputs": { "persona": "helpful but will *always* tell a joke first before calling tools" } } ``` **Response** ```json { "data": [ { "id": "data_lwWadasRw0vT4XDarZuNQ", "index": 0, "output": "Why did the weather report go to school? To become a little brighter!\n\nLet me check the current weather in San Francisco for you.", "raw_output": "Why did the weather report go to school? To become a little brighter!\n\nLet me check the current weather in San Francisco for you.", "model_config_id": "prv_rlwVnPhRsiMKfnTusferP", "output_message": { "role": "assistant", "content": "Why did the weather report go to school? To become a little brighter!\n\nLet me check the current weather in San Francisco for you.", "name": null, "tool_call_id": null, "tool_calls": [ { "id": "call_1b6yHTGiB51P2I75T6yZrm63", "type": "function", "function": { "name": "get_current_weather", "arguments": "{\"location\":\"San Francisco, CA\"}" } } ], "tool_call": null }, "inputs": { "persona": "helpful but will *always* tell a joke first before calling tools" }, "finish_reason": "tool_call", "tool_results": [], "messages": [ { "role": "system", "content": "You are a helpful assistant with persona helpful but will *always* tell a joke first before calling tools. \nUse tools to respond to user's queries.\n", "name": null, "tool_call_id": null, "tool_calls": null, "tool_call": null }, { "role": "user", "content": "What is the weather in SF?", "name": null, "tool_call_id": null, "tool_calls": null, "tool_call": null } ], "tool_calls": [ { "id": "call_1b6yHTGiB51P2I75T6yZrm63", "type": "function", "function": { "name": "get_current_weather", "arguments": "{\"location\":\"San Francisco, CA\"}" } } ], "tool_call": { "name": "get_current_weather", "arguments": "{\"location\":\"San Francisco, CA\"}" } } ], "provider_responses": [ { "id": "chatcmpl-9TbcxkvnK9Q0VTCO89GGPsgWiI7LY", "choices": [ { "finish_reason": "tool_calls", "index": 0, "logprobs": null, "message": { "content": "Why did the weather report go to school? To become a little brighter!\n\nLet me check the current weather in San Francisco for you.", "role": "assistant", "function_call": null, "tool_calls": [ { "id": "call_1b6yHTGiB51P2I75T6yZrm63", "function": { "arguments": "{\"location\":\"San Francisco, CA\"}", "name": "get_current_weather" }, "type": "function" } ] } } ], "created": 1716843179, "model": "gpt-4o-2024-05-13", "object": "chat.completion", "system_fingerprint": "fp_43dfabdef1", "usage": { "completion_tokens": 46, "prompt_tokens": 137, "total_tokens": 183 } } ], "project_id": "pr_TfhDgggIsPi3cgmhq2yeA", "num_samples": 1, "logprobs": null, "suffix": null, "user": null, "usage": { "prompt_tokens": 137, "generation_tokens": 46, "total_tokens": 183 }, "metadata": null, "provider_request": { "messages": [ { "content": "You are a helpful assistant with persona helpful but will *always* tell a joke first before calling tools. \nUse tools to respond to user's queries.\n", "role": "system" }, { "content": "What is the weather in SF?", "role": "user" } ], "stream": false, "n": 1, "model": "gpt-4o", "temperature": 0.7, "top_p": 1, "presence_penalty": 0, "frequency_penalty": 0, "tools": [ { "type": "function", "function": { "name": "get_current_weather", "description": "Get the current weather in a given location", "parameters": { "type": "object", "properties": { "location": { "type": "string", "name": "Location", "description": "The city and state, e.g. San Francisco, CA" }, "unit": { "type": "string", "name": "Unit", "enum": [ "celsius", "fahrenheit" ] } }, "required": [ "location" ] } } }, { "type": "function", "function": { "name": "get_stock_price", "description": "Get current stock price", "parameters": { "type": "object", "properties": { "ticker_symbol": { "type": "string", "name": "Ticker Symbol", "description": "Ticker symbol of the stock" } }, "required": [] } } } ], "tool_choice": "auto" }, "session_id": null, "tool_choice": null } ``` **SDK Code** ```python ai-assistant-with-tools import requests url = "https://api.humanloop.com/v4/chat-deployed" payload = { "stream": False, "messages": [ { "role": "user", "content": "What is the weather in SF?" } ], "project": "ai-assistant", "inputs": { "persona": "helpful but will *always* tell a joke first before calling tools" } } headers = { "X-API-KEY": "", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json()) ``` ```javascript ai-assistant-with-tools const url = 'https://api.humanloop.com/v4/chat-deployed'; const options = { method: 'POST', headers: {'X-API-KEY': '', 'Content-Type': 'application/json'}, body: '{"stream":false,"messages":[{"role":"user","content":"What is the weather in SF?"}],"project":"ai-assistant","inputs":{"persona":"helpful but will *always* tell a joke first before calling tools"}}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); } ``` ```go ai-assistant-with-tools package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.humanloop.com/v4/chat-deployed" payload := strings.NewReader("{\n \"stream\": false,\n \"messages\": [\n {\n \"role\": \"user\",\n \"content\": \"What is the weather in SF?\"\n }\n ],\n \"project\": \"ai-assistant\",\n \"inputs\": {\n \"persona\": \"helpful but will *always* tell a joke first before calling tools\"\n }\n}") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("X-API-KEY", "") req.Header.Add("Content-Type", "application/json") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby ai-assistant-with-tools require 'uri' require 'net/http' url = URI("https://api.humanloop.com/v4/chat-deployed") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["X-API-KEY"] = '' request["Content-Type"] = 'application/json' request.body = "{\n \"stream\": false,\n \"messages\": [\n {\n \"role\": \"user\",\n \"content\": \"What is the weather in SF?\"\n }\n ],\n \"project\": \"ai-assistant\",\n \"inputs\": {\n \"persona\": \"helpful but will *always* tell a joke first before calling tools\"\n }\n}" response = http.request(request) puts response.read_body ``` ```java ai-assistant-with-tools import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.humanloop.com/v4/chat-deployed") .header("X-API-KEY", "") .header("Content-Type", "application/json") .body("{\n \"stream\": false,\n \"messages\": [\n {\n \"role\": \"user\",\n \"content\": \"What is the weather in SF?\"\n }\n ],\n \"project\": \"ai-assistant\",\n \"inputs\": {\n \"persona\": \"helpful but will *always* tell a joke first before calling tools\"\n }\n}") .asString(); ``` ```php ai-assistant-with-tools request('POST', 'https://api.humanloop.com/v4/chat-deployed', [ 'body' => '{ "stream": false, "messages": [ { "role": "user", "content": "What is the weather in SF?" } ], "project": "ai-assistant", "inputs": { "persona": "helpful but will *always* tell a joke first before calling tools" } }', 'headers' => [ 'Content-Type' => 'application/json', 'X-API-KEY' => '', ], ]); echo $response->getBody(); ``` ```csharp ai-assistant-with-tools using RestSharp; var client = new RestClient("https://api.humanloop.com/v4/chat-deployed"); var request = new RestRequest(Method.POST); request.AddHeader("X-API-KEY", ""); request.AddHeader("Content-Type", "application/json"); request.AddParameter("application/json", "{\n \"stream\": false,\n \"messages\": [\n {\n \"role\": \"user\",\n \"content\": \"What is the weather in SF?\"\n }\n ],\n \"project\": \"ai-assistant\",\n \"inputs\": {\n \"persona\": \"helpful but will *always* tell a joke first before calling tools\"\n }\n}", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ``` ```swift ai-assistant-with-tools import Foundation let headers = [ "X-API-KEY": "", "Content-Type": "application/json" ] let parameters = [ "stream": false, "messages": [ [ "role": "user", "content": "What is the weather in SF?" ] ], "project": "ai-assistant", "inputs": ["persona": "helpful but will *always* tell a joke first before calling tools"] ] as [String : Any] let postData = JSONSerialization.data(withJSONObject: parameters, options: []) let request = NSMutableURLRequest(url: NSURL(string: "https://api.humanloop.com/v4/chat-deployed")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers request.httpBody = postData as Data let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```