> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://humanloop.com/docs/v5/api/prompts/call-stream/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/_mcp/server. # Call Prompt POST https://api.humanloop.com/v5/prompts/call Content-Type: application/json Call a Prompt. Calling a Prompt calls the model provider before logging the request, responses and metadata to Humanloop. You can use query parameters `version_id`, or `environment`, to target an existing version of the Prompt. Otherwise the default deployed version will be chosen. Instead of targeting an existing version explicitly, you can instead pass in Prompt details in the request body. In this case, we will check if the details correspond to an existing version of the Prompt. If they do not, we will create a new version. This is helpful in the case where you are storing or deriving your Prompt details in code. Reference: https://humanloop.com/docs/api/prompts/call ## Authentication - `X-API-KEY` header (required) — API Key authentication via header ## Request ### Query parameters - `version_id` (string, optional) — A specific Version ID of the Prompt to log to. - `environment` (string, optional) — Name of the Environment identifying a deployed version to log to. ### Body (application/json) This endpoint expects an object. - `stream` (true, required) — If true, tokens will be sent as data-only server-sent events. If num_samples > 1, samples are streamed back independently. - `path` (string, optional) — Path of the Prompt, including the name. This locates the Prompt in the Humanloop filesystem and is used as as a unique identifier. For example: `folder/name` or just `name`. - `id` (string, optional) — ID for an existing Prompt. - `messages` (list of ChatMessage, optional) — The messages passed to the to provider chat endpoint. - `tool_choice` (PromptsCallStreamRequestToolChoice, optional) — Controls how the model uses tools. The following options are supported: * `'none'` means the model will not call any tool and instead generates a message; this is the default when no tools are provided as part of the Prompt. * `'auto'` means the model can decide to call one or more of the provided tools; this is the default when tools are provided as part of the Prompt. * `'required'` means the model must call one or more of the provided tools. * `{'type': 'function', 'function': {name': }}` forces the model to use the named function. - `prompt` (PromptsCallStreamRequestPrompt, optional) — The Prompt configuration to use. Two formats are supported: - An object representing the details of the Prompt configuration - A string representing the raw contents of a .prompt file A new Prompt version will be created if the provided details do not match any existing version. - `inputs` (map from string to any, optional) — The inputs passed to the prompt template. - `source` (string, optional) — Identifies where the model was called from. - `metadata` (map from string to any, optional) — Any additional metadata to record. - `start_time` (datetime, optional) — When the logged event started. - `end_time` (datetime, optional) — When the logged event ended. - `source_datapoint_id` (string, optional) — Unique identifier for the Datapoint that this Log is derived from. This can be used by Humanloop to associate Logs to Evaluations. If provided, Humanloop will automatically associate this Log to Evaluations that require a Log for this Datapoint-Version pair. - `trace_parent_id` (string, optional) — The ID of the parent Log to nest this Log under in a Trace. - `user` (string, optional) — End-user ID related to the Log. - `environment` (string, optional) — The name of the Environment the Log is associated to. - `save` (boolean, optional, default: true) — Whether the request/response payloads will be stored on Humanloop. - `log_id` (string, optional) — This will identify a Log. If you don't provide a Log ID, Humanloop will generate one for you. - `provider_api_keys` (ProviderApiKeys, optional) — API keys required by each provider to make API calls. The API keys provided here are not stored by Humanloop. If not specified here, Humanloop will fall back to the key saved to your organization. - `num_samples` (integer, optional, default: 1) — The number of generations. - `return_inputs` (boolean, optional, default: true) — Whether to return the inputs in the response. If false, the response will contain an empty dictionary under inputs. This is useful for reducing the size of the response. Defaults to true. - `logprobs` (integer, optional) — Include the log probabilities of the top n tokens in the provider_response - `suffix` (string, optional) — The suffix that comes after a completion of inserted text. Useful for completions that act like inserts. ## Response ### 200 - Streaming response of `PromptCallStreamResponse`. - `index` (integer, required) — The index of the sample in the batch. - `id` (string, required) — ID of the log. - `prompt_id` (string, required) — ID of the Prompt the log belongs to. - `version_id` (string, required) — ID of the specific version of the Prompt. - `output` (string, optional) — Generated output from your model for the provided inputs. Can be `None` if logging an error, or if creating a parent Log with the intention to populate it later. - `created_at` (datetime, optional) — User defined timestamp for when the log was created. - `error` (string, optional) — Error message if the log is an error. - `provider_latency` (double, optional) — Duration of the logged event in seconds. - `stdout` (string, optional) — Captured log and debug statements. - `output_message` (ChatMessage, optional) — The message returned by the provider. - `prompt_tokens` (integer, optional) — Number of tokens in the prompt used to generate the output. - `reasoning_tokens` (integer, optional) — Number of reasoning tokens used to generate the output. - `output_tokens` (integer, optional) — Number of tokens in the output generated by the model. - `prompt_cost` (double, optional) — Cost in dollars associated to the tokens in the prompt. - `output_cost` (double, optional) — Cost in dollars associated to the tokens in the output. - `finish_reason` (string, optional) — Reason the generation finished. ## Errors ### 422 Prompts Call Stream Request Unprocessable Entity Error Validation Error - `detail` (list of ValidationError, optional) ## Types ### ChatMessage - `role` (enum, required) — Role of the message author. - Allowed values: `user`, `assistant`, `system`, `tool`, `developer` - `content` (ChatMessageContent, optional) — The content of the message. - `name` (string, optional) — Optional name of the message author. - `tool_call_id` (string, optional) — Tool call that this message is responding to. - `tool_calls` (list of ToolCall, optional) — A list of tool calls requested by the assistant. - `thinking` (list of ChatMessageThinkingItem, optional) — Model's chain-of-thought for providing the response. Present on assistant messages if model supports it. ### PromptsCallStreamRequestToolChoice Controls how the model uses tools. The following options are supported: * `'none'` means the model will not call any tool and instead generates a message; this is the default when no tools are provided as part of the Prompt. * `'auto'` means the model can decide to call one or more of the provided tools; this is the default when tools are provided as part of the Prompt. * `'required'` means the model must call one or more of the provided tools. * `{'type': 'function', 'function': {name': }}` forces the model to use the named function. ### PromptsCallStreamRequestPrompt The Prompt configuration to use. Two formats are supported: - An object representing the details of the Prompt configuration - A string representing the raw contents of a .prompt file A new Prompt version will be created if the provided details do not match any existing version. ### ProviderApiKeys - `openai` (string, optional) - `mock` (string, optional) - `anthropic` (string, optional) - `deepseek` (string, optional) - `bedrock` (string, optional) - `cohere` (string, optional) - `openai_azure` (string, optional) - `openai_azure_endpoint` (string, optional) - `google` (string, optional) ### ValidationError - `loc` (list of ValidationErrorLocItem, required) - `msg` (string, required) - `type` (string, required) ### ChatMessageContent The content of the message. ### ToolCall A tool call to be made. - `id` (string, required) - `type` ("function", required) — The type of tool to call. - `function` (FunctionTool, required) — A function tool to be called by the model where user owns runtime. ### ChatMessageThinkingItem ### ToolChoice Tool choice to force the model to use a tool. - `type` ("function", required) — The type of tool to call. - `function` (FunctionToolChoice, required) — A function tool to be called by the model where user owns runtime. ### PromptKernelRequest Base class used by both PromptKernelRequest and AgentKernelRequest. Contains the consistent Prompt-related fields. - `model` (string, required) — The model instance used, e.g. `gpt-4`. See [supported models](https://humanloop.com/docs/reference/supported-models) - `endpoint` (enum, optional) — The provider model endpoint used. - Allowed values: `complete`, `chat`, `edit` - `template` (PromptKernelRequestTemplate, optional) — The template contains the main structure and instructions for the model, including input variables for dynamic values. For chat models, provide the template as a ChatTemplate (a list of messages), e.g. a system message, followed by a user message with an input variable. For completion models, provide a prompt template as a string. Input variables should be specified with double curly bracket syntax: `{{input_name}}`. - `template_language` (enum, optional) — The template language to use for rendering the template. - Allowed values: `default`, `jinja` - `provider` (enum, optional) — The company providing the underlying model service. - Allowed values: `anthropic`, `bedrock`, `cohere`, `deepseek`, `google`, `groq`, `mock`, `openai`, `openai_azure`, `replicate` - `max_tokens` (integer, optional, default: -1) — The maximum number of tokens to generate. Provide max_tokens=-1 to dynamically calculate the maximum number of tokens to generate given the length of the prompt - `temperature` (double, optional, default: 1) — What sampling temperature to use when making a generation. Higher values means the model will be more creative. - `top_p` (double, optional, default: 1) — An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass. - `stop` (PromptKernelRequestStop, optional) — The string (or list of strings) after which the model will stop generating. The returned text will not contain the stop sequence. - `presence_penalty` (double, optional, default: 0) — Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the generation so far. - `frequency_penalty` (double, optional, default: 0) — Number between -2.0 and 2.0. Positive values penalize new tokens based on how frequently they appear in the generation so far. - `other` (map from string to any, optional) — Other parameter values to be passed to the provider call. - `seed` (integer, optional) — If specified, model will make a best effort to sample deterministically, but it is not guaranteed. - `response_format` (ResponseFormat, optional) — The format of the response. Only `{"type": "json_object"}` is currently supported for chat. - `reasoning_effort` (PromptKernelRequestReasoningEffort, optional) — Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget. - `tools` (list of ToolFunction, optional) — The tool specification that the model can choose to call if Tool calling is supported. - `linked_tools` (list of string, optional) — The IDs of the Tools in your organization that the model can choose to call if Tool calling is supported. The default deployed version of that tool is called. - `attributes` (map from string to any, optional) — Additional fields to describe the Prompt. Helpful to separate Prompt versions from each other with details on how they were created or used. ### ValidationErrorLocItem ### FunctionTool A function tool to be called by the model where user owns runtime. - `name` (string, required) - `arguments` (string, optional) ### AnthropicThinkingContent - `type` ("thinking", required) - `thinking` (string, required) — Model's chain-of-thought for providing the response. - `signature` (string, required) — Cryptographic signature that verifies the thinking block was generated by Anthropic. ### AnthropicRedactedThinkingContent - `type` ("redacted_thinking", required) - `data` (string, required) — Thinking block Anthropic redacted for safety reasons. User is expected to pass the block back to Anthropic ### FunctionToolChoice A function tool to be called by the model where user owns runtime. - `name` (string, required) ### PromptKernelRequestTemplate The template contains the main structure and instructions for the model, including input variables for dynamic values. For chat models, provide the template as a ChatTemplate (a list of messages), e.g. a system message, followed by a user message with an input variable. For completion models, provide a prompt template as a string. Input variables should be specified with double curly bracket syntax: `{{input_name}}`. ### PromptKernelRequestStop The string (or list of strings) after which the model will stop generating. The returned text will not contain the stop sequence. ### ResponseFormat Response format of the model. - `type` (enum, required) - Allowed values: `json_object`, `json_schema` - `json_schema` (map from string to any, optional) — The JSON schema of the response format if type is json_schema. ### PromptKernelRequestReasoningEffort Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget. ### ToolFunction - `name` (string, required) — Name for the tool referenced by the model. - `description` (string, required) — Description of the tool referenced by the model - `strict` (boolean, optional, default: false) — If true, forces the model to output json data in the structure of the parameters schema. - `parameters` (map from string to any, optional) — Parameters needed to run the Tool, defined in JSON Schema format: https://json-schema.org/ ## Examples **Request** ```json { "stream": true } ``` **Response** ```text event: data: {"output":"output","created_at":"2024-01-15T09:30:00Z","error":"error","provider_latency":1.1,"stdout":"stdout","output_message":{"content":"content","name":"name","tool_call_id":"tool_call_id","role":"user","tool_calls":[{"id":"id","type":"function","function":{"name":"name"}}],"thinking":[{"type":"thinking","thinking":"thinking","signature":"signature"}]},"prompt_tokens":1,"reasoning_tokens":1,"output_tokens":1,"prompt_cost":1.1,"output_cost":1.1,"finish_reason":"finish_reason","index":1,"id":"id","prompt_id":"prompt_id","version_id":"version_id"} ``` **SDK Code** ```python import requests url = "https://api.humanloop.com/v5/prompts/call" payload = { "stream": True } headers = { "X-API-KEY": "", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json()) ``` ```typescript import { HumanloopClient } from "humanloop"; const client = new HumanloopClient({ apiKey: "YOUR_API_KEY" }); const response = await client.prompts.callStream({}); for await (const item of response) { console.log(item); } ``` ```go package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.humanloop.com/v5/prompts/call" payload := strings.NewReader("{\n \"stream\": true\n}") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("X-API-KEY", "") req.Header.Add("Content-Type", "application/json") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby require 'uri' require 'net/http' url = URI("https://api.humanloop.com/v5/prompts/call") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["X-API-KEY"] = '' request["Content-Type"] = 'application/json' request.body = "{\n \"stream\": true\n}" response = http.request(request) puts response.read_body ``` ```java import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.humanloop.com/v5/prompts/call") .header("X-API-KEY", "") .header("Content-Type", "application/json") .body("{\n \"stream\": true\n}") .asString(); ``` ```php request('POST', 'https://api.humanloop.com/v5/prompts/call', [ 'body' => '{ "stream": true }', 'headers' => [ 'Content-Type' => 'application/json', 'X-API-KEY' => '', ], ]); echo $response->getBody(); ``` ```csharp using RestSharp; var client = new RestClient("https://api.humanloop.com/v5/prompts/call"); var request = new RestRequest(Method.POST); request.AddHeader("X-API-KEY", ""); request.AddHeader("Content-Type", "application/json"); request.AddParameter("application/json", "{\n \"stream\": true\n}", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ``` ```swift import Foundation let headers = [ "X-API-KEY": "", "Content-Type": "application/json" ] let parameters = ["stream": true] as [String : Any] let postData = JSONSerialization.data(withJSONObject: parameters, options: []) let request = NSMutableURLRequest(url: NSURL(string: "https://api.humanloop.com/v5/prompts/call")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers request.httpBody = postData as Data let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```