> This page is for version v4.0.
> For other versions, use one of these documentation indexes:
> - v5.0 (default): https://humanloop.com/docs/v5/llms.txt
> - v4.0: https://humanloop.com/docs/v4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

# Create

POST https://api.humanloop.com/v4/completion
Content-Type: application/json

Create a completion by providing details of the model configuration in the request.

Reference: https://humanloop.com/docs/v4/api/completions/create

## Authentication

- `X-API-KEY` header (required) — API Key authentication via header

## Request

### Body (application/json)

This endpoint expects an object.

- `stream` (true, required) — If true, tokens will be sent as data-only server-sent events. If num_samples > 1, samples are streamed back independently.
- `model_config` (ModelConfigCompletionRequest, required) — The model configuration used to generate.
- `project` (string, optional) — Unique project name. If no project exists with this name, a new project will be created.
- `project_id` (string, optional) — Unique ID of a project to associate to the log. Either this or `project` must be provided.
- `session_id` (string, optional) — ID of the session to associate the datapoint.
- `session_reference_id` (string, optional) — A unique string identifying the session to associate the datapoint to. Allows you to log multiple datapoints to a session (using an ID kept by your internal systems) by passing the same `session_reference_id` in subsequent log requests. Specify at most one of this or `session_id`.
- `parent_id` (string, optional) — ID associated to the parent datapoint in a session.
- `parent_reference_id` (string, optional) — A unique string identifying the previously-logged parent datapoint in a session. Allows you to log nested datapoints with your internal system IDs by passing the same reference ID as `parent_id` in a prior log request. Specify at most one of this or `parent_id`. Note that this cannot refer to a datapoint being logged in the same request.
- `inputs` (map from string to any, optional) — The inputs passed to the prompt template.
- `source` (string, optional) — Identifies where the model was called from.
- `metadata` (map from string to any, optional) — Any additional metadata to record.
- `save` (boolean, optional, default: true) — Whether the request/response payloads will be stored on Humanloop.
- `source_datapoint_id` (string, optional) — ID of the source datapoint if this is a log derived from a datapoint in a dataset.
- `provider_api_keys` (ProviderApiKeys, optional) — API keys required by each provider to make API calls. The API keys provided here are not stored by Humanloop. If not specified here, Humanloop will fall back to the key saved to your organization.
- `num_samples` (integer, optional, default: 1) — The number of generations.
- `template_language` (enum, optional) — The template language to use for rendering the template.
  - Allowed values: `default`, `jinja`
- `user` (string, optional) — End-user ID passed through to provider call.
- `return_inputs` (boolean, optional, default: true) — Whether to return the inputs in the response. If false, the response will contain an empty dictionary under inputs. This is useful for reducing the size of the response. Defaults to true.
- `logprobs` (integer, optional) — Include the log probabilities of the top n tokens in the provider_response
- `suffix` (string, optional) — The suffix that comes after a completion of inserted text. Useful for completions that act like inserts.
- `seed` (integer, optional, deprecated) — Deprecated field: the seed is instead set as part of the request.config object.

## Response

### 200

- Streaming response of `CompletionResponse`.
- `data` (list of DataResponse, required) — Array containing the generation responses.
- `provider_responses` (list of any, required) — The raw responses returned by the model provider.
- `project_id` (string, optional) — Unique identifier of the parent project. Will not be provided if the request was made without providing a project name or id
- `num_samples` (integer, optional, default: 1) — How many completions to make for each set of inputs.
- `logprobs` (integer, optional) — Include the log probabilities of the top n tokens in the provider_response
- `suffix` (string, optional) — The suffix that comes after a completion of inserted text. Useful for completions that act like inserts.
- `user` (string, optional) — End-user ID passed through to provider call.
- `usage` (Usage, optional) — Counts of the number of tokens used and related stats.
- `metadata` (map from string to any, optional) — Any additional metadata to record.
- `provider_request` (map from string to any, optional) — The raw request sent to the model provider.
- `session_id` (string, optional) — ID of the session if it belongs to one.

## Errors

### 422 Completions Create Stream Request Unprocessable Entity Error

Validation Error

- `detail` (list of ValidationError, optional)

## Types

### ModelConfigCompletionRequest

Completion model config request

- `model` (string, required) — The model instance used. E.g. text-davinci-002.
- `name` (string, optional) — A friendly display name for the model config. If not provided, a name will be generated.
- `description` (string, optional) — A description of the model config.
- `provider` (enum, optional) — The company providing the underlying model service.
  - Allowed values: `anthropic`, `bedrock`, `cohere`, `deepseek`, `google`, `groq`, `mock`, `openai`, `openai_azure`, `replicate`
- `max_tokens` (integer, optional, default: -1) — The maximum number of tokens to generate. Provide max_tokens=-1 to dynamically calculate the maximum number of tokens to generate given the length of the prompt
- `temperature` (double, optional, default: 1) — What sampling temperature to use when making a generation. Higher values means the model will be more creative.
- `top_p` (double, optional, default: 1) — An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass.
- `stop` (ModelConfigCompletionRequestStop, optional) — The string (or list of strings) after which the model will stop generating. The returned text will not contain the stop sequence.
- `presence_penalty` (double, optional, default: 0) — Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the generation so far.
- `frequency_penalty` (double, optional, default: 0) — Number between -2.0 and 2.0. Positive values penalize new tokens based on how frequently they appear in the generation so far.
- `other` (map from string to any, optional) — Other parameter values to be passed to the provider call.
- `seed` (integer, optional) — If specified, model will make a best effort to sample deterministically, but it is not guaranteed.
- `response_format` (ResponseFormat, optional) — The format of the response. Only type json_object is currently supported for chat.
- `reasoning_effort` (ModelConfigCompletionRequestReasoningEffort, optional) — Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget.
- `template_language` (enum, optional) — The template language to use for rendering the template.
  - Allowed values: `default`, `jinja`
- `endpoint` (enum, optional) — The provider model endpoint used.
  - Allowed values: `complete`, `chat`, `edit`
- `prompt_template` (string, optional) — Prompt template that will take your specified inputs to form your final request to the model. Input variables within the prompt template should be specified with syntax: `{{input_name}}`.

### ProviderApiKeys

- `openai` (string, optional)
- `mock` (string, optional)
- `anthropic` (string, optional)
- `deepseek` (string, optional)
- `bedrock` (string, optional)
- `cohere` (string, optional)
- `openai_azure` (string, optional)
- `openai_azure_endpoint` (string, optional)
- `google` (string, optional)

### DataResponse

- `id` (string, required) — Unique ID for the model inputs and output logged to Humanloop. Use this when recording feedback later.
- `index` (integer, required) — The index for the sampled generation for a given input. The num_samples request parameter controls how many samples are generated.
- `output` (string, required) — Output text returned from the provider model with leading and trailing whitespaces stripped.
- `raw_output` (string, required) — Raw output text returned from the provider model.
- `inputs` (map from string to any, required) — The inputs passed to the prompt template.
- `model_config_id` (string, required) — The model configuration used to create the generation.
- `finish_reason` (string, optional) — Why the generation ended. One of 'stop' (indicating a stop token was encountered), or 'length' (indicating the max tokens limit has been reached), or 'tool_call' (indicating that the model has chosen to call a tool - in which case the tool_call parameter of the response will be populated). It will be set as null for the intermediary responses during a stream, and will only be set as non-null for the final streamed token.
- `tool_results` (list of ToolResultResponse, optional) — Results of any tools run during the generation.

### Usage

- `prompt_tokens` (integer, required) — Number of tokens used in the prompt.
- `generation_tokens` (integer, required) — Number of tokens produced by the generation.
- `total_tokens` (integer, required) — Total number of tokens used by the prompt and generation combined.
- `reasoning_tokens` (integer, optional) — Number of tokens used for model used to 'think' before creating a response to the prompt.

### ValidationError

- `loc` (list of ValidationErrorLocItem, required)
- `msg` (string, required)
- `type` (string, required)

### ModelConfigCompletionRequestStop

The string (or list of strings) after which the model will stop generating. The returned text will not contain the stop sequence.

### ResponseFormat

Response format of the model.

- `type` (enum, required)
  - Allowed values: `json_object`, `json_schema`
- `json_schema` (map from string to any, optional) — The JSON schema of the response format if type is json_schema.

### ModelConfigCompletionRequestReasoningEffort

Guidance on how many reasoning tokens it should generate before creating a response to the prompt. OpenAI reasoning models (o1, o3-mini) expect a OpenAIReasoningEffort enum. Anthropic reasoning models expect an integer, which signifies the maximum token budget.

### ToolResultResponse

A result from a tool used to populate the prompt template

- `id` (string, required)
- `name` (string, required)
- `signature` (string, required)
- `result` (string, required)

### ValidationErrorLocItem

## Examples

**Request**

```json
{
  "stream": true,
  "model_config": {
    "model": "model"
  }
}
```

**Response**

```json
[
  {
    "project_id": "project_id",
    "num_samples": 1,
    "logprobs": 1,
    "suffix": "suffix",
    "user": "user",
    "data": [
      {
        "id": "id",
        "index": 1,
        "output": "output",
        "raw_output": "raw_output",
        "inputs": {
          "key": "value"
        },
        "finish_reason": "finish_reason",
        "model_config_id": "model_config_id",
        "tool_results": [
          {
            "id": "id",
            "name": "name",
            "signature": "signature",
            "result": "result"
          }
        ]
      }
    ],
    "usage": {
      "prompt_tokens": 1,
      "generation_tokens": 1,
      "total_tokens": 1,
      "reasoning_tokens": 1
    },
    "metadata": {
      "key": "value"
    },
    "provider_responses": [
      {
        "key": "value"
      }
    ],
    "provider_request": {
      "key": "value"
    },
    "session_id": "session_id"
  }
]
```

**SDK Code**

```python
import requests

url = "https://api.humanloop.com/v4/completion"

payload = {
    "stream": True,
    "model_config": { "model": "model" }
}
headers = {
    "X-API-KEY": "<apiKey>",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
```

```javascript
const url = 'https://api.humanloop.com/v4/completion';
const options = {
  method: 'POST',
  headers: {'X-API-KEY': '<apiKey>', 'Content-Type': 'application/json'},
  body: '{"stream":true,"model_config":{"model":"model"}}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
```

```go
package main

import (
	"fmt"
	"strings"
	"net/http"
	"io"
)

func main() {

	url := "https://api.humanloop.com/v4/completion"

	payload := strings.NewReader("{\n  \"stream\": true,\n  \"model_config\": {\n    \"model\": \"model\"\n  }\n}")

	req, _ := http.NewRequest("POST", url, payload)

	req.Header.Add("X-API-KEY", "<apiKey>")
	req.Header.Add("Content-Type", "application/json")

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby
require 'uri'
require 'net/http'

url = URI("https://api.humanloop.com/v4/completion")

http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true

request = Net::HTTP::Post.new(url)
request["X-API-KEY"] = '<apiKey>'
request["Content-Type"] = 'application/json'
request.body = "{\n  \"stream\": true,\n  \"model_config\": {\n    \"model\": \"model\"\n  }\n}"

response = http.request(request)
puts response.read_body
```

```java
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.post("https://api.humanloop.com/v4/completion")
  .header("X-API-KEY", "<apiKey>")
  .header("Content-Type", "application/json")
  .body("{\n  \"stream\": true,\n  \"model_config\": {\n    \"model\": \"model\"\n  }\n}")
  .asString();
```

```php
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('POST', 'https://api.humanloop.com/v4/completion', [
  'body' => '{
  "stream": true,
  "model_config": {
    "model": "model"
  }
}',
  'headers' => [
    'Content-Type' => 'application/json',
    'X-API-KEY' => '<apiKey>',
  ],
]);

echo $response->getBody();
```

```csharp
using RestSharp;

var client = new RestClient("https://api.humanloop.com/v4/completion");
var request = new RestRequest(Method.POST);
request.AddHeader("X-API-KEY", "<apiKey>");
request.AddHeader("Content-Type", "application/json");
request.AddParameter("application/json", "{\n  \"stream\": true,\n  \"model_config\": {\n    \"model\": \"model\"\n  }\n}", ParameterType.RequestBody);
IRestResponse response = client.Execute(request);
```

```swift
import Foundation

let headers = [
  "X-API-KEY": "<apiKey>",
  "Content-Type": "application/json"
]
let parameters = [
  "stream": true,
  "model_config": ["model": "model"]
] as [String : Any]

let postData = JSONSerialization.data(withJSONObject: parameters, options: [])

let request = NSMutableURLRequest(url: NSURL(string: "https://api.humanloop.com/v4/completion")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "POST"
request.allHTTPHeaderFields = headers
request.httpBody = postData as Data

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```