> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> Experiments allow you to set up A/B tests between multiple model configs.

Experiments can be used to compare different prompt templates, parameter combinations (such as temperature and presence penalties), and even base models.

### Prerequisites

* You already have a Prompt — if not, please follow our [Prompt creation](/docs/guides/create-prompt) guide first.
* You have integrated `humanloop.complete_deployed()` or the `humanloop.chat_deployed()` endpoints, along with the `humanloop.feedback()` with the [API](https://www.postman.com/humanloop/workspace/humanloop) or [Python SDK](./generate-and-log-with-the-sdk).

> **Info**
>
> This guide assumes you're using an OpenAI model. If you want to use other
> providers or your model, refer to the [guide for running an experiment with your model provider](./use-your-own-model-provider).

## Create an experiment

### Navigate to the **Experiments** tab of your Prompt

### Click the **Create new experiment** button

1. Give your experiment a descriptive name.
2. Select a list of feedback labels to be considered as positive actions - this will be used to calculate the performance of each of your model configs during the experiment.
3. Select which of your project’s model configs to compare.
4. Then click the **Create** button.

![](/docs/_fern-img/9cde9f7b2489950a5da2c2fad8ce04010ffb2b41e04069df430b88e922b92697.webp)

## Set the experiment live

Now that you have an experiment, you need to set it as the project’s active experiment:

### Navigate to the **Experiments** tab.

Of a Prompt go to the **Experiments** tab.

### Choose the **Experiment** card you want to deploy.

### Click the **Deploy** button

Next to the Environments label, click the **Deploy** button.

### Select the environment to deploy the experiment.

We only have one environment by default so select the 'production' environment.

![](/docs/_fern-img/ff73e8b1b44e12331ba01c1661fe01d292928bfedc439c05e5d021ddc19dc367.webp)

> **Check**
>
> Now that your experiment is active, any SDK or API calls to generate will
> sample model configs from the list you provided when creating the experiment
> and any subsequent feedback captured using feedback will contribute to the
> experiment performance.

## Monitor experiment progress

Now that an experiment is live, the data flowing through your generate and feedback calls will update the experiment progress in real-time:

### Navigate back to the **Experiments** tab.

### Select the **Experiment** card

Here you will see the performance of each model config with a measure of confidence based on how much feedback data has been collected so far:

![](/docs/_fern-img/ff73e8b1b44e12331ba01c1661fe01d292928bfedc439c05e5d021ddc19dc367.webp)![You can toggle on and off existing model configs and choose to add new model configs from your project over the lifecycle of an experiment](/docs/_fern-img/b7995b55f885453b1216dbc7107ea65324257d80d6cb63e07a9cf105818bd373.webp)

🎉 Your experiment can now give you insight into which of the model configs your users prefer.

> **Tip**
>
> How quickly you can draw conclusions depends on how much traffic you have flowing through your project.
>
> Generally, you should be able to draw some initial conclusions after on the order of hundreds of examples.