> This page is for version v4.0.
> For other versions, use one of these documentation indexes:
> - v5.0 (default): https://humanloop.com/docs/v5/llms.txt
> - v4.0: https://humanloop.com/docs/v4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> How do you evaluate your large language model use case using a dataset and an evaluator on Humanloop?

> **Paid Feature**
>
> This feature is not available for the Free tier. Please contact us if you wish
> to learn more about our [Enterprise plan](https://humanloop.com/pricing)

## Create an offline evaluator

### Prerequisites

* You need to have access to Evaluations
* You also need to have a Prompt – if not, please follow our [Prompt creation](/docs/v4/guides/create-prompt) guide.
* Finally, you need at least a few Logs in your prompt. Use the **Editor** to generate some logs if you have none.

> **Info**
>
> You need logs for your project because we will use these as a source of test
> datapoints for the dataset we create. If you want to make arbitrary test
> datapoints from scratch, see our guide to doing this from the API. We will
> soon update the app to enable arbitrary test datapoint creation from your
> browser.

For this example, we will evaluate a model responsible for extracting critical information from a customer service request and returning this information in JSON. In the image below, you can see the model config we've drafted on the left and an example of it running against a customer query on the right.

![](/docs/_fern-img/7a98b9714497eb6ea1992996cea3659eb2c48ad9cd51125c500dd01224791f9c.webp)

### Set up a dataset

We will create a dataset based on existing logs in the project.

### Navigate to the **Logs** tab

### Select the logs you would like to convert into test datapoints

### From the dropdown menu in the top right (see below), choose **Add to Dataset**

![Creating test datapoints from a selection of existing project datapoints.](/docs/_fern-img/8319576e3ddd32d8092f0589f01e27ac9580bc098e252eafd1f6f1d4dce5148c.webp)

### In the dialog box, give the new dataset a name and provide an optional description. Click **Create dataset**.

![](/docs/_fern-img/7a3ce56a6b618a3e5a8845b45ffceed3263c1f28d04b9ac642eaea42f8b419cf.webp)

> **Info**
>
> You can add more datapoints to the same dataset later by clicking the 'add to
> existing dataset' button at the top.

### Go to the **Datasets** tab.

### Click on the newly created dataset. One datapoint will be present for each log you selected in Step 3

![The newly created dataset, containing datapoints converted from existing logs in the project.](/docs/_fern-img/840124397192bf5d5f7fed82129b6c8be2bd222c8ae31622658a2f46b3b808e8.webp)

### Click on a datapoint to inspect its parameters.

> **Tip**
>
> A test datapoint contains inputs (the variables passed into your model config template), an optional sequence of messages (if used for a chat model) and a target representing the desired output.
>
> When existing logs are converted to datapoints, the datapoint target defaults to the output of the source Log.

In our example, we created datapoints from existing logs. The default behaviour is that the original log's output becomes an output field in the target JSON.

To access the `feature` field more efficiently in our evaluator, we'll modify the datapoint targets to be a raw JSON with a feature key.

![The original log was an LLM generation which outputted a JSON value. The conversion process has placed this into the \`output\` field of the testcase target.](/docs/_fern-img/74d505949cbe063586bd3592a42ea53c9fb7a43a3e98461446cfdf06fa778d0b.webp)

### Modify the datapoint if you need to make refinements

You can provide an arbitrary JSON object as the target.

![After editing, we have a clean JSON object recording the salient characteristics of the datapoint's expected output.](/docs/_fern-img/c608579da10d3cb8a4f61ff21417b925d755ebdf279e36921bf49c6fe2cb4f7f.webp)

## Create an offline evaluator

Having set up a dataset, we'll now create the evaluator. As with online evaluators, it's a Python function but for offline mode, it also takes a `testcase` parameter alongside the generated log.

### Navigate to the evaluations section, and then the Evaluators tab

### Select **+ New Evaluator** and choose **Offline Evaluation**

### Choose **Start from scratch**

For this example, we'll use the code below to compare the LLM generated output with what we expected for that testcase.

**`Python`**

```python Python
import json
from json import JSONDecodeError

def it_extracts_correct_feature(log, testcase):
    expected_feature = testcase["target"]["feature"]

    try:
        # The model is expected to produce valid JSON output
        # but it could fail to do so.
        output = json.loads(log["output"])
        actual_feature = output.get("feature", None)
        return expected_feature == actual_feature

    except JSONDecodeError:
        # If the model didn't even produce valid JSON, then
        # we evaluate the output as bad.
        return False
```

### Use the Debug Console

In the debug console at the bottom of the dialog, click **Load data** and then **Datapoints from dataset**. Select the dataset you created in the previous section. The console will be populated with its datapoints.

![The debug console. Use this to load test datapoints from a dataset and perform debug runs with any model config in your project.](/docs/_fern-img/7a9db9aa92ec7897125f69ac47ebcae9822f7f4c2e4254fe831bf96021654c6a.webp)

#### Choose a model config from the dropdown menu.

#### Click the run button at the far right of one of the test datapoints.

A new debug run will be triggered, which causes an LLM generation using that datapoint's inputs and messages parameters. The generated log and the test datapoint will be passed to the evaluator, and the resulting evaluation will be displayed in the **Result** column.

### Click **Create** when you are happy with the evaluator.

## Trigger an offline evaluation

Now that you have an offline evaluator and a dataset, you can use them to evaluate the performance of any model config in your project.

### Go to the **Evaluations** section.

### In the **Runs** tab, click **Run Evaluation**

### In the dialog box, choose a model config to evaluate and select your newly created dataset and evaluator.

![](/docs/_fern-img/f48a63ef97c90ca4c28625f0c7cf84e3aace0a0bce060061388709a4ebb26bf1.webp)

### Click **Batch Generate**

### A new evaluation is launched. Click on the card to inspect the results.

A batch generation has now been triggered. This means that the model config you selected will be used to generate a log for each datapoint in the dataset. It may take some time for the evaluation to complete, depending on how many test datapoints are in your dataset and what model config you are using. Once all the logs have been generated, the evaluator will execute for each in turn.

### Inspect the results of the evaluation.

![](/docs/_fern-img/a834bbc007d64cb9aa68f0e4ae69c5c2047afeced885cf49b1e699fe0755da55.webp)