> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> Learn how to create a code Evaluators in Humanloop to assess the performance of your AI applications. This guide covers setting up an offline evaluator, writing evaluation logic, and using the debug console.

A code [Evaluator](/docs/explanation/evaluators) is a Python function that takes a generated [Log](/docs/explanation/logs) (and optionally a testcase [Datapoint](/docs/explanation/datasets) if comparing to expected results) as input and returns a **judgement**.
The judgement is in the form of a boolean or number that measures some criteria of the generated Log defined within the code.

Code Evaluators provide a flexible way to evaluate the performance of your AI applications, allowing you to re-use existing evaluation packages as well as define custom evaluation heuristics.

We support a fully featured Python environment; details on the supported packages can be found in the [environment reference](/docs/v5/reference/python-environment)

### Prerequisites

You should have an existing [Prompt](/docs/explanation/prompts) to evaluate and already generated some [Logs](/docs/explanation/logs).
Follow our guide on [creating a Prompt](/docs/guides/prompts/create-prompt).

In this example, we'll reference a Prompt that categorises a user query about Humanloop's product and docs by which feature it relates to.

![An example Prompt with a variable \`\{\{query}}\`.](/docs/_fern-img/c472e1a18df70a05953765e200c66067d5acac3d2a87b44fe73c3b3592483421.webp)

## Create a code Evaluator

### Create a new Evaluator

* Click the **New** button at the bottom of the left-hand sidebar, select **Evaluator**, then select **Code**.

![Create code evaluator.](/docs/_fern-img/e1d3ea67d54681edcf429a89b2da68b5f864f81a6451059a144c69a984fb7dd2.webp)

* Give the Evaluator a name when prompted in the sidebar, for example `Category Validator`.

### Define the Evaluator code

After creating the Evaluator, you will automatically be taken to the code editor.
For this example, our Evaluator will check that the feature category returned by the Prompt is from the list of allowed feature categories. We want to ensure our categoriser isn't hallucinating new features.

* Make sure the **Mode** of the Evaluator is set to **Online** in the options on the left.
* Copy and paste the following code into the code editor:

**`Python`**

```python Python

ALLOWED_FEATURES = [
    "Prompt Editor",
    "Model Integrations",
    "Online Monitoring",
    "Offline Evaluations",
    "Dataset Management",
    "User Management",
    "Roles Based Access Control",
    "Deployment Options",
    "Collaboration",
    "Agents and chaining"
]

def validate_feature(log):
    print(f"Full log output: \n {log['output']}")
    # Parse the final line of the log output to get the returned category
    feature = log["output"].split("\n")[-1]
    return feature in ALLOWED_FEATURES
```

> **Code Organisation**
>
> You can define multiple functions in the code Editor to organize your
> evaluation logic. The final function defined is used as the main Evaluator
> entry point that takes the Log argument and returns a valid judgement.

### Debug the code with Prompt Logs

* In the debug console beneath where you pasted the code, click **Select Prompt or Dataset** and find and select the Prompt you're evaluating.
  The debug console will load a sample of Logs from that Prompt.

![The debug console for testing the code.](/docs/_fern-img/23fb727c1f14215ef73e7db1ff64688ce907a79171f62471f832393736dd1e56.webp)

* Click the **Run** button at the far right of one of the loaded Logs to trigger a debug run. This causes the code to be executed with the selected Log as input and populates the **Result** column.
* Inspect the output of the executed code by selecting the arrow to the right of **Result**.

![Inspect evaluator log in debug console.](/docs/_fern-img/1a61ae0b778f34f5e5347b65ecb2c2c2f4440453174d0ea229166ce7019fe520.webp)

### Save the code

Now that you've validated the behaviour, save the code by selecting the **Save** button at the top right of the Editor and optionally provide a suitable version name and description.

### Inspect Evaluator logs

Navigate to the **Logs** tab of the Evaluator to see and debug all the historic usages of this Evaluator.

![Evaluator logs table.](/docs/_fern-img/ca9b1cfba3682ebd6e1d3f4c2a430bf0acc04a5107203d5b6179a915a494cc76.webp)

## Monitor a Prompt

Now that you have an Evaluator, you can use it to monitor the performance of your Prompt by linking it so that it is automatically run on new Logs.

### Link the Evaluator to the Prompt

* Navigate to the **Dashboard** of your Prompt
* Select the **Monitoring** button above the graph and select **Connect Evaluators**.
* Find and select the Evaluator you just created and click **Chose**.

![Select Evaluator for monitoring.](/docs/_fern-img/c051741592944e93c023ae1153468cf6273aa78fd122a19bb6631b95c4398008.webp)

> **Linking Evaluators for Monitoring**
>
> You can link to a deployed version of the Evaluator by choosing the
> environment such as `production`, or you can link to a specific version of the
> Evaluator. If you want changes deployed to your Evaluator to be automatically
> reflected in Monitoring, link to the environment, otherwise link to a specific
> version.

This linking results in: - An additional graph on your Prompt dashboard showing the Evaluator results over time. - An additional column in your Prompt Versions table showing the aggregated Evaluator results for each version. - An additional column in your Logs table showing the Evaluator results for each Log.

### Generate new Logs

Navigate to the **Editor** tab of your Prompt and generate a new Log by entering a query and clicking **Run**.

### Inspect the Monitoring results

Navigate to the **Logs** tab of your Prompt and see the result of the linked Evaluator against the new Log. You can filter on this value in order to [create a Dataset](/docs/guides/evals/create-dataset-api) of interesting examples.

![See the results of monitoring on your logs.](/docs/_fern-img/8e0d221da83ae3924ca5e111928b9ec31f5731f0409e09d2cf0c56619f6bf874.webp)

## Evaluating a Dataset

When running a code Evaluator on a [Dataset](/docs/explanation/datasets), you can compare a generated [Log](/docs/explanation/logs) to each Datapoint's target. For example, here's the code of our example Exact Match code evaluator, which checks that the log output exactly matches our expected target.

**`Python`**

```python Python
def exact_match(log, testcase):
    target = testcase["target"]["output"]
    generation = log["output"]

    return target == generation
```

## Next steps

* Explore [AI Evaluators](/docs/evaluation/guides/llm-as-a-judge) and [Human Evaluators](/docs/evaluation/guides/human-evaluators) to complement your code-based judgements for more qualitative and subjective criteria.
* Combine your Evaluator with a [Dataset](/docs/explanation/datasets) to run [Evaluations](/docs/guides/evals/run-evaluation-ui) to systematically compare the performance of different versions of your AI application.