> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> Learn how to create and use online evaluators to observe the performance of your models.

> **Paid Feature**
>
> This feature is not available for the Free tier. Please contact us if you wish
> to learn more about our [Enterprise plan](https://humanloop.com/pricing)

## Create an online evaluator

### Prerequisites

* You need to have access to evaluations.
* You also need to have a Prompt – if not, please follow our [Prompt creation](/docs/v4/guides/create-prompt) guide.
* Finally, you need at least a few logs in your project. Use the **Editor** to generate some logs if you don't have any yet.

To set up an online Python evaluator:

### Go to the **Evaluations** page in one of your projects and select the **Evaluators** tab

### Select **+ New Evaluator** and choose **Code Evaluator** in the dialog

![Selecting the type of a new evaluator](/docs/_fern-img/971d836d04768272c0436f3fbb2966af1a514fdf84866e582d57925712bbcf80.webp)

### From the library of presets on the left-hand side, we'll choose **Valid JSON** for this guide. You'll see a pre-populated evaluator with Python code that checks the output of our model is valid JSON grammar.

![The evaluator editor after selecting \*\*Valid JSON\*\* preset](/docs/_fern-img/f8dd14b241a451186542c2adb4c86160c944ab77f5620a39626878aed578cde8.webp)

### In the debug console at the bottom of the dialog, click **Random logs from project**. The console will be populated with five datapoints from your project.

![The debug console (you can resize this area to make it easier to view the logs)](/docs/_fern-img/ee14dd6374b9292000d6dc58f17cac4eacb657724e10c2380f81dfc672977c8f.webp)

### Click the **Run** button at the far right of one of the log rows. After a moment, you'll see the **Result** column populated with a `True` or `False`.

![The \*\*Valid JSON\*\* evaluator returned \`True\` for this particular log, indicating the text output by the model was grammatically correct JSON.](/docs/_fern-img/131bc603efd4c61b3c10b66b9b1d22cb796027dde394129886b0154ba6eeec57.webp)

### Explore the `log` dictionary in the table to help understand what is available on the Python object passed into the evaluator.

### Click **Create** on the left side of the page.

## Activate an evaluator for a project

### On the new \*\*Valid JSON \*\* evaluator in the Evaluations tab, toggle the switch to **on** - the evaluator is now activated for the current project.

![Activating the new evaluator to run automatically on your project.](/docs/_fern-img/c193bbe73b9601d3558a9d7bd44e99d51cdd1f880e44bd1739718e4d9b0d9e12.webp)

### Go to the **Editor**, and generate some fresh logs with your model.

### Over in the **Logs** tab you'll see the new logs. The **Valid JSON** evaluator runs automatically on these new logs, and the results are displayed in the table.

![The \*\*Logs\*\* table includes a column for each activated evaluator in your project. Each activated evaluator runs on any new logs in the project.](/docs/_fern-img/8250d9b1a79b1b16fe6f6006c63b4820b3e18f5d2aa0aa62ee41fbdfb8ab5abe.webp)

## Track the performance of models

### Prerequisites

* A Humanloop project with a reasonable amount of data.
* An Evaluator activated in that project.

To track the performance of different model configs in your project:

### Go to the **Dashboard** tab.

In the table of model configs at the
bottom, choose a subset of the project's model configs.

### Use the graph controls

At the top of the page to select the date range and time granularity
of interest.

### Review the relative performance

For each activated Evaluator shown in the graphs, you can see the relative performance of the model configs you selected.

![](/docs/_fern-img/64338d3886d43a3f41b412537b73eadc29368935660af95832887e545829e903.webp)

> **Available Modules**
>
> The following Python modules are available to be imported in your code evaluators:
>
> * `re`
> * `math`
> * `random`
> * `datetime`
> * `json` (useful for validating JSON grammar as per the example above)
> * `jsonschema` (useful for more fine-grained validation of JSON output - see the in-app example)
> * `sqlglot` (useful for validating SQL query grammar)
> * `requests` (useful to make further LLM calls as part of your evaluation - see the in-app example for a suggestion of how to get started).