> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://humanloop.com/docs/v4/guides/evaluation/monitoring/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/_mcp/server. > Learn how to create and use online evaluators to observe the performance of your models. > **Paid Feature** > > This feature is not available for the Free tier. Please contact us if you wish > to learn more about our [Enterprise plan](https://humanloop.com/pricing) ## Create an online evaluator ### Prerequisites * You need to have access to evaluations. * You also need to have a Prompt – if not, please follow our [Prompt creation](/docs/v4/guides/create-prompt) guide. * Finally, you need at least a few logs in your project. Use the **Editor** to generate some logs if you don't have any yet. To set up an online Python evaluator: ### Go to the **Evaluations** page in one of your projects and select the **Evaluators** tab ### Select **+ New Evaluator** and choose **Code Evaluator** in the dialog ![Selecting the type of a new evaluator](/docs/_fern-img/971d836d04768272c0436f3fbb2966af1a514fdf84866e582d57925712bbcf80.webp) ### From the library of presets on the left-hand side, we'll choose **Valid JSON** for this guide. You'll see a pre-populated evaluator with Python code that checks the output of our model is valid JSON grammar. ![The evaluator editor after selecting \*\*Valid JSON\*\* preset](/docs/_fern-img/f8dd14b241a451186542c2adb4c86160c944ab77f5620a39626878aed578cde8.webp) ### In the debug console at the bottom of the dialog, click **Random logs from project**. The console will be populated with five datapoints from your project. ![The debug console (you can resize this area to make it easier to view the logs)](/docs/_fern-img/ee14dd6374b9292000d6dc58f17cac4eacb657724e10c2380f81dfc672977c8f.webp) ### Click the **Run** button at the far right of one of the log rows. After a moment, you'll see the **Result** column populated with a `True` or `False`. ![The \*\*Valid JSON\*\* evaluator returned \`True\` for this particular log, indicating the text output by the model was grammatically correct JSON.](/docs/_fern-img/131bc603efd4c61b3c10b66b9b1d22cb796027dde394129886b0154ba6eeec57.webp) ### Explore the `log` dictionary in the table to help understand what is available on the Python object passed into the evaluator. ### Click **Create** on the left side of the page. ## Activate an evaluator for a project ### On the new \*\*Valid JSON \*\* evaluator in the Evaluations tab, toggle the switch to **on** - the evaluator is now activated for the current project. ![Activating the new evaluator to run automatically on your project.](/docs/_fern-img/c193bbe73b9601d3558a9d7bd44e99d51cdd1f880e44bd1739718e4d9b0d9e12.webp) ### Go to the **Editor**, and generate some fresh logs with your model. ### Over in the **Logs** tab you'll see the new logs. The **Valid JSON** evaluator runs automatically on these new logs, and the results are displayed in the table. ![The \*\*Logs\*\* table includes a column for each activated evaluator in your project. Each activated evaluator runs on any new logs in the project.](/docs/_fern-img/8250d9b1a79b1b16fe6f6006c63b4820b3e18f5d2aa0aa62ee41fbdfb8ab5abe.webp) ## Track the performance of models ### Prerequisites * A Humanloop project with a reasonable amount of data. * An Evaluator activated in that project. To track the performance of different model configs in your project: ### Go to the **Dashboard** tab. In the table of model configs at the bottom, choose a subset of the project's model configs. ### Use the graph controls At the top of the page to select the date range and time granularity of interest. ### Review the relative performance For each activated Evaluator shown in the graphs, you can see the relative performance of the model configs you selected. ![](/docs/_fern-img/64338d3886d43a3f41b412537b73eadc29368935660af95832887e545829e903.webp) > **Available Modules** > > The following Python modules are available to be imported in your code evaluators: > > * `re` > * `math` > * `random` > * `datetime` > * `json` (useful for validating JSON grammar as per the example above) > * `jsonschema` (useful for more fine-grained validation of JSON output - see the in-app example) > * `sqlglot` (useful for validating SQL query grammar) > * `requests` (useful to make further LLM calls as part of your evaluation - see the in-app example for a suggestion of how to get started). > In this guide, we will demonstrate how to create and use online evaluators to observe the performance of your models.