> This page is for version v5.0 (default).
> For other versions, use one of these documentation indexes:
> - v5.0 (default): https://humanloop.com/docs/v5/llms.txt
> - v4.0: https://humanloop.com/docs/v4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> How to use Humanloop to evaluate multiple different Prompts across a Dataset.

An **Evaluation** on Humanloop leverages a [Dataset](/docs/explanation/datasets), a set of [Evaluators](/docs/explanation/evaluators) and different versions of a [Prompt](/docs/explanation/prompts) to compare.

The Dataset contains datapoints describing the inputs (and optionally the expected results) for a given task. The Evaluators define the criteria for judging the performance of the Prompts when executed using these inputs.

[comment]: <> "The above should be in Explanation section but we haven't move it there"

Prompts, when evaluated, produce [Logs](/docs/explanation/logs). These Logs are then judged by the Evaluators. You can see the summary of Evaluators judgments to systematically compare the performance of the different Prompt versions.

### Prerequisites

* A set of [Prompt](/docs/explanation/prompts) versions you want to compare - see the guide on [creating Prompts](/docs/guides/prompts/create-prompt).
* A [Dataset](/docs/explanation/datasets) containing datapoints for the task - see the guide on [creating a Dataset](/docs/evaluation/guides/upload-dataset-csv).
* At least one [Evaluator](/docs/explanation/evaluators) to judge the performance of the Prompts - see the guides on creating [Code](/docs/evaluation/guides/code-based-evaluator), [AI](/docs/evaluation/guides/llm-as-a-judge) and [Human](/docs/evaluation/guides/human-evaluators) Evaluators.

## Run an Evaluation via UI

For this example, we're going to evaluate the performance of a Support Agent that responds to user queries about Humanloop's product and documentation.
Our goal is to understand which base model between `gpt-4o`, `gpt-4o-mini` and `claude-3-5-sonnet-20241022` is most appropriate for this task.

### Navigate to the Evaluations tab of your Prompt

* Go to the Prompt you want to evaluate and then click on **Evaluations** tab at the top of the page.
* Click the **Evaluate** button top right to create a new Evaluation.
* Click the **+Run** button top right to create a new Evaluation Run.

![Prompt Evaluations Run tab.](/docs/_fern-img/b70033e5c58c4618ec2ce272235c0f78685454ee051d854f16b860b3800519f7.webp)

### Set up an Evaluation Run

* Select a Dataset using **+Dataset**.
* Add the Prompt versions you want to compare using **+Prompt**.
* Add the Evaluators you want to use to judge the performance of the Prompts using **+Evaluator**.

> **Log Caching**
>
> By default the system will re-use Logs if they exist for the chosen Dataset, Prompts and Evaluators. This makes it easy to extend Evaluation Run without paying the cost of re-running your Prompts and Evaluators.
>
> If you want to force the system to re-run the Prompts against the Dataset producing a new batch of Logs, you can click on regenerate button next to the Logs count

* Click **Save**. Humanloop will start generating Logs for the Evaluation.

![In progress Evaluation run](/docs/_fern-img/23d29aaf9ab289f631b5b039605fd7691380bfd3721d9b4cb61917214a297e29.webp)

> **Using your Runtime**
>
> This guide assumes both the Prompt and Evaluator Logs are generated using the
> Humanloop runtime. For certain use cases where more flexibility is required,
> the runtime for producing Logs instead lives in your code - see our guide on
> [Logging](/docs/v5/guides/observability/logging-through-api), which also works with our
> Evaluations feature. We have a guide for how to run Evaluations with Logs
> generated in your code coming soon!

### Review the results

Once the Logs are produced, you can review the performance of the different Prompt versions by navigating to the **Stats** tab.

* The top spider plot provides you with a summary of the average Evaluator performance across all the Prompt versions.
  In our case, `gpt-4o`, although on average slightly slower and more expensive on average, is significantly better when it comes to **User Satisfaction**.

![Evaluation Spider plot](/docs/_fern-img/a0eeb4f39975c753b980c3c6e05374e386c7661e5bfa71b7116861fea99c364a.webp)

* Below the spider plot, you can see the breakdown of performance per Evaluator.

![Evaluation Evaluator stats breakdown](/docs/_fern-img/881f18d26f903cbeaab29d65d31ffa44edd08ac4ab7e612db8baf873c5ec9124.webp)

* To drill into and debug the Logs that were generated, navigate to the **Review** tab at top left of the Run page.
  The Review view allows you to better understand performance and replay logs in our Prompt Editor.

![Drill down to Evaluatoin Logs.](/docs/_fern-img/4641a672adb8af14007e4151145bd0778fa55a530be7eedcff2b373a62b22097.webp)

### Next Steps

* Incorporate this Evaluation process into your Prompt engineering and deployment workflow.
* Setup Evaluations where the runtime for producing Logs lives in your code - see our guide on [Logging](/docs/development/guides/log-to-a-prompt).
* Utilise Evaluations as part of your [CI/CD pipeline](/docs/evaluation/guides/cicd-integration)