> This page is for version v5.0 (default).
> For other versions, use one of these documentation indexes:
> - v5.0 (default): https://humanloop.com/docs/v5/llms.txt
> - v4.0: https://humanloop.com/docs/v4/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> This tutorial will explain how to set up evaluations through the Humanloop UI and use them to iteratively improve your Agent.

> **Note**
>
> This tutorial expands on the Agent we created in the [Quickstart Guide](/docs/v5/quickstart/agent-evals-in-ui).
>
> If you've already completed the Quickstart, you can jump directly to the [Run second Evaluation](/docs/v5/tutorials/evaluate-agent-in-ui#run-second-evaluation) section.
> If you haven't completed the Quickstart yet, please begin with the first step below.

For this tutorial, we're going to evaluate **Outreach Agent**, which is designed to compose personalized outbound messages to potential customers. The Agent uses a [Tool](/docs/explanation/tools) to research information about the lead before writing a message.

We'll show how to assess the quality of the Agent and compare two Agent versions side by side.

#### Account setup

Create a Humanloop Account

If you haven't already, [create an account](https://app.humanloop.com/signup) or [log in](https://app.humanloop.com/login) to Humanloop

Add an OpenAI API Key

If you're the first person in your organization, you'll need to add an API key to a model provider.

1. Go to OpenAI and [grab an API key](https://platform.openai.com/api-keys).
2. In Humanloop [Organization Settings](https://app.humanloop.com/account/api-keys) set up OpenAI as a model provider.

> **Info**
>
> Using the Prompt Editor will use your OpenAI credits in the same way that the
> OpenAI playground does. Keep your API keys for Humanloop and the model
> providers private.

## Clone an Agent

### Clone Outreach Agent from Library

In this quickstart, we will use a pre-configured Agent from the Humanloop Library.

Navigate to the Library by clicking the Library button in the upper-left corner. Select the **Outreach Agent** and click the **Clone to Workspace** button in the upper-right corner.

![](/docs/_fern-img/5eb455c9af00590e41d229c45693483b8e327f19e2bb51c36f39b10e5d244506.webp)

This will create an **Outreach Agent** folder in your workspace. Inside the folder, you'll find:

* The **Outreach Agent**.
* **Hacker News Search** Tool used by the Agent to research potential customers.
* A **Customer Outreach** Dataset and three **Evaluators** that we will use to assess the Agent.

### Try out the Agent

The **Outreach Agent** looks up information about the lead on Hacker News and composes an outbound message to them.

Before we kick off the first evaluation, run the Agent in the Editor to get a feel for how it works:

1. Click on the **Outreach Agent** file.
2. On the left-hand side, you can configure your Agent's Parameters, Instructions, and add Tools.
3. On the right-hand side, you can specify Inputs and run the Agent on demand.
4. Enter *Coca-Cola* as the organization and *Humanloop* as the lead in the Inputs section in the top right-hand side.
5. Click the Run button.

![](/docs/_fern-img/7192532e25b53f5d9b6cae719e4edbd0e15d22ffb1d6af24f03aebbd0b5bdaa5.webp)

### Run Eval

Evaluations are an efficient way to improve your Agent iteratively. You can test versions of the Agent against a [Dataset](https://humanloop.com/docs/evaluation/guides/create-dataset) and see how changing the Agent's configuration impacts the performance.

To test the **Outreach Agent**, navigate to the **Evals** tab and click on the **+ Evaluation** button.

Create a new Run by clicking on the **+ Run** button. Then, follow these steps:

1. Click on the **Dataset** button and select **Customer Outreach** Dateset.
2. Click on the **Agent** button and select the "v1" version.
3. Click on **Evaluators** and add the three Evaluators included in the Output Agent folder: **Friendly Tone**, **Tool Call** and **Message Length**

The first two Evaluators will check if the message is friendly and if the Tool was used.
The **Message Length** Evaluator will show the number of words in the output, providing a baseline value for all further evaluations.

Click **Save**. Humanloop will start generating Logs for the Evaluation.

![](/docs/_fern-img/48c64f972b5458324bd9771da5d7389d7ae290687e389b094cb663b47eb362ed.webp)

### Review results

After the Run is completed, you can review the Logs produced and corresponding judgments in the **Review** tab.

![](/docs/_fern-img/904856454a7f1bbd57496f2c110c1b05e7b2451e685749e3aa2712f7c8a78bed.webp)

The summary of all Evaluators is displayed in the **Stats** tab.

![](/docs/_fern-img/c11e3447a6bb8530c31b4a4569d9384f7c985ccf4784cf9a8df6e6b6153dd4cc.webp)

## Run second Evaluation

### Iterate on the Agent

HackerNews is a limited resource because it lacks background information about potential customers and does not include all recent news articles related to them.

To enhance the search phase, connect Google Search Tool that enables our Agent to traverse through more sophisticated Google search results.

Additionally, add a dedicated Write Personalized Message Prompt that is solely responsible for writing outbound messages. This approach allows for separate iteration on the writing block and the use of different LLM parameters specifically for the writing step.

### Clone new Tools

Navigate back to **Library** and clone the **Google Search** Tool and the **Write Personalized Message** Prompt.

![](/docs/_fern-img/2d8af4ef9924cfcca09fd13957308aa7774cab1903a5a0f901707fbb4226fd03.webp)

### Setup Google Search tool

> **Note**
>
> To use the Google Search tool, you need to obtain an API key from the third-party Serper. Connecting the Agent to a
> third party makes the Agent much more powerful. Serper offers a free API tier that you can use for this tutorial - to
> obtain an API key, sign up at [https://serper.dev/](https://serper.dev/)

Click on Google Search file inside Outreach Agent folder

Add the API key:

1. Open the **Google Search** Tool in your workspace.
2. Click the **Environment Variables** button in the upper-right corner.
3. In the **Name** field, enter: `SERPER_API_KEY`
4. In the **Value** field, paste your API key.

![](/docs/_fern-img/7075fbfa96f77cbf603e927f797375450d76d5022ff4b7f715a51456d1b90b7a.webp)

### Add Tools to the Agent

Click on Outreach Agent File, then click on **+ Tools** button on left bottom corner and choose **Google Search** Tool and **Write Personalized Message** Prompt from the list. Remove HackerNews Search Tool as it's no longer needed.

![](/docs/_fern-img/97bee798a80e4d996d36a28212e9e1d5e22f4e797757a4355fa0dd7bc7b5eb68.webp)

Save the Agent and name it "v2".

### Run another Evaluation

We can now create a new Run with the new Agent version. Click on the **+ Run** button and select the newly created Agent version.

![](/docs/_fern-img/b34c4c35ae910a5ffcbe5f17903758053562d932757e75ba54dea10ef93c23a5.webp)

Navigate to the **Stats** tab to see how the two versions compare to each other.

![](/docs/_fern-img/423f04c0938f0540358aa90ae93b8a3a2ba7d1daf35767ad0ed57f7188d5dd7b.webp)

To see the two versions side by side, click on the **Review** tab.

![](/docs/_fern-img/1d47185e9e32c54f4d17259fd2fe9217db1ae8aaae75c23f3a0667acc72173c6.webp)

The second version of the Agent we evaluated used the Google Search Tool to extract more relevant information. It also searched for background information on each lead, something the initial version of the Agent lacked.

Adding a new Tools to the Agent resulted in a more personalized message and improved outreach.

In this tutorial, you've created an Agent that can help your organization compose personalized messages for your prospects. You've evaluated the initial version, made changes to the Agent, and compared the newly created version with the initial one.

## Next steps

Now that you've successfully run your first Eval, you can explore customizing it for your use case:

* Learn how your subject-matter experts (such as your sales team) can [review and evaluate model outputs](/docs/evaluation/guides/manage-multiple-reviewers) in Humanloop to help improve your AI product.
* Explore how you can set up [Human Evaluators](/docs/evaluation/guides/human-evaluators) to get human feedback on your Agent outputs.