> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> Integrate your existing evaluation process with Humanloop.

LLM and code Evaluators generally live on the Humanloop runtime environment. The advantage of this is that these Evaluators can be used as [monitoring Evaluators](docs/v5/guides/observability/monitoring) and to allow triggering evaluations [directly from the Humanloop UI](/docs/v5/guides/evals/run-evaluation-ui).

Your setup however can be more complex: your Evaluator has library dependencies that are not present in the [runtime environment](/docs/v5/reference/python-environment), your LLM evaluator has multiple reasoning steps, or you prefer managing the logic yourself.

External Evaluators address this: they are registered with Humanloop but their code definition remains in your environment. In order to evaluate a Log, you call the logic yourself and send the judgment to Humanloop.

In this tutorial, we will build a chat agent that answers questions asked by children, and evaluate its performance using an external Evaluator.

## Create the agent

We reuse the chat agent from our [evaluating an agent tutorial](/docs/v5/tutorials/agent-evaluation).

**`main.py`**

```python title="main.py" maxLines=35
from humanloop import Humanloop
from openai import OpenAI
from openai.types.chat.chat_completion_message import ChatCompletionMessage as Message
import wikipedia
import json


openai = OpenAI(api_key="ADD YOUR KEY HERE")
humanloop = Humanloop(api_key="ADD YOUR KEY HERE")


def search_wikipedia(query: str) -> dict:
    """Search Wikipedia to get up-to-date information for a query."""
    try:
        page = wikipedia.page(query)
        return {
            "title": page.title,
            "content": page.content,
            "url": page.url,
        }
    except Exception as _:
        return {
            "title": "",
            "content": "No results found",
            "url": "",
        }


def call_model(messages: list[Message]) -> Message:
    """Calls the model with the given messages"""
    system_message = {
        "role": "system",
        "content": (
            "You are an assistant that helps to answer user questions. "
            "You should leverage wikipedia to answer questions so that "
            "the information is up to date. If the response from "
            "Wikipedia does not seem relevant, rephrase the question "
            "and call the tool again. Then finally respond to the user."
        ),
    }
    response = openai.chat.completions.create(
        model="gpt-4o",
        messages=[system_message] + messages,
        tools=[
            {
                "type": "function",
                "function": {
                    "name": "search_wikipedia",
                    "description": "Search the internet to get up to date answers for a query.",
                    "parameters": {
                        "type": "object",
                        "required": ["query"],
                        "properties": {
                            "query": {"type": "string"},
                        },
                        "additionalProperties": False,
                    },
                },
            }
        ],
    )
    return response.choices[0].message.to_dict(exclude_unset=False)


def call_agent(question: str) -> str:
    """Calls the main agent loop and returns the final result"""
    messages = [{"role": "user", "content": question}]
    # Retry for a relevant response 3 times at most
    for _ in range(3):
        response = call_model(messages)
        messages.append(response)
        if response["tool_calls"]:
            # Call wikipedia to get up-to-date information
            for tool_call in response["tool_calls"]:
                source = search_wikipedia(
                    **json.loads(tool_call["function"]["arguments"])
                )
                messages.append(
                    {
                        "role": "tool",
                        "content": json.dumps(source),
                        "tool_call_id": tool_call["id"],
                    }
                )
        else:
            # Respond to the user
            return response["content"]


if __name__ == "__main__":
    result = call_agent("Where does the sun go at night?")
    print(result)
```

**`main.ts`**

```typescript title="main.ts" maxLines=35
import { HumanloopClient } from "humanloop";
import OpenAI from "openai";
import type { ChatCompletionMessageParam as Message } from "openai/resources";
import wikipedia from "wikipedia";
import fs from "fs";
import readline from "readline";

const openai = new OpenAI({ apiKey: "<OPENAI_API_KEY>" });
const humanloop = new HumanloopClient({ apiKey: "<HUMANLOOP_API_KEY>" });

type WikiResult = {
  title: string;
  content: string;
  url: string;
};

const searchWikipedia = async ({ query }: { query: string }) => {
  const NO_RESULT_FOUND: WikiResult = {
    title: "",
    content: "No results found",
    url: "",
  };

  try {
    const page = await wikipedia.page(query);

    if (page) {
      return {
        title: page?.title || "",
        content: (await page?.content()) || "",
        url: `https://en.wikipedia.org/wiki/${encodeURIComponent(
          page?.title || ""
        )}`,
      } as WikiResult;
    }

    return NO_RESULT_FOUND;
  } catch (error) {
    return NO_RESULT_FOUND;
  }
};

const callModel = async ({ messages }: { messages: Array<Message> }) => {
  const systemMessage: Message = {
    role: "system",
    content:
      "You are an assistant that helps to answer user questions. " +
      "You should leverage wikipedia to answer questions so that " +
      "the information is up to date. If the response from " +
      "Wikipedia does not seem relevant, rephrase the question " +
      "and call the tool again. Then finally respond to the user.",
  };

  const response = await openai.chat.completions.create({
    model: "gpt-4o",
    messages: [systemMessage, ...messages],
    tools: [
      {
        type: "function",
        function: {
          name: "search_wikipedia",
          description:
            "Search the internet to get up to date answers for a query.",
          parameters: {
            type: "object",
            required: ["query"],
            properties: {
              query: { type: "string" },
            },
            additionalProperties: false,
          },
        },
      },
    ],
  });

  return response.choices[0].message;
}

const callAgent = async ({ question }: { question: string }) => {
  const messages: Message[] = [{ role: "user", content: question }];

  for (let _ = 0; _ < 3; _++) {
    const response = await callModel({ messages });
    messages.push(response);

    if (response.tool_calls) {
      for (const toolCall of response.tool_calls) {
        const args = JSON.parse(toolCall.function.arguments);
        const source = await searchWikipedia(args);

        messages.push({
          role: "tool",
          content: JSON.stringify(source),
          tool_call_id: toolCall.id,
        });
      }
    } else {
      return response.content || "";
    }
  }

  return "Could not get a relevant response after multiple attempts.";
}

const main = async () => {
  const result = await callAgent({
    question: "Where does the sun go at night?",
  });
  console.log(result);
}

main();
```

Run the agent and check that it works as expected:

#### Python

```bash
python main.py
```

```plaintext
Okay! Imagine the Earth is like a big ball, and we live on it.
The sun doesn't really “go” anywhere—it stays in the same spot,
shining all the time. But our Earth is spinning like a top!
```

#### TypeScript

```bash
npx tsx main.ts
```

```plaintext
Okay! Imagine the Earth is like a big ball, and we live on it.
The sun doesn't really “go” anywhere—it stays in the same spot,
shining all the time. But our Earth is spinning like a top!
```

## Evaluate the agent

#### How Evaluators work

Evaluators are callables that take the Log's dictionary representation as input and return a judgment. The Evaluator's judgment should respect the `return_type` present in the Evaluator's [specification](/docs/v5/api/evaluators/upsert#request.body.spec).

The Evaluator can take an additional `target` argument to compare the Log against. The target is provided in an Evaluation context by the validation [Dataset](/docs/v5/explanation/datasets).

For more details, check out our [Evaluator explanation](/docs/v5/explanation/evaluators).

### Define external Evaluator

The Evaluator takes a `log` argument, which represents the Log created by calling `call_agent`.

Let's add a simple Evaluator that checks if the agent's answers are too long. Add this in the `agent.py` file:

```python
if __name__ == "__main__":
    def easy_to_understand(log):
        return len(log["output"]) < 100
```

### Add dataset

Create a file called `dataset.jsonl` and add the following:

**`dataset.jsonl`**

```jsonl title="dataset.jsonl" maxLines=5
{"inputs": {"question": "Why is the sky blue?"}}
{"inputs": {"question": "Where does the sun go at night?"}}
{"inputs": {"question": "Why do birds fly?"}}
{"inputs": {"question": "What makes rainbows?"}}
{"inputs": {"question": "Why do we have to sleep?"}}
{"inputs": {"question": "How do fish breathe underwater?"}}
{"inputs": {"question": "Why do plants need water?"}}
{"inputs": {"question": "How does the moon stay in the sky?"}}
{"inputs": {"question": "What are stars made of?"}}
{"inputs": {"question": "Why do we have seasons?"}}
{"inputs": {"question": "How does the TV work?"}}
{"inputs": {"question": "Why do dogs wag their tails?"}}
{"inputs": {"question": "What makes cars go?"}}
{"inputs": {"question": "Why do we need to brush our teeth?"}}
{"inputs": {"question": "What do ants eat?"}}
{"inputs": {"question": "Why does the wind blow?"}}
{"inputs": {"question": "How do airplanes stay in the air?"}}
{"inputs": {"question": "Why does the ocean look so big?"}}
{"inputs": {"question": "What makes the grass green?"}}
{"inputs": {"question": "Why do we have to eat vegetables?"}}
{"inputs": {"question": "How do butterflies fly?"}}
{"inputs": {"question": "Why do some animals live in the zoo?"}}
{"inputs": {"question": "How do magnets stick to the fridge?"}}
{"inputs": {"question": "What makes fire hot?"}}
{"inputs": {"question": "Why do leaves change color?"}}
{"inputs": {"question": "What happens when we flush the toilet?"}}
{"inputs": {"question": "Why do we have belly buttons?"}}
{"inputs": {"question": "What makes the clouds move?"}}
{"inputs": {"question": "Why do we have eyebrows?"}}
{"inputs": {"question": "How do seeds turn into plants?"}}
{"inputs": {"question": "Why does the moon change shape?"}}
{"inputs": {"question": "Why do bees make honey?"}}
{"inputs": {"question": "What makes ice melt?"}}
{"inputs": {"question": "Why do we sneeze?"}}
{"inputs": {"question": "How do trains stay on the tracks?"}}
{"inputs": {"question": "Why do stars twinkle?"}}
{"inputs": {"question": "Why can't we see air?"}}
{"inputs": {"question": "What makes the Earth spin?"}}
{"inputs": {"question": "Why do frogs jump?"}}
{"inputs": {"question": "Why do cats purr?"}}
{"inputs": {"question": "How do phones let us talk to people far away?"}}
{"inputs": {"question": "Why does the moon follow us?"}}
{"inputs": {"question": "What makes lightning?"}}
{"inputs": {"question": "Why does it snow?"}}
{"inputs": {"question": "Why do we have shadows?"}}
{"inputs": {"question": "Why do boats float?"}}
{"inputs": {"question": "What makes our heart beat?"}}
{"inputs": {"question": "Why do some animals sleep all winter?"}}
{"inputs": {"question": "Why do we have to wear shoes?"}}
{"inputs": {"question": "What makes music?"}}
```

### Add Evaluation

Instantiate an Evaluation using the client's \[]`evaluations.run`]\(/docs/v5/sdk/run-evaluation) utility. `easy_to_understand` is an external Evaluator, so we provide its definition via the `callable` argument. At runtime, `evaluations.run` will call the function and submit the judgment to Humanloop.

**`agent.py`**

```python title="agent.py" maxLines=100 highlight={5-28}
if __name__ == "__main__":
    def easy_to_understand(log):
        return len(log["output"]) < 100

    # Read the evaluation dataset
    with open("dataset.jsonl", "r") as fp:
        dataset = [json.loads(line) for line in fp]

    humanloop.evaluations.run(
        name="QA Agent Answer Comprehensiveness",
        file={
            "path": "QA Agent/Agent",
            "callable": call_agent,
        },
        evaluators=[
            {
                "path": "QA Agent/Comprehension",
                "callable": easy_to_understand,
                "args_type": "target_free",
                "return_type": "boolean",
            }
        ],
        dataset={
            "path": "QA Agent/Children Questions",
            "datapoints": dataset,
        },
        workers=8,
    )
```

### Run the evaluation

#### Python

**`Terminal`**

```bash title="Terminal"
python main.py
```

**`Terminal`**

```bash title="Terminal" maxLines=50
Navigate to your Evaluation:
https://app.humanloop.com/project/fl_9CCIoTzySPfUFeIxfYE6g/evaluations/evr_67tEc2DiR83fy9iTaqyPA/stats

Flow Version ID: flv_9ECTrfeZYno2OIj9KAqlz
Run ID: rn_67tEcDYV6mqUS86hD8vrP

Running 'Agent' over the Dataset 'Children Questions' using 8 workers 
[##############--------------------------] 15/50 (30.00%) | ETA: 14

...

📊 Evaluation Results for QA Agent/Agent
+------------------------+---------------------+
|                        |        Latest       |
+------------------------+---------------------+
|                 Run ID |        67tEc        |
+------------------------+---------------------+
|             Version ID |        9ECTr        |
+------------------------+---------------------+
|                  Added | 2024-11-19 21:49:02 |
+------------------------+---------------------+
|             Evaluators |                     |
+------------------------+---------------------+
| QA Agent/Comprehension |         3.24        |
+------------------------+---------------------+
```

#### TypeScript

**`Terminal`**

```bash title="Terminal" maxLines=50
npx tsx main.ts
```

**`Terminal`**

```bash title="Terminal" maxLines=50
Navigate to your Evaluation:
https://app.humanloop.com/project/fl_9CCIoTzySPfUFeIxfYE6g/evaluations/evr_67tEc2DiR83fy9iTaqyPA/stats

Flow Version ID: flv_9ECTrfeZYno2OIj9KAqlz
Run ID: rn_67tEcDYV6mqUS86hD8vrP

Running 'Agent' over the Dataset 'Children Questions' using 8 workers 
[##############--------------------------] 15/50 (30.00%) | ETA: 14

...

📊 Evaluation Results for QA Agent/Agent
+------------------------+---------------------+
|                        |        Latest       |
+------------------------+---------------------+
|                 Run ID |        67tEc        |
+------------------------+---------------------+
|             Version ID |        9ECTr        |
+------------------------+---------------------+
|                  Added | 2024-11-19 21:49:02 |
+------------------------+---------------------+
|             Evaluators |                     |
+------------------------+---------------------+
| QA Agent/Comprehension |         3.24        |
+------------------------+---------------------+
```

Click on the link to see the results when the Evaluation is complete.

## Add detailed logging

> **Info**
>
> If you use a programming language not supported by the SDK, or want more control, see our guide on [logging through the API](/docs/v5/guides/observability/logging-through-api) for an alternative to decorators.

Up to this point, we have treated the agent as a black box, reasoning about its behavior by looking at the inputs and outputs.

Let's use Humanloop logging to observe the step-by-step actions taken by the agent.

#### Python

Modify `main.py`:

**`main.py`**

```python title="main.py" maxLines=100 highlight={1,5,10,15}
@humanloop.tool(path="QA Agent/Search Wikipedia")
def search_wikipedia(query: str) -> dict:
    ...

@humanloop.prompt(path="QA Agent/Prompt")
def call_model(messages: list[Message]) -> Message:
    ...

@humanloop.flow(path="QA Agent/Agent")
def call_agent(question: str) -> str:
    ...
```

#### TypeScript

> **Note**
>
> To auto-instrument calls to OpenAI, pass the module in the Humanloop constructor:
>
> ```typescript
> const humanloop = new HumanloopClient({
>     apiKey: "<HUMANLOOP_API_KEY>",
>     instrumentProviders: {
>         // Pass the OpenAI module, not the initialized client
>         OpenAI
>     }
> });
> ```

Modify `main.ts`:

**`main.ts`**

```typescript title="main.ts" maxLines=100 highlight={20-21,26-27,32-33}
const searchWikipedia = humanloop.tool({
    path: "QA Agent/Search Wikipedia",
    version: {
        function: {
            name: "Search Wikipedia",
            description: "Search Wikipedia for the best article to answer a question",
            strict: true,
            parameters: {
                type: "object",
                properties: {
                    query: {
                        type: "string",
                        description: "The question to search Wikipedia for",
                    },
                },
                required: ["query"],
            },
        },
    },
    // Wraps the initial function body
    callable: async ({ query }) => { ... },
});

const callModel = humanloop.prompt({
    path: "QA Agent/Prompt",
    // Wraps the initial function body
    callable: async ({ messages }) => { ... },
});

const callAgent = humanloop.flow({
    path: "QA Agent/Agent",
    // Wraps the initial function body
    callable: async ({ question }) => { ... },
});
```

Evaluate the agent again. When it's done, head to your workspace and click the **Agent** [Flow](/docs/v5/guides/explanations/flows) on the left. Select the Logs tab from the top of the page.

![](/docs/_fern-img/5b84bec69140c251a42867c8a2da5699a7f62eb82eebe25a3c613ccc7ca4c703.webp)

The decorators divide the code in logical components, allowing you to observe the steps taken to answer a question. Every step taken by the agent creates a Log.

## Next steps

You've learned how to integrate your existing evaluation process with Humanloop.

Learn more about Humanloop's features in these guides:

* Learn how to use Evaluations to improve on your feature's performance in our [tutorial on evaluating a chat agent](/docs/v5/tutorials/agent-evaluation).

* Evals work hand in hand with logging. Learn how to log detailed information about your AI project in [logging setup guide](/docs/v5/quickstart/set-up-logging).