> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> Learn how to use Humanloop to systematically test and improve your LLM agent systems.

> **Note**
>
> Humanloop offers first-class support for agents on its runtime. This tutorial provides guidance on how to instrument existing code-based agentic systems with Humanloop.
>
> Check out our [Evaluating Agents in UI](/docs/v5/tutorials/evaluate-agent-in-ui) tutorial for a more streamlined experience, including tool calling on the Humanloop runtime, and autonomous agents.

Working with LLMs is daunting: you are dealing with a black box that outputs unpredictable results.

Humanloop provides tools to make your development process systematic, bringing it closer to traditional software testing and quality assurance.

In this tutorial, we’ll build an agentic question-and-answer system and use Humanloop to iterate on its performance. The agent will provide simple, child-friendly answers using Wikipedia as its source of factual information.

## Prerequisites

#### Account setup

Create a Humanloop Account

If you haven't already, [create an account](https://app.humanloop.com/signup) or [log in](https://app.humanloop.com/login) to Humanloop

Add an OpenAI API Key

If you're the first person in your organization, you'll need to add an API key to a model provider.

1. Go to OpenAI and [grab an API key](https://platform.openai.com/api-keys).
2. In Humanloop [Organization Settings](https://app.humanloop.com/account/api-keys) set up OpenAI as a model provider.

> **Info**
>
> Using the Prompt Editor will use your OpenAI credits in the same way that the
> OpenAI playground does. Keep your API keys for Humanloop and the model
> providers private.

#### Install dependencies

#### Python

Install the project's dependencies:

```python
pip install humanloop openai wikipedia
```

Humanloop SDK requires Python 3.9 or higher. Optionally, create a virtual environment to keep dependencies tidy.

#### TypeScript

Install the project's dependencies:

```typescript
npm install humanloop openai wikipedia
```

## Initial agent code

Let's build the first iteration of the agent. We'll use OpenAI's function calling to connect the agent to the Wikipedia API. The agent will also be allowed to refine its answer through multiple iterations if the initial tool call doesn't yield relevant information.

**`main.py`**

```python title="main.py" maxLines=35
from humanloop import Humanloop
from openai import OpenAI
from openai.types.chat.chat_completion_message import ChatCompletionMessage as Message
import wikipedia
import json


openai = OpenAI(api_key="ADD YOUR KEY HERE")
humanloop = Humanloop(api_key="ADD YOUR KEY HERE")


def search_wikipedia(query: str) -> dict:
    """Search Wikipedia to get up-to-date information for a query."""
    try:
        page = wikipedia.page(query)
        return {
            "title": page.title,
            "content": page.content,
            "url": page.url,
        }
    except Exception as _:
        return {
            "title": "",
            "content": "No results found",
            "url": "",
        }


def call_model(messages: list[Message]) -> Message:
    """Calls the model with the given messages"""
    system_message = {
        "role": "system",
        "content": (
            "You are an assistant that helps to answer user questions. "
            "You should leverage wikipedia to answer questions so that "
            "the information is up to date. If the response from "
            "Wikipedia does not seem relevant, rephrase the question "
            "and call the tool again. Then finally respond to the user."
        ),
    }
    response = openai.chat.completions.create(
        model="gpt-4o",
        messages=[system_message] + messages,
        tools=[
            {
                "type": "function",
                "function": {
                    "name": "search_wikipedia",
                    "description": "Search the internet to get up to date answers for a query.",
                    "parameters": {
                        "type": "object",
                        "required": ["query"],
                        "properties": {
                            "query": {"type": "string"},
                        },
                        "additionalProperties": False,
                    },
                },
            }
        ],
    )
    return response.choices[0].message.to_dict(exclude_unset=False)


def call_agent(question: str) -> str:
    """Calls the main agent loop and returns the final result"""
    messages = [{"role": "user", "content": question}]
    # Retry for a relevant response 3 times at most
    for _ in range(3):
        response = call_model(messages)
        messages.append(response)
        if response["tool_calls"]:
            # Call wikipedia to get up-to-date information
            for tool_call in response["tool_calls"]:
                source = search_wikipedia(
                    **json.loads(tool_call["function"]["arguments"])
                )
                messages.append(
                    {
                        "role": "tool",
                        "content": json.dumps(source),
                        "tool_call_id": tool_call["id"],
                    }
                )
        else:
            # Respond to the user
            return response["content"]


if __name__ == "__main__":
    result = call_agent("Where does the sun go at night?")
    print(result)
```

**`main.ts`**

```typescript title="main.ts" maxLines=35
import { HumanloopClient } from "humanloop";
import OpenAI from "openai";
import type { ChatCompletionMessageParam as Message } from "openai/resources";
import wikipedia from "wikipedia";
import fs from "fs";
import readline from "readline";

const openai = new OpenAI({ apiKey: "<OPENAI_API_KEY>" });
const humanloop = new HumanloopClient({ apiKey: "<HUMANLOOP_API_KEY>" });

type WikiResult = {
  title: string;
  content: string;
  url: string;
};

const searchWikipedia = async ({ query }: { query: string }) => {
  const NO_RESULT_FOUND: WikiResult = {
    title: "",
    content: "No results found",
    url: "",
  };

  try {
    const page = await wikipedia.page(query);

    if (page) {
      return {
        title: page?.title || "",
        content: (await page?.content()) || "",
        url: `https://en.wikipedia.org/wiki/${encodeURIComponent(
          page?.title || ""
        )}`,
      } as WikiResult;
    }

    return NO_RESULT_FOUND;
  } catch (error) {
    return NO_RESULT_FOUND;
  }
};

const callModel = async ({ messages }: { messages: Array<Message> }) => {
  const systemMessage: Message = {
    role: "system",
    content:
      "You are an assistant that helps to answer user questions. " +
      "You should leverage wikipedia to answer questions so that " +
      "the information is up to date. If the response from " +
      "Wikipedia does not seem relevant, rephrase the question " +
      "and call the tool again. Then finally respond to the user.",
  };

  const response = await openai.chat.completions.create({
    model: "gpt-4o",
    messages: [systemMessage, ...messages],
    tools: [
      {
        type: "function",
        function: {
          name: "search_wikipedia",
          description:
            "Search the internet to get up to date answers for a query.",
          parameters: {
            type: "object",
            required: ["query"],
            properties: {
              query: { type: "string" },
            },
            additionalProperties: false,
          },
        },
      },
    ],
  });

  return response.choices[0].message;
}

const callAgent = async ({ question }: { question: string }) => {
  const messages: Message[] = [{ role: "user", content: question }];

  for (let _ = 0; _ < 3; _++) {
    const response = await callModel({ messages });
    messages.push(response);

    if (response.tool_calls) {
      for (const toolCall of response.tool_calls) {
        const args = JSON.parse(toolCall.function.arguments);
        const source = await searchWikipedia(args);

        messages.push({
          role: "tool",
          content: JSON.stringify(source),
          tool_call_id: toolCall.id,
        });
      }
    } else {
      return response.content || "";
    }
  }

  return "Could not get a relevant response after multiple attempts.";
}

const main = async () => {
  const result = await callAgent({
    question: "Where does the sun go at night?",
  });
  console.log(result);
}

main();
```

Run the agent and check that it works as expected:

#### Python

```bash
python main.py
```

```plaintext
Okay! Imagine the Earth is like a big ball, and we live on it.
The sun doesn't really “go” anywhere—it stays in the same spot,
shining all the time. But our Earth is spinning like a top!
```

#### TypeScript

```bash
npx tsx main.ts
```

```plaintext
Okay! Imagine the Earth is like a big ball, and we live on it.
The sun doesn't really “go” anywhere—it stays in the same spot,
shining all the time. But our Earth is spinning like a top!
```

## Add Humanloop to the agent

Humanloop offers first-class support for agentic systems, plus the ability to effortlessly switch between providers.

## Evaluate the agent

#### How Evaluators work

Evaluators are callables that take the Log's dictionary representation as input and return a judgment. The Evaluator's judgment should respect the `return_type` present in the Evaluator's [specification](/docs/v5/api/evaluators/upsert#request.body.spec).

The Evaluator can take an additional `target` argument to compare the Log against. The target is provided in an Evaluation context by the validation [Dataset](/docs/v5/explanation/datasets).

For more details, check out our [Evaluator explanation](/docs/v5/explanation/evaluators).

Let's check if the agent respects the requirement of providing easy-to-understand answers.

We will create an [Evaluation](/docs/v5/guides/evals/run-evaluation-ui) to benchmark the performance of the agent. An Evaluation requires a [Dataset](/docs/v5/guides/explanations/datasets) and at least one [Evaluator](/docs/v5/guides/explanations/evaluators).

### Create LLM judge

We will use an LLM judge to automatically evaluate the agent's responses.

We will define the Evaluator in code, but you can also [manage Evaluators in the UI](/docs/v5/guides/evals/llm-as-a-judge).

Replace your `main` function with the following:

**`main.py`**

```python title="main.py" maxLines=50
if __name__ == "__main__":
    humanloop.evaluators.upsert(
        path="QA Agent/Comprehension",
        spec={
            "arguments_type": "target_free",
            "return_type": "number",
            "evaluator_type": "llm",
            "prompt": {
                "model": "gpt-4o",
                "endpoint": "complete",
                "template": (
                    "You must decide if an explanation is simple "
                    "enough to be understood by a 5-year old. "
                    "A better explanation is shorter and uses less jargon. "
                    "Rate the answer from 1 to 10, where 10 is the best.\n"
                    "\n<Question>\n{{log.inputs.question}}\n</Question>\n\n"
                    "\n<Answer>\n{{log.output}}</Answer>\n\n"
                    "First provide your rationale, then on a newline, "
                    "output your judgment."
                ),
                "provider": "openai",
                "temperature": 0,
            },
        },
    )
```

**`main.ts`**

```typescript title="main.ts" maxLines=50
const main = async () => {
  await humanloop.evaluators.upsert({
    path: "QA Agent/Comprehension",
    spec: {
      argumentsType: "target_free",
      returnType: "number",
      evaluatorType: "llm",
      prompt: {
        model: "gpt-4o",
        endpoint: "complete",
        template:
          "You must decide if an explanation is simple " +
          "enough to be understood by a 5-year old. " +
          "A better explanation is shorter and uses less jargon. " +
          "Rate the answer from 1 to 10, where 10 is the best.\n" +
          "\n<Question>\n{{log.inputs.question}}\n</Question>\n\n" +
          "\n<Answer>\n{{log.output}}</Answer>\n\n" +
          "First provide your rationale, then on a newline, " +
          "output your judgment.",
        provider: "openai",
        temperature: 0,
      },
    },
  });
}
```

### Add Dataset

Create a file called `dataset.jsonl` and add the following:

**`dataset.jsonl`**

```jsonl title="dataset.jsonl" maxLines=5
{"inputs": {"question": "Why is the sky blue?"}}
{"inputs": {"question": "Where does the sun go at night?"}}
{"inputs": {"question": "Why do birds fly?"}}
{"inputs": {"question": "What makes rainbows?"}}
{"inputs": {"question": "Why do we have to sleep?"}}
{"inputs": {"question": "How do fish breathe underwater?"}}
{"inputs": {"question": "Why do plants need water?"}}
{"inputs": {"question": "How does the moon stay in the sky?"}}
{"inputs": {"question": "What are stars made of?"}}
{"inputs": {"question": "Why do we have seasons?"}}
{"inputs": {"question": "How does the TV work?"}}
{"inputs": {"question": "Why do dogs wag their tails?"}}
{"inputs": {"question": "What makes cars go?"}}
{"inputs": {"question": "Why do we need to brush our teeth?"}}
{"inputs": {"question": "What do ants eat?"}}
{"inputs": {"question": "Why does the wind blow?"}}
{"inputs": {"question": "How do airplanes stay in the air?"}}
{"inputs": {"question": "Why does the ocean look so big?"}}
{"inputs": {"question": "What makes the grass green?"}}
{"inputs": {"question": "Why do we have to eat vegetables?"}}
{"inputs": {"question": "How do butterflies fly?"}}
{"inputs": {"question": "Why do some animals live in the zoo?"}}
{"inputs": {"question": "How do magnets stick to the fridge?"}}
{"inputs": {"question": "What makes fire hot?"}}
{"inputs": {"question": "Why do leaves change color?"}}
{"inputs": {"question": "What happens when we flush the toilet?"}}
{"inputs": {"question": "Why do we have belly buttons?"}}
{"inputs": {"question": "What makes the clouds move?"}}
{"inputs": {"question": "Why do we have eyebrows?"}}
{"inputs": {"question": "How do seeds turn into plants?"}}
{"inputs": {"question": "Why does the moon change shape?"}}
{"inputs": {"question": "Why do bees make honey?"}}
{"inputs": {"question": "What makes ice melt?"}}
{"inputs": {"question": "Why do we sneeze?"}}
{"inputs": {"question": "How do trains stay on the tracks?"}}
{"inputs": {"question": "Why do stars twinkle?"}}
{"inputs": {"question": "Why can't we see air?"}}
{"inputs": {"question": "What makes the Earth spin?"}}
{"inputs": {"question": "Why do frogs jump?"}}
{"inputs": {"question": "Why do cats purr?"}}
{"inputs": {"question": "How do phones let us talk to people far away?"}}
{"inputs": {"question": "Why does the moon follow us?"}}
{"inputs": {"question": "What makes lightning?"}}
{"inputs": {"question": "Why does it snow?"}}
{"inputs": {"question": "Why do we have shadows?"}}
{"inputs": {"question": "Why do boats float?"}}
{"inputs": {"question": "What makes our heart beat?"}}
{"inputs": {"question": "Why do some animals sleep all winter?"}}
{"inputs": {"question": "Why do we have to wear shoes?"}}
{"inputs": {"question": "What makes music?"}}
```

### Run an Evaluation

Add this to your `main` function:

**`main.py`**

```python title="main.py" maxLines=100 highlight={4-24}
if __name__ == "__main__":
    # ...

    # Read the evaluation dataset
    with open("dataset.jsonl", "r") as fp:
        dataset = [json.loads(line) for line in fp]

    humanloop.evaluations.run(
        name="QA Agent Answer Check",
        file={
            "path": "QA Agent/Agent",
            "callable": call_agent,
        },
        evaluators=[{"path": "QA Agent/Comprehension"}],
        dataset={
            "path": "QA Agent/Dataset",
            "datapoints": dataset,
        },
        workers=8,
    )
```

**`main.ts`**

```typescript title="main.ts" maxLines=100 highlight={4-29}
const main = async () => {
  // ...
  
  // Read the evaluation dataset
  const dataset: any[] = [];
  const fileStream = fs.createReadStream("dataset.jsonl");
  const rl = readline.createInterface({
    input: fileStream,
    crlfDelay: Infinity,
  });

  for await (const line of rl) {
    dataset.push(JSON.parse(line));
  }

  // Run the evaluation
  await humanloop.evaluations.run({
    name: "QA Agent Answer Check",
    file: {
      path: "QA Agent/Agent",
      callable: callAgent,
    },
    evaluators: [{ path: "QA Agent/Comprehension" }],
    dataset: {
      path: "QA Agent/Dataset",
      datapoints: dataset,
    },
    concurrency: 8,
  });
}
```

Run your file and let the Evaluation finish:

#### Python

**`Terminal`**

```bash title="Terminal"
python main.py
```

**`Terminal`**

```bash title="Terminal" maxLines=50
Navigate to your Evaluation:
https://app.humanloop.com/project/fl_9CCIoTzySPfUFeIxfYE6g/evaluations/evr_67tEc2DiR83fy9iTaqyPA/stats

Flow Version ID: flv_9ECTrfeZYno2OIj9KAqlz
Run ID: rn_67tEcDYV6mqUS86hD8vrP

Running 'Agent' over the Dataset 'Children Questions' using 8 workers 
[##############--------------------------] 15/50 (30.00%) | ETA: 14

...

📊 Evaluation Results for QA Agent/Agent
+------------------------+---------------------+
|                        |        Latest       |
+------------------------+---------------------+
|                 Run ID |        67tEc        |
+------------------------+---------------------+
|             Version ID |        9ECTr        |
+------------------------+---------------------+
|                  Added | 2024-11-19 21:49:02 |
+------------------------+---------------------+
|             Evaluators |                     |
+------------------------+---------------------+
| QA Agent/Comprehension |         3.24        |
+------------------------+---------------------+
```

#### TypeScript

**`Terminal`**

```bash title="Terminal" maxLines=50
npx tsx main.ts
```

**`Terminal`**

```bash title="Terminal" maxLines=50
Navigate to your Evaluation:
https://app.humanloop.com/project/fl_9CCIoTzySPfUFeIxfYE6g/evaluations/evr_67tEc2DiR83fy9iTaqyPA/stats

Flow Version ID: flv_9ECTrfeZYno2OIj9KAqlz
Run ID: rn_67tEcDYV6mqUS86hD8vrP

Running 'Agent' over the Dataset 'Children Questions' using 8 workers 
[##############--------------------------] 15/50 (30.00%) | ETA: 14

...

📊 Evaluation Results for QA Agent/Agent
+------------------------+---------------------+
|                        |        Latest       |
+------------------------+---------------------+
|                 Run ID |        67tEc        |
+------------------------+---------------------+
|             Version ID |        9ECTr        |
+------------------------+---------------------+
|                  Added | 2024-11-19 21:49:02 |
+------------------------+---------------------+
|             Evaluators |                     |
+------------------------+---------------------+
| QA Agent/Comprehension |         3.24        |
+------------------------+---------------------+
```

## Iterate and evaluate again

The score of the initial setup is quite low. Click the Evaluation link from the terminal and switch to the Logs view. You will see that the model tends to provide elaborate answers.

![](/docs/_fern-img/5151b0f22899a2ae3d255ec7a89eb0a2892af12ad9748113b0177d1ee4a24d9a.webp)

#### Python

Let's modify the LLM prompt inside `call_model`:

**`main.py`**

```python title="main.py" maxLines=100 highlight={11-12}

def call_model(messages: list[Message]) -> Message:
    """Calls the model with the given messages"""
    system_message = {
        "role": "system",
        "content": (
          "You are an assistant that help to answer user questions. "
          "You should leverage wikipedia to answer questions so that "
          "the information is up to date. If the response from Wikipedia "
          "does not seem relevant, rephrase the question and call the "
          "tool again. Then finally respond to the user. "
          "Formulate the response so that it is easy to understand "
          "for a 5 year old."
        )
    }
    response = openai.chat.completions.create(
        model="gpt-4o",
        messages=[system_message] + messages,
        tools=[
            {
                "type": "function",
                "function": {
                    "name": "search_wikipedia",
                    "description": "Search the internet to get up to date answers for a query.",
                    "parameters": {
                        "type": "object",
                        "required": ["query"],
                        "properties": {
                            "query": {"type": "string"},
                        },
                        "additionalProperties": False,
                    },
                }
            }
        ],
    )
    return response.choices[0].message.to_dict(exclude_unset=False)
```

Run the agent again and let the Evaluation finish:

```bash
python main.py
```

#### TypeScript

Let's modify the LLM prompt inside `callModel`:

**`main.ts`**

```typescript title="main.ts" maxLines=100 highlight={10-11}
const callModel = async ({ messages }: { messages: Array<Message> }) => {
  const systemMessage: Message = {
    role: "system",
    content:
      "You are an assistant that helps to answer user questions. " +
      "You should leverage wikipedia to answer questions so that " +
      "the information is up to date. If the response from " +
      "Wikipedia does not seem relevant, rephrase the question " +
      "and call the tool again. Then finally respond to the user. "+
      "Formulate the response so that it is easy to understand " + 
      "for a 5 year old.",
  };

  const response = await openai.chat.completions.create({
    model: "gpt-4o",
    messages: [systemMessage, ...messages],
    tools: [
      {
        type: "function",
        function: {
          name: "search_wikipedia",
          description:
            "Search the internet to get up to date answers for a query.",
          parameters: {
            type: "object",
            required: ["query"],
            properties: {
              query: { type: "string" },
            },
            additionalProperties: false,
          },
        },
      },
    ],
  });

  return response.choices[0].message;
}
```

Run the agent again and let the Evaluation finish:

```typescript
npx tsx main.ts
```

**`Terminal`**

```bash title="Terminal" maxLines=50
Flow Version ID: flv_9ECTrfeZYno2OIj9KAqlz
Run ID: rn_WnIwPSI7JFKEtwTS0l3mj

Navigate to your Evaluation:
https://app.humanloop.com/project/fl_9CCIoTzySPfUFeIxfYE6g/evaluations/rn_WnIwPSI7JFKEtwTS0l3mj/stats

Running 'Agent' over the Dataset 'Children Questions' using 8 workers 
[######################------------------] 34/50 (68.00%) | ETA: 14

...

+------------------------+---------------------+---------------------+
|                        |       Control       |        Latest       |
+------------------------+---------------------+---------------------+
|                 Run ID |        67tEc        |        WnIwP        |
+------------------------+---------------------+---------------------+
|             Version ID |        9ECTr        |        9ECTr        |
+------------------------+---------------------+---------------------+
|                  Added | 2024-11-19 22:05:17 | 2024-11-19 22:24:13 |
+------------------------+---------------------+---------------------+
|             Evaluators |                     |                     |
+------------------------+---------------------+---------------------+
| QA Agent/Comprehension |         3.24        |         8.04        |
+------------------------+---------------------+---------------------+

Change of [4.80] for Evaluator QA Agent/Comprehension
```

Click the Evaluation link again. The agent's performance has improved significantly.

![](/docs/_fern-img/88de21feccf4678040d38351c9c03acd72776d6c0a9742236026da74383c46c3.webp)

## Add detailed logging

> **Info**
>
> If you use a programming language not supported by the SDK, or want more control, see our guide on [logging through the API](/docs/v5/guides/observability/logging-through-api) for an alternative to decorators.

Up to this point, we have treated the agent as a black box, reasoning about its behavior by looking at the inputs and outputs.

Let's use Humanloop logging to observe the step-by-step actions taken by the agent.

#### Python

Modify `main.py`:

**`main.py`**

```python title="main.py" maxLines=100 highlight={1,5,10,15}
@humanloop.tool(path="QA Agent/Search Wikipedia")
def search_wikipedia(query: str) -> dict:
    ...

@humanloop.prompt(path="QA Agent/Prompt")
def call_model(messages: list[Message]) -> Message:
    ...

@humanloop.flow(path="QA Agent/Agent")
def call_agent(question: str) -> str:
    ...
```

#### TypeScript

> **Note**
>
> To auto-instrument calls to OpenAI, pass the module in the Humanloop constructor:
>
> ```typescript
> const humanloop = new HumanloopClient({
>     apiKey: "<HUMANLOOP_API_KEY>",
>     instrumentProviders: {
>         // Pass the OpenAI module, not the initialized client
>         OpenAI
>     }
> });
> ```

Modify `main.ts`:

**`main.ts`**

```typescript title="main.ts" maxLines=100 highlight={20-21,26-27,32-33}
const searchWikipedia = humanloop.tool({
    path: "QA Agent/Search Wikipedia",
    version: {
        function: {
            name: "Search Wikipedia",
            description: "Search Wikipedia for the best article to answer a question",
            strict: true,
            parameters: {
                type: "object",
                properties: {
                    query: {
                        type: "string",
                        description: "The question to search Wikipedia for",
                    },
                },
                required: ["query"],
            },
        },
    },
    // Wraps the initial function body
    callable: async ({ query }) => { ... },
});

const callModel = humanloop.prompt({
    path: "QA Agent/Prompt",
    // Wraps the initial function body
    callable: async ({ messages }) => { ... },
});

const callAgent = humanloop.flow({
    path: "QA Agent/Agent",
    // Wraps the initial function body
    callable: async ({ question }) => { ... },
});
```

Evaluate the agent again. When it's done, head to your workspace and click the **Agent** [Flow](/docs/v5/guides/explanations/flows) on the left. Select the Logs tab from the top of the page.

![](/docs/_fern-img/5b84bec69140c251a42867c8a2da5699a7f62eb82eebe25a3c613ccc7ca4c703.webp)

The decorators divide the code in logical components, allowing you to observe the steps taken to answer a question. Every step taken by the agent creates a Log.

## Next steps

We've built a complex agentic workflow and learned how to use Humanloop to add logging to it and evaluate its performance.

Take a look at these resources to learn more about evals on Humanloop:

* Learn how to [create a custom dataset](/docs/v5/guides/evals/create-dataset) for your project.

* Learn more about using [LLM Evaluators](/docs/v5/guides/evals/llm-as-a-judge) on Humanloop.