> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://humanloop.com/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server.

> Getting up and running with Humanloop is quick and easy. This guide will explain how to set up evaluations on Humanloop and use them to iteratively improve your applications.

## Prerequisites

#### Install and initialize the SDK

First you need to install and initialize the SDK. If you have already done this, skip to the next section.

Open up your terminal and follow these steps:

1. Install the Humanloop SDK:

```python
pip install humanloop
```

```typescript
npm install humanloop
```

2. Initialize the SDK with your Humanloop API key (you can get it from the [Organization Settings page](https://app.humanloop.com/account/api-keys)).

```python
from humanloop import Humanloop
humanloop = Humanloop(api_key="<YOUR HUMANLOOP KEY>")

# Check that the authentication was successful
print(humanloop.prompts.list())
```

```typescript
import { HumanloopClient, Humanloop } from "humanloop";

const humanloop = new HumanloopClient({ apiKey: "YOUR_API_KEY" });

// Check that the authentication was successful
console.log(await humanloop.prompts.list());
```

This quickstart will take you through running your first Eval with Humanloop.

You'll learn how to trigger an evaluation from code, interpret an eval-report on Humanloop and use it to improve your AI features.

## Create an evals script

Add the following code in a file:

```python maxLines=100
from humanloop import Humanloop

humanloop = Humanloop(api_key="<YOUR HUMANLOOP KEY>")

checks = humanloop.evaluations.run(
  name="Initial Test",
  file={
      "path": "Scifi/App",
      # Replace with your AI model
      "callable": lambda messages: (
          "I'm sorry, Dave. I'm afraid I can't do that."
          if messages[-1]["content"].lower() == "hal"
          else "Beep boop!"
      )
  },
  # Replace with your test dataset
  dataset={
      "path": "Scifi/Tests",
      "datapoints": [
        {
          "messages": [
            {
              "role": "system",
              "content": "You are an AI that responds like famous sci-fi AIs."
            },
            {
              "role": "user",
              "content": "HAL"
            }
          ],
          "target": {
            "output": "I'm sorry, Dave. I'm afraid I can't do that."
          }
        },
        {
          "messages": [
            {
              "role": "system",
              "content": "You are an AI that responds like famous sci-fi AIs."
            },
            {
              "role": "user",
              "content": "R2D2"
            }
          ],
          "target": {
            "output": "Beep boop beep!"
          }
        }
      ]
  }, 
  evaluators=[
      {"path": "Example Evaluators/Code/Exact match"},
      {"path": "Example Evaluators/Code/Latency"},
      {"path": "Example Evaluators/AI/Semantic similarity"},
  ],
)
```

```typescript maxLines=100
import { HumanloopClient } from "humanloop";

const humanloop = new HumanloopClient({
  apiKey: "<YOUR HUMANLOOP KEY>",
});

const checks = humanloop.evaluations.run({
  name: "Initial Test",
  file: {
    path: "Scifi/App",
    // Replace with your AI model
    callable: (inputs, messages) =>
      messages && messages[messages.length - 1].content.toLowerCase() === "hal"
        ? "I'm sorry, Dave. I'm afraid I can't do that."
        : "Beep boop!",
  },
  // Replace with your test dataset
  dataset: {
    path: "Scifi/Tests",
    datapoints: [
      {
        messages: [
          {
            role: "system",
            content: "You are an AI that responds like famous sci-fi AIs.",
          },
          {
            role: "user",
            content: "HAL",
          },
        ],
        target: { output: "I'm sorry, Dave. I'm afraid I can't do that." },
      },
      {
        messages: [
          {
            role: "system",
            content: "You are an AI that responds like famous sci-fi AIs.",
          },
          {
            role: "user",
            content: "R2D2",
          },
        ],
        target: { output: "Beep boop beep!" },
      },
    ],
  },
  evaluators: [
    { path: "Example Evaluators/Code/Exact match" },
    { path: "Example Evaluators/Code/Latency" },
    { path: "Example Evaluators/AI/Semantic similarity" },
  ],
});
```

This sets up the basic structure of an [Evaluation](/docs/guides/evals/overview):

1. A **callable** function that you want to evaluate. The callable should take your inputs and/or messages and returns a string. The `file` argument defines the callable as well as the location of where the evaluation results will appear on Humanloop.
2. A test [Dataset](/docs/explanation/datasets) of inputs and/or messages to run your function over and optional expected targets to evaluate against.
3. A set of [Evaluators](/docs/explanation/evaluators) to provide judgments on the output of your function. This example uses default evaluators that come with every Humanloop workspace. Evaluators can also be defined locally and pushed to the Humanloop runtime.

It returns a `checks` object that contains the results of the eval per Evaluator.

## Run your script

Run your script with the following command:

```python
python main.py
```

```typescript
npx tsx index.ts
```

You will see a URL to view your evals on Humanloop. A summary of progress and the final results will be displayed directly in your terminal:

![Eval progress and url in terminal](/docs/_fern-img/32e39d973b3b0f50bc360161ba29b2e25beb54cbf1cfcc692ed183ee78eac191.webp)![Eval results in terminal](/docs/_fern-img/5de4e223930791e53111e086a42f143aa6ecdb05efe83af703c360fc9441b4b2.webp)

## View the results

Navigate to the URL provided in your terminal to see the result of running your script on Humanloop.
This `Stats` view will show you the live progress of your local eval runs as well summary statistics of the final results.
Each new run will add a column to your `Stats` view, allowing you to compare the performance of your LLM app over time.

The `Logs` and `Review` tabs allow you to drill into individual datapoints and view the outputs of different runs side-by-side to understand how to improve your LLM app.

![Eval results in terminal](/docs/_fern-img/2b54c3400657342825fbad870b341f85c492fe03df3fedb35daf3de229f657a8.webp)

## Make a change and re-run

Your first run resulted in a `Semantic similarity` score of 3 (out of 5) and an `Exact match` score of 0. Try and make a change to your `callable` to improve
the output and re-run your script. A second run will be added to your `Stats` view and the difference in performance will be displayed.

![Re-run eval results in terminal](/docs/_fern-img/3625dbdd26ca14184b99b27c9f92616c3a6b3a95ce8141f2e18dc6e0c4544366.webp)![Re-run eval results in UI](/docs/_fern-img/f3be766e51038a90dd524d37928cc14069c3d682df913c9263a1c454eeee7da2.webp)

## Next steps

Now that you've run your first eval on Humanloop, you can:

* Explore our detailed [tutorial](/docs/tutorials/rag-evaluation) on evaluating a real RAG app where you'll learn about versioning your app, customizing logging, adding Evaluator thresholds and more.
* Create your own [Dataset](/docs/explanation/datasets) of test cases to evaluate your LLM app against.
* Create your own [Evaluators](/docs/explanation/evaluators) to provide judgments on the output of your LLM app.