> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://humanloop.com/docs/v5/guides/evals/run-evaluation-api/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/_mcp/server. An **Evaluation** on Humanloop leverages a [Dataset](/docs/explanation/datasets), a set of [Evaluators](/docs/explanation/evaluators) and different versions of a [Prompt](/docs/explanation/prompts) to compare. > **Note** > > In this guide, we use a Dataset to evaluate the performance of different Prompt versions. To learn how to evaluate Prompts > without a Dataset, see the guide on [Spot-check your Logs](./spot-check-logs). ### Prerequisites * A set of [Prompt](/docs/explanation/prompts) versions you want to compare - see the guide on [creating Prompts](/docs/guides/prompts/create-prompt). * A [Dataset](/docs/explanation/datasets) containing datapoints for the task - see the guide on [creating a Dataset via API](/docs/v5/guides/evals/create-dataset-api). * At least one [Evaluator](/docs/explanation/evaluators) to judge the performance of the Prompts - see the guides on creating [Code](/docs/evaluation/guides/code-based-evaluator), [AI](/docs/evaluation/guides/llm-as-a-judge) and [Human](/docs/evaluation/guides/human-evaluators) Evaluators. ## Run an Evaluation For this guide, we're going to evaluate the performance of a Support Agent that responds to user queries about Humanloop's product and documentation. Our goal is to understand which base model between `gpt-4o`, `gpt-4o-mini` and `claude-3-5-sonnet-20241022` is most appropriate for this task. ### Create a Prompt Create a Support Agent Prompt with three versions each using a different base model. ```python from humanloop import Humanloop humanloop = Humanloop(api_key="YOUR_API_KEY") system_message = "You are a helpful assistant. Your job is to respond to FAQ style queries about the Humanloop documentation and platform. Be polite and succinct." gpt_4o = humanloop.prompts.upsert( path="Run Evaluation via API/Support Agent", model="gpt-4o", endpoint="chat", template=[ { "content": system_message, "role": "system", } ], provider="openai", version_name="gpt-4o", version_description="FAQ style support agent using gpt-4o", ) gpt_4o_mini = humanloop.prompts.upsert( path="Run Evaluation via API/Support Agent", model="gpt-4o-mini", endpoint="chat", template=[ { "content": system_message, "role": "system", } ], provider="openai", version_name="gpt-4o-mini", version_description="FAQ style support agent using gpt-4o-mini", ) sonnet = humanloop.prompts.upsert( path="Run Evaluation via API/Support Agent", model="claude-3-5-sonnet-20241022", endpoint="chat", template=[ { "content": system_message, "role": "system", } ], provider="anthropic", version_name="claude-3-5-sonnet", version_description="FAQ style support agent using Claude 3.5 Sonnet", ) # store prompt versions for later use prompt_versions = [gpt_4o.version_id, gpt_4o_mini.version_id, sonnet.version_id] ``` ```typescript import { HumanloopClient } from "humanloop"; const humanloop = new HumanloopClient({ apiKey: "HUMANLOOP_API_KEY" }); const systemMessage = "You are a helpful assistant. Your job is to respond to FAQ style queries about the Humanloop documentation and platform. Be polite and succinct."; const gpt_4o = await humanloop.prompts.upsert({ path: "Run Evaluation via API/Support Agent", model: "gpt-4o", endpoint: "chat", template: [ { "content": systemMessage, "role": "system", } ], provider: "openai", version_name: "gpt-4o", version_description: "FAQ style support agent using gpt-4o" }) const gpt_4o_mini = await humanloop.prompts.upsert({ path: "Run Evaluation via API/Support Agent", model: "gpt-4o-mini", endpoint: "chat", template: [ { "content": systemMessage, "role": "system", } ], provider: "openai", version_name: "gpt-4o-mini", version_description: "FAQ style support agent using gpt-4o-mini" }) const sonnet = await humanloop.prompts.upsert({ path: "Run Evaluation via API/Support Agent", model: "claude-3-5-sonnet-20241022", endpoint: "chat", template: [ { "content": systemMessage, "role": "system", } ], provider: "anthropic", version_name: "claude-3-5-sonnet", version_description: "FAQ style support agent using Claude 3.5 Sonnet" }) // store prompt versions for later use const promptVersions = [gpt_4o.versionId, gpt_4o_mini.versionId, sonnet.versionId] ``` ### Create a Dataset We defined sample data that contains user messages and desired responses for the Support Agent Prompt. We will now create a Dataset with these datapoints. ```python humanloop.datasets.upsert( path="Run Evaluation via API/Dataset with user questions", datapoints=[ { "messages": [{ "role": "user", "content": "How do I manage my organization's API keys?", }], "target": {"answer": "Hey, thanks for your questions. Here are steps for how to achieve: 1. Log in to the Humanloop Dashboard \n\n2. Click on \"Organization Settings.\"\n If you do not see this option, you might need to contact your organization admin to gain the necessary permissions.\n\n3. Within the settings or organization settings, select the option labeled \"API Keys\" on the left. Here you will be able to view and manage your API keys.\n\n4. You will see a list of existing API keys. You can perform various actions, such as:\n - **Generate New API Key:** Click on the \"Generate New Key\" button if you need a new API key.\n - **Revoke an API Key:** If you need to disable an existing key, find the key in the list and click the \"Revoke\" or \"Delete\" button.\n - **Copy an API Key:** If you need to use an existing key, you can copy it to your clipboard by clicking the \"Copy\" button next to the key.\n\n5. **Save and Secure API Keys:** Make sure to securely store any new or existing API keys you are using. Treat them like passwords and do not share them publicly.\n\nIf you encounter any issues or need further assistance, it might be helpful to engage with an engineer or your IT department to ensure you have the necessary permissions and support.\n\nWould you need help with anything else?"}, }, { "messages":[{ "role": "user", "content": "Can I use my code evaluator for monitoring my legal-copilot prompt?", }], "target": {"answer": "Hey, thanks for your questions. Here are steps for how to achieve: 1. Navigate to your Prompt dashboard. \n 2. Select the `Monitoring` button on the top right of the Prompt dashboard \n 3. Within the model select the Version of the Evaluator you want to turn on for monitoring. \n\nWould you need help with anything else?"}, }, ], action="set", version_name="User questions", version_description="Add two new questions and answers", ) ``` ```typescript await humanloop.datasets.upsert({ path: "Run Evaluation via API/Dataset with user questions", datapoints: [{ "messages": [{ "role": "user", "content": "How do i manage my organizations API keys?", }], "target": {"answer": "Hey, thanks for your questions. Here are steps for how to achieve: 1. Log in to the Humanloop Dashboard \n\n2. Click on \"Organization Settings.\"\n If you do not see this option, you might need to contact your organization admin to gain the necessary permissions.\n\n3. Within the settings or organization settings, select the option labeled \"API Keys\" on the left. Here you will be able to view and manage your API keys.\n\n4. You will see a list of existing API keys. You can perform various actions, such as:\n - **Generate New API Key:** Click on the \"Generate New Key\" button if you need a new API key.\n - **Revoke an API Key:** If you need to disable an existing key, find the key in the list and click the \"Revoke\" or \"Delete\" button.\n - **Copy an API Key:** If you need to use an existing key, you can copy it to your clipboard by clicking the \"Copy\" button next to the key.\n\n5. **Save and Secure API Keys:** Make sure to securely store any new or existing API keys you are using. Treat them like passwords and do not share them publicly.\n\nIf you encounter any issues or need further assistance, it might be helpful to engage with an engineer or your IT department to ensure you have the necessary permissions and support.\n\nWould you need help with anything else?"}, }, { "messages":[{ "role": "user", "content": "Hey, can do I use my code evaluator for monitoring my legal-copilot prompt?", }], "target": {"answer": "Hey, thanks for your questions. Here are steps for how to achieve: 1. Navigate to your Prompt dashboard. \n 2. Select the `Monitoring` button on the top right of the Prompt dashboard \n 3. Within the model select the Version of the Evaluator you want to turn on for monitoring. \n\nWould you need help with anything else?"}, }], action: "set", version_name: "User questions", version_description: "Add two new questions and answers" }); ``` ### Create an Evaluation We create an Evaluation Run to compare the performance of the different Prompts using the Dataset we just created. For this guide, we selected *Semantic Similarity*, *Cost* and *Latency* Evaluators. You can find these Evaluators in the **Example Evaluators** folder in your workspace. > **Note** > > "Semantic Similarity" Evaluator measures the degree of similarity between the model's response and the expected output. The similarity is rated on a scale from 1 to 5, where 5 means very similar. ```python evaluation = humanloop.evaluations.create( name="Evaluation via API", file={ "path": "Run Evaluation via API/Support Agent", }, evaluators=[{"path": "Example Evaluators/AI/Semantic Similarity"}, {"path": "Example Evaluators/Code/Cost"}, {"path": "Example Evaluators/Code/Latency"}], ) # Create a Run for each prompt version for prompt_version in prompt_versions: humanloop.evaluations.create_run( id=evaluation.id, dataset={"path": "Run Evaluation via API/Dataset with user questions"}, version={"version_id": prompt_version}, ) ``` ```typescript const evaluation = await humanloop.evaluations.create({ name: "Evaluation via API", file: { "path": "Run Evaluation via API/Support Agent", }, evaluators: [{"path": "Example Evaluators/AI/Semantic Similarity"}, {"path": "Example Evaluators/Code/Cost"}, {"path": "Example Evaluators/Code/Latency"}], }); for (const promptVersion of promptVersions) { await humanloop.evaluations.createRun(evaluation.id, { dataset: { path: "Run Evaluation via API/Dataset with user questions" }, version: { versionId: promptVersion }, }); } ``` ### Inspect the Evaluation stats When Runs are completed, you can inspect the Evaluation Stats to see the summary of the Evaluators judgments. ```python evaluation_stats = humanloop.evaluations.get_stats( id=evaluation.id, ) print(evaluation_stats.report) ``` ```typescript const evaluationStats = await humanloop.evaluations.getStats(evaluation.id); console.log(evaluationStats.report); ``` ![Drill down to Evaluatoin Logs.](/docs/_fern-img/a6fcdc40a0b1a9a6ce5f37b7ab2e3b9d91ea1d63ff9e8ca2597f5ac85ae57bdb.webp) Alternatively you can see detailed stats in the Humanloop UI. Navigate to the Prompt, click on the **Evaluations** tab at the top of the page and select the Evaluation you just created. The stats are displayed in the **Stats** tab. ![Drill down to Evaluatoin Logs.](/docs/_fern-img/3e2a2a13d6a29b68a767e6fc0c9c9ca7b4a7cb8da13edc7d08e105adb60c023e.webp) # Run an Evaluation using your runtime If you choose to execute Prompts using your own Runtime, you still can benefit from Humanloop Evaluations. In code snippet below, we run Evaluators hosted on Humanloop using logs produced by the OpenAI client. ```python # create new Humanloop prompt prompt = humanloop.prompts.upsert( path="Run Evaluation via API/Support Agent my own runtime", model="gpt-4o", endpoint="chat", template=[ { "content": "You are a helpful assistant. Your job is to respond to FAQ style queries about the Humanloop documentation and platform. Be polite and succinct.", "role": "system", } ], provider="openai", ) # create the evaluation evaluation = humanloop.evaluations.create( name="Evaluation via API using my own runtime", file={ "path": "Run Evaluation via API/Support Agent my own runtime", }, evaluators=[{"path": "Example Evaluators/AI/Semantic Similarity"}, {"path": "Example Evaluators/Code/Cost"}, {"path": "Example Evaluators/Code/Latency"}], ) # use dataset created in previous steps datapoints = humanloop.datasets.list_datapoints(dataset.id) import openai openai_client = openai.OpenAI(api_key="OPENAI_API_KEY") # create a run run = humanloop.evaluations.create_run( id=evaluation.id, dataset={"version_id": dataset.version_id}, version={"version_id": prompt.version_id}, ) # for each datapoint in the dataset, create a chat completion for datapoint in datapoints: # create a run chat_completion = openai_client.chat.completions.create( messages=datapoint.messages, model=prompt.model ) # log the prompt humanloop.prompts.log( id=prompt.id, run_id=run.id, version_id=prompt.version_id, source_datapoint_id=datapoint.id, output_message=chat_completion.choices[0].message, messages=datapoint.messages, ) ``` ```typescript // create a new Humanloop prompt const prompt = await humanloop.prompts.upsert({ path: "Run Evaluation via API/Support Agent my own runtime", model: "gpt-4o", endpoint: "chat", template: [ { content: "You are a helpful assistant. Your job is to respond to FAQ style queries about the Humanloop documentation and platform. Be polite and succinct.", role: "system", }, ], provider: "openai", }); // create the evaluation const evaluation = await humanloop.evaluations.create({ name: "Evaluation via API using my own runtime", file: { path: "Run Evaluation via API/Support Agent my own runtime", }, evaluators: [ { path: "Example Evaluators/AI/Semantic Similarity" }, { path: "Example Evaluators/Code/Cost" }, { path: "Example Evaluators/Code/Latency" }, ], }); // use dataset created in previous steps const datapoints = await humanloop.datasets.listDatapoints(dataset.id); // import openai const openaiClient = new OpenAI({ apiKey: "OPENAI_API_KEY", }); // create a run const run = await humanloop.evaluations.createRun(evaluation.id, { dataset: { versionId: dataset.versionId }, version: { versionId: prompt.versionId }, }); // for each datapoint in the dataset, create a chat completion for (const datapoint of datapoints.data) { // create a run const chatCompletion = await openaiClient.chat.completions.create({ messages: datapoint.messages, model: prompt.model, }); // log the prompt await humanloop.prompts.log({ id: prompt.id, runId: run.id, versionId: prompt.versionId, sourceDatapointId: datapoint.id, outputMessage: chatCompletion.choices[0].message, messages: datapoint.messages, }); } ``` ## Next steps * Learn how to [set up LLM as a Judge](./llm-as-a-judge) to evaluate your AI applications. > In this guide, we will walk through how to programmatically evaluate multiple different Prompts to compare the quality and performance of each version.