> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://humanloop.com/docs/v5/guides/evals/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server. # Evaluation ## Docs - [Run an Evaluation via the UI](https://humanloop.com/docs/guides/evals/run-evaluation-ui.md): How to use Humanloop to evaluate multiple different Prompts across a Dataset. - [Run an Evaluation via the API](https://humanloop.com/docs/guides/evals/run-evaluation-api.md): In this guide, we will walk through how to programmatically evaluate multiple different Prompts to compare the quality and performance of each version. - [Upload a Dataset from CSV](https://humanloop.com/docs/guides/evals/upload-dataset-csv.md): Learn how to create Datasets in Humanloop to define fixed examples for your projects, and build up a collection of input-output pairs for evaluation and fine-tuning. - [Create a Dataset via the API](https://humanloop.com/docs/guides/evals/create-dataset-api.md): Learn how to create Datasets in Humanloop to define fixed examples for your projects, and build up a collection of input-output pairs for evaluation and fine-tuning. - [Create a Dataset from existing Logs](https://humanloop.com/docs/guides/evals/create-dataset-from-logs.md): Learn how to create Datasets in Humanloop to define fixed examples for your projects, and build up a collection of input-output pairs for evaluation and fine-tuning. - [Set up a code Evaluator](https://humanloop.com/docs/guides/evals/code-based-evaluator.md): Learn how to create a code Evaluators in Humanloop to assess the performance of your AI applications. This guide covers setting up an offline evaluator, writing evaluation logic, and using the debug console. - [Set up LLM as a Judge](https://humanloop.com/docs/guides/evals/llm-as-a-judge.md): Learn how to use LLM as a judge to check for PII in Logs. - [Set up a Human Evaluator](https://humanloop.com/docs/guides/evals/human-evaluators.md): Learn how to set up a Human Evaluator in Humanloop. Human Evaluators allow your subject-matter experts and end-users to provide feedback on Prompt Logs. - [Run a Human Evaluation](https://humanloop.com/docs/guides/evals/run-human-evaluation.md): Collect judgments from subject-matter experts (SMEs) to better understand the quality of your AI product. - [Manage multiple reviewers](https://humanloop.com/docs/guides/evals/manage-multiple-reviewers.md): Learn how to split the work between your SMEs - [Compare and Debug Prompts](https://humanloop.com/docs/guides/evals/comparing-prompts.md): In this guide, we will walk through comparing the outputs from multiple Prompts side-by-side using the Humanloop Editor environment and using diffs to help debugging. - [Set up CI/CD Evaluations](https://humanloop.com/docs/guides/evals/cicd-integration.md): Learn how to automate LLM evaluations as part of your CI/CD pipeline using Humanloop and GitHub Actions. - [Spot-check your Logs](https://humanloop.com/docs/guides/evals/spot-check-logs.md): Learn how to use the Humanloop Python SDK to sample a subset of your Logs and create an Evaluation Run to spot-check them. - [Use external Evaluators](https://humanloop.com/docs/guides/evals/use-external-evaluators.md): Integrate your existing evaluation process with Humanloop. - [Evaluate external logs](https://humanloop.com/docs/guides/evals/evaluate-external-logs.md): Run an Evaluation on Humanloop with your own