> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://humanloop.com/docs/v4/guides/evaluation/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/docs/_mcp/server. # Evaluation and Monitoring ## Docs - [Overview](https://humanloop.com/docs/v4/guides/evaluation/overview.md): Learn how to set up and use Humanloop's evaluation framework to test and track the performance of your prompts. - [Run an evaluation](https://humanloop.com/docs/v4/guides/evaluation/evaluate-models-offline.md): How do you evaluate your large language model use case using a dataset and an evaluator on Humanloop? - [Set up evaluations using API](https://humanloop.com/docs/v4/guides/evaluation/evaluations-using-api.md): How to use Humanloop to evaluate your large language model use-case, using a dataset and an evaluator. - [Use LLMs to evaluate logs](https://humanloop.com/docs/v4/guides/evaluation/use-llms-to-evaluate-logs.md): Learn how to use LLM as a judge to check for PII in Logs. - [Self-hosted evaluations](https://humanloop.com/docs/v4/guides/evaluation/self-hosted-evaluations.md): Learn how to run an evaluation in your own infrastructure and post the results to Humanloop. - [Evaluating externally generated Logs](https://humanloop.com/docs/v4/guides/evaluation/evaluating-externally-generated-logs.md): Learn how to use the Humanloop Python SDK to create an evaluation run and post-generated logs. - [Evaluating with human feedback](https://humanloop.com/docs/v4/guides/evaluation/evaluating-with-human-feedback.md): Learn how to set up a human evaluator to collect feedback on the output of your model. - [Set up Monitoring](https://humanloop.com/docs/v4/guides/evaluation/monitoring.md): Learn how to create and use online evaluators to observe the performance of your models.