> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://humanloop.com/docs/v5/explanation/datasets/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/_mcp/server. > Discover how Humanloop manages datasets, with version control and collaboration to enable you to evaluate and fine-tune your models. ![](/docs/_fern-img/2268eb4bd02172d9d4c2f797f0f9642c9592a21cd59238c1b4b1ce54c53557b7.webp) Datasets on Humanloop are collections of Datapoints used for evaluation and fine-tuning. You can think of a Datapoint as a test case for your AI application, which contains the following fields: * **Inputs**: a collection of prompt variable values that replace the `{{variables}}` defined in your prompt template during generation. * **Messages**: for chat models, you can have a history of chat messages that are appended to the prompt during generation. * **Target**: a value that in its simplest form describes the desired output string for the given inputs and messages history. For more advanced use cases, you can define a JSON object containing whatever fields are necessary to evaluate the model's output. ![](/docs/_fern-img/0780b6b0c582d4a3875112b34be96fb22ab2cd30a3dc8725b7b4e8e111761b01.webp) ## Versioning A Dataset will have multiple Versions as you iterate on the test cases for your task. This tends to be an evolving process as you learn how your [Prompts](./prompts) behave and how users interact with your AI application in the wild. Dataset Versions are immutable and are uniquely defined by the contents of the Datapoints. When you change, add, or remove Datapoints, this constitutes a new Version. Each [Evaluation](/docs/guides/evals/run-evaluation-ui) is linked to a specific Dataset Version, ensuring that your evaluation results are always traceable to the exact set of test cases used. ## Creating a Dataset A Dataset can be created in the following ways: * [Upload a CSV in the UI](/docs/guides/evals/upload-dataset-csv) * [Create a Datapoint from an existing Log](/docs/guides/evals/create-dataset-from-logs) * [Create a Dataset via the API](/docs/guides/evals/create-dataset-api) ## Using Datasets for Evaluations Datasets are foundational for Evaluations on Humanloop. Evaluations are run by iterating over the Datapoints in a Dataset, generating output from different versions of your AI application for each one. The Datapoints provide the specific test cases to evaluate, with each containing the input variables and optionally a target output that defines the desired behavior. When a target is specified, Evaluators can compare the generated outputs to the targets to assess how well each version performed. #### [Run an Evaluation](/docs/guides/evals/run-evaluation-ui) Get started with using Datasets for Evaluation via UI #### [Run an Evaluation](/docs/guides/evals/run-evaluation-api) Get started with using Datasets for Evaluation via code > Datasets are collections of Datapoints used for evaluation and fine-tuning.