> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://humanloop.com/docs/v5/changelog/2024/12/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://humanloop.com/_mcp/server. # December ## Improved TypeScript SDK Evals *December 18th, 2024* We've enhanced our TypeScript SDK with an evaluation utility, similar to our [Python SDK](https://humanloop.com/docs/v5/changelog/2024/10#evaluations-sdk-improvements). The utility can run evaluations on either your runtime or Humanloop's. To use your local runtime, you need to provide: * A callable function that takes your inputs/ messages * A Dataset of inputs/ messages to evaluate the function against * A set of Evaluators to use to provide judgments on the outputs of your function Here's how our [evals in code guide](https://humanloop.com/docs/v5/quickstart/evals-in-code) looks in the new TypeScript SDK: ```typescript maxLines=50 import { Humanloop } from "humanloop"; // Get API key at https://app.humanloop.com/account/api-keys const hl = new Humanloop({ apiKey: "", }); const checks = hl.evaluations.run({ name: "Initial Test", file: { path: "Scifi/App", callable: (messages: { role: string; content: string }[]) => { // Replace with your AI model logic const lastMessageContent = messages[messages.length - 1].content.toLowerCase(); return lastMessageContent === "hal" ? "I'm sorry, Dave. I'm afraid I can't do that." : "Beep boop!"; }, }, dataset: { path: "Scifi/Tests", // Replace with your own dataset datapoints: [ { messages: [ { role: "system", content: "You are an AI that responds like famous sci-fi AIs." }, { role: "user", content: "HAL" }, ], target: { output: "I'm sorry, Dave. I'm afraid I can't do that.", }, }, { messages: [ { role: "system", content: "You are an AI that responds like famous sci-fi AIs." }, { role: "user", content: "R2D2" }, ], target: { output: "Beep boop beep!", }, }, ], }, // Replace with your own Evaluators evaluators: [ { path: "Example Evaluators/Code/Exact match" }, { path: "Example Evaluators/Code/Latency" }, { path: "Example Evaluators/AI/Semantic similarity" }, ], }); console.log("Evaluation checks:", checks); ``` > **Tip** > > Check out this [cookbook example](https://github.com/humanloop/humanloop-cookbook/tree/main/node-evaluate-medqa) to learn how to evaluate a RAG pipeline with the new SDK. ## SDK Decorators in Typescript \[beta] *December 17th, 2024* We're excited to announce the beta release of our [TypeScript SDK](https://www.npmjs.com/package/humanloop/v/0.8.9-beta5), aligning with the Python logging utilities we introduced [last month](/docs/changelog/2024/11#logging-with-decorators). The new utilities help you integrate Humanloop with minimal changes to your existing code base. Take this basic chat agent instrumented through Humanloop: ```typescript maxLines=50 const callModel = (traceId: string, messages: MessageType[]) => { const response = await openAIClient.chat.completions.create({ model: "gpt-4o", temperature: 0.8, messages: messages, }); const output = response.choices[0].message.content || ""; await humanloop.prompts.log({ path: "Chat Agent/Call Model", prompt: { model: "gpt-4o", messages: [...messages, { role: "assistant", content: output }], temperature: 0.8, }, traceParentId: traceId, }) return output; } const chatAgent = () => { const traceId = humanloop.flows.log( path: "Chat Agent/Agent", ).id const messages = [{ role: "system", content: "You are a helpful assistant." }]; while (true) { const userMessage = await getCLIInput(); if (userMessage === "exit") { break; } messages.push({ role: "user", content: userMessage }); const response = await callModel(traceId, messages); messages.push({ role: "assistant", content: response }); } humanloop.flows.updateLog( traceId, { traceStatus: "complete", messages: messages } ) } ``` Using the new logging utilities, the SDK will automatically manage the Files and logging for you. Through them you can integrate Humanloop to your project with less changes to your existing codebase. Calling a function wrapped in an utility will create a Log on Humanloop. Furthermore, the SDK will detect changes to the LLM hyperparameters and create a new version automatically. The code below is equivalent to the previous example: ```typescript maxLines=50 const callModel = (messages: MessageType[]) => humanloop.prompt({ path: "Chat Agent/Call Model", callable: async (inputs: any, messages: MessageType[]) => { const response = await openAIClient.chat.completions.create({ model: "gpt-4o", temperature: 0.8, messages: messages, }); return response.choices[0].message.content || ""; }, })(undefined, messages); const chatAgent = () => humanloop.flow({ path: "Chat Agent/Agent", callable: async (inputs: any, messages: MessageType[]) => { const messages = [{ role: "system", content: "You are a helpful assistant." }]; while (true) { const userMessage = await getCLIInput(); if (userMessage === "exit") { break; } messages.push({ role: "user", content: userMessage }); const response = await callModel(messages); messages.push({ role: "assistant", content: response }); } return messages; }, })(undefined, []); ``` This release introduces three decorators: * **`flow()`**: Serves as the entry point for your AI features. Use it to call other decorated functions and trace your feature's execution. * **`prompt()`**: Monitors LLM client library calls to version your Prompt Files. Supports **OpenAI**, **Anthropic**, and **Replicate** clients. Changing the provider or hyperparameters creates a new version in Humanloop. * **`tool()`**: Versions tools using their source code. Includes a `jsonSchema` decorated to streamline function calling. > **Tip** > > Explore our [cookbook example](https://github.com/humanloop/humanloop-cookbook/tree/main/node-instrument-chat-agent) to see a simple chat agent instrumented with the new logging utilities. ## Function-calling AI Evaluators *December 15th, 2024* We've updated our AI Evaluators to use function calling by default, improving their reliability and performance. We've also updated the AI Evaluator Editor to support this change. ![AI Evaluator Editor with function calling](/docs/_fern-img/b2035f21f695e5dc699bfa8cc3493dd161d6965978cde14b1429731409c2c269.webp) New AI Evaluators will now use function calling by default. When you create an AI Evaluator in Humanloop, you will now create an AI Evaluator with a `submit_judgment(judgment, reasoning)` tool that takes `judgment` and `reasoning` as arguments. When you run this Evaluator on a Log, Humanloop will force the model to call the tool. The model will then return an appropriate judgment alongside its reasoning. You can customize the AI Evaluator in its Editor tab. Here, Humanloop displays a "Parameters" and a "Template" section, similar to the Prompt Editor, allowing you to define the messages and parameters used to call the model. In the "Judgment" section below those, you can customize the function descriptions and disable the `reasoning` argument. To test the AI Evaluator, you can load Logs from a Prompt with the **Select a Prompt or Dataset** button in the **Debug console** panel. After Logs are loaded, click the **Run** button to run the AI Evaluator on the Logs. The resulting judgments will be shown beside the Logs. If reasoning is enabled, you can view the reasoning by hovering over the judgment or by clicking the **Open in drawer** button next to the judgment. ![AI Evaluator Editor with function calling](/docs/_fern-img/0a579db1ac36c48253802e119d04c6a5aaa56261a747f7eb1350445714de2190.webp) ## New models: Gemini 2.0 Flash, Llama 3.3 70B *December 12th, 2024* To support you in adopting the latest models, we've added support for more new models, including the [latest experimental models for Gemini](https://developers.googleblog.com/en/the-next-chapter-of-the-gemini-era-for-developers/). These include `gemini-2.0-flash-exp` with better performance than Gemini 1.5 Pro and tool use, and [`gemini-exp-1206`](https://blog.google/feed/gemini-exp-1206/), the latest experimental advanced model. We've also added support for [Llama 3.3 70B](https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md) on Groq, Meta's latest model with performance comparable to their largest Llama 3.1 405B model. You can start using these models in your Prompts by going to the Editor and selecting the model from the dropdown. (To use the Gemini models, you need to have a Google API key saved in your Humanloop [account settings](https://app.humanloop.com/hl-demo/account/api-keys).) ![Gemini 2.0 Flash in Prompt Editor](/docs/_fern-img/6a1dd659eadd7a7b8e6897b60957616bc050d8e34b7cab81625cb44cb0cdc0d7.webp) ## Drag and Drop in the Sidebar *December 9th, 2024* You can now drag and drop files into the sidebar to organize your Prompts, Evaluators, Datasets, and Flows into Directories. With this much requested feature, you can easily reorganize your workspace hierarchy without having to use the 'Move...' modals. This improvement makes it easier to maintain a clean and organized workspace. We recommend using a Directory per project to group together related files. ## Logs with user-defined IDs *December 6th, 2024* We’ve added the ability to create Logs with your own unique ID, which you can then use to reference the Log when making API calls to Humanloop. ```python highlight={19} my_id = "my_very_own_and_unique_id" # create Log with "my_very_own_and_unique_id" id humanloop.prompts.call( path="path_to_the_prompt", prompt={ "model": "gpt-4", "template": [ { "role": "system", "content": "You are a helpful assistant. Tell the truth, the whole truth, and nothing but the truth", }, ], }, log_id=my_id, messages=[{"role": "user", "content": "Is it acceptable to put pineapples on pizza?"}], ) # add evaluator judgment to this Log using your own id humanloop.evaluators.log( parent_id=my_id, path="path_to_my_evaluator", judgment="good", spec={ "arguments_type": "target_free", "return_type": "select", "evaluator_type": "human", "options": [{"name": "bad", "valence": "negative"}, {"name": "good", "valence": "positive"}] }) ``` This is particularly useful for providing judgments on the Logs without requiring you to store Humanloop-generated IDs in your application. ## Flow Trace in Review View *December 3rd, 2024* We've added the ability to see the full Flow trace directly in the Review view. This is useful to get the full context of what was called during the execution of a Flow. ![](/docs/_fern-img/10d55645e54ffbbfe8c1b2f8d6d44c7df7ca67319e0e9d36f054489eb664d071.webp) To open the Log drawer side panel, click on the Log ID above the Log output in the Review view.