We all know how to run unit tests on our apps in CI/CD workflows, but we're now in a new world of AI apps with non-deterministic output. How do we evaluate changes to AI apps to ensure that the LLM answer quality remains high? What tools can we use in our workflows, and how do we avoid unnecessary LLM costs?
In this interactive session, Pamela Fox, cloud advocate in Python at Microsoft, will demonstrate automated tools for AI chat app evaluation and give you a chance to try out the workflows for yourself. You'll experience the difference between both a failed and successful evaluation.
Note: This stage has limited capacity and will seat attendees at a first come first serve basis. Adding this session to your schedule does not reserve a seat. Please attend early.
On demand catalog
Select a session below to watch.
Save the pass/fail for unit tests—there’s a better way to evaluate AI apps
, Cloud Advocate in Python, Microsoft
Session Type: Sandbox Session
Key Takeaway 1: Discover multiple types of metrics to help evaluate AI applications.
Key Takeaway 2: Learn how to configure a GitHub Action with access to powerful LLMs so that it can calculate LLM-based metrics.
Key Takeaway 3: Understand the complexity involved in evaluating AI apps and why the same simple pass/fail used in unit tests doesn't apply.
Topic: AI, Automated Infrastructure Deployment
Target Audience: Enterprise - Developers, Startups
Industry: Applicable to all
Level: Level 100: Introductory
GitHub Product: Actions
Delivery Format: In-person