AI Evals Course

AI Evals For Engineers & PMs

Pedro Javier Martos Velasco

Principal Software Engineer @ The Workshop

I've learned a lot

About three weeks ago, I started the course "AI Evals For Engineers & PMs" on maven.com. In all honesty, I wasn't sure what I was getting myself into. Sure, I'm a seasoned, product-minded engineer, and I was certain that the material would prove useful, especially considering that proper evaluation is still quite niche, even more so in Spain. Becoming something of a local go-to person in that particular field sounded rather appealing. Oh boy, was I not prepared for this. I should have brought a bigger boat. The amount of material, thought processes, tips, tricks, war stories, tools, recommendations and anecdotes is astonishing. There’s just so much to unpack. I found myself rewatching every session several times to make sure I didn’t miss even the tiniest detail. When it all clicks in your brain, it suddenly feels so intuitive. Guess what? Thinking in terms of evals reinforces your product mentality. And it paves the road to continuous testing AKA true observability. I believe I'm a much more complete engineer now than I was before starting this course. I will never look at logs, traces or assertions the same way again. Thanks to the titans Hamel H. and Shreya Shankar for leading this course. I've learned a lot from you folks.

About three weeks ago, I started the course "AI Evals For Engineers & PMs" on maven.com. In all honesty, I wasn't sure what I was getting myself into. Sure, I'm a seasoned, product-minded engineer, and I was certain that the material would prove useful, especially considering that proper evaluation is still quite niche, even more so in Spain. Becoming something of a local go-to person in that particular field sounded rather appealing. Oh boy, was I not prepared for this. I should have brought a bigger boat. The amount of material, thought processes, tips, tricks, war stories, tools, recommendations and anecdotes is astonishing. There’s just so much to unpack. I found myself rewatching every session several times to make sure I didn’t miss even the tiniest detail. When it all clicks in your brain, it suddenly feels so intuitive. Guess what? Thinking in terms of evals reinforces your product mentality. And it paves the road to continuous testing AKA true observability. I believe I'm a much more complete engineer now than I was before starting this course. I will never look at logs, traces or assertions the same way again. Thanks to the titans Hamel H. and Shreya Shankar for leading this course. I've learned a lot from you folks.

Pastor Soto

Independent

One of the best, if not the best, courses on AI.

This course provided a hands-on path to understand how AI products can be improved and how critical evaluations are for any business. The best lesson from me was having to understand the nuances of evaluation systems, and how a deep understanding of your data can make massive improvements on your results. One of the best, if not the best, courses on AI!!

This course provided a hands-on path to understand how AI products can be improved and how critical evaluations are for any business. The best lesson from me was having to understand the nuances of evaluation systems, and how a deep understanding of your data can make massive improvements on your results. One of the best, if not the best, courses on AI!!

Mert Ozgun

Churnkey

Concrete steps I can apply right away

What I loved about this course is that it didn’t just explain concepts, it gave me concrete steps I could apply right away. Shreya and Hamel focus on what actually works in practice, not just theory. I now have a clear, structured process for evaluating and improving our AI systems, and I’ve already started using it in our workflow. If you’re tired of vague advice and want something you can actually do, this is it.

What I loved about this course is that it didn’t just explain concepts, it gave me concrete steps I could apply right away. Shreya and Hamel focus on what actually works in practice, not just theory. I now have a clear, structured process for evaluating and improving our AI systems, and I’ve already started using it in our workflow. If you’re tired of vague advice and want something you can actually do, this is it.

Nihit N

CEO, Yuuki

This course saved me precious time.

As a CEO of an AI agent company, the quality of our product is of highest priority. And without this course I was ending up spending too much time. This course easily saved my 50 hours of research and trial and error.

As a CEO of an AI agent company, the quality of our product is of highest priority. And without this course I was ending up spending too much time. This course easily saved my 50 hours of research and trial and error.

Dan Treasure

@dan__treasure

For 3 weeks I’ve had the privilege to take the AI Evals course from @HamelHusain & @sh_reya. My team has known we need to skill up on evals and we all read what we can but there is no comparison to lecture + course reader + homework + guest presentations + office hours
For 3 weeks I’ve had the privilege to take the AI Evals course from @HamelHusain & @sh_reya. My team has known we need to skill up on evals and we all read what we can but there is no comparison to lecture + course reader + homework + guest presentations + office hours

Abhilash N L

HackerRank, Principal Product Manager

Highly valuable for PMs, not just Engineers

Starting this course as a Product Manager I was a bit worried that I might lag behind other folks when learning about evals but I was pleasantly surprised at how the whole course along with its reading material and homework made things so simple for me to learn and grasp the key concepts. In addition to the right process the course also teaches one about the tools to use to implement the process which is of great practical value. Highly highly recommend this course for all the AI practitioners.
Starting this course as a Product Manager I was a bit worried that I might lag behind other folks when learning about evals but I was pleasantly surprised at how the whole course along with its reading material and homework made things so simple for me to learn and grasp the key concepts. In addition to the right process the course also teaches one about the tools to use to implement the process which is of great practical value. Highly highly recommend this course for all the AI practitioners.

NS
Nabeel Seedat

Researcher, Cambridge

I learned how to "look at the data" systematically

Fantastic insights into making LLM evals work. Without Evals, you are just in the dark about whether your AI systems actually work. Hamel & Shreya's course does a fantastic job of both systematically and thoroughly covering different important dimensions to evals. The biggest takeaway is the immense value of just looking at the data. I highly recommend it, for the great content! A huge bonus is the companion reader.
Fantastic insights into making LLM evals work. Without Evals, you are just in the dark about whether your AI systems actually work. Hamel & Shreya's course does a fantastic job of both systematically and thoroughly covering different important dimensions to evals. The biggest takeaway is the immense value of just looking at the data. I highly recommend it, for the great content! A huge bonus is the companion reader.

Jaume

Senior Data Scientist at New Work SE

Worth the time investment

Great course! Main take away is "look at your data", but there's much more like, how to automate your evals, how to implement CI/CD for LLM apps, how handle dataset labeling and more. It's great time investment
Great course! Main take away is "look at your data", but there's much more like, how to automate your evals, how to implement CI/CD for LLM apps, how handle dataset labeling and more. It's great time investment

Mike Powers

AI Engineer

Exceptionally clear with a relentlessly practical focus

Exceptionally clear with a relentlessly practical focus. Hamel and Shreya don't sell you on fancy tools or generic benchmarks. Instead, they distill hard-won experience from hundreds of client engagements into what actually works for building reliable AI applications. They show you exactly what steps you must never skip, how and where you can and can't automate, and how to move from hoping your AI works to systematic understanding of your pipeline's real behavior. The focus is on looking at your actual data and identifying the failure modes that matter for your specific use case. The spicy takes and real-world case studies make it clear this isn't theoretical, and the accompanying course reader is invaluable. No fluff, no jargon, just the systematic framework you need to build AI systems that people can actually depend on in production.

Exceptionally clear with a relentlessly practical focus. Hamel and Shreya don't sell you on fancy tools or generic benchmarks. Instead, they distill hard-won experience from hundreds of client engagements into what actually works for building reliable AI applications. They show you exactly what steps you must never skip, how and where you can and can't automate, and how to move from hoping your AI works to systematic understanding of your pipeline's real behavior. The focus is on looking at your actual data and identifying the failure modes that matter for your specific use case. The spicy takes and real-world case studies make it clear this isn't theoretical, and the accompanying course reader is invaluable. No fluff, no jargon, just the systematic framework you need to build AI systems that people can actually depend on in production.

Taylor Holland

Head of DSE Weights and Biases

The most practical course I’ve taken on AI

Evaluations are the fundamental building block for building AI systems (at least if you want them to actually work). Hamel & Shreya’s course is the best out there to learn build AI applications that actually work.  I highly recommend this course To anyone working in AI.

Evaluations are the fundamental building block for building AI systems (at least if you want them to actually work). Hamel & Shreya’s course is the best out there to learn build AI applications that actually work.  I highly recommend this course To anyone working in AI.

Oualid Nouri

CEO, UnaMind Informatics

Plenty of real-world examples and thoughtful exercises

This course reframes evaluation not as a final step, but as a structured process that’s central to building reliable AI applications. For practitioners, this shift in perspective is incredibly valuable—it provides a clear path to understanding system behavior, tracking progress, and making informed decisions across the AI lifecycle. Shreya and Hamel combine clarity with practicality. Through real-world examples and thoughtful exercises, they show how to identify failure modes, perform targeted error analysis, and align evaluation with product goals—without relying on complex jargon. A key takeaway: asking “Is our system improving?” requires more than metrics—it requires a repeatable, insight-driven process. This course gives you the tools to build that process. Highly recommended for engineers, PMs, and anyone serious about delivering trustworthy, high-impact AI systems.

This course reframes evaluation not as a final step, but as a structured process that’s central to building reliable AI applications. For practitioners, this shift in perspective is incredibly valuable—it provides a clear path to understanding system behavior, tracking progress, and making informed decisions across the AI lifecycle. Shreya and Hamel combine clarity with practicality. Through real-world examples and thoughtful exercises, they show how to identify failure modes, perform targeted error analysis, and align evaluation with product goals—without relying on complex jargon. A key takeaway: asking “Is our system improving?” requires more than metrics—it requires a repeatable, insight-driven process. This course gives you the tools to build that process. Highly recommended for engineers, PMs, and anyone serious about delivering trustworthy, high-impact AI systems.

Shaw Talebi

AI Educator & Builder | PhD, Physics

99% of AI engineers don’t do this one thing… So if you do, you can beat them and build apps that actually make an impact. Over the past weeks, I’ve been taking Hamel and Shreya’s course on AI evals. My biggest insight from the course (so far) is that: Looking at your data is the highest-leverage thing you can do as an engineer! "What does that mean?" That means you block an hour on your calendar to read user inputs and your AI system’s corresponding output. Then, for each example where the system made a mistake, write detailed notes on why the output is bad. "That sounds like a lot of work. Is it really worth it?" YES! I was also skeptical (or just lazy 😅), but after doing this for a LinkedIn ghostwriter project, my only regret is that I didn’t make this part of my development workflow sooner. Here’s the process: 1) Generate 49 LinkedIn posts using GPT-4.1 based on unstructured notes/ideas 2) Vibe code a custom data annotator using Streamlit 3) Read each input-output pair and write notes on model mistakes 4) Categorize mistakes and evaluate their frequency Step 4 is the key result of this effort. It allowed me to triage the most common/severe errors and focus my development effort on resolving those errors. Thus, my improvements weren't random, ad hoc fixes but a clear path toward getting the greatest lift in performance from the least effort (i.e. leverage). Excited to share more about this project in upcoming YouTube videos :) 🙏 A huge thanks to Hamel and Shreya for putting this course together. It’s taken a tremendous amount of guesswork out of the LLM application development process for me. -- ♻️ If this post was helpful, repost it!
99% of AI engineers don’t do this one thing… So if you do, you can beat them and build apps that actually make an impact. Over the past weeks, I’ve been taking Hamel and Shreya’s course on AI evals. My biggest insight from the course (so far) is that: Looking at your data is the highest-leverage thing you can do as an engineer! "What does that mean?" That means you block an hour on your calendar to read user inputs and your AI system’s corresponding output. Then, for each example where the system made a mistake, write detailed notes on why the output is bad. "That sounds like a lot of work. Is it really worth it?" YES! I was also skeptical (or just lazy 😅), but after doing this for a LinkedIn ghostwriter project, my only regret is that I didn’t make this part of my development workflow sooner. Here’s the process: 1) Generate 49 LinkedIn posts using GPT-4.1 based on unstructured notes/ideas 2) Vibe code a custom data annotator using Streamlit 3) Read each input-output pair and write notes on model mistakes 4) Categorize mistakes and evaluate their frequency Step 4 is the key result of this effort. It allowed me to triage the most common/severe errors and focus my development effort on resolving those errors. Thus, my improvements weren't random, ad hoc fixes but a clear path toward getting the greatest lift in performance from the least effort (i.e. leverage). Excited to share more about this project in upcoming YouTube videos :) 🙏 A huge thanks to Hamel and Shreya for putting this course together. It’s taken a tremendous amount of guesswork out of the LLM application development process for me. -- ♻️ If this post was helpful, repost it!

Danielle Dirks

Product Strategy & Systems | Acquisition to Activation, User-Centered Platforms, AI & Automation | Building @ It’s Kinda Magic

I just want to give a shout out to the AI Evals for PMs and Engineers course I've been taking from Shreya (lnkd.in/gnnzxuhT) and Hamel (lnkd.in/g4Fb78M4) over on Maven. it has been one of the most impactful learning experiences I’ve had in the AI space. The course offers a deep dive into what is essentially test-driven development for AI systems. We covered everything from frameworks for identifying failure modes and building our own eval tooling to designing collaborative workflows and analyzing failure funnels in agentic systems. The best part, is it’s not just theory. We've had hands-on assignments each week, paired with readings and live walkthroughs from over a dozen guest experts who shared real-world implementations. What stood out most was how practical and rigorous the approach was. This isn’t surface-level prompting advice or a high-level overview. It’s a grounded, methodical framework for making AI products actually work, and keeping them working as they evolve. In fact, I found it so valuable that I’m taking it again in July to go even deeper. If you’re a product manager, engineer, or builder working with AI and you want a solid foundation for evaluating and improving your systems, I can’t recommend this course enough. They'll be doing their last live course starting July 21 and I highly recommend signing up. lnkd.in/gRZxpGYm I’ll be writing more about evals soon, but this was the spark I needed to take things seriously. #AIProduct #AIevals #ProductManagement #ParlanceLabs
I just want to give a shout out to the AI Evals for PMs and Engineers course I've been taking from Shreya (lnkd.in/gnnzxuhT) and Hamel (lnkd.in/g4Fb78M4) over on Maven. it has been one of the most impactful learning experiences I’ve had in the AI space. The course offers a deep dive into what is essentially test-driven development for AI systems. We covered everything from frameworks for identifying failure modes and building our own eval tooling to designing collaborative workflows and analyzing failure funnels in agentic systems. The best part, is it’s not just theory. We've had hands-on assignments each week, paired with readings and live walkthroughs from over a dozen guest experts who shared real-world implementations. What stood out most was how practical and rigorous the approach was. This isn’t surface-level prompting advice or a high-level overview. It’s a grounded, methodical framework for making AI products actually work, and keeping them working as they evolve. In fact, I found it so valuable that I’m taking it again in July to go even deeper. If you’re a product manager, engineer, or builder working with AI and you want a solid foundation for evaluating and improving your systems, I can’t recommend this course enough. They'll be doing their last live course starting July 21 and I highly recommend signing up. lnkd.in/gRZxpGYm I’ll be writing more about evals soon, but this was the spark I needed to take things seriously. #AIProduct #AIevals #ProductManagement #ParlanceLabs

Sam Julien

@samjulien

This course is unbelievably good, and the accompanying reader is gold. I’m getting a ton out of it so far. Highly recommend you take advantage of a chance to learn from the best! twitter.com/hamelhusain/status/1928507728508629312
This course is unbelievably good, and the accompanying reader is gold. I’m getting a ton out of it so far. Highly recommend you take advantage of a chance to learn from the best! twitter.com/hamelhusain/status/1928507728508629312

Inspiration

@tespry

Currently taking "AI Evals for Engineers & PMs" course with @sh_reya and @HamelHusain. And it's tremendously helpful!!! As a DS in my organization, I often noticed confusion around best practices and available approaches for measuring the objective quality of LLM responses.
Currently taking "AI Evals for Engineers & PMs" course with @sh_reya and @HamelHusain. And it's tremendously helpful!!! As a DS in my organization, I often noticed confusion around best practices and available approaches for measuring the objective quality of LLM responses.

Tomás Lucas @tomaslucas.bsky.social

@tomaslucas_

One of the many things I'm learning in the evaluations course is creating a continuous improvement flywheel cycle. I've tried to reflect this in the accompanying diagram, based on a chapter from @HamelHusain and @sh_reya's book, provided for the Evaluation course for GenAI apps
One of the many things I'm learning in the evaluations course is creating a continuous improvement flywheel cycle. I've tried to reflect this in the accompanying diagram, based on a chapter from @HamelHusain and @sh_reya's book, provided for the Evaluation course for GenAI apps

Gang Rui

@limgangrui

Learnt so much from the AI Evals course by @HamelHusain & @sh_reya Prior to the course, I struggled a lot trying to build production grade AI apps using evals. Thankfully, this course has equipped me loads of knowledge to tackle them. Here's what I've learnt: 1. A systematic
Learnt so much from the AI Evals course by @HamelHusain & @sh_reya Prior to the course, I struggled a lot trying to build production grade AI apps using evals. Thankfully, this course has equipped me loads of knowledge to tackle them. Here's what I've learnt: 1. A systematic

Pastor Soto

Machine Learning Engineer | Mentor at @DeepLearningAI & Code in Place | Medical Student | Helping Businesses Make Smarter, Data-Driven Decisions

The picture illustrates the challenges you'll face in your AI applications. The course by Shreya Shankar and Hamel H. shows how this failure manifests in AI applications and what to do about it. They did a great job of creating a systematic path to understand the process that evaluates your product and answers the question: Is my application good? Are we making progress? Is Model A better than Model B for my use case? These are some of my takeaways from the course: - Looking at your data is the biggest ROI thing on anything! - Evaluation is a continuous process - A systematic approach helps to break down the complexity of the problems The most impressive thing about the course was how they handled this topic and made it manageable for anyone, Product Managers, Engineers, and everyone who joined. This course was good in itself, but Shreya and Hamel went a step further to answer every single question from almost 700 students. This was a massive task that they mastered. The next cohort will be the last live cohort, so this is a one-time opportunity to learn from them live! By far one of the best courses I've taken, if you want to learn how to build systems based on AI applications, this is the place to go. Thank you, Shreya Shankar and Hamel H., for putting this together and sharing with everyone!
The picture illustrates the challenges you'll face in your AI applications. The course by Shreya Shankar and Hamel H. shows how this failure manifests in AI applications and what to do about it. They did a great job of creating a systematic path to understand the process that evaluates your product and answers the question: Is my application good? Are we making progress? Is Model A better than Model B for my use case? These are some of my takeaways from the course: - Looking at your data is the biggest ROI thing on anything! - Evaluation is a continuous process - A systematic approach helps to break down the complexity of the problems The most impressive thing about the course was how they handled this topic and made it manageable for anyone, Product Managers, Engineers, and everyone who joined. This course was good in itself, but Shreya and Hamel went a step further to answer every single question from almost 700 students. This was a massive task that they mastered. The next cohort will be the last live cohort, so this is a one-time opportunity to learn from them live! By far one of the best courses I've taken, if you want to learn how to build systems based on AI applications, this is the place to go. Thank you, Shreya Shankar and Hamel H., for putting this together and sharing with everyone!

Radek Osmulski 🇺🇦

@radekosmulski

I started the LLM Evals companion book. What a phenomenal read! @HamelHusain and @sh_reya are experienced explorers taking you on a journey into the fascinating new world of putting LLMs to good use. A world that they started to chart expertly! My first impressions:
I started the LLM Evals companion book. What a phenomenal read! @HamelHusain and @sh_reya are experienced explorers taking you on a journey into the fascinating new world of putting LLMs to good use. A world that they started to chart expertly! My first impressions:

Tonny Ouma

@devopsdream

I have spent the past year assisting customers in building production-scale LLM applications. @sh_reya and @HamelHusain have done an excellent job of conceptually explaining the path to production, which involves navigating the "Three Gulfs."
I have spent the past year assisting customers in building production-scale LLM applications. @sh_reya and @HamelHusain have done an excellent job of conceptually explaining the path to production, which involves navigating the "Three Gulfs."

Alex Strick van Linschoten

@strickvl

Just completed a session on systematic failure mode analysis for LLM applications as part of the @HamelHusain / @sh_reya evals course. The approach reminded me of my historian background—turns out the skills for wrestling with unstructured data translate well to LLM evaluation.
Just completed a session on systematic failure mode analysis for LLM applications as part of the @HamelHusain / @sh_reya evals course. The approach reminded me of my historian background—turns out the skills for wrestling with unstructured data translate well to LLM evaluation.

Taka 高 Gang Gang

@JeromeDawgYo

I want to give a huge thanks to @sh_reya and @HamelHusain for leading an amazing course on AI Evals! They've laid out in such simple yet intuitive steps how to improve AI apps w/ data-driven evidence. The 3 Gulphs has been a real game changer in my way of AI problem-solving!
I want to give a huge thanks to @sh_reya and @HamelHusain for leading an amazing course on AI Evals! They've laid out in such simple yet intuitive steps how to improve AI apps w/ data-driven evidence. The 3 Gulphs has been a real game changer in my way of AI problem-solving!

David Aronchick

CEO, Expanso

Value Packed Course

Hamel and Shreya teach absolutely must-have skills for AI that is rarely taught elsewhere.   There are many courses that teach how to use tools and frameworks, but there aren't many courses that teach you the right process to follow.   The things they teach in this course are timeless and apply to any AI problem that you might work on now or in the future. 

Hamel and Shreya teach absolutely must-have skills for AI that is rarely taught elsewhere.   There are many courses that teach how to use tools and frameworks, but there aren't many courses that teach you the right process to follow.   The things they teach in this course are timeless and apply to any AI problem that you might work on now or in the future. 

J
JB

AI Engineer

One of the most practical courses I've ever taken

This is one of the most practical courses I've ever taken and the LLM Evals companion book is fantastic! I was able to put the concepts directly into action throughout the course, which has already improved our AI evals dramatically.

This is one of the most practical courses I've ever taken and the LLM Evals companion book is fantastic! I was able to put the concepts directly into action throughout the course, which has already improved our AI evals dramatically.

Jeremy Lewi

Software Engineer, Doghouse Labs

Foundational course for moving beyond vibes

This course is the basis for how we are designing evals for our AI SRE. The course let us go from vibes to quantitative metrics that let us systematically improve the AI.

This course is the basis for how we are designing evals for our AI SRE. The course let us go from vibes to quantitative metrics that let us systematically improve the AI.

Aditya Kabra

Co-founder, Applicative AI

I now have a systematic process to make my AI better.

So I've been building these AI tools and chatbots for a while. They worked okay most of the time, maybe 80%, and I thought that was pretty good. But then I started wondering, what if I actually needed something reliable - production level? So, I took this course by Hamel and Shreya on AI evals. Turns out the entire point is to look at the data. It's actually about sitting down and really looking at what your AI is doing, like, actually examining the input, output data. They taught me this whole process: test your AI with different types of questions, write down everything that goes wrong, then figure out patterns. Sounds boring, but here's the thing, I was missing so many errors I didn't even know existed. Now I actually know the process to identify when my AI tools work and when they don't. Instead of just crossing my fingers and hoping, I have a real process to make them better. It's a lot of manual work, but honestly, it's the only way to build something people can actually depend on.
So I've been building these AI tools and chatbots for a while. They worked okay most of the time, maybe 80%, and I thought that was pretty good. But then I started wondering, what if I actually needed something reliable - production level? So, I took this course by Hamel and Shreya on AI evals. Turns out the entire point is to look at the data. It's actually about sitting down and really looking at what your AI is doing, like, actually examining the input, output data. They taught me this whole process: test your AI with different types of questions, write down everything that goes wrong, then figure out patterns. Sounds boring, but here's the thing, I was missing so many errors I didn't even know existed. Now I actually know the process to identify when my AI tools work and when they don't. Instead of just crossing my fingers and hoping, I have a real process to make them better. It's a lot of manual work, but honestly, it's the only way to build something people can actually depend on.

Ryan Lingo

@RyanLingo

Wanted to share my experience with @sh_reya @HamelHusain's maven.com/parlance-labs/evals. It is the most in-depth, practical content on LLMs I've encountered. Lectures and exclusive book are 🔥. Essential if you're building AI applications.
Wanted to share my experience with @sh_reya @HamelHusain's maven.com/parlance-labs/evals. It is the most in-depth, practical content on LLMs I've encountered. Lectures and exclusive book are 🔥. Essential if you're building AI applications.

Kashyap Amin

@kashyapramin

Participating in the meticulously designed and comprehensive “AI Evals For Engineers & PMs” course, led by @HamelHusain and @sh_reya, was a significant learning and enlightening experience.
Participating in the meticulously designed and comprehensive “AI Evals For Engineers & PMs” course, led by @HamelHusain and @sh_reya, was a significant learning and enlightening experience.

adi

@adidoit

Error Analysis as taught by @sh_reya and @HamelHusain in maven.com/parlance-labs/evals is literally 1000x ROI for AI product builders Biggest learning is applying an ML mindset - train, test, holdout data sets and improving through structured error analysis (i.e., LOOK AT YOUR DATA)
Error Analysis as taught by @sh_reya and @HamelHusain in maven.com/parlance-labs/evals is literally 1000x ROI for AI product builders Biggest learning is applying an ML mindset - train, test, holdout data sets and improving through structured error analysis (i.e., LOOK AT YOUR DATA)

Dmitry Labazkin

@labdmitriy

I'm currently in the middle of the fantastic "AI Evals For Engineers & PMs" course by @HamelHusain and @sh_reya! maven.com/parlance-labs/evals It's genuinely one of the most impactful courses I've ever taken. Biggest takeaways? In 🧵
I'm currently in the middle of the fantastic "AI Evals For Engineers & PMs" course by @HamelHusain and @sh_reya! maven.com/parlance-labs/evals It's genuinely one of the most impactful courses I've ever taken. Biggest takeaways? In 🧵