Evaluate Models and Apps Quiz

5 questions Pass: 70% +25 pts

Quiz covering Build Generative Applications

Evaluate Models and Apps Quiz

5 questions | Pass: 70% | Earn 25 points

Questions in this quiz

A preview of the 5 questions covered. Start the quiz above to answer them, check your score, and read the explanations.

  1. 1

    Which of the following metrics is most commonly used to evaluate the factual accuracy of a generative AI application's output against a ground-truth reference?

  2. 2

    You are building a RAG (Retrieval-Augmented Generation) application and notice the model frequently hallucinates when the retrieved context is irrelevant. Which evaluation approach best addresses this?

  3. 3

    When evaluating an agentic workflow, why is 'Model-based Evaluation' (using an LLM to grade another LLM) often preferred over static unit tests?

  4. 4

    In the context of evaluating a production-grade generative application, what is the primary purpose of 'A/B Testing'?

  5. 5

    You are performing a 'Golden Dataset' evaluation. If your model achieves high performance on the golden dataset but performs poorly in production, what is the most likely cause?