Evaluation Metrics Quiz

5 questions Pass: 70% +25 pts

Quiz covering Model Evaluation

Evaluation Metrics Quiz

5 questions | Pass: 70% | Earn 25 points

Questions in this quiz

A preview of the 5 questions covered. Start the quiz above to answer them, check your score, and read the explanations.

  1. 1

    Which evaluation metric is commonly used to measure the similarity between a model-generated summary and a reference summary based on n-gram overlap?

  2. 2

    When evaluating a generative model using Perplexity (PPL), what does a lower score indicate?

  3. 3

    Why is it often insufficient to rely solely on automated metrics like BLEU or ROUGE for evaluating large language models (LLMs) in creative writing tasks?

  4. 4

    You are evaluating a classifier for a highly imbalanced dataset where the positive class is rare. Which metric would provide the most misleading assessment of model performance?

  5. 5

    In the context of 'LLM-as-a-judge', what is the primary risk associated with using a stronger model (e.g., GPT-4) to evaluate the outputs of a smaller model?