Auto Scaling Endpoints Quiz

5 questions Pass: 70% +25 pts

Quiz covering ML Infrastructure

Auto Scaling Endpoints Quiz

5 questions | Pass: 70% | Earn 25 points

Questions in this quiz

A preview of the 5 questions covered. Start the quiz above to answer them, check your score, and read the explanations.

  1. 1

    What is the primary purpose of implementing auto-scaling for machine learning model endpoints?

  2. 2

    Which metric is most commonly used as a trigger for scaling ML inference endpoints when latency is a critical business requirement?

  3. 3

    What is the main risk of setting a 'scale-in' cooldown period that is too short?

  4. 4

    When deploying a deep learning model that requires a GPU, why is 'target tracking' scaling often preferred over simple 'step scaling'?

  5. 5

    You observe that your auto-scaled endpoint frequently fails to handle sudden 'burst' traffic despite having a high maximum instance limit. What is the most likely architectural bottleneck?