Question 1
A company has been modernizing on cloud-native platforms for the past few years and has been running some small consumer support utilities on their production NKP cluster. After a thorough testing and QA cycle with simulated workloads on a development cluster, the company is ready to bring their online retail application into the fold. While they have sufficient system resources to scale the NKP cluster properly from a performance standpoint, they also want to ensure they properly scale their monitoring stack's resource settings to retain a sufficient amount of data to see how overall system resource utilization trends for the NKP cluster over several months' time with the added workloads. Which NKP Platform Application component should the company be most concerned with adjusting, and how should their Platform Engineer adjust it?
The NKP Platform Application component the company should be most concerned with adjusting is Prometheus, by increasing its storage retention and compute resources.
Why this is correct:
- Metrics Retention: Prometheus stores time-series metrics data; to retain several months of historical data, the data storage capacity must be significantly increased
- Data Volume Increase: With the new retail application workloads added to the cluster, Prometheus will collect substantially more metrics, requiring additional disk storage
- Resource Scaling: Prometheus itself (CPU and memory) needs to be scaled up to handle the increased volume of metrics collection and queries across multiple workloads
- Long-Term Trending: To see system resource utilization trends over several months, retention policies must be configured to keep data for that duration
- Monitoring Stack Capacity: Prometheus is the core component of the monitoring stack that handles metrics storage and retrieval
Prometheus is the critical component to adjust when scaling monitoring infrastructure for increased workload and longer data retention requirements.
