Cloud Scale and the Architecture of Real-Time Monitoring
How automated cloud resources expand on live demand and the continuous diagnostic tools required to keep enterprise systems stable
In my August column, I explored cloud resource monitoring—placing emphasis on private clouds, data centers and cloud structure—and introduced the needs and requirements for scalability and monitoring. As a brief review and to set the tone for the conclusion, the definition of scale has multiple interpretations.
Scale
“Scale” can mean “size,” or the dimensional physical property usually associated with ground space. Scalability refers to that space’s ability to (easily) expand to house more space, equipment or growth.
For instance, in computing, Lenovo describes scale as a system’s ability to handle growing amounts of work, data or users without losing performance or stability.
When discussing storage—such as disk drives or arrays—core scaling is described as vertical scaling, or “scaling up.” To scale up is to add more power, such as faster CPUs or GPUs, more memory or larger storage capacity, especially on a single machine or server.
Horizontal scaling, or “scaling out,” means to add more individual computers or servers to share the workload.
Other meanings in computing include “UI scaling,” image scaling or using a mathematical multiplier to convert a range of numbers or values to fit a specific digital format.
Cloud Scaling
Cloud scaling is the ability of cloud computing infrastructure to dynamically adjust resources—like computing power, storage or network bandwidth—either up or down to match changing workloads.
The professional video industry's #1 source for news, trends and product and tech information. Sign up below.
Unlike traditional, chassis-centric physical servers (or frameworks), which typically require manually adding physical components as in hardware upgrades, cloud environments scale almost instantly using automated software.
Fig. 1 shows a breakdown of how clouds scale, the mechanisms they use and how that differs from traditional infrastructure.
Clouds scale, in part, like computers or storage (i.e., horizontally, vertically or diagonally)—as described earlier—with diagonal scaling adding a dynamic, real-time element.
Dynamic Scaling
All clouds need to behave dynamically—that is, to expand or contract in reaction to the needs or demands of clients using their services. Diagonal scaling is a hybrid of vertical and horizontal scaling in which a system first scales horizontally, adding more instances until it hits a limit.
Those instances are then upgraded vertically to larger sizes to handle even more intensive workloads. Often, the larger size means bringing on additional equipment that waits in standby until the management system calls upon it, or when the software that requires those services demands it.
Road maps, associated in part with artificial intelligence, will be a system’s ability to “self-scale.” Here, a cloud enables DevOps pipelines to automatically allocate computing resources based on real-time traffic and, as demands spike, the system employs “auto scaling” to ensure the application doesn’t crash.
Monitoring
Both human and machine-based monitoring systems must interchange and, in turn, support commercial and technical scalability to accommodate rapid growth and evolving needs. Bugs will trigger other routines that can automatically change flows, spin up alternative code sets and even halt the testing process while another fix or solution can be deployed to serve as corrective actions.
This further enables the active system to avoid large up-front commitments and to prefer solutions that scale easily as needed. Alternative solutions can then demonstrate immediate value without extensive proof of concept.
“Technical scalability” promotes easy expansion without significant additional effort or cost—up front or downstream.
Service Mechanisms
Automated systems are engineered into the cloud “package” specific to the service provider and are typically key marketing and performance factors. Such systems depend on automation to make these transitions “seamless” and usually “hands-off” for users, except in specific cases in which the cloud model has been contracted to scale based upon various factors.
Those factors can include overruns or the need to meet time-constrained deadlines or demands, using tools such as “load balancers,” which act as traffic directors that sit in front of the application. As new servers are added (via horizontal scaling), the load balancer (Fig. 2) automatically starts routing a portion of the incoming user traffic to them, so that no single server gets overwhelmed.
Another tool available is serverless computing. This form of scaling (such as AWS Lambda or Google Cloud Functions) is where developers write code and the cloud provider handles everything else. If zero people use the app, zero servers run. If 10,000 people trigger the app at the same second, the cloud instantly runs 10,000 parallel copies of the code.
Remember that scaling is the system’s capability to handle growth, focused on long-term capacity planning or meeting a specific high-demand baseline.
However, while often used interchangeably with scaling, elasticity is the speed and fluidity with which the system handles both spikes and drops in real time (focused on immediate adaptation to fluctuating demands). Fluidity means the smooth, seamless and automated movement of resources as they expand and contract in real time. Zero human intervention is the action where cloud (and AI) systems need to shrink and expand automatically based on live demand, without needing a human to manually approve or configure new servers.
Instant adaptability refers to “continual availability,” “load predictability/reactivity” and “reactional monitoring”—each are key parts of a cloud’s enterprise solution. For example, when an app goes viral on social media, compute resources must flow into the system within seconds to handle the spike. When the “viral” rush ends, those resources must drain away just as fast.
Cost efficiency is where cloud providers advertise that “you only pay for what you use.” The system is engineered with a fluid infrastructure and automatically changes the scale of service, with the cost dynamics (aka billing) remaining flexible. As a result, the cloud provider seldom leaves expensive, unused servers sitting idle when traffic is modest, but it is ready to provision those servers and bring them online nearly instantly.
A robust monitoring platform should also have a large ecosystem of free, vendor-backed integrations. Open-source solutions often require ongoing maintenance and custom integration efforts. Many startups lack comprehensive, supported integrations. They prefer vendor-backed integrations from stable, established companies to ensure support and reliability.
Up-to-Date Documentation
Good and current documentation is essential for quick setup and ongoing maintenance, whether for cloud-only or AI applications. Clear, up-to-date, software-centric documentation supported in the cloud will allow for both manual and automated diagnostics reporting and analysis. Complete documentation should be detailed, accessible, and include human-readable screenshots. It should be linked to those subroutines for quick reporting that can be married into an AI training process.
By employing a large language model (LLM), testing and monitoring systems can quickly locate and isolate both the code and the instructions to speed diagnosis and corrective identification. By using an LLM library of diagnostic processes, anyone can be tagged into the testing process, alleviating the need for specialized subject matter experts during nominal testing, unnecessarily spinning up changes or manipulating core systems.
Configurable Monitors and Alerts
Monitoring tools should allow for easy configuration of persistent, reliable monitors and alerts, i.e.:
- Monitors should adapt to infrastructure changes without breaking.
- Tag-based queries should enable comprehensive and flexible alerting.
- Persistent alerts should ensure continuous monitoring regardless of environment changes.
Real-Time and Collaboration Integration
Alerts must be efficiently routed to the right stakeholders, consisting of project managers, SMEs and software owners. Avoid alerts buried in email or unsupported notification systems. They must integrate with downstream tools like PagerDuty, ServiceNow or customized workflows. Centralized monitoring should support combined infrastructure and DevOps process alerts for faster troubleshooting.
Developer Access Without Infrastructure Exposure
Developers often need access to performance data without risking security or compliance violations. Key practices include:
- Providing secure, read-only access to logs, metrics and traces.
- Giving developers access to relevant troubleshooting data without having to wade through cloud infrastructure access policies.
- Creating safe environments where developers can monitor code and deployment integration securely to speed infrastructure development and flow, reducing risks from failure or security compromises.
Monitoring tools should integrate relevant data from all phases of the DevOps cycle—often in real time—rather than requiring teams to linger on lengthy reports that require multiple sign-offs before moving through each step of the diagnostic model set.
Comprehensive enterprise DevOps monitoring requirements for the cloud will likely evolve through each cycle and should be thought of in future uses and applications. This is especially useful where a system can self-train cloud servers to move through effective procedures, minimizing no results and unnecessary repetitive processes.
Karl Paulsen recently retired as a CTO and has regularly contributed to TV Tech on topics related to media, networking, workflow, cloud and systemization for the media and entertainment industry. He is a SMPTE Fellow with more than 50 years of engineering and managerial experience in commercial TV and radio broadcasting. For over 25 years he has written on featured topics in TV Tech magazine—penning the magazine’s “Storage and Media Technologies” and “Cloudspotter’s Journal” columns.