Consolidation and automation of CERN's cloud monitoring dashboards
Abstract
This project addresses the challenges in managing CERN’s extensive cloud monitoring infrastructure, which relies on over 100 Grafana dashboards to track the health of thousands of virtual and physical machines. To overcome limitations in version control, styling consistency, and bulk editing, we implemented a "dashboards as code" solution using JSONnet and Grafonnet. This approach is enhanced by a custom abstraction layer that simplifies specification and enforces consistency through defaults. A CI/CD pipeline automates the compilation and deployment of dashboards to development and production instances, establishing a robust Git-based versioning system that improves maintainability and scalability for CERN’s cloud monitoring.
Full text
CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS August 2025 AUTHOR(S): Dennis Alexander Mertens Velasquez Department of Advanced Computing Sciences (DACS), Maastricht University SUPERVISOR(S): Richard Bachmann Varsha Bhat Ulrich Schwickerath
CERN openlab Report 2025 PROJECT SPECIFICATION The CERN Cloud Infrastructure service administrates CERN’s private cloud, which provides computation resources to physics analysis and IT services for the whole organisation. At present, this includes around 14.000 virtual machines and more than 10.000 physical ones. The service is built upon OpenStack, a leading open-source software used around the world. As part of daily operations, it is vital to continuously monitor the health and performance of our systems. This not only gives insight to operators but also provides our users with information about their resources. This monitoring is accomplished using CERN’s monitoring platform, with the collected data being displayed by well over 100 Grafana dashboards1. The goal of this project is to consolidate our many monitoring dashboards into a single software-defined source, which will be easier to update and expand to accommodate future needs. You will write and deploy the Continuous Integration (CI) setup for the creation of such dashboards, and start the development effort of replicating our existing dashboards with JSONnet/Grafonnet code. 1Example of a Grafana dashboard: https://monit-grafana-open.cern.ch/d/8f4TgzF7z/cern-openstackoverview?orgId=16 CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS 1
CERN openlab Report 2025 ABSTRACT This project addresses the challenges in managing CERN’s extensive cloud monitoring infrastructure, which relies on over 100 Grafana dashboards to track the health of thousands of virtual and physical machines. To overcome limitations in version control, styling consistency, and bulk editing, we implemented a "dashboards as code" solution using JSONnet and Grafonnet. This approach is enhanced by a custom abstraction layer that simplifies specification and enforces consistency through defaults. A CI/CD pipeline automates the compilation and deployment of dashboards to development and production instances, establishing a robust Git-based versioning system that improves maintainability and scalability for CERN’s cloud monitoring. CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS 2
CERN openlab Report 2025 TABLE OF CONTENTS 1 INTRODUCTION 4 1.1 DataSources ..................................... 4 1.2 Panels ......................................... 5 1.3 Versioning System, Consistent Styling, & Bulk Editing . . . . . . . . . . . . . . 6 2 DASHBOARDS AS CODE 7 2.1 Abstractions With Grafonnet . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 2.2 TheCI/CDPipeline ................................. 8 3 DISCUSSION & CONCLUSION 9 3.1 AlternativeApproach................................. 9 3.2 FutureWork...................................... 10 4 Special Thanks 10 CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS 3
CERN openlab Report 2025 1 INTRODUCTION Grafana [5] is a visualization and monitoring platform, able to connect to multiple data sources and produce interactive and non-interactive plots both in real time or static ones. This way, aspects of a complex system can be broken down into relevant constituents and monitored. It is possible to run one’s own Grafana instance. This project in particular is concerned with two instances hosted at CERN: monit-grafana.cern.ch and monit-grafana-dev.cern.ch, which we refer to as monit-instance and dev-instance respectively. We use the dev-instance for testing and the monit-instance for deployment. Everything that gets deployed in the monit-instance, has to be tested in the dev-instance first. This separation ensures we can test changes thoroughly before they become visible to users. Grafana instances follow a hierarchical structure, where the root is the instance itself, followed by the organizations hosted within, and finally, arbitrarily deep nested interleaves of folders and/or dashboards. In principle, this hierarchy works like any file system, except that we have two novel concepts: instance and organization. Figure 1illustrates this idea. Figure 1: Left, a conceptual diagram of Grafana’s file system hierarchy. Right, an arbitrary example of an instance with two organizations and some folder under each. A Grafana instance may only contain organizations, and an organization may only contain folders or dashboards. Finally, a folder may contain either folders or dashboards. The reason for this separation is mainly for security and management. Organizations offer complete separation in terms of permissions, configuration, and activity tracking. For example, each organization has access to its own list of data sources. 1.1 Data Sources Though any Grafana instance is mainly concerned with plotting, it needs to interact with multiple sources providing data about the system being monitored. To that end, the platform allows users to connect data sources via a URL. Once connected, users are able to specify what kind of data to retrieve by providing a query. Even though each data source follows its own CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS 4
CERN openlab Report 2025 query language, this layer of abstraction allows for standardized integration with the rest of the platform. Any data source can be identified by a unique name.2For example, Figure 2shows a data source listed in the dev-instance. The name is associated with a type, in this case, Prometheus, and a URL. Figure 2: The monit-prom-lts - openstack Prometheus data source. Naturally, for any given data source, it is possible to specify a myriad of different queries by combining atomic statements. For instance, Listing 1and Listing 2show how to query different data from the same data source. sum( irate( node_network_receive_bytes_total{ submitter_hostgroup="cloud_networking/controller/backend/ovn_sb_relay/$region" }[$__rate_interval] )* 8 ) Listing 1: Example of a query for the monit-prom-lts - openstack data source. sum( irate( node_network_transmit_bytes_total{ submitter_hostgroup="cloud_networking/controller/backend/ovn_sb_relay/$region" }[$__rate_interval] )* 8 ) Listing 2: Example of another query for the monit-prom-lts - openstack data source. Through queries, one can specify the kind and shape of data to pull from a given data source to be displayed in a panel. 1.2 Panels A Grafana dashboard is made of panels, and a panel is simply an umbrella term for a component that can display some data pulled from a data source. For example, the queries in Listing 1 and Listing 2are from the panel shown in Figure 3; a timeseries panel. 2By far, the most frequently used data source during development was monit-prom-lts - openstack. CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS 5
CERN openlab Report 2025 Figure 3: A timeseries panel displaying data pulled from monit-prom-lts - openstack using the queries in Listing 1and Listing 2—one for each timeseries. There are different types of panels, each designed to plot different types of data. In particular, timeseries, tabular, and stat3panels were the most relevant for this project. A static panel merely displays the data retrieved through a user-specified query on load. If a refresh rate is specified, the panel may become dynamic if the data returned by the query changes over time. By contrast, a panel becomes interactive when a user declares template variables to be used in queries. For example, if we look at Listing 1and Listing 2, we can see that both refer to a region variable by "$region". If we look at Figure 3, at the top-left corner, we find said region variable assigned to "pdc". When a viewer assigns a different value, all panels referring to region automatically update their plots by replacing the referenced variable name with its assigned value and running the query again. 1.3 Versioning System, Consistent Styling, & Bulk Editing Grafana’s versioning system allows tracking the chronology of changes applied to each panel in a dashboard, and their authorship. However, unlike specialized versioning systems, it does not provide a means for branching and merging. For example, Git [2] provides these commodities, enabling teams to make contributions in parallel and manage the complexity that comes with this. Figure 4: A dashboard’s version history in Grafana. Furthermore, upon creation, each panel in a dashboard is independent in the sense that its styling parameters4are free and completely up to the author’s preference. For example, if it 3Astat panel only shows a number, which may be static, interactive, or dynamic. 4e.g., color pallete, color thresholds, font, etc. CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS 6
CERN openlab Report 2025 is decided that plots showing availability counts of some resource, such as memory, processors, etc., should change colors as they approach critical limits, the only way to enforce consistency is through policy. Grafana does not provide a way to ensure every panel and dashboard uses the same color palette. Similarly, since each panel is independent, bulk editing is not possible through Grafana’s user interface.5For instance, if it is decided that a particular color for signaling (visually) needs to be changed, the only way of effecting said change is by going over each panel—one by one. More importantly, if multiple dashboards have to be migrated to a different organization or instance, it is not guaranteed that the data sources connected to each dashboard will match. Hence, it may be necessary to update them.6 2 DASHBOARDS AS CODE Considering the three aforementioned main limitations, the concept of dashboards as code was proposed in order to circumvent Grafana’s limitations by using Git for versioning by Ewoud Ketele [8]. Since Grafana stores dashboards and all components within, as nested JSON strings, it is in principle possible to simply export all dashboards and manage them with Git. However, the issue is that a single JSON string describing a dashboard is exceedingly long7and difficult to interpret, since such strings are generated automatically by user interface (GUI). Hence, in order to keep the process ergonomic, it is necessary to add a layer of abstraction. 2.1 Abstractions With Grafonnet Grafonnet [6] is the official library for writing dashboards as code. It is built on JSONnet [3]; a superset of JSON. With this library, we can—in principle—write a dashboard’s specification using human-readable terms, and then unroll it into a JSON string that Grafana can understand. In practice, dashboard specifications written in Grafonnet still incur a significant overhead in terms of the number of characters that need to be typed and the level of detail the user is required to provide. Furthermore, the library does not enable us to enforce consistent styling, nor does it provide any advantages regarding bulk editing not already present in JSON. In light of the aforementioned shortcomings, we implemented yet another layer of abstraction atop Grafonnet. Figure 5shows the specification for a single panel in JSON, in Grafonnet, and finally in our abstraction layer. In our library of abstractions, consistent styling is not enforced per se, but encouraged. By design, it is possible to deviate from the agreed standards. However, one would have to make the extra effort of specifying parameters different from defaults. Similarly, bulk editing is facilitated by the fact that many features and/or characteristics of dashboards are inherited from said defaults. Hence, if a change must be effected that is supposed to be applied everywhere, changing the defaults is often enough. 5While Grafana does provide some tools for bulk editing, these are tailored for specific aspects and barely cover the full extent of changes a user may apply. 6Another example is when a data source itself becomes deprecated. 7For example, we have one dashboard with 19 panels that results in a JSON string with 2365 lines, 5315 words, and 89266 characters. CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS 7
CERN openlab Report 2025 Figure 5: An example of the evolution from JSON to Grafonnet, and finally to our own abstraction layer. Naturally, since the original dashboards are exported as raw JSON strings by Grafana, we have to port them to our library by hand. This is an arduous task, considering the number of dashboards. Nonetheless, this is a fruitful dynamic process. If we find a pattern that cannot be mapped to our abstractions, we have to implement new ones. Throughout this process, our library of abstractions becomes richer and thus is able to implement an increasingly wider range of panels. It was a serendipitous discovery that, as we added more abstractions to our library, it became easier for a large language model (LLM) assistant to port raw JSON strings to code.8Combined with our CI/CD pipeline’s ability to check and report errors at various stages, in principle, it should be possible to automate the whole process.9 2.2 The CI/CD Pipeline The conversion from our layer of abstraction to JSON is carried out automatically by a CI/CD pipeline in our GitLab repo10. We have two branches: master and qa. In principle, the master branch is connected to the monit-instance, and the qa branch to the dev-instance. Every commit and push to either branch triggers the pipeline, which then proceeds to compile and deploy the dashboard specifications to the appropriate Grafana instance. Figure 6shows a high-level diagram of the processing steps within the pipeline. Figure 6: Illustration of the CI/CD pipeline when processing dashboard specifications. 8We used Claude Sonnet 3.7 Thinking, and provided JSONnet and Grafonnet documentation in the context. 9However, this is out of scope for this project. 10You can find our repo at https://gitlab.cern.ch/dmertens/damvcipj. CONSOLIDATION & AUTOMATION OF CERN’S CLOUD MONITORING DASHBOARDS 8