A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment
Abstract
The Tomcat service has been operating in Kubernetes for the past four years. With the advent of new technologies, deployment methods have also evolved. In particular, some applications and components of the current infrastructure are already managed declaratively via Terraform and a GitOps controller. This project aims to explore how ClusterAPI and Crossplane can employ a declarative approach to manage and deploy Kubernetes clusters in a cloud-native way. Another important aspect of the project is interoperability; the proposed solution should be developed and tested both on-premise and on Oracle Cloud Infrastructure, further enhancing disaster recovery.
Full text
A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment September 2024 AUTHOR: Mouad El Haouari (moua[email protected]) SUPERVISORS: Antonio Nappi ([email protected]) Artur Wiecek ([email protected])
Mouad El Haouari - CERN Openlab Report // 2024 2 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Acknowledgement I would like to express my sincere gratitude to CERN Openlab for providing invaluable opportunities to students from around the world, including those from non-member state countries. Their commitment to boosting students’ careers with valuable internships is truly commendable. I extend my heartfelt thanks to my supervisors, Antonio and Arthur, for their support, expert guidance, and patience throughout my internship. Their mentorship has been a huge plus to my learning and professional growth. I am also grateful to my other teammates, Hubert, Adrian, and Daniel, who created a welcoming and supportive environment. Their kindness and willingness to assist made my internship experience both enjoyable and enriching. Lastly, I want to acknowledge my fellow summer students, with a special mention to Tomasz and Victor, who were in the same team with me, and with whom I shared many memorable moments. The camaraderie we developed added a special dimension to this internship, creating friendships and connections that I will cherish long after this experience. This summer at CERN has been an extraordinary journey of learning and personal growth. Hopefully, our paths will cross somewhere in the future. ~ Mouad ~
Mouad El Haouari - CERN Openlab Report // 2024 3 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment ABSTRACT The Tomcat service has been operating in Kubernetes for the past four years. With the advent of new technologies, deployment methods have also evolved. In particular, some applications and components of the current infrastructure are already managed declaratively via Terraform and a GitOps controller. This project aims to explore how ClusterAPI and Crossplane can employ a declarative approach to manage and deploy Kubernetes clusters in a cloud-native way. Another important aspect of the project is interoperability; the proposed solution should be developed and tested both on-premise and on Oracle Cloud Infrastructure, further enhancing disaster recovery.
Mouad El Haouari - CERN Openlab Report // 2024 4 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment TABLE OF CONTENTS 1. Introduction ..........................................................................................................................................5 2. Existing workflows...............................................................................................................................5 a. Kubernetes clusters management......................................................................................................5 b. Packages and Applications deployment on Kubernetes clusters .......................................................7 3. Objectives .............................................................................................................................................8 4. Investigations .......................................................................................................................................9 a. Overview ..........................................................................................................................................9 b. CERN’s OpenStack ........................................................................................................................ 10 i. ClusterAPI & Magnum ............................................................................................................... 10 ii. ClusterAPI & the vanilla approach: building the image ............................................................. 10 iii. ClusterAPI & the vanilla approach: installing ClusterAPI ...................................................... 11 iv. ClusterAPI & the vanilla approach: testing cluster creation ................................................... 11 v. Crossplane .................................................................................................................................. 14 c. Oracle Cloud Infrastructure ............................................................................................................ 15 i. ClusterAPI .................................................................................................................................. 15 ii. Crossplane .................................................................................................................................. 18 5. Proposed Approach ............................................................................................................................ 18 6. The one button deployment ................................................................................................................ 19 7. Migration ............................................................................................................................................ 21 a. Translating Terraform to Crossplane .............................................................................................. 21 b. Secret management ......................................................................................................................... 24 8. Conclusion & Perspectives ................................................................................................................. 27 9. References .......................................................................................................................................... 29 10. Appendix ........................................................................................................................................ 32
Mouad El Haouari - CERN Openlab Report // 2024 5 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment 1. Introduction With the widespread and rapid adoption of Kubernetes as the framework for deploying applications, there has been a growing need for tools and services to manage Kubernetes clusters and deploy applications on them, especially at a large scale. Both on-premise and public cloud providers have had to adapt to this shift, offering both self-managed and managed services specifically tailored for provisioning and managing Kubernetes clusters. Infrastructure as Code (IaC) tools also play a significant role in this new era by supporting these Kubernetesspecific services. Some IaC tools were designed from the ground up to be cloud-native, such as Crossplane, and run within Kubernetes, offering a unique approach to managing both infrastructure and applications. Others have focused specifically on the lifecycle management of Kubernetes clusters, such as ClusterAPI. During my internship at CERN as an Openlab summer student, I had the privilege to work within the IT department in the team responsible for providing hosting infrastructure based on Kubernetes for SSO (Single Sign On) and many others critical java applications for Finance and Administrative Processes (FAP) and engineering (EN) departments such as EDH (Electronic Document Handling) and EDMS (Engineering & Equipment Data Management Service). The team’s Kubernetes clusters are all currently running within CERN on-premise OpenStack cloud and the current workflow for their management leverages Terraform as the IaC tool in addition to Gitlab CI/CD to automate cluster operations. My project consisted of investigating how we can use Kubernetes-native IaC tools (ClusterAPI and Crossplane) in addition to GitOps practices (ArgoCD) for managing Kubernetes clusters on both on-premise CERN OpenStack cloud and Oracle public cloud (OCI) as a replacement to the current terraform and GitLab CI/CD approach. The aim of my supervisor was to have a unified way of managing both infrastructure an application through Kubernetes object. Another aspect is effective disaster recovery, with Kubernetes-native approach we can eliminate a lot of dependency present in the current approach and recover our clusters faster through a KinD or MiniKube cluster that can be run locally. 2. Existing workflows a. Kubernetes clusters management Our Kubernetes clusters are managed in CERN Openstack cloud through the Magnum [1] module which makes container orchestration engines (COE) such as Kubernetes available as first-class resources in OpenStack.
Mouad El Haouari - CERN Openlab Report // 2024 6 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment A cluster is created using a ClusterTemplate [2] which is a collection of parameters that describe how a cluster can be constructed. Some parameters are relevant to the infrastructure of the cluster, while others are for the particular COE. Labels [3] are one of these parameters. They are a set of key/value pairs used for specifying additional parameters such us the packages that need to be installed on the cluster and their versions. In CERN's OpenStack cloud, cluster templates are already defined and provided by the Kubernetes team. These templates include CERN-specific labels [4] in addition to those defined in Magnum's documentation. For each recent Kubernetes minor version, the Kubernetes team offers two types of templates: • single-AZ: clusters created from this type of templates will have all their nodes located in one availability zone that user can define through a label. • multi-AZ: used for creating a highly available cluster where nodes are spread across multiple availability zones. In addition to multi-AZ templates, Nodegroups [5] are also used to enable high availability. Nodegroups are another first-class resource provided by the Magnum module. They represent a subset of nodes belonging to the same cluster that have similar properties and they can be assigned to a specific availability zone. By default, when a cluster is created it already has two Nodegroups, ‘default-master’ for master nodes and ‘default-worker’ for worker nodes. In order to have a highly available cluster, user needs to explicitly create additional Nodegroups and spread them across multiple availability zone. My team’s approach consisted of using the templates provided by the Kubernetes team and overriding their labels with custom ones in order to provide Kubernetes environments that allow the execution of the applications we are responsible for. Since there were multiple applications under my team’s responsibilities, and for each of these applications there were test, dev, and prod clusters. Managing clusters manually through the OpenStack CLI was not an option, especially as these clusters were spread across multiple OpenStack projects. This is why the management system in place relied primarily on: • Terraform, an IaC tool that allows cluster provisioning and management using a declarative configuration language (HCL - HashiCorp Configuration Language). • GitLab CI/CD pipeline, which automates operations and reflect clusters specifications declared in the GitLab repository in the OpenStack cloud.
Mouad El Haouari - CERN Openlab Report // 2024 7 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Figure 1. Current workflow for managing Kubernetes clusters lifecycle As demonstrated in the figure above, the workflow functions as follows: • when user pushes Terraform code describing a cluster specifications (see Figure 10 in Appendix for an example) to the Git repository, the GitLab CI/CD pipeline will run on each push event and executes some Terraform commands (validate, plan, apply) triggering by that Terraform to send an API request to Openstack Magnum module which takes care of provisioning/upgrading the cluster. • Once the clusters are created, the magnum module will respond with cluster’s state which is then persisted by terraform in the backend (Postgres database) and the kubeconfig file of the cluster is saved in CERN secret store (Tbag/Teigi). Despite the advantages of this approach, such as declarative cluster management and version control, there were two main drawbacks: • Terraform is not Kubernetes-native and is not straightforward to maintain, as it uses HCL as a configuration language, while my team's work primarily involves Kubernetes objects/resources (YAML). • It is not effective for disaster recovery, as there is a strong dependency on external CI/CD tools and a database. b. Packages and Applications deployment on Kubernetes clusters In addition to the infrastructure layer, the application layer management has its own workflow. Applications, along with additional packages like the monitoring stack, are managed using ArgoCD (a GitOps implementation), which automates deployment by continuously monitoring changes in Git repositories that host applications specifications (YAML code). ArgoCD ensures consistent and automated updates by synchronizing the desired applications state with Kubernetes clusters whenever changes are detected. To further enhance maintainability, YAML code defining applications and packages is organized in a Kustomize [6] format. For each application, there is a base containing the common YAML code among all
Mouad El Haouari - CERN Openlab Report // 2024 8 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment the clusters and an overlay for each cluster that patches the base by adding, replacing, or removing code. The code ultimately deployed in each cluster is the result of applying the overlay patches to the base. To connect the infrastructure and application layers, clusters created using the Terraform & GitLab CI/CD workflow must be registered as target clusters in the ArgoCD instance so that each application gets deployed to its designated cluster. Since both layers involve fundamentally different workflows, this process of registering target clusters is performed manually using the ArgoCD CLI tool. 3. Objectives With the drawbacks of the current approach in mind, our main objective was to propose a unified, cloudnative method for managing Kubernetes clusters across both on-premises (OpenStack) and public cloud environments (Oracle Cloud Infrastructure). By unified, we mean using Kubernetes objects to manage applications and infrastructure on both on-premises and public cloud environments. By cloudnative, we mean that our IaC tool will run within Kubernetes and execute a reconciliation loop, constantly trying to align the current state of the clusters with their desired state specified in the Kubernetes resources. Figure 2. A unified and cloud-native approach for managing Kubernetes clusters on hybrid cloud Since GitOps (ArgoCD) has already been adopted for managing applications, we also wanted to integrate the Kubernetes-native approach with the current ArgoCD workflow. This way, when new clusters are created, they are automatically registered as target clusters within ArgoCD, making it easy for us to hydrate them with the required packages. This setup will allow us to manage everything using Kubernetes and will eliminate many dependencies present in the current approach, leading to a more effective disaster recovery process that only requires: • A Kubernetes cluster: which can be running on a local machine using KinD [7] or Minikube [8]
Mouad El Haouari - CERN Openlab Report // 2024 9 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment • An ArgoCD instance: running on that cluster After that, all that is left is to point to the repository (which can be hosted locally, on GitLab, or GitHub) containing the YAML files necessary to deploy the IaC tool, the required providers, and the Kubernetes resources defining the clusters. 4. Investigations a. Overview To achieve the objectives outlined earlier, we conducted some investigations to explore the available cloudnative IaC tools and assess how they integrate with CERN Openstack and Oracle Cloud Infrastructure (OCI). We primarily focused on two tools that seemed to align with our use case: • ClusterAPI: which is focused solely on the lifecycle management of Kubernetes clusters. • Crossplane: a general-purpose IaC tool similar to Terraform but designed to run on Kubernetes. In fact, it is actually based on Terraform under the hood (providers are created using a code generation tool called Upjet [9] that allows code generation of Crossplane providers from Terraform providers). The table below highlights the results of our investigations. IaC tool / Cloud OpenStack Oracle Cloud Infrastructure ClusterAPI - Does not support creating Kubernetes clusters through Magnum - It can be used as a backend for Magnum instead of Heat [10] service - Creating clusters using services other than Magnum does not work with CERN Openstack cloud due to networking constraints Supports creating both self-managed and managed clusters Crossplane Supports creating clusters through Magnum No provider available (there is one, but it has been archived and it is not officially recognized by Crossplane)
Mouad El Haouari - CERN Openlab Report // 2024 16 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment This taint was still present and not removed from the nodes because the component responsible for this, OCI Cloud Controller Manager, was not installed on the cluster after its creation. CCMs [27] (Cloud Controller Managers) are very common when dealing with Kubernetes on public cloud. Since I mostly worked with Kubernetes in on-premise environments, I missed it. Test 2 Creating a self-managed cluster with addons (using cluster-template-ociaddons.yaml) Results Success Details This template could be really helpful for automating the deployment of CSI 2, CNI 3, CCM, and any other required packages. This approach for creating clusters on OCI is the most suitable for CERN, as it is fundamentally similar to what is being used on-premises (clustertemplate-oci-addons.yaml ⇔ Magnum template, addons ⇔ Magnum labels). Test 3 Creating a managed cluster from OKE (using cluster-templatemanaged.yaml) Results Nodepool was not created in OCI. Details This template relies on the concept of machine pools, which is still an experimental feature in ClusterAPI. A feature flag must be enabled before installing CAPI providers (export EXP_MACHINE_POOL=true) and I missed that step. 2 Container Storage Interface (CSI): CSI is a standardized interface that enables the integration of storage systems with Kubernetes. 3 Container Network Interface (CNI): CNI provides a standardized way to manage the networking aspects of containerized applications, allowing different networking solutions (such as Calico, Flannel, or Weave) to integrate seamlessly.
Mouad El Haouari - CERN Openlab Report // 2024 17 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Test 4 Using a kubeconfig file retrieved from ClusterAPI (clusterctl get kubconfig) for accessing managed clusters. Results The cluster was accessible for few minutes and then access was suddenly lost and the following error message was shown: you must be logged in to the server (Unauthorized). Details The cluster’s authentication token was changing every couple of minutes. The token in the kubeconfig I was using had expired, which explains the error message. OKE clusters have a special kubeconfig, retrieved only through the OCI CLI, that does not contain a static token for user authentication. Instead, it uses exec-based authentication, where every time a kubectl command is executed, a command from the OCI CLI is run to retrieve the token, which is then injected into the user request [28]. clusterctl, in this case, was not retrieving the kubeconfig with exec-based authentication. It was only retrieving a kubeconfig with the current valid authentication token. The kubeconfig with exec-based authentication is only used for granting access to users who will interact with the cluster through kubectl. However, the kubeconfig generated by clusterctl is suitable for other purposes, such as CI/CD pipelines [28]. A Clusterctl plugin could be implemented to support the retrieval of kubeconfig with exec-based authentication. Test 5 Creating a managed cluster with self-managed nodes (using clustertemplate-managed-self-managed-nodes.yaml) Results worker machine is created and in running state in OCI but worker node never reaches a ready state in Kubernetes although all the requirements are fulfilled. Details /
Mouad El Haouari - CERN Openlab Report // 2024 18 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Test 6 Testing reconciliation by scaling a managed cluster from the OCI UI. Results The state changed in the CRs but the controller didn’t try to take down new node. The status of the cluster in the OCI UI was alternating between ‘updating’ and ‘active’ all the time. Details When scaling up, the ‘status.phase’ of the MachinePool CR is set to ‘scaling’, ‘status.NodepoolLifecycleState’ is set to ‘updating’ in OCIManagedMachinePool CR (see Figure 15 and Figure 16 in Appendix). When scaling down, MachinePool CR went back to its normal state (‘status.phase’ set to ‘running’) but “status.NodepoolLifecycleState” OCIManagedMachinePool CR was still set to UPDATING and the cluster state in the OCI UI was still constantly switching between “running” and “updating”. ii. Crossplane During my search through the Upbound marketplace for Crossplane providers [29], I noticed that the OCI provider was notably absent from the extensive collection of over 100 providers, which included all major public cloud vendors. Upon further investigation, I found a blog post from Oracle [30] offering a pre-release preview of a Crossplane provider, but it didn't link to any code source. Later on in my search, I stumbled across a project for the provider on GitHub [31] where development efforts had begun but were discontinued 6 months later, leading to the repository being recently archived. Having an OCI Crossplane provider would be very beneficial for people who use OKE and want to manage everything from Kubernetes. Hybrid cloud is another scenario where the provider would be helpful, and the CERN use case is a perfect example of that. In addition to these compelling use cases, Upjet can significantly reduce the development effort and time required for the Crossplane provider, as it can be leveraged to generate boilerplate code from the existing Terraform provider. 5. Proposed Approach Our investigation results indicated that neither ClusterAPI nor Crossplane is capable of provisioning Kubernetes clusters on both CERN OpenStack and Oracle Cloud Infrastructure. Each tool's compatibility is mutually exclusive (if one tool supports a particular provider, it lacks support for the other).
Mouad El Haouari - CERN Openlab Report // 2024 19 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Despite Crossplane not being the optimal solution due to its lack of strong reconciliation, we opted for it anyway because all Kubernetes clusters under my team’s responsibility are currently running in CERN OpenStack, and we don’t have any clusters running on Oracle yet. In addition to Crossplane, ArgoCD is also leveraged as a Kubernetes-native CD pipeline to automate clusters lifecycle management following the GitOps approach. ArgoCD and Crossplane are deployed in the management cluster, which can be hosted on any cloud provider or even locally using KinD or Minikube. The management cluster is responsible only for executing cluster operations and hosting clusters specifications as Kubernetes objects, contrary to workload clusters where applications will be actually running. Figure 3. Newly proposed workflow for managing Kubernetes clusters lifecycle The workflow, as described in the figure above, functions as follows: • The user will push YAML code defining cluster specifications to a Git repository that can be hosted on GitLab, GitHub, or even locally. • ArgoCD, running in the management cluster, will constantly watch for newly added cluster specifications in the Git repository and deploy them as Crossplane custom resources (Kubernetes Objects) within the management cluster. • The reconciliation loop executed by the Crossplane controllers will detect these newly added custom resources and trigger the creation of a workload cluster on the CERN OpenStack cloud. 6. The one button deployment In the previous section, we explained how our approach works, its internal architecture, and how the components within this architecture communicate with each other to achieve the intended functionalities. Now, we will dive deeper into the details and explain our automated approach for setting up the management cluster with all the necessary components, as well as how to configure them to collaborate effectively.
Mouad El Haouari - CERN Openlab Report // 2024 20 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Before discussing automation, let me first walk you through the scenario we would need to follow if we wanted to do everything manually. The table below outlines each step of this scenario and explains some of its details. Step N° Actions Explanation 0 • Provision a Kubernetes cluster • Install ArgoCD on that cluster The Kubernetes cluster will serve as the management cluster and will host ArgoCD, Crossplane and the custom resources that will describe our clusters. 1 • Install Crossplane core [32] Crossplane has a plugin architecture where there is the core component and a plugin for each cloud often called providers. 2 • Install openstack-provider [33] for Crossplane • Install argocd-provider [34] for Crossplane The openstack-provider will allow us to create or manage already existing cluster on the CERN OpenStack cloud. The argocd-provider is used for automating the registrations of these clusters in ArgoCD as target clusters for application deployments. 3 • Install openstack-providerconfig [35] • Install argocd-provider-config [36] In this step, we will configure the openstackprovider with our OpenStack cloud credentials and the argocd-provider with the credentials of an ArgoCD admin user needed for registering clusters created/imported using Crossplane as target clusters within ArgoCD. 4 • Deploy Crossplane custom resource that will represent our clusters This will include custom resources for both openstack-provider (Clusterv1, Nodegroupv1) and argocd-provider (Cluster) (see Figure 12, Figure 13 and Figure 14 for in Appendix for examples) 5 • Deploy packages/applications This step is optional but can be executed to hydrate newly created clusters with the packages/dependencies needed by applications or deploy applications themselves directly. To automate this process as much as possible, we leveraged the app of apps pattern [37] in addition to the sync waves feature [38] of ArgoCD, where each step (from 1 to 5) is represented through an ArgoCD
Mouad El Haouari - CERN Openlab Report // 2024 21 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment application with the step number set as its sync wave value (see Figure 18, Figure 19 and Figure 20 in Appendix for examples). It is important to note, however, that using sync waves with Crossplane custom resources does not work by default and requires additional configuration to function properly [39]. The main app, in the app-of-apps architecture, will be responsible for deploying the five applications that represent the five steps, ensuring that the logical order of deployment is respected: applications with higher sync wave values will not be deployed until those with lower sync wave values are deployed and in a healthy state (see Figure 17 in Appendix). As a result, deploying the main app (which can happen with a one-button click) will lead to the setup of the management cluster, the provisioning/import of OpenStack clusters, and deployment of packages/applications on top of these clusters. This one-button deployment is particularly useful for disaster recovery, as it automates the entire process from end to end, without relying on any external dependencies that could also be affected by the disaster. 7. Migration a. Translating Terraform to Crossplane Now that we have validated our new cluster management approach and the one-button deployment method, the final step will be to translate our cluster specifications written in HCL into Crossplane custom resources, which will be deployed through ArgoCD in the management cluster. Since we already use Kustomize in conjunction with ArgoCD for deploying applications and packages, we decided to do the same with clusters. Another reason behind this decision is to allow for easy maintenance of cluster artifacts, as they share many common fields, especially between ones created from the same template. With Kustomize, we will have a base for each cluster template and an overlay for each cluster created from that template, which will customize the base by updating and adding new properties to match the specifications defined in Terraform. However, since cluster specifications and OpenStack credentials are managed separately in Crossplane (each defined in a separate CR), unlike in Terraform, where everything is included in the tfvars file, we decided to maintain a dedicated base for OpenStack configurations (see Figure 21 and Figure 22 in Appendix) and multiple overlays for each OpenStack project, patching the base with the project’s credentials.
Mouad El Haouari - CERN Openlab Report // 2024 22 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment To automate the translation process, we wrote a Python script that parses cluster specifications written in HCL (tfvars file) and generates both config and cluster overlays. The base, on the other hand, is created manually in both cases. Creating the base for the config was relatively straightforward compared to the cluster one, where we needed to identify the fields that are common in all clusters created from the same template. One emerging challenge here was labels. The labels are almost identical between clusters created using the same template, with one to a maximum of three labels changing from one cluster to another. By performing ‘diff’ operations between clusters created using the same template, we identified the lowest common denominator of labels corresponding to a cluster template and included these labels in its base (see Figure 23 and Figure 24 in Appendix for an example). For the remaining labels, we identified two categories: • Optional labels: These are not present in all cluster specifications, so they are not included in the base but might be added in overlays as a patch to the base using an “add” operation. • Changing labels: These labels have a common value across most clusters, except for a small subset where the value differs. They are included in the base with their common value and might be patched in the overlay using a “replace” operation if the value differs from the common one. Optional and changing labels are defined for each cluster template and represented using a Python dictionary as showcased in the figure below. This dictionary is leveraged by the translation algorithm for generating cluster overlays. TEMPLATES = { "v1.29":{ "v1.29.2-1":{ "optional_labels": ["availability_zone"], "changing_labels": { "cephfs_csi_enabled": "false", "manila_csi_enabled": "false", "manila_enabled": "false" } }, "v1.29.2-1-multi":{ "optional_labels": [], "changing_labels": { "cephfs_csi_enabled": "false", "manila_csi_enabled": "false", "manila_enabled": "false"
Mouad El Haouari - CERN Openlab Report // 2024 23 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment } }, "v1.29.2-2-multi":{ "optional_labels": [], "changing_labels": {} } } } Figure 4. Excerpt from the changing and optional labels python dictionary Both cluster and config overlays consist of a set of YAML files. To simplify the generation of these files, we created a template for each overlay file using Jinja [40] as the templating engine (see Figure 25 in Appendix for an example). The translation algorithm, as shown in the figure below, loads and customizes these templates based on the parsed cluster specifications and the identified changing/optional labels (the parsed cluster specifications are passed as a parameter in both render_config_overlay() and render_cluster_overlay() function calls. However, only the render_cluster_overlay() function requires the changing/optional labels and utilizes the TEMPLATES dictionary, which is declared as a global variable). in_dir = sys.argv[1] out_dir = sys.argv[2] # loop over communities for community_dir in os.listdir(in_dir): # get input and output paths for communities in_community_dir_path = os.path.join(in_dir, community_dir) out_community_dir_path = os.path.join(out_dir, community_dir) # create output community os.makedirs(out_community_dir_path, exist_ok=True) # create config under output community config_path = os.path.join(out_community_dir_path, "config") os.makedirs(config_path, exist_ok=True) # copy config overlay kustomization.yaml shutil.copy2( "./overlay-templates/config/kustomization.yaml", os.path.join(config_path,"kustomization.yaml") ) first_dir = 0 # loop over clusters for cluster_dir in os.listdir(in_community_dir_path):
Mouad El Haouari - CERN Openlab Report // 2024 24 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment # get input and output path for clusters in_cluster_dir_path = os.path.join(in_community_dir_path, cluster_dir) out_cluster_dir_path = os.path.join(out_community_dir_path, cluster_dir) # create output cluster os.makedirs(out_cluster_dir_path, exist_ok=True) # render overlays with open(os.path.join(in_cluster_dir_path,"terraform.tfvars"), 'r') as file: # translate tfvars file to python dict tfvars_dict = hcl2.load(file) # render config overlay if first_dir == 0: first_dir = 1 patch_provider_config, patch_secret = render_config_overlay(tfvars_dict) with open(os.path.join(config_path,'patch-provider-config.yaml'),'w') as file: file.write(patch_provider_config) with open(os.path.join(config_path,'patch-secret.yaml'),'w') as file: file.write(patch_secret) # render cluster overlay kustomization,patch_cluster,patch_clusterv1,nodegroupv1=render_cluster_overlay(tfvars_dict) with open(os.path.join(out_cluster_dir_path,'kustomization.yaml'),'w') as file: file.write(kustomization) with open(os.path.join(out_cluster_dir_path,'patch-cluster.yaml'),'w') as file: file.write(patch_cluster) with open(os.path.join(out_cluster_dir_path,'patch-clusterv1.yaml'),'w') as file: file.write(patch_clusterv1) if nodegroupv1: with open(os.path.join(out_cluster_dir_path,'nodegroupv1.yaml'),'w') as file: file.write(nodegroupv1) Figure 5. Terraform to Crossplane translation script b. Secret management As we mentioned earlier, saving the kubeconfig files for newly created clusters was the final step in the Terraform & GitLab CI/CD workflow. To ensure that the entire workflow is migrated to ArgoCD & Crossplane, we had to implement a new approach for secret management. When clusters are created through Crossplane openstack provider, their kubeconfig is saved in a Secret resource of type ‘connection.crossplane.io/v1alpha1’, specifically as a value to the ‘attribute.kubeconfig.raw_config’ key. Based on this default Crossplane behavior for kubeconfig
Mouad El Haouari - CERN Openlab Report // 2024 25 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment management, we created a CronJob that will be deployed on the management cluster and will constantly monitor for newly created kubeconfig secrets, retrieve the value of the kubeconfig key, decode it, and save it in CERN's secret store using its dedicated CLI. The figures below demonstrate the cronjob code and the YAML artifacts used for its deployment on Kubernetes. from kubernetes import client, config import base64 import os def main(): namespace = os.environ.get('NAMESPACE', 'crossplane-clusters') secret_type = os.environ.get('SECRET_TYPE', 'connection.crossplane.io/v1alpha1') kubeconfig_key = os.environ.get('KUBECONFIG_KEY', 'attribute.kubeconfig.raw_config') config.load_incluster_config() v1 = client.CoreV1Api() secrets = v1.list_namespaced_secret(namespace) filtered_secrets = [ secret for secret in secrets.items if secret.type == secret_type ] for s in filtered_secrets: kubeconfig_encoded = s.data.get(kubeconfig_key) kubeconfig_decoded = base64.b64decode(kubeconfig_encoded).decode('utf-8') print(kubeconfig_decoded) # save kubeconfig in the secret store # ... if __name__ == '__main__': main() Figure 6. python code for the kubeconfig cronjob FROM python:3.9-slim WORKDIR /app
Mouad El Haouari - CERN Openlab Report // 2024 32 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment 10. Appendix # Openrc parameters auth_url = "<auth_url>" identity_api_version = "<identity_api_version>" interface = "<interface>" keypair = "<keypair>" project_id = "<project_id>" project_name = "JEEDY AIS" region_name = "cern" # Cluster parameters cluster_name = "k8s-ais-batch-dev-multi-v25" cluster_template_name = "kubernetes-1.25.3-3-multi" create_timeout = "60" flavor = "m2.xlarge" master_count = "1" master_flavor = "m2.xlarge" node_count = "1" # Nodegroups definition nodegroups_list = [ { flavor_id = "m2.xlarge", merge_labels = "true", name = "zone-a-xlarge", node_count = "1", labels = { availability_zone = "cern-geneva-a" } }, { flavor_id = "m2.xlarge", merge_labels = "true", name = "zone-b-xlarge", node_count = "2", labels = { availability_zone = "cern-geneva-b" } }, { flavor_id = "m2.xlarge", merge_labels = "true" name = "zone-c-xlarge", node_count = "2", labels = {
Mouad El Haouari - CERN Openlab Report // 2024 33 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment availability_zone = "cern-geneva-c" } } ] # Labels for kubernetes cluster labels = { "admission_control_list" : "ExtendedResourceToleration,NamespaceLifecycle,LimitRanger,ServiceAccount,Defau ltStorageClass,DefaultTolerationSeconds,MutatingAdmissionWebhook,ValidatingAdmi ssionWebhook,ResourceQuota,Priority", "autoscaler_tag" : "v1.21.0-cern.0", "calico_ipv4pool" : "10.100.0.0/16", "calico_ipv4pool_ipip" : "Always", "calico_tag" : "v3.24.5", "cephfs_csi_enabled" : "false", "cephfs_csi_version" : "cern-csi-1.0-3", "cern_chart_enabled" : "true", "cern_chart_version" : "0.12.0", "cert_manager_api" : "true", "cgroup_driver" : "cgroupfs", "cloud_provider_enabled" : "true", "cloud_provider_tag" : "v1.24.5", "container_infra_prefix" : "<container_infra_prefix>", "container_runtime" : "containerd", "containerd_tarball_sha256" : "<containerd_tarball_sha256>", "containerd_tarball_url" : "<containerd_tarball_url>", "coredns_tag" : "1.8.7", "cvmfs_csi_enabled" : "false", "cvmfs_csi_version" : "v2.0.0", "eos_enabled" : "false", "etcd_tag" : "v3.4.13", "heapster_enabled" : "false", "heat_container_agent_tag" : "train-stable-6", "helm_client_tag" : "v2.16.6", "ignition_version" : "3.3.0", "influx_grafana_dashboard_enabled" : "false", "ingress_controller" : "", "ip_family_policy" : "single_stack", "keystone_auth_enabled" : "true", "kube_csi_enabled" : "false", "kube_csi_version" : "cern-csi-1.0-2", "kube_dashboard_enabled" : "false", "kube_tag" : "v1.25.3-cern.0", "kubeapi_options" : "--feature-gates=CSIMigration=true",
Mouad El Haouari - CERN Openlab Report // 2024 34 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment "kubecontroller_options" : "--feature-gates=", "kubelet_options" : "--feature-gates= --system-reserved=memory=500Mi -- resolv-conf=/run/systemd/resolve/resolv.conf", "logging_installer" : "helm", "manila_csi_enabled" : "false", "manila_enabled" : "false", "manila_version" : "v0.3.0", "metrics_server_enabled" : "true", "monitoring_enabled" : "false", "nginx_ingress_controller_tag" : "v1.0.4", "nvidia_gpu_enabled" : "false", "nvidia_gpu_tag" : "35-5.16.13-200.fc35.x86_64-470.82.00", "oidc_enabled" : "false", "oidc_groups_claim" : "<oidc_groups_claim>", "oidc_groups_prefix" : "<oidc_groups_prefix>", "oidc_issuer_url" : "<oidc_issuer_url>", "oidc_username_claim" : "<oidc_username_claim>", "oidc_username_prefix" : "<oidc_username_prefix>", "snapshot_controller_enabled" : "false", "tiller_enabled" : "true", "tiller_tag" : "v2.16.6", "traefik_ingress_controller_tag" : "2.5.4", "use_podman" : "true" } Figure 10. Tfvars file example apiVersion: infrastructure.cluster.x-k8s.io/v1beta1 kind: OpenStackCluster metadata: name: capi-test namespace: default spec: apiServerLoadBalancer: enabled: true network: id: <network_id> subnets: - id: <subnet_1_id> identityRef: cloudName: openstack name: capi-test-cloud-config network: id: <network_id>
Mouad El Haouari - CERN Openlab Report // 2024 35 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment subnets: - id: <subnet_2_id> disableExternalNetwork: true disableAPIServerFloatingIP: true Figure 11. OpenstackCluster CR used for testing CAPI with CERN openstack apiVersion: containerinfra.openstack.crossplane.io/v1alpha1 kind: ClusterV1 metadata: annotations: crossplane.io/external-name: k8s-ais-batch-dev-multi-v25 name: k8s-ais-batch-dev-multi-v25 spec: deletionPolicy: Orphan forProvider: clusterTemplateId: kubernetes-1.25.3-3-multi createTimeout: 60 flavor: m2.xlarge keypair: shared-jeedy-key labels: admission_control_list: ExtendedResourceToleration,NamespaceLifecycle,LimitRanger,ServiceAccount,Defaul tStorageClass,DefaultTolerationSeconds,MutatingAdmissionWebhook,ValidatingAdmis sionWebhook,ResourceQuota,Priority autoscaler_tag: v1.21.0-cern.0 calico_ipv4pool: 10.100.0.0/16 calico_ipv4pool_ipip: Always calico_tag: v3.24.5 cephfs_csi_enabled: "false" cephfs_csi_version: cern-csi-1.0-3 cern_chart_enabled: "true" cern_chart_version: 0.12.0 cert_manager_api: "true" cgroup_driver: cgroupfs cloud_provider_enabled: "true" cloud_provider_tag: v1.24.5 container_infra_prefix: <container_infra_prefix> container_runtime: containerd containerd_tarball_sha256: <containerd_tarball_sha256> containerd_tarball_url: <containerd_tarball_url> coredns_tag: 1.8.7 cvmfs_csi_enabled: "false" cvmfs_csi_version: v2.0.0
Mouad El Haouari - CERN Openlab Report // 2024 36 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment eos_enabled: "false" etcd_tag: v3.4.13 heapster_enabled: "false" heat_container_agent_tag: train-stable-6 helm_client_tag: v2.16.6 ignition_version: 3.3.0 influx_grafana_dashboard_enabled: "false" ingress_controller: "" ip_family_policy: single_stack keystone_auth_enabled: "true" kube_csi_enabled: "false" kube_csi_version: cern-csi-1.0-2 kube_dashboard_enabled: "false" kube_tag: v1.25.3-cern.0 kubeapi_options: --feature-gates=CSIMigration=true kubecontroller_options: --feature-gates= kubelet_options: --feature-gates= --system-reserved=memory=500Mi -- resolv-conf=/run/systemd/resolve/resolv.conf logging_installer: helm manila_csi_enabled: "false" manila_enabled: "false" manila_version: v0.3.0 metrics_server_enabled: "true" monitoring_enabled: "false" nginx_ingress_controller_tag: v1.0.4 nvidia_gpu_enabled: "false" nvidia_gpu_tag: 35-5.16.13-200.fc35.x86_64-470.82.00 oidc_enabled: "false" oidc_groups_claim: <oidc_groups_claim> oidc_groups_prefix: <oidc_groups_prefix> oidc_issuer_url: <oidc_issuer_url> oidc_username_claim: <oidc_username_claim> oidc_username_prefix: <oidc_username_prefix> snapshot_controller_enabled: "false" tiller_enabled: "true" tiller_tag: v2.16.6 traefik_ingress_controller_tag: 2.5.4 use_podman: "true" masterCount: 1 masterFlavor: m2.xlarge mergeLabels: true name: k8s-ais-batch-dev-multi-v25 nodeCount: 1 providerConfigRef:
Mouad El Haouari - CERN Openlab Report // 2024 37 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment name: jeedy-ais writeConnectionSecretToRef: name: k8s-ais-batch-dev-multi-v25 namespace: crossplane-clusters Figure 12. Clusterv1 CR example (crossplane-openstack-provider) apiVersion: containerinfra.openstack.crossplane.io/v1alpha1 kind: NodegroupV1 metadata: annotations: crossplane.io/external-name: k8s-ais-batch-dev-multi-v25/zone-a-xlarge name: zone-a-xlarge spec: deletionPolicy: Orphan forProvider: clusterId: k8s-ais-batch-dev-multi-v25 flavorId: m2.xlarge labels: availability_zone: cern-geneva-a mergeLabels: true name: zone-a-xlarge nodeCount: 1 providerConfigRef: name: jeedy-ais Figure 13. Nodegroupv1 CR example (crossplane-openstack-provider) apiVersion: cluster.argocd.crossplane.io/v1alpha1 kind: Cluster metadata: name: k8s-ais-batch-dev-multi-v25 spec: forProvider: config: kubeconfigSecretRef: key: attribute.kubeconfig.raw_config name: k8s-ais-batch-dev-multi-v25 namespace: crossplane-clusters name: k8s-ais-batch-dev-multi-v25 providerConfigRef: name: argocd-provider Figure 14. Cluster CR example (crossplane-argocd-provider)
Mouad El Haouari - CERN Openlab Report // 2024 38 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Name: oci-capi-managed-mp-0 Namespace: default Labels: cluster.x-k8s.io/cluster-name=oci-capi-managed Annotations: cluster.x-k8s.io/replicas-managed-by: API Version: cluster.x-k8s.io/v1beta1 Kind: MachinePool Metadata: Creation Timestamp: 2024-08-13T09:21:07Z Finalizers: machinepool.cluster.x-k8s.io Generation: 4 Owner References: API Version: cluster.x-k8s.io/v1beta1 Kind: Cluster Name: oci-capi-managed UID: d858ab51-27df-42a0-9a80-51f0c1c90f62 Resource Version: 7108296 UID: 32801290-c721-4e89-a18a-5594e705f574 Spec: Cluster Name: oci-capi-managed Min Ready Seconds: 0 Provider ID List: <ocid1.instance.oc1.eu-frankfurt-1…> <ocid1.instance.oc1.eu-frankfurt-1…> Replicas: 1 Template: Metadata: Spec: Bootstrap: Data Secret Name: Cluster Name: oci-capi-managed Infrastructure Ref: API Version: infrastructure.cluster.x-k8s.io/v1beta1 Kind: OCIManagedMachinePool Name: oci-capi-managed-mp-0 Namespace: default Version: v1.30.1 Status: Available Replicas: 2 Bootstrap Ready: true Conditions: Last Transition Time: 2024-08-13T09:41:18Z Status: True Type: Ready
Mouad El Haouari - CERN Openlab Report // 2024 39 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Last Transition Time: 2024-08-13T09:21:07Z Status: True Type: BootstrapReady Last Transition Time: 2024-08-13T09:32:38Z Status: True Type: InfrastructureReady Last Transition Time: 2024-08-13T09:41:18Z Status: True Type: ReplicasReady Infrastructure Ready: true Node Refs: API Version: v1 Kind: Node Name: 10.0.78.220 UID: df0b8645-ee4d-4b2a-9413-48b8ea7fa038 API Version: v1 Kind: Node Name: 10.0.69.111 UID: 75df6b54-76e9-4da0-8e8b-40edc0e5ad51 Observed Generation: 4 Phase: Scaling Ready Replicas: 2 Replicas: 2 Unavailable Replicas: -1 Events: <none> Figure 15. MachinePool CR after reconciliation test of CAPI on OCI Name: oci-capi-managed-mp-0 Namespace: default Labels: cluster.x-k8s.io/cluster-name=oci-capi-managed Annotations: <none> API Version: infrastructure.cluster.x-k8s.io/v1beta2 Kind: OCIManagedMachinePool Metadata: Creation Timestamp: 2024-08-13T09:21:07Z Finalizers: ocimanagedmachinepool.infrastructure.cluster.x-k8s.io Generation: 4 Owner References: API Version: cluster.x-k8s.io/v1beta1 Block Owner Deletion: true Controller: true Kind: MachinePool
Mouad El Haouari - CERN Openlab Report // 2024 40 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment Name: oci-capi-managed-mp-0 UID: 32801290-c721-4e89-a18a-5594e705f574 Resource Version: 7553555 UID: 13c6eaca-28de-4756-b70e-6eb76e19b896 Spec: Id: <ocid1.nodepool.oc1.eu-frankfurt-1…> Node Pool Node Config: Node Pool Pod Network Option Details: Cni Type: OCI_VCN_IP_NATIVE Vcn Ip Native Pod Network Options: Nsg Names: pod Subnet Names: pod Nsg Names: worker Placement Configs: Availability Domain: BKrI:EU-FRANKFURT-1-AD-2 Fault Domains: FAULT-DOMAIN-1 FAULT-DOMAIN-2 FAULT-DOMAIN-3 Subnet Name: worker Availability Domain: BKrI:EU-FRANKFURT-1-AD-3 Fault Domains: FAULT-DOMAIN-1 FAULT-DOMAIN-2 FAULT-DOMAIN-3 Subnet Name: worker Availability Domain: BKrI:EU-FRANKFURT-1-AD-1 Fault Domains: FAULT-DOMAIN-1 FAULT-DOMAIN-2 FAULT-DOMAIN-3 Subnet Name: worker Node Shape: VM.Standard.E4.Flex Node Shape Config: Ocpus: 1 Node Source Via Image: Boot Volume Size In G Bs: 50 Image Id: <ocid1.image.oc1.eu-frankfurt-1…> Provider ID: <oci://ocid1.nodepool.oc1.eu-frankfurt-1…> Provider ID List: <ocid1.instance.oc1.eu-frankfurt-1…>
Mouad El Haouari - CERN Openlab Report // 2024 41 A cloud-native approach for managing Kubernetes clusters in a hybrid cloud environment <ocid1.instance.oc1.eu-frankfurt-1…> Version: v1.30.1 Status: Conditions: Last Transition Time: 2024-08-13T09:32:37Z Status: True Type: NodePoolReady Infrastructure Machine Kind: OCIMachinePoolMachine Nodepool Lifecycle State: UPDATING Ready: true Replicas: 2 Events: Type Reason Age From Message ---- ------ ---- --- - ------- Normal NodePoolReady 58s (x2396 over 22h) ocimanagedmachinepoolcontroller Node pool is in ready state Figure 16. OCIManagedMachinePool CR after reconciliation test of CAPI on OCI apiVersion: argoproj.io/v1alpha1 kind: Application metadata: name: main-app namespace: argocd finalizers: - resources-finalizer.argocd.argoproj.io spec: destination: namespace: default server: https://kubernetes.default.svc project: default source: path: sub-apps repoURL: <repoURL> targetRevision: HEAD syncPolicy: automated: prune: true Figure 17. ArgoCD main app for the one button deployment apiVersion: argoproj.io/v1alpha1