scieee AI-readable full text Open interactive document viewer

Qubernetes : Towards a unified cloud-native execution platform for hybrid classic-quantum computing

Stirbu, Vlad,Kinanen, Otso,Haghparast, Majid,Mikkonen, Tommi

Full text

This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY 4.0 https://creativecommons.org/licenses/by/4.0/ Qubernetes : Towards a unified cloud-native execution platform for hybrid classicquantum computing © 2024 the Authors Published version Stirbu, Vlad; Kinanen, Otso; Haghparast, Majid; Mikkonen, Tommi Stirbu, V., Kinanen, O., Haghparast, M., & Mikkonen, T. (2024). Qubernetes : Towards a unified cloud-native execution platform for hybrid classic-quantum computing. Information and Software Technology, 175, Article 107529. https://doi.org/10.1016/j.infsof.2024.107529 2024 Information and Software Technology 175 (2024) 107529 Available online 26 July 2024 0950-5849/© 2024 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/). Contents lists available at ScienceDirect Information and Software Technology journal homepage: www.elsevier.com/locate/infsof Qubernetes: Towards a unified cloud-native execution platform for hybrid classic-quantum computing Vlad Stirbu∗ , Otso Kinanen, Majid Haghparast, Tommi Mikkonen University of Jyväskylä, Jyväskylä, Finland ARTICLE INFO Keywords: Quantum software Hybrid classical-quantum software Containers Quantum software development lifecycle Cloud-native computing ABSTRACT Context: The emergence of quantum computing proposes a revolutionary paradigm that can radically transform numerous scientific and industrial application domains. The ability of quantum computers to scale computations beyond what the current computers are capable of implies better performance and efficiency for certain algorithmic tasks. Objective: However, to benefit from such improvement, quantum computers must be integrated with existing software systems, a process that is not straightforward. In this paper, we propose a unified execution model that addresses the challenges that emerge from building hybrid classical-quantum applications at scale. Method: Following the Design Science Research methodology, we proposed a convention for mapping quantum resources and artifacts to Kubernetes concepts. Then, in an experimental Kubernetes cluster, we conducted experiments for scheduling and executing quantum tasks on both quantum simulators and hardware. Results: The experimental results demonstrate that the proposed platform Qubernetes (or Kubernetes for quantum) exposes the quantum computation tasks and hardware capabilities following established cloud-native principles, allowing seamless integration into the larger Kubernetes ecosystem. Conclusion: The quantum computing potential cannot be realized without seamless integration into classical computing. By validating that it is practical to execute quantum tasks in a Kubernetes infrastructure, we pave the way for leveraging the existing Kubernetes ecosystem as an enabler for hybrid classical-quantum computing. 1. Introduction Quantum computers have demonstrated the potential to revolutionize various fields, including cryptography, drug discovery, materials science, and machine learning, by leveraging the principles of quantum mechanics. However, the current generation of quantum computers, known as noisy intermediate-scale quantum (NISQ) computers [1], suffer from noise and errors, making them challenging to operate. Additionally, the development of quantum algorithms requires specialized knowledge in the field of quantum mechanics and mathematics, which is not readily available to the majority of software professionals. These factors pose a significant entry barrier to leveraging the unique capabilities of quantum systems. For the existing base of business applications, classical computing has already proven its capabilities across a diverse range of solutions. However, some of the computations they must perform can be accelerated with quantum computing, much like graphical processing units (GPUs) are used today. Therefore, quantum systems should not function in isolation, but they must coexist and interoperate with ∗Corresponding author. E-mail address: [email protected] (V. Stirbu). 1https://kubernetes.io/. classical systems. To this end, the current way of building and operating quantum computers hinders their adoption, as application developers have to learn the bespoke way in which their programs are executed on the hardware. To make matters worse, the quantum simulator of the hardware target used for execution has to be explicitly selected, which blurs the line between the development and the operational phase in a product or software development lifecycle. This paper proposes an approach where the focus is placed on the orchestration of classical and quantum computations. Kubernetes,1a widely used system for automating deployment, scaling, and management of containerized applications, is used as the underlying infrastructure. In this approach, the quantum computations are packaged as containers that are executed on quantum-capable nodes alongside classical computations. Constructed in this way, Qubernetes – the quantum-enhanced Kubernetes – is tailored to fit hybrid classicalquantum applications. The rest of this paper is organized as follows. In Section 2, we present the fundamental concepts of quantum computing, the quantum https://doi.org/10.1016/j.infsof.2024.107529 Received 6 November 2023; Received in revised form 13 May 2024; Accepted 16 July 2024 Information and Software Technology 175 (2024) 107529 2 V. Stirbu et al. software development and key challenges faced by the developers and the hardware operators of hybrid classic-quantum systems. In Section 3, we introduce the methodological background of this research. In Section 4, we introduce the objectives of the solution, crystallized as requirements that need to be satisfied by a unified cloud-native hybrid classical-quantum computing execution platform. In Section 5, we introduce Qubernetes, a Kubernetes platform extension that enables the execution of heterogeneous classic-quantum computing tasks. In Section 6we describe the experimental setup and the application scenarios used to validate the Qubernetes concept. In Section 7we discuss how Qubernetes addresses the requirements and the needs of software developers. In Section 8we address the threats to validity. Concluding remarks are provided in Section 9. 2. Background and motivation 2.1. Quantum computing fundamentals Qubits, which stands for quantum bits, are the fundamental units of quantum information in quantum computing. Unlike conventional bits, which can exist in one of two states (0 or 1), qubits can exist in multiple states simultaneously, thanks to the principles of superposition and entanglement, which are unique to quantum mechanics [2]. This new computing paradigm enables the development of a new breed of algorithms [3] that leverage the qubit capabilities to speed up the performance of computational tasks beyond what is possible with the existing classical computers [4]. For example, factoring large numbers using classical algorithms has exponential complexity, while using Shor’s algorithm has polynomial complexity. The physical implementation of quantum computers can be split into two categories: specialized (e.g., special purpose computers designed to solve optimization problems using annealing programming approach) or general-purpose (e.g., allowing programming of individual qubits using pulses or gate programming approaches). The current technological candidates for building gate-based general-purpose quantum computers fit within one of the following categories: superconducting – tiny superconducting materials are cooled to extremely low temperatures to manifest their quantum properties, trapped ion – ions are trapped within electromagnetic fields, or photonic – quantum information stored in photons can be manipulated and transmitted over long distances. In the longer term, the topological quantum computers, leveraging the collective properties of ensembles of particles, will overcome the current NISQ limitations and achieve fault-tolerant operations [5]. Although these quantum computers are not yet advanced enough to achieve fault-tolerance or reach the scale required for quantum advantage [6,7], they provide an experimentation platform to develop new generations of hardware and quantum algorithms and validate quantum technology in real-world use cases. Whether a quantum computer is general-purpose or specialized, the selection of quantum qubit implementation technology can enhance hardware efficiency for specific problem classes [8,9]. To use the hardware effectively, application developers must consider these differences when designing and optimizing the software’s functionality and operations. Further, the concept of distributed quantum computers [10], which interlink multiple distinct quantum machines through quantum communication networks, emerges as a potential solution to amplify the available quantum volume [11], beyond what is possible using a single quantum computer. Nevertheless, the intricacies inherent in the distributed quantum computers remain hidden from users, as compilers aware of the distributed architecture of the target system shield them from such complexities. In essence, the quantum compiler plays a vital role in achieving the effective execution of generic quantum circuits on existing physical hardware platforms, making the compilers an active research area in quantum computing [12]. 2.2. Quantum development kits A typical hybrid classic-quantum software system is understood as a classical program that has one or more software components that are implemented using quantum technology, as depicted in Fig. 1. A quantum component relies on quantum algorithms [3], which are transformed into quantum circuits. The quantum circuit describes quantum computations in a machine-independent language, such as quantum assembly (QASM) [13]. This circuit is translated by a computer that controls the quantum computer in a machine-specific circuit and a sequence of operations, such as pulses [14], that control the operation on individual hardware qubits. The translation process, implemented using quantum compilers, encompasses supplementary actions like breaking down quantum gates, optimizing quantum circuits, and providing fault-tolerant iterations of the circuit. Application developer use tools like Qiskit2and Cirq3for writing, manipulating and optimizing quantum circuits. These Python libraries allow researchers and application developers to interact with nowadays’ NISQ computers, allowing them to run quantum programs on a variety of simulators and hardware designs, abstracting away the complexities of low-level operations and allowing researchers and developers to focus on algorithm design and optimization. Tools like TensorFlow Quantum4and PennyLane5play a crucial role in facilitating the development of machine learning quantum software. These frameworks provide high-level abstractions and interfaces that bridge the gap between quantum computing and classical machine learning. They allow researchers and developers to integrate quantum algorithms seamlessly into the machine learning development process by providing access to quantum simulators and hardware, as well as offering a range of quantum-friendly classical optimization techniques. TensorFlow Quantum leverages the power of Google’s TensorFlow ecosystem, enabling the combination of classical and quantum neural networks for hybrid quantum–classical machine learning models. PennyLane offers a unified framework for developing quantum machine learning algorithms, supporting various quantum devices and seamlessly integrating them with classical machine learning libraries. These tools provide a foundation for researchers to explore and experiment with quantum machine learning, accelerating the progress and adoption of quantum computing in the field of machine learning. 2.3. Notebooks, simulators, and proxy access to quantum hardware Jupyter6notebooks and quantum simulators play a vital role in supporting developers of quantum programs. Jupyter provides an interactive and collaborative environment where developers can write, execute, and visualize their quantum code in an accessible manner. They allow for the combination of code, explanatory text, and visualizations, making it easier to experiment, iterate, and document the development process. Quantum simulators, on the other hand, enable developers to simulate the behavior of quantum systems without the need for physical quantum hardware. These simulators provide a valuable testing ground for verifying and debugging quantum algorithms, allowing developers to gain insights into their performance and behavior before running them on actual quantum devices. Developers can iterate quickly, gain a deeper understanding of quantum concepts, and refine their quantum programs efficiently. 2https://qiskit.org. 3https://quantumai.google/cirq. 4https://www.tensorflow.org/quantum. 5https://pennylane.ai. 6https://jupyter.org. Information and Software Technology 175 (2024) 107529 3 V. Stirbu et al. Fig. 1. Quantum computing model: components and interfaces. Traditional cloud computing providers, such as AWS Braket,7Azure Quantum,8Google Quantum AI9or IBM Quantum,10 offer comprehensive quantum development services. These services are designed to optimize the development process with integrated tools like Jupyter notebooks and task schedulers. Developers can create quantum applications and algorithms across multiple hardware platforms simultaneously. This approach ensures flexibility, allowing fine-tune algorithms for specific systems while maintaining the ability to develop applications that are compatible with various quantum hardware platforms. 2.4. Hybrid classical-quantum computing approaches High-Performance Computing (HPC) is the mainstream approach for running scientific and engineering simulations at scale. Integrating the quantum computing and HPC software stacks enables quantum technology to accelerate parts of the simulations. Two notable approaches for integrating the two software stacks are HPC-QC [15], which leverages the Open Message Passing Interface (OpenMPI11) compatible architectures, and XACC [16] approach based on the OSGi12 architecture. Similarly, the existing base of cloud applications can benefit from using quantum computing to accelerate the appropriate computational tasks, a trend that is not overlooked by the major quantum development toolkit providers. For example, Qiskit’s quantum-serverless [17] proposes a cloud-based approach for running hybrid classical-quantum programs. The proposed programming model, conforming to the RAY13 computing framework, makes it easy to scale Python workloads on a Kubernetes cluster in which the quantum execution environment is represented by a distributed Qiskit runtime that allows transparent access to multiple QPUs. 2.5. Development process The software development life-cycle (SDLC) of hybrid classicquantum applications consists of a multi-faceted approach [18], as depicted in Fig. 2. At the top level, the classical software development process starts by identifying user needs and deriving them into system requirements. These requirements are transformed into a design and implemented. The result is verified against the requirements and validated against user needs. Once the software system enters the operational phase, any detected anomalies are used to identify potential new system requirements, if necessary. A dedicated track for quantum 7https://aws.amazon.com/braket/. 8https://learn.microsoft.com/en-us/azure/quantum/. 9https://quantumai.google. 10 https://quantum-computing.ibm.com. 11 https://www.open-mpi.org. 12 https://www.osgi.org. 13 https://www.ray.io. components is followed within the SDLC [19], specific to the implementation of quantum technology. The requirements for these components are converted into a design, which is subsequently implemented on classic computers, verified on simulators or real quantum hardware, and integrated into the larger software system. During the operational phase, the quantum software components are executed on actual quantum hardware. The scheduling ensures efficient utilization of the scarce quantum hardware resources, while monitoring capabilities enable the detection of anomalies throughout the operational stage. As quantum computers are a scarce resource, it is not practical to develop quantum software components directly on hardware. Instead, developers can use simulators that use commonly available and less expensive classical resources (e.g., CPUs and GPUs) for the early stages of development and testing. As simulators become more sophisticated, being able to simulate the noise of actual hardware, developers can perform fast iterations with confidence. Only when the components are mature enough the development can be continued on actual the hardware that will be used during the execution phase. This approach ensures that the use of quantum resources is effective. Commercial entities, like QuantumPath [20], provide an integrated offering that covers multiple developments, including requirements management, editing and source code version control, and remote execution via proxy to quantum hardware. The integrated approach has near-term advantages as it lowers the entry barrier into a technologically complex environment. However, in the long term, as quantum technology is integrated into existing classical applications, the development methodologies and the tooling that support them will be inherited from what is already used for classical software development by the respective organizations. This is a particularly important concern for regulated industries (e.g., finance or medical [21]) where regulatoryrelated automation is implemented in tools like JIRA14/Polarion15 – project and requirements management, and GitHub16/GitLab17 – version control and code level change management. 2.6. Towards cloud-native quantum computing Quantum technology has the ability to deliver quantum advantage for an array of applications (e.g., machine learning [22] or optimizations [23]), that can be implemented in cloud-native environments. When used within this context, the quantum technology must be properly integrated into the larger technological ecosystem (e.g. Kubernetes), and using modern DevOps practices [24], leveraging containers as the standard way of packaging software artifacts, and a high degree of automation employed at every stage of the SDLC. 14 https://www.atlassian.com/software/jira. 15 https://polarion.plm.automation.siemens.com. 16 https://github.com. 17 https://about.gitlab.com. Information and Software Technology 175 (2024) 107529 4 V. Stirbu et al. Fig. 2. The software development lifecycle model for hybrid classical-quantum systems. Fig. 3. Design science research methodology applied to Qubernetes development. 3. Methodology The Qubernetes concept was developed using the problem-centric approach of the Design Science Research Methodology (DSRM) [25]. The starting point was to answer the research question How to effectively implement hybrid classical-quantum computing? The research question was translated into a set of six objectives that need to be met by the solution to enable cloud-native integration. Further, in the Design and Development phase we introduced how quantum computing concepts like quantum computer and computation tasks are exposed in to Kubernetes as quantum nodes and jobs. Then, for the Demonstration phase, we described a test cluster, conforming to the Qubernetes convention, and provided an example quantum jobs developed using the Qiskit toolkit that is executed on all targets: CPU and GPU simulators (e.g., Qiskit-Aer), and quantum hardware (e.g., HELMI). For the Evaluation phase, we have discussed how the Qubernetes solution addresses the objectives. The process is depicted in Fig. 3. We have carefully considered the reproducibility of the test environment and opted for an approach in which the essential software artifacts are included in the paper using the established conventions for each technology: YAML specifications for serialized Kubernetes objects,18 Dockerfile for the container descriptions,19 and Python source for the simple quantum test program developed using Qiskit toolkit. The steps describing setting up a Kubernetes cluster, configuring the internal container registry, build an publish container images to registry or interacting with the cluster using the kubectl,20 have been 18 https://kubernetes.io/docs/concepts/overview/working-with-objects/. 19 https://docs.docker.com/reference/dockerfile/. 20 https://kubernetes.io/docs/reference/kubectl/. omitted for brevity, as they are covered by ample documentation on the respective projects’ websites. Nevertheless, we have provided throughout the manuscript, whenever necessary, footnotes with the links that lead to the relevant online documentation. We acknowledge that our access to the HELMI quantum computer is attributed to our university’s membership in the consortia that owns the hardware, a circumstance that is not easily replicable. However, the CPU and GPU capabilities of the test cluster can be replicated by anyone with access to general-purpose computing and an Nvidia-compatible GPU, which are commercially available off-the-shelf products. We believe that this approach strikes the right balance between completeness and brevity, allowing the reader not only to replicate our results but to continue experimentation. 4. Objectives The shift to cloud computing has simplified the process of developing scalable applications. However, to fully harness the benefits of cloud computing, applications must adhere to cloud-native architectural principles [26]. This entails designing applications as small, loosely coupled components that can be bundled with their dependencies into portable containers and deployed on the immutable infrastructure. By leveraging the service discovery, load-balancing, and selfhealing capabilities inherent in cloud platforms, development teams, comprising both software development and operations expertise, can automate the software development lifecycle and streamline delivery processes. Furthermore, emphasizing observability through integrated monitoring and logging offers valuable insights into performance, health, and behavior, empowering teams to swiftly respond to potential anomalies. Information and Software Technology 175 (2024) 107529 5 V. Stirbu et al. Fig. 4. The solution boundaries within the hybrid classical-quantum application domain. Kubernetes is the industry-standard container orchestration platform for automating deployment, scaling, and management of containerized cloud-native applications. Developed as an open-source solution by Cloud Native Computing Foundation (CNCF),21 together with the myriad of projects that offer supporting functionality, it allows users to deploy applications on the managed infrastructure of the major cloud providers (e.g., AWS EKS,22 Azure AKS,23 or GCP GKE24), smaller or regional cloud providers, or on-prem – using own infrastructure. The reach functionality and wide industry adoption make Kubernetes the prime candidate for developing a cloud-native execution platform for hybrid classical-quantum computing. Quantum computing technology holds the potential to enhance the performance of cloud applications, particularly in domains such as machine learning and optimizations [27]. To facilitate seamless integration, the implementation of quantum components should align with existing development conventions and practices established in classical applications whenever possible. It is crucial to acknowledge that cloudnative applications are developed using a diverse array of programming languages and frameworks. In the realm of machine learning alone, there are various tools such as KubeFlow,25 Seldon Core,26 and RAY, to name a few. Consequently, a cloud-native solution for exposing quantum computing resources needs to focus on the low-level interface between containerized workloads and simulators/hardware. Simultaneously, it should maintain an open high-level interface between the classical and the quantum components, allowing for flexibility and interworking with different programming languages and frameworks, as illustrated in Fig. 4. The following objectives crystallize the focus on the low-level interface described above. O1 - Design control and SDLC: The design controls are part of a comprehensive quality system that covers the lifetime of a product or 21 https://www.cncf.io/. 22 https://aws.amazon.com/eks/. 23 https://azure.microsoft.com/en-us/products/kubernetes-service. 24 https://cloud.google.com/kubernetes-engine. 25 https://www.kubeflow.org. 26 https://www.seldon.io/solutions/seldon-core. Fig. 5. Design controls. service. The process ensures that the user needs are met by the resulting product or service and that the design inputs and outputs on which the design process is based are verified through a rigorous review process, see Fig. 5. They are based upon established quality assurance and engineering principles [28], covering changes to the product, service, or manufacturing process design, including those occurring long after a device has been introduced to the market. From a quantum software perspective, the software component developed using quantum technology needs to be validated and packaged in a format that is appropriate for execution during the quantum execution phase. O2 - Runtime support: The quantum programming frameworks (e.g., Qiskit or Cirq) employ distinct methods for exposing the quantum hardware as backends. As the framework includes a runtime for running the code, they are responsible for converting the input circuits, which are machine-independent, into machine-specific configurations using an internal representation expressed in QASM. Alternatively, an open and extensible toolchain and runtime based on intermediate representations for quantum programs that extend the LLVM compiler framework [29] are currently under development in the QIR Alliance.27 The QIR compiler has the ability not only to convert between the machine-independent and the machine-dependent circuits but also to mix intermediate representations originating from different quantum programming languages expressed as QIR. Further, the QIR ecosystem enables developers to create programs with complex classical and quantum instructions via its interoperability with LLVM. These aspects of the execution environment have to be exposed at the platform level so that users can execute their quantum software on the appropriate hardware. O3 - Programming model: Gate-level and pulse-level quantum programming are two distinct approaches used to control and manipulate quantum computers. In gate-level programming, quantum operations are expressed as a sequence of quantum gates that act on qubits. These gates are akin to logic gates in classical computing and are specified in a quantum circuit. Gate-level programming provides a highlevel, hardware-independent representation of quantum algorithms. Most quantum programming frameworks support gate-level programming, e.g., Qiskit, Cirq, or TKET.28 Similarly, machine learning-oriented quantum programming (e.g., Pennylane29) are gate-based [30]. On the other hand, pulse-level quantum programming involves direct manipulation of the microwave or laser pulses that drive the qubits. This level of programming is hardware-centric and enables fine-grained control over the quantum operations, providing opportunities for optimizing 27 https://www.qir-alliance.org. 28 https://www.quantinuum.com/developers/tket. 29 https://pennylane.ai. Information and Software Technology 175 (2024) 107529 6 V. Stirbu et al. quantum algorithms. Pulse-level programming is well-suited for practitioners who want to harness the full potential of quantum hardware via specialized programming languages (e.g., Jaqal,30 Qiskit Pulse,31 or SimuQ [31]). O4 - Scheduling: The scheduler is a software component that has the responsibility to find the appropriate resources required for executing correctly a quantum software component. Besides the basic functionality, the scheduler might consider additional inputs that affect its decisions. For example, the energy requirements for completing the job vs the cost of the energy can play a significant role in deciding the time when to schedule the execution. Similarly, from a time perspective, the scheduler can do more than act as a queue so that quantum executions that need to be completed fast are prioritized first, while the others are scheduled when the quantum hardware utilization decreases. O5 - Execution: The execution is the phase during which the quantum software component is run on the actual hardware. The execution typically involves the preparation of the hardware, a step performed by the control software that runs on a classical computer. Following the preparation, the quantum program is executed a number of times, with the results being collected and aggregated into a data structure that includes a probability distribution of the results. O6 - Monitoring: The monitoring component performs comprehensive observation of the system performance targeted to the users and to the operators of the platform. Monitoring the execution allows the users to determine if there are anomalies in the execution that can lead to modification of the program. Similarly, monitoring allows operators to determine how the quantum hardware is utilized and detect how to improve resource utilization. Monitoring also fulfills the enabling layer of billing. 5. Qubernetes: design and concepts Qubernetes (Q8s) is a quantum computing-aware extension of Kubernetes. In this section, we describe how the quantum computing resources are mapped to the Kubernetes native concepts, serving as the foundation for building cloud-native hybrid classical-quantum applications. 5.1. Quantum resource mapping overview In contrast to traditional Kubernetes, Q8s introduces the following pivotal additions: the quantum-capable node definition and the quantum job definition that facilitates execution of quantum computations on quantum-capable nodes. Quantum nodes seamlessly integrate quantum hardware and its associated control circuit capabilities into the Kubernetes cluster, while the quantum-aware scheduler is able to schedule jobs that instantiate the pods that need access to quantum hardware on the corresponding quantum nodes, as depicted in Fig. 6. 5.2. Quantum node The quantum capable node joining the cluster is identified using specific labels (e.g., accelerator), and the QPU’s capacity in their Node specification (e.g., vendor.example.com/qpu), see Listing 1. The capacity indicated by the node is used by the scheduler to allocate pods on compatible nodes. As current quantum hardware is typically able to execute one task at a time, the value 1 means that the node is able to execute a task, while the value 0 indicates that it is not. 30 https://www.sandia.gov/quantum/quantum-information-sciences/ projects/qscout-jaqal/. 31 https://qiskit.org/documentation/apidoc/pulse.html. apiVersion: v1 kind: Node metadata: labels: accelerator: qpu status: capacity: vendor.example.com/qpu: 1 Listing 1: Quantum computing capable node specification. 5.3. Quantum job AJob in Kubernetes is a workload resource designed to spawn a single Pod and ensure its reliable execution until completion. Given that quantum programs typically adhere to a batch execution model, reusing the Job workload is a well-suited choice. 1apiVersion: batch/v1 2kind: Job 3metadata: 4name: quantum-job 5spec: 6template: 7spec: 8nodeSelector: 9accelerator: qpu 10 containers: 11 -name: quantum-task 12 image: registry.example.com/program:v1.2.3 13 command: ["./extrypoint.sh"] 14 resources: 15 requests: 16 vendor.example.com/qpu: 1 17 limits: 18 vendor.example.com/qpu: 1 Listing 2: Quantum job specification. The specific quantum task that needs to be executed as part of the Job is described by the spec.template key that includes a cue for the scheduler that the pod needs to be executed on a quantum capable node (e.g., nodeSelector), and it needs one slice of the specific hardware capacity (e.g., vendor.example.com/qpu). 1apiVersion: v1 2kind: Pod 3metadata: 4name: quantum-pod 5spec: 6nodeSelector: 7accelerator: qpu 8containers: 9-name: quantum-task 10 image:"registry.example.com/program:v1.2.3" 11 resources: 12 requests: 13 vendor.example.com/qpu: 1 14 limits: 15 vendor.example.com/qpu: 1 Listing 3: Pod specification created from the template described in the Job. 5.4. Scheduling and execution Kubernetes has sophisticated scheduling capabilities for classical computing that are able to handle heterogeneous computing capabilities like CPUs with different architectures (e.g., amd64 or arm64), Information and Software Technology 175 (2024) 107529 7 V. Stirbu et al. Fig. 6. Qubernetes: quantum aware Kubernetes. GPUs (e.g., AMD, Intel, Nvidia), or even more exotic accelerators like TPUs (e.g., on Google Kubernetes Engine) or FPGAs. Using the labels and capabilities exposed by the quantum capable nodes, and the node selection preferences and the computing needs requested by pods, the default scheduler (e.g., kube-scheduler), without being aware of quantum computing internals, can create the pods, move them in Pending state, and wait till the appropriate nodes become available. Once scheduled, a Pod moves into Running state, during which the quantum circuit is actually executed on the quantum hardware. Once the execution ends successfully, the pod state changes to Succeeded, and the corresponding Job becomes Completed. In case the execution fails, the pod status changes to Failed. The Job output can be fetched using kubectl logs jobs/quantum-job, as for any Kubernetes jobs. 5.5. Logging and monitoring Logging is the process of capturing, storing, and analyzing the data generated by containers, applications, and infrastructure within a Kubernetes cluster. It plays a crucial role in monitoring, troubleshooting, and maintaining the health and performance of containerized applications and the underlying infrastructure. Kubernetes logging typically involves the collection of log data from various sources, such as containers, pods, and nodes, and centralizing it for analysis and visualization. Effective logging at the quantum node and pod level helps Kubernetes administrators and developers gain valuable insights into the application’s behavior, diagnose issues, and ensure the reliability and security of the containerized quantum workloads. Monitoring is an essential aspect of managing containerized applications within Kubernetes clusters. It involves the continuous collection, analysis, and visualization of data related to the performance, health, and resource utilization of both the applications and the underlying infrastructure. Kubernetes monitoring provides real-time insights into the behavior of containers, pods, nodes, and other resources, enabling administrators to proactively identify and resolve issues, optimize resource allocation, and ensure the reliability and scalability of the entire environment. Administrators can leverage tools such as Prometheus,32 Grafana,33 or other Kubernetes-native monitoring solutions to enable operators to gain a comprehensive understanding of the cluster’s operational status by tracking metrics, setting up alerts, or creating detailed dashboards. This data-driven approach is fundamental for maintaining the availability and performance of applications in dynamic, containerized environments. 6. Demonstration This section describes the environment used to demonstrate the use of the Qubernetes platform. We start with a description of the experimental cluster in which the demonstration was conducted. Then we describe the scenarios used for running quantum programs inside the test Qubernetes cluster. 32 https://prometheus.io. 33 https://grafana.com. 6.1. Experimental cluster setup The evaluation of Qubernetes was performed on a Kubernetes cluster containing both classical and quantum computing resources (see Fig. 7). The classical nodes had CPU and GPU capabilities, allowing quantum computations to be executed in simulators, including the ones supported by Nvidia’s cuQuantum.34 The quantum node exposed the QPU functionality as a a virtual QPU, implemented by a classical program (e.g., the entrypoint.sh script included in the container) that sends commands over secure shell (ssh) to the IQM 5-qubit computer attached to the LUMI supercomputer operated by CSC35 in Finland. 1from qiskit import QuantumCircuit, transpile 2from qiskit_aer import AerSimulator 3 4# Use Aer’s AerSimulator 5simulator =AerSimulator() 6 7# Create a Quantum Circuit acting on the q register 8circuit =QuantumCircuit(2,2) 9 10 # Add a H gate on qubit 0 11 circuit.h(0) 12 13 # Add a CX (CNOT) gate on control qubit 0 and target qubit 1 14 circuit.cx(0,1) 15 16 # Map the quantum measurement to the classical bits 17 circuit.measure([0,1], [0,1]) 18 19 # Compile the circuit for the support instruction set (basis_gates)↪ 20 # and topology (coupling_map) of the backend 21 compiled_circuit =transpile(circuit, simulator) 22 23 # Execute the circuit on the aer simulator 24 job =simulator.run(compiled_circuit, shots=shotsAmount) 25 26 # Grab results from the job 27 result =job.result() 28 29 # Returns counts 30 counts =result.get_counts(compiled_circuit) 31 print("\nTotal count for 00 and 11 are:", counts) Listing 4: Simplified test program intended to run on CPU. The test application was a simple quantum program developed using the Qiskit framework, depicted in Listing 4. The program contains all the structural elements expected in a typical quantum program regardless of the programming framework used (e.g., Cirq, PennyLine, etc.): backend selection (line 5), quantum circuit definition (lines 34 https://developer.nvidia.com/cuquantum-sdk. 35 https://docs.csc.fi/computing/quantum-computing/overview/. Information and Software Technology 175 (2024) 107529 8 V. Stirbu et al. Fig. 7. Experimental Qubernetes cluster setup. Fig. 8. The representation of the quantum circuit used in the experiment. 8–17), transpilation of the machine-independent circuit to the backendspecific circuit (line 21), execution on the backend (line 24), and using the results (lines 27–31). The simple quantum circuit consisting of two qubits and a 2-qubit gate (depicted in Fig. 8) is light enough in terms of gate complexity that can be executed in all target environments (e.g., CPU or GPU-based simulators or actual quantum computers), but still demonstrates a measurable result of a quantum computation task. The program is packaged as a container, together with the appropriate dependencies and the entrypoint.sh script, then published to the cluster’s internal container registry. The blueprint of the container specification is presented in Listing 5. 1FROM --platform=amd64 nvidia/cuda:11.6.2-base-ubuntu20.04↪ 2 3COPY requirements.txt . 4RUN pip install - r requirements.txt 5 6COPY test.py . 7COPY entrypoint.sh . 8 9CMD ["./entrypoint.sh"] Listing 5: The container blueprint for executing the quantum task in a Pod. The program is executed in the cluster as a Job that requires the execution completion of one Pod following the Kubernetes conventions. The quantum jobs are submitted, and the results of the execution are fetched using kubectl commands apply and logs, as expected in a Kubernetes cluster. 6.2. Execute the quantum computation task in simulator The experiment’s objective is to run a test program on a classical node within the cluster, utilizing the high-performance quantum computing simulator qiskit-aer,36 which includes realistic noise models. Initially, the program is executed on a node that solely relies on CPU resources, as evident in the Job specification by the absence of resource requests (e.g., as seen in lines 14–18 in Listing 2). 36 https://github.com/Qiskit/qiskit-aer. Subsequently, the program is adapted to employ qiskit-aergpu, the GPU-accelerated version of the simulator. This modified execution takes place on a GPU-enabled node within the cluster, as indicated by the necessary hardware specified in the Job configuration (e.g., lines 16 and 18 in Listing 2are altered to nvidia.com/gpu: 1). 6.3. Execute the quantum computation task on quantum hardware The aim of the experiment is to run the test program on the HELMI quantum computer. The test program is adjusted to utilize the HELMI backend.37 An entrypoint.sh script that communicates with HELMI via SSH, executes the required commands, and waits for their completion is added to the container image. The Job submission is scheduled to run on a designated node configured as described in Listing 1. The Job description has additional configuration that exposes the needed ssh keys in the running Pod, enabling entrypoint.sh script to communicate securely with HELMI. 7. Discussion In this section, we first discuss how Qubernetes meets the objectives for a hybrid classical-quantum cloud native execution platform. Additionally, we compare how Qubernetes compares with alternative approaches, and propose future research directions. 7.1. QPU-capable node implementation Within the experimental setup, the role of the quantum computer is assumed by the HELMI computer, operated by CSC. Our approach involves accessing the HELMI computer and executing the necessary commands to run the quantum program through an SSH session. Given that HELMI is an older system, this method of integrating its functionality into the Kubernetes cluster serves as a proof of concept. Fortunately, recent developments in quantum computing have seen new hardware vendors and cloud providers offering remote APIs for their quantum computers (e.g., Atos QML38 or AWS Braket). Further, ongoing research initiatives like European High-Performance Computing Joint Undertaking39 (EuroHPC JU) are working on defining Universal Quantum Access [32], a concept that would not only enable access to various local and remote quantum computers via standardized interfaces and protocols, but would also facilitate the effective use of these quantum resources. These advancements will facilitate a more straightforward implementation of quantum resources at the node level. Overall, Kubernetes has the ability to expose the runtime and hardware capabilities using node labels, fulfilling the intent of objectives O2and O3, and collect the logs entries from the Pods to a centralized drain (e.g., Prometheus), enabling monitoring according to objective O6. 37 https://docs.csc.fi/computing/quantum-computing/helmi/running-onhelmi/. 38 https://pypi.org/project/qlmaas/. 39 https://eurohpc-ju.europa.eu/.