Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [1] CONTRASTIVE LEARNING FOR WORKLOAD CHARACTERIZATION AND PREDICTIVE SCHEDULING IN HETEROGENEOUS COMPUTING Hiroto Tanaka* Ayaka Saito Graduate School of Information Science and Technology, The University of Tokyo, Japan * Corresponding author:
[email protected] ABSTRACT Heterogeneous computing systems, integrating diverse processing units such as central processing units (CPUs), graphics processing units (GPUs), and field-programmable gate arrays (FPGAs), have emerged as dominant architectures for high-performance computing applications. However, effective workload characterization and task scheduling remain critical challenges due to the complex interplay between diverse hardware capabilities and dynamic application requirements. This paper presents a novel framework leveraging contrastive learning (CL) for automated workload characterization and predictive scheduling in heterogeneous computing environments. Our approach employs self-supervised representation learning to extract meaningful workload features from historical execution traces without requiring extensive labeled datasets. By maximizing agreement between augmented views of similar workload patterns while maintaining separation from dissimilar patterns, the proposed contrastive learning framework learns robust representations that capture essential workload characteristics including resource utilization patterns, communication overhead, and computational intensity. These learned representations enable accurate prediction of task execution times across different processing units and facilitate intelligent scheduling decisions that optimize system throughput and resource utilization. Experimental evaluation on real-world heterogeneous computing workloads demonstrates that our contrastive learning-based approach achieves superior performance compared to traditional scheduling heuristics, reducing average task completion time by 28% and improving resource utilization efficiency by 34%. The framework shows particular effectiveness in handling workload diversity and adapting to dynamic system conditions, making it suitable for production heterogeneous computing clusters. Keywords: Contrastive Learning, Heterogeneous Computing, Workload Characterization, Predictive Scheduling, SelfSupervised Learning, Resource Management 1. INTRODUCTION The landscape of high-performance computing has undergone a fundamental transformation with the widespread adoption of heterogeneous computing architectures. Modern computing systems increasingly integrate multiple types of processing units, each optimized for specific computational patterns, to achieve superior performance and energy efficiency compared to homogeneous systems [1]. Central processing units excel at sequential and control-intensive tasks, graphics processing units provide massive parallelism for data-intensive operations, and specialized accelerators such as tensor processing units and field-programmable gate arrays offer customized computational capabilities for domain-specific applications. This architectural heterogeneity enables computing systems to match diverse application requirements with appropriate hardware resources, potentially achieving significant performance improvements and energy savings. Despite the compelling advantages of heterogeneous computing, realizing their full potential requires solving complex resource management challenges [2]. The fundamental difficulty lies in making intelligent decisions about which tasks should execute on which processing units to optimize system-wide objectives such as throughput, latency, energy consumption, or fairness. Traditional scheduling approaches for homogeneous systems, which assume uniform processing capabilities across all resources, fail to capture the nuanced performance characteristics of heterogeneous architectures [3]. Each processing unit exhibits distinct computational capabilities, memory hierarchies, communication patterns, and energy profiles, creating a vast space of possible task-to-resource mappings with dramatically different performance outcomes. Furthermore, modern computing workloads demonstrate significant diversity in their computational characteristics, ranging
Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [2] from control-intensive applications with irregular memory access patterns to highly parallel data-intensive computations with predictable execution behavior [4]. Effective scheduling in heterogeneous systems fundamentally depends on accurate workload characterization, which involves understanding and predicting how different applications will behave when executed on different processing units [5]. Traditional workload characterization approaches rely on either static code analysis or profiling-based methods, both of which have significant limitations. Static analysis techniques examine program source code to identify computational patterns, but they struggle with dynamic behaviors that only manifest at runtime, such as input-dependent execution paths, dynamic memory access patterns, and runtime dependencies [6]. Profiling-based approaches execute applications on target hardware platforms to measure actual performance characteristics, but they incur substantial overhead from exhaustive profiling across all possible hardware configurations, particularly problematic as the number of heterogeneous processing units grows [7]. Moreover, both approaches face challenges when dealing with emerging applications or novel workload patterns for which historical performance data is unavailable or limited. Recent advances in machine learning, particularly self-supervised representation learning through contrastive learning, offer promising solutions to these challenges [8]. Contrastive learning has demonstrated remarkable success in computer vision, natural language processing, and other domains by learning meaningful representations from unlabeled data through maximizing agreement between augmented views of the same instance while minimizing agreement between different instances [9]. The key advantage of contrastive learning lies in its ability to discover useful features automatically without requiring extensive manually-labeled training datasets, which are expensive and time-consuming to obtain in heterogeneous computing contexts [10]. By treating different execution traces of similar workloads as positive pairs and traces from dissimilar workloads as negative pairs, contrastive learning can extract robust representations that capture essential workload characteristics relevant for scheduling decisions. This paper makes several key contributions to address the workload characterization and predictive scheduling challenges in heterogeneous computing environments. First, we develop a novel contrastive learning framework specifically designed for heterogeneous computing workload analysis, incorporating domain-specific data augmentation strategies that preserve semantically meaningful workload properties while introducing sufficient variation to enable effective contrastive learning [11]. Second, we design a specialized encoder architecture that efficiently processes execution traces and system telemetry data to generate compact yet informative workload representations suitable for downstream scheduling tasks [12]. Third, we propose a predictive scheduling algorithm that leverages the learned workload representations to make intelligent task-to-resource assignment decisions, considering both historical performance patterns and current system state [13]. Finally, we conduct comprehensive experimental evaluation using real-world heterogeneous computing workloads and demonstrate significant performance improvements over state-of-the-art scheduling heuristics across multiple metrics including makespan reduction, resource utilization efficiency, and adaptability to dynamic workload conditions. 2. LITERATURE REVIEW The challenge of effective task scheduling in heterogeneous computing environments has attracted significant research attention over the past decades [14]. Early approaches focused on list-based scheduling heuristics that assign priorities to tasks based on various metrics such as critical path length, computational weight, or data dependencies [15]. The Heterogeneous Earliest Finish Time (HEFT) algorithm, proposed as one of the most influential list-based heuristics, prioritizes tasks based on upward rank values and assigns each task to the processor that minimizes its earliest finish time [16]. While HEFT and its variants have demonstrated good performance for many application scenarios, they rely on simplifying assumptions about workload behavior and system characteristics that may not hold in dynamic production environments [17]. More sophisticated scheduling approaches have incorporated machine learning techniques to improve decision quality [18]. Early machine learning-based schedulers employed supervised learning methods that train predictive models using historical execution data to estimate task execution times on different processors [19]. However, these approaches face significant challenges related to data labeling requirements, as obtaining ground-truth execution time measurements for all possible task-processor combinations is prohibitively expensive [20]. Furthermore, supervised learning models trained on specific workload distributions often exhibit poor generalization when encountering novel application patterns or system configurations not represented in the training data [21].
Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [3] Recent developments in self-supervised learning have opened new possibilities for learning from unlabeled data [22]. Contrastive learning, in particular, has emerged as a powerful paradigm for representation learning across diverse domains [23]. The fundamental principle of contrastive learning involves constructing positive pairs of similar instances and negative pairs of dissimilar instances, then training an encoder to maximize agreement between positive pairs while minimizing agreement between negative pairs [24]. SimCLR introduced a simple framework for contrastive learning of visual representations that achieved state-of-the-art performance on image classification tasks using only self-supervised pretraining [25-27]. MoCo proposed a momentum-based contrastive learning approach that maintains a dynamic dictionary of encoded representations to enable more effective negative sampling [28-30]. Applications of contrastive learning have expanded beyond computer vision to include natural language processing, graph neural networks, and time series analysis [31]. In the context of system optimization and resource management, contrastive learning has been applied to problems such as program optimization, performance prediction, and anomaly detection [32]. However, the application of contrastive learning specifically to heterogeneous computing workload characterization and scheduling remains relatively unexplored [33, 34]. Some recent work has investigated using deep learning for workload prediction in cloud computing environments, but these approaches typically rely on supervised learning with labeled training data rather than self-supervised representation learning [35]. 3. METHODOLOGY 3.1 Problem Formulation and Task Graph Representation We formulate the heterogeneous computing scheduling problem using a directed acyclic graph (DAG) representation, which naturally captures both computational requirements and inter-task dependencies. A workload is represented as G = (V, E), where V is the set of tasks and E represents communication dependencies between tasks. Each task v_i in V has associated execution times on different processor types, forming a computation cost matrix W where w_i,j denotes the execution time of task v_i on processor p_j. The edges in E are weighted by communication costs c_i,j representing data transfer overhead between dependent tasks. Figure 1: the task graph representation As illustrated in Figure 1, a typical heterogeneous workload consists of multiple interdependent tasks with varying computational characteristics. The computation cost matrix reveals significant performance heterogeneity across different processor types, with execution time ratios ranging from 1.75 to 10.75 for different tasks. For instance, task T4 shows minimal variation across processors (execution times of 7, 10, and 4 time units), indicating it is relatively insensitive to processor selection, while task T3 exhibits substantial heterogeneity (execution times of 32, 27, and 43 time units), suggesting careful processor assignment is critical for this task. This heterogeneity in execution behavior motivates our contrastive learning approach, which aims to automatically discover workload
Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [4] patterns that correlate with processor affinity without requiring manual feature engineering or exhaustive profiling. The scheduling objective is to assign each task to a processor and determine execution order such that precedence constraints are satisfied and overall makespan is minimized. Traditional approaches compute static priority rankings based on task attributes such as upward rank, critical path position, or estimated finish times. However, these heuristics struggle to capture complex interactions between workload characteristics and hardware capabilities that determine actual performance in heterogeneous systems. Our contrastive learning framework addresses this limitation by learning representations that encode nuanced workload-hardware affinity patterns from historical execution data. 3.2 Contrastive Learning Framework Architecture Our proposed contrastive learning framework for workload characterization consists of several key components that work together to learn meaningful representations from unlabeled execution traces. The overall architecture follows the general contrastive learning paradigm but incorporates domain-specific modifications tailored to heterogeneous computing workloads. The framework takes as input raw execution traces collected from heterogeneous computing systems and produces compact vector representations that encode essential workload characteristics relevant for scheduling decisions. The data preprocessing pipeline begins by collecting comprehensive execution telemetry from the heterogeneous computing cluster. For each task execution, we record a rich set of features including CPU utilization patterns, memory access frequencies, cache miss rates, instruction mix statistics, inter-processor communication volumes, and actual execution times on different processor types. These raw measurements are then normalized and structured into fixed-length feature vectors that serve as input to the contrastive learning framework. The preprocessing stage also includes filtering outliers, handling missing values through interpolation, and applying dimensionality reduction techniques to manage the high-dimensional nature of execution telemetry data. A crucial aspect of our framework is the data augmentation strategy specifically designed for heterogeneous computing workloads. Unlike image augmentations in computer vision that apply geometric transformations, our augmentations preserve the semantic meaning of workload characteristics while introducing controlled variations. We employ three primary augmentation techniques including temporal jittering that randomly shifts execution trace timestamps within small windows to account for minor timing variations, feature masking that randomly masks subsets of telemetry features to encourage the model to learn robust representations not dependent on any single feature, and noise injection that adds small Gaussian noise to numerical features to improve generalization. By applying pairs of randomly selected augmentations to each execution trace, we generate positive pairs that represent the same underlying workload pattern despite surface-level variations. 3.3 Neural Network Encoder Architecture Design The encoder network architecture is designed to effectively process the sequential nature of execution traces while capturing both local and global patterns. Drawing inspiration from recent advances in deep learning for time series and sequential data, we develop a hybrid architecture that balances representational capacity with computational efficiency. The architecture evolution reflects the progression from simple feedforward networks to more sophisticated recurrent models capable of capturing temporal dependencies in workload behavior. Figure 2: the architecture comparison between ANN and LSTM
Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [5] As shown in Figure 2, the architecture comparison reveals the effectiveness of different neural network configurations for encoding workload characteristics. We begin with a simple Artificial Neural Network (ANN) as a reference model, which serves as a baseline for evaluating more sophisticated architectures. The simple ANN processes execution trace features through fully connected layers but lacks the ability to capture temporal dependencies inherent in workload execution patterns. To address this limitation, we introduce Long Short-Term Memory (LSTM) models that maintain hidden state across time steps, enabling the network to learn sequential patterns in resource utilization and performance metrics. The single-layer LSTM model demonstrates improved performance over the simple ANN by capturing short-term temporal correlations in workload behavior. However, for complex workloads with multi-scale temporal patterns, deeper architectures prove beneficial. The two-layer deep LSTM model adds an additional recurrent layer followed by a feedforward projection layer, allowing the network to learn hierarchical temporal representations where lower layers capture fine-grained execution patterns and higher layers abstract to workload-level characteristics. For the most demanding workloads, we employ a three-layer deep LSTM architecture that further extends the representational capacity through additional recurrent processing layers followed by multiple feedforward layers for final representation projection. The encoder architecture processes execution traces by first applying the LSTM layers to extract temporal features from the sequential telemetry data. The recurrent layers use gated mechanisms to selectively retain or forget information across time steps, enabling long-range dependencies to be captured while avoiding vanishing gradient problems during training. The final LSTM hidden state is then passed through fully connected projection layers that map the high-dimensional LSTM output to a lower-dimensional embedding space suitable for contrastive learning. This compressed representation serves as the workload signature used for downstream scheduling decisions. 3.4 Contrastive Loss Function and Training Procedure The training objective for our contrastive learning framework is formulated using the InfoNCE loss function, which has been shown to be effective for learning discriminative representations. For a given batch of execution traces, we construct positive pairs by applying different augmentations to the same trace and negative pairs by considering traces from different workloads. The contrastive loss encourages the encoder to produce similar representations for positive pairs while pushing apart representations from negative pairs in the embedding space. Mathematically, for an anchor trace representation z_i and its positive pair z_j, the loss maximizes their cosine similarity while minimizing similarity to all negative samples in the batch. The training procedure follows an iterative optimization process where we sample mini-batches of execution traces, apply augmentations to generate positive pairs, compute representations using the encoder network, calculate the contrastive loss, and update encoder parameters using stochastic gradient descent with momentum. We employ several training techniques to improve convergence and representation quality including temperature scaling in the loss function to control the concentration of the learned representations, large batch sizes to provide more negative samples per training iteration which improves the quality of learned representations, and learning rate scheduling with warmup and cosine decay to ensure stable training dynamics. The encoder is trained for multiple epochs until convergence, typically measured by monitoring the contrastive loss on a validation set of execution traces. 3.5 Predictive Scheduling Algorithm Design Once the contrastive learning framework has been trained to produce high-quality workload representations, we leverage these representations in a predictive scheduling algorithm. The scheduling algorithm operates in two stages including an offline precomputation phase where we encode a library of historical execution traces to build a database of workload representations paired with their observed performance characteristics on different processor types, and an online scheduling phase where incoming tasks are encoded using the trained encoder, matched against the historical database to find similar workload patterns, and assigned to processors based on predicted execution times derived from similar historical executions. This approach combines the benefits of learned representations with instance-based reasoning to make scheduling decisions. The matching process in the online scheduling phase uses cosine similarity in the learned embedding space to identify the top-k most similar historical workloads for each incoming task. For each candidate processor type, we aggregate the execution time statistics from these similar historical workloads to predict the expected execution time of the new task. The aggregation uses inverse-distance weighting where more similar historical workloads
Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [6] contribute more strongly to the prediction. This approach provides robustness against individual outliers while still allowing the predictor to adapt to fine-grained variations in workload characteristics. The final scheduling decision considers both predicted execution times and current processor availability to minimize overall makespan while balancing load across the heterogeneous system. 4. RESULTS AND DISCUSSION 4.1 Experimental Setup and Performance Heterogeneity Analysis We evaluate our proposed contrastive learning-based scheduling framework using a comprehensive experimental setup that includes both synthetic and real-world heterogeneous computing workloads. The evaluation platform consists of a heterogeneous cluster with diverse processing units including eight Intel Xeon CPUs with different generations and core counts, four NVIDIA GPUs ranging from V100 to A100 architectures, and two Xilinx FPGAs with varying logic capacities. This hardware diversity creates a challenging scheduling environment with substantial performance variations across processor types for different workload categories. Figure 3: the performance comparison among different benchmarks As shown in Figure 3, the performance comparison reveals dramatic heterogeneity in workload-processor affinity across different benchmark applications. Several workloads demonstrate extreme GPU speedups exceeding 20x compared to CPU execution, particularly for data-parallel computations such as dense matrix operations and signal processing kernels. For instance, the leftmost benchmarks in Figure 3 show GPU speedups of 82x and 20x, indicating these workloads are highly suitable for GPU acceleration due to their massive data parallelism and regular memory access patterns. Conversely, control-intensive workloads and those with irregular memory accesses show minimal GPU benefit or even CPU superiority, with speedup values close to or below 1.0x. This heterogeneity pattern underscores the critical importance of intelligent workload characterization and processor assignment. Simple heuristics that always prefer GPU execution would achieve excellent performance for data-parallel workloads but severely degrade performance for control-intensive applications. Similarly, conservative policies that default to CPU execution would miss substantial acceleration opportunities for GPUamenable workloads. The wide performance variance across workloads motivates our contrastive learning approach, which automatically learns to distinguish workload characteristics that correlate with processor affinity through self-supervised representation learning rather than relying on manually-engineered features or exhaustive profiling. Our evaluation employs multiple metrics to assess scheduling performance from different perspectives. The primary metric is makespan, defined as the total time from the submission of the first task until the completion of the last task in a given workload set. We also measure resource utilization efficiency by calculating the percentage
Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [7] of time that each processor spends executing tasks versus idling, with higher utilization indicating better resource exploitation. Additionally, we evaluate prediction accuracy by comparing predicted execution times against actual measured times using mean absolute percentage error. Finally, we assess the scheduling algorithm's ability to adapt to workload diversity by measuring performance consistency across workload categories with different computational characteristics. 4.2 Performance Comparison and Analysis The experimental results demonstrate that our contrastive learning-based scheduling framework achieves substantial performance improvements over traditional scheduling heuristics. Compared to the baseline HEFT algorithm, our approach reduces average makespan by 28 percent across the complete workload suite, with particularly strong improvements of 35 to 40 percent for workloads exhibiting high computational diversity and moderate task parallelism. The performance gains stem from the framework's ability to accurately characterize workload behavior and predict execution times on different processor types, enabling more informed scheduling decisions that better match tasks to appropriate hardware resources. The learned representations effectively capture the workload characteristics that determine processor affinity. For GPU-amenable workloads with high data parallelism, the contrastive learning encoder produces representations that cluster closely in the embedding space and are accurately matched to historical GPU-accelerated executions. When these workloads are scheduled, the predictor correctly identifies GPU assignment as optimal based on similar historical patterns. Conversely, for control-intensive workloads with irregular execution behavior, the encoder generates distinct representations that match to historical CPU executions, leading to appropriate CPU assignment decisions. This automatic discovery of workload-hardware affinity patterns eliminates the need for manual feature engineering or domain expertise in workload classification. Resource utilization analysis reveals that our framework achieves 34 percent higher average processor utilization compared to baseline methods. This improvement results from better load balancing across heterogeneous resources and reduced idle time through more accurate execution time predictions. The contrastive learning representations enable the scheduler to identify subtle patterns in workload behavior that traditional heuristics miss, such as memory-intensive tasks that benefit from processors with larger cache hierarchies or communication-intensive applications that perform better when co-located on processing units with low-latency interconnects. By capturing these nuanced characteristics, our framework makes more effective use of available hardware resources. The prediction accuracy analysis shows that our contrastive learning-based execution time predictor achieves a mean absolute percentage error of 12 percent, compared to 31 percent for baseline regression models and 45 percent for simple average-based predictions. The improved prediction accuracy directly translates to better scheduling decisions, as the scheduler can more confidently assign tasks to processors with genuinely lower execution times rather than making decisions based on noisy or inaccurate estimates. Notably, the prediction accuracy remains relatively stable across different workload categories, demonstrating that the learned representations generalize effectively to diverse application types rather than overfitting to specific workload patterns seen during training. Performance consistency across workload diversity is measured by examining the coefficient of variation in makespan across different workload categories. Our framework exhibits a coefficient of variation of 0.18, substantially lower than 0.34 for HEFT and 0.41 for random scheduling, indicating more predictable and reliable performance across diverse workload types. This consistency is particularly valuable in production environments where workload characteristics vary dynamically and scheduling algorithms must maintain acceptable performance without requiring manual tuning for each workload category. The self-supervised nature of contrastive learning enables the framework to automatically discover relevant workload features rather than relying on hand-crafted heuristics that may work well for some applications but poorly for others. 4.3 Ablation Studies and Component Analysis To understand the contribution of different components in our framework, we conduct ablation studies that systematically remove or modify key design choices. First, we examine the impact of the contrastive learning pretraining stage by comparing against a variant that uses the same encoder architecture but trains directly on the scheduling task using supervised learning with labeled execution time data. The contrastive pretraining version achieves 18 percent lower makespan, demonstrating that the self-supervised representations learned through
Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [8] contrastive learning capture more useful features for scheduling than representations learned purely through supervised training on limited labeled data. Second, we evaluate the importance of domain-specific data augmentations by comparing our tailored augmentation strategy against generic augmentation techniques borrowed from other domains. Using our heterogeneous computing-specific augmentations yields 11 percent better makespan compared to using standard augmentation techniques naively applied to execution traces. This improvement confirms that carefully designing augmentations to preserve semantic workload properties while introducing meaningful variations is crucial for effective contrastive learning in this domain. The temporal jittering and feature masking augmentations prove particularly valuable for capturing workload invariances, while noise injection provides marginal but consistent improvements in representation robustness. Third, we analyze the impact of the encoder architecture by comparing our LSTM-based design against simpler alternatives including a pure multilayer perceptron baseline and different recurrent network configurations. The deep LSTM architecture achieves the best performance with 8 percent lower makespan than single-layer LSTM and 14 percent lower than the multilayer perceptron. The recurrent layers effectively capture temporal patterns in resource utilization over execution windows, while the deeper architecture enables the model to learn hierarchical representations where lower layers capture fine-grained execution dynamics and higher layers abstract to workload-level characteristics. This architectural choice proves essential for learning high-quality representations from sequential execution traces that encode both short-term execution patterns and long-term workload behavior. 5. CONCLUSION This paper has presented a novel framework leveraging contrastive learning for workload characterization and predictive scheduling in heterogeneous computing environments. Our approach addresses fundamental limitations of traditional scheduling heuristics by automatically learning meaningful workload representations from unlabeled execution traces through self-supervised contrastive learning. The framework incorporates domain-specific data augmentations that preserve semantic workload properties while introducing controlled variations, a deep LSTM encoder architecture that effectively processes sequential execution telemetry to capture temporal dependencies, and a predictive scheduling algorithm that exploits learned representations to make informed task-to-resource assignment decisions. Extensive experimental evaluation demonstrates substantial performance improvements over state-of-the-art scheduling methods, including 28 percent reduction in average makespan and 34 percent improvement in resource utilization efficiency. The experimental analysis of performance heterogeneity across CPU and GPU processors reveals dramatic variations in workload-hardware affinity, with GPU speedups ranging from negligible to over 80x depending on workload characteristics. This heterogeneity underscores the critical importance of accurate workload characterization, which our contrastive learning framework addresses through automatic feature discovery rather than manual engineering. The learned representations exhibit strong generalization across diverse workload categories and maintain consistent performance without requiring manual parameter tuning for different application types. The success of our approach demonstrates that self-supervised contrastive learning provides an effective paradigm for learning workload representations that capture the nuanced patterns determining performance in heterogeneous systems. By formulating workload characterization as a contrastive learning problem, we enable the automatic discovery of features that correlate with processor affinity without requiring exhaustive profiling or domain expertise. The use of DAG-based task graph representations naturally captures both computational requirements and inter-task dependencies, providing a structured input format that the neural encoder can effectively process. The progression from simple feedforward networks to deep LSTM architectures reflects the increasing complexity required to accurately model temporal execution patterns and hierarchical workload characteristics. Ablation studies confirm that each component of the framework contributes meaningfully to overall performance, with contrastive pretraining, domain-specific augmentations, and the deep LSTM encoder architecture all playing essential roles. The contrastive pretraining stage enables learning from abundant unlabeled execution data rather than requiring expensive labeled datasets, while the carefully designed augmentations ensure that learned representations capture workload-invariant features rather than spurious execution artifacts. The deep LSTM architecture provides the representational capacity needed to encode complex temporal patterns in workload behavior that simpler models cannot capture. The success of contrastive learning for heterogeneous computing scheduling suggests several promising directions for future research. First, extending the framework to handle dynamic workload arrivals and online adaptation
Volume-09 Issue 12, December -2025 ISSN: 2456-9348 Impact Factor: 8.232 International Journal of Engineering Technology Research & Management (IJETRM) https://ijetrm.com/ IJETRM (http://ijetrm.com/) [9] could further improve performance in production environments where workload characteristics evolve over time. The current offline-online two-stage approach could be enhanced with continual learning techniques that update the encoder and predictor as new execution data becomes available, enabling the system to adapt to shifting workload distributions or hardware configurations. Second, incorporating multi-objective optimization criteria beyond makespan minimization, such as energy efficiency or fairness constraints, would broaden the applicability of the approach. The contrastive learning framework could be extended to learn multi-faceted representations that encode both performance and energy characteristics, enabling the scheduler to make Pareto-optimal decisions that balance competing objectives. Third, investigating transfer learning techniques to enable rapid adaptation to new hardware platforms or workload types with minimal retraining could reduce deployment overhead. The learned representations might transfer across different heterogeneous systems, allowing a model trained on one cluster to be fine-tuned for another with limited additional data. Fourth, combining contrastive learning with reinforcement learning might enable the development of fully adaptive scheduling policies that continuously improve through interaction with the heterogeneous computing environment. The learned workload representations could serve as state features for a reinforcement learning agent that learns scheduling policies through trial and error, potentially discovering novel strategies that outperform human-designed heuristics. These future directions hold significant potential for advancing the state-of-the-art in heterogeneous computing resource management and enabling more efficient utilization of increasingly complex computing infrastructures. As heterogeneous systems continue to proliferate across cloud datacenters, edge computing platforms, and highperformance computing facilities, the need for intelligent workload characterization and scheduling will only grow. Self-supervised learning techniques like contrastive learning offer a scalable and effective approach to this challenge, automatically adapting to the specific characteristics of each deployment without requiring extensive manual configuration or domain expertise. REFERENCES 1) Zhou, W., Yang, L., Zhao, L., Zhang, R., Cui, Y., Huang, H., ... & Wang, C. (2025). Vision technologies with applications in traffic surveillance systems: A holistic survey. ACM Computing Surveys, 58(3), 1-47. 2) Hong, C. H., & Varghese, B. (2019). Resource management in fog/edge computing: a survey on architectures, infrastructure, and algorithms. ACM computing surveys (csur), 52(5), 1-37. 3) Jiang, B., Cao, J., Tan, Y., & Qiu, S. (2025). Deep Learning Architectures for Sequential Decision-Making in Financial Systems: From Fraud Detection to Risk Management. Journal of Banking and Financial Dynamics, 9(9), 1-11. 4) Guo, Z., Tang, Y., Zhai, J., Yuan, T., Jin, J., Wang, L., ... & Li, R. (2024). A Survey on Performance Modeling and Prediction for Distributed DNN Training. IEEE Transactions on Parallel and Distributed Systems. 5) Wang, X., Tang, Z., Guo, J., Meng, T., Wang, C., Wang, T., & Jia, W. (2025). Empowering edge intelligence: A comprehensive survey on on-device ai models. ACM Computing Surveys, 57(9), 1-39. 6) Qian, C., Zhang, M., Nie, Y., Lu, S., & Cao, H. (2023). A survey on bug deduplication and triage methods from multiple points of view. Applied Sciences, 13(15), 8788. 7) St John, T. (2021). Performance analysis and optimization for extreme scale systems. University of Delaware. 8) Yang, Y., Ding, G., Chen, Z., & Yang, J. (2025). GART: Graph Neural Network-based Adaptive and Robust Task Scheduler for Heterogeneous Distributed Computing. IEEE Access. 9) He K, Fan H, Wu Y. Momentum contrast for unsupervised visual representation learning. IEEE Conference on Computer Vision and Pattern Recognition. 2020:9729-9738. 10) Jing, L., & Tian, Y. (2020). Self-supervised visual feature learning with deep neural networks: A survey. IEEE transactions on pattern analysis and machine intelligence, 43(11), 4037-4058. 11) Robinson, J., Chuang, C. Y., Sra, S., & Jegelka, S. (2020). Contrastive learning with hard negative samples. arXiv preprint arXiv:2010.04592. 12) Grill, J. B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., ... & Valko, M. (2020). Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems, 33, 21271-21284. 13) Zbontar J, Jing L, Misra I. Barlow twins: Self-supervised learning via redundancy reduction. International Conference on Machine Learning. 2021;139:12310-12320. 14) Kocot, B., Czarnul, P., & Proficz, J. (2023). Energy-aware scheduling for high-performance computing systems: A survey. Energies, 16(2), 890.