scieee AI-readable full text Open interactive document viewer

Biclustering in bioinformatics using big data and High Performance Computing applications: challenges and perspectives, a review

López Fernández, Aurelio; Gómez-Vela, Francisco A.; Rodríguez–Baena, Domingo S.; Delgado-Chaves, Fernando M.; Gonzalez‑Dominguez, Jorge

Abstract

Biclustering is a powerful machine learning technique that simultaneously groups rows and columns in matrix-based datasets. Applied to gene expression data in bio informatics, its use has expanded alongside the rapid growth of high-throughput sequencing technologies, leading to massive and complex biological datasets. This review aims to examine how biclustering methods and their validation strategies are evolving to meet the demands of High Performance Computing (HPC) and Big Data environments. We present a structured classification of existing approaches based on the computational paradigms they employ, including MPI/OpenMP, Apache Hadoop/Spark, and GPU/CUDA. By synthesising these developments, we highlight current trends and outline key research challenges. The knowledge gathered in this work may support researchers in adapting and scaling biclustering algorithms to analyse large-scale biomedical data more efficiently. Our contribution is intended to bridge the gap between algorithmic innovation and computational scalability in the context of bioinformatics and data-intensive applications.

Full text

Vol.:(0123456789) The Journal of Supercomputing (2025) 81:1123 https://doi.org/10.1007/s11227-025-07563-6 Biclustering inbioinformatics using big data andHigh Performance Computing applications: challenges andperspectives, areview AurelioLópez‑Fernández4· FranciscoA.Gomez‑Vela1· DomingoS.Rodriguez‑Baena1· FernandoM.Delgado‑Chaves2· JorgeGonzalez‑Dominguez3 Accepted: 6 June 2025 © The Author(s) 2025 Abstract Biclustering is a powerful machine learning technique that simultaneously groups rows and columns in matrix-based datasets. Applied to gene expression data in bioinformatics, its use has expanded alongside the rapid growth of high-throughput sequencing technologies, leading to massive and complex biological datasets. This review aims to examine how biclustering methods and their validation strategies are evolving to meet the demands of High Performance Computing (HPC) and Big Data environments. We present a structured classification of existing approaches based on the computational paradigms they employ, including MPI/OpenMP, Apache Hadoop/Spark, and GPU/CUDA. By synthesising these developments, we highlight current trends and outline key research challenges. The knowledge gathered in this work may support researchers in adapting and scaling biclustering algorithms to analyse large-scale biomedical data more efficiently. Our contribution is intended to bridge the gap between algorithmic innovation and computational scalability in the context of bioinformatics and data-intensive applications. Keywords Biclustering· Big data· High Performance Computing· Bioinformatics Abbreviations ACV Average correlation value AWS Amazon web services CC Cheng–Church CCS The condition-dependent correlation subgroups CPU Central processing unit CUDA Compute unified device architecture EBI European bioinformatics institute EBIC Evolutionary-based bIClustering ENCODE Encyclopedia of DNA elements Extended author information available on the last page of the article A.López-Fernández et al. 1123 Page 2 of 52 FCA Formal concept analysis FLOC FLexible overlapped biClustering FPGA Field-programmable gate array GB Gigabyte GBC Geometric biclustering algorithm GMQL GenoMetric query language GPGPU General-purpose computing on graphics processing units GPU Graphics processing unit GTEX Genotype-tissue expression project HDFS Hadoop distributed file system HPC High Performance Computing KEGG Kyoto encyclopedia of genes and genomes LCS Longest common sequence MATLAB MATrix laboratory MCC Matthews correlation coefficient MPI Message passing interface MR MapReduce MSR Mean square residue NGS Next-generation sequencing NMF Non-negative matrix factorisation NMI Normalised mutual information PBD-SPEA2 Parallel biclustering detection PC Parallel coordinate PCR Polymerase chain reaction PGAS Partitioned global address space POSIX Portable operating system interface PPI Protein–protein interaction RAM Random access memory RDD Robust distributed datasets RNA Ribonucleic acid SM NVIDIA Streaming multi-processor SMP Shared memory parallelism SPEA2 Strength pareto front evolutionary algorithm2 SPMD Single program, multiple data SPUs Stream processor units SPs Streaming processors SQL Structured query language STRING Search tool for the retrieval of interacting genes/protein TCGA The cancer genome atlas UM Unified memory Biclustering inbioinformatics using big data andHigh… Page 3 of 52 1123 1 Introduction Biclustering techniques have been successfully applied to several research areas, including energy consumption [1], economics [2], trading forecasting [3], marketing [4], or recommendation systems [5]. In bioinformatics, these techniques are often used for applications such as analysing gene expression data, discovering and annotating new functionalities for unclassified genes, finding expression modules, reconstructing biological networks, elucidating disease mechanisms, and stratifying patients [6, 7]. The application of biclustering techniques on gene expression datasets has two main advantages over clustering: (a) it groups both genes and experimental conditions, which is much closer to biology since a subset of genes may have common biological behaviour only under a subset of experimental conditions or samples [8], and (b) it considers group overlapping, allowing genes to contribute to more than one biological activity. Despite these advantages, several challenges arise, mainly due to the exponential growth of biological datasets—in both volume and complexity—generated by next-generation sequencing (NGS) technologies [9] and the creation of large-scale genomic consortiums [10]. The complete annotation and quantification of the expression level of all genes and their isoforms in a specific sample [11], along with the data volume that represents gene sequencing information from 100 million to 2 billion individual patients in the year 2025 [10], demonstrate some challenges faced in this field. Also, it is essential to consider that this data may originate from various sources, be stored in several formats, and typically contain noise and artefacts that require resolution [12]. This increases the computational complexity of biclustering methods used to evaluate and extract meaningful information from large biomedical and biological datasets. To address these challenges, High Performance Computing (HPC) approaches, such as General-Purpose Computing on Graphics Processing Units (GPGPU), Big Data frameworks, and traditional parallel and distributed computing paradigms, have been adopted to improve performance and scalability [13]. Despite having proven highly useful in multiple contexts, conventional biclustering techniques exhibit a number of technical and computational limitations when dealing with extremely large datasets. These limitations can be summarised as follows: • High computational complexity: Conventional techniques often aim to find optimal solutions at the cost of high computational complexity. This becomes unsustainable when input matrices contain millions of rows and columns. For example, the first true biclustering algorithm, the Cheng–Church algorithm [14], has approximately quadratic complexity in the number of rows and columns, making it poorly scalable. • Memory limitations: Many algorithms store entire matrices or multiple auxiliary structures in memory, which becomes unfeasible with very large datasets. For instance, algorithms like SAMBA [15] may encounter memory management issues when handling gene expression datasets with more than 100,000 genes and conditions. A.López-Fernández et al. 1123 Page 4 of 52 • Sensitivity to noise and redundancy: In large datasets, noise is inevitable and can lead to many spurious patterns. Classical techniques are not always designed to effectively filter out this noise. Algorithms such as Bimax [16] search for exact biclusters (i.e. noise-free), which makes them less useful for large-scale real-world data. In addition to these three main issues, there is also a lack of adaptability to distributed structures and an inability to handle heterogeneous or dynamic data. The importance of studying biclustering under the Big Data and HPC paradigms lies in the growing need to extract biologically meaningful patterns from massive and complex datasets. For example, in the work presented by Liu etal. [17], five gene expression datasets were processed, generating a total of 8280536 biclusters. From a genomic consortium point of view, another example can be found in The Cancer Genome Atlas (TCGA) [18], which provides multi-omics data for thousands of tumour samples across numerous cancer types. The identification of biclusters in this context can help discover cancer-specific gene expression modules, stratify patients based on molecular profiles, and support the design of targeted therapies. However, the volume and complexity of these types of dataset means that traditional biclustering methods are insufficient. So a systematic examination of new biclustering approaches that integrate robust and efficient computational strategies with biological interpretability is required. Since the publication of the work by Sara C. Madeira etal. [19], many studies have been presented recently whose objective is to present a classification or comparative study. Despite this wealth of information, to the best of our knowledge, none of these papers has focused their study on how new or even existing biclustering techniques have met the challenge of being able to adapt to a Big Data or HPC ecosystem to process the biological and biomedical datasets that are currently available. There is also no related work on the ability of biclustering techniques to validate the large number of results they generate or how bicluster validation methods address this problem. Moreover, in order to cope with a Big Data ecosystem, these techniques should also take into account the various challenges they must face when trying to adapt to some specific HPC methods. However, there are also no works that address the challenges involved in this type of adaptation. This paper presents an overview for researchers and practitioners to discover suitable biclustering solutions for large-scale bioinformatics applications while promoting future research that connects biological knowledge with computational scalability. This work explores the evolution of biclustering techniques in response to the increasing scale and complexity of bioinformatic datasets. It examines the incorporation of High Performance Computing (HPC) and Big Data frameworks as key enablers for improving the scalability and applicability of these methods in real-world biomedical scenarios. Furthermore, it analyses the principal limitations and ongoing challenges that persist in this interdisciplinary field, highlighting the need for continued methodological innovation to ensure both computational efficiency and biological interpretability. Therefore, we can summarise the contributions of this work as follows: Biclustering inbioinformatics using big data andHigh… Page 5 of 52 1123 • A systematic and up-to-date survey of biclustering techniques that address the challenges posed by large-scale biomedical data. • A classification of these methods based on their computational paradigms and practical implementations. • A critical analysis of their biological applications and the reported computational performance based on a comprehensive review of the existing literature. • A summary of current challenges and future directions in this field. The structure of the rest of the paper is as follows: Section3 presents a list of the most relevant biclustering reviews. Section4 groups the problems that HPC and Big Data applications used by biclustering techniques must address to support huge datasets and obtain their results in the shortest possible time. Biclustering techniques based on Big Data and HPC applications are examined in Section 5. Section 6 reviews the different bicluster validation methods and validation processes implemented in biclustering techniques that can validate large amounts of results in the shortest possible execution time. Section7 describes the challenges that biclustering and its validation measures must face to adapt to a Big Data ecosystem. Finally, Section8 summarises the main conclusions derived from this work. 2 Review methodology This study employs a structured process based on the PRISMA 2020 principles to guarantee scientific rigour, transparency, and repeatability. Two distinct yet interconnected analyses were conducted to encapsulate the recent advancements in biclustering methodologies employed in bioinformatics, particularly in light of Big Data and high-performance computing issues. The initial study concentrated on discovering review articles and surveys that investigate biclustering techniques used in bioinformatics. The aim was to evaluate the depth, breadth, and emphasis of previous evaluations, specifically with their handling of scalability, validation, and practical usefulness in bioinformatics. The second study focused on original research articles that proposed or implemented biclustering methods inside High Performance Computing (HPC) and Big Data environments. This encompasses biclustering methods using traditional parallel and distributed models (e.g. MPI, OpenMP), Big Data frameworks (e.g. Apache Spark, Hadoop), and GPGPU acceleration (e.g. CUDA, OpenCL). The objective is to delineate trends, problems, and the technological deficiencies in scalable bicluster deployments and their applications within the bioinformatics domain. 2.1 Search strategy anddata sources Both analyses employed systematic search strategies across three major databases: Scopus, PubMed, and Google Scholar. The searches were carried out in May 2025 and were limited to peer-reviewed articles published between 1 January 2009 and 25 May 2025. Only articles written in English were included. A.López-Fernández et al. 1123 Page 6 of 52 For the review and survey works, queries combined terms such as: ”biclustering”, ”co-clustering”, ”survey”, ”review”, ”bioinformatics”, and ”gene expression”. For the HPC/Big Data implementations, queries included combinations of: ”biclustering”, ”co-clustering”, ”parallel computing”, ”distributed computing”, ”high-performance computing”, ”HPC”, ”Big Data”, ”Apache Spark”, ”Hadoop”, ”GPU”, ”GPGPU”, ”CUDA”, and ”bioinformatics”. In all cases, Boolean operators and platform-specific syntax were used. Although all analyses were performed for all platforms, some illustrative queries for each database are shown below. PubMed query (survey and review identification): ((biclustering[Title/Abstract] OR ”coclustering”[Title/Abstract]) AND (review[Title/ Abstract] OR survey[Title/Abstract] OR benchmark*[Title/Abstract]) AND (bioinformatics[MeSH Terms] OR ”gene expression”[Title/Abstract])) AND (”2009/01/01”[Date - Publication] : ”3000”[Date - Publication]) Scopus query (Big Data-based biclustering algorithms): (TITLE-ABS-KEY(biclustering OR ”co-clustering”) AND TITLE-ABS-KEY(”Big Data” OR ”Spark” OR ”Apache Spark” OR ”Hadoop” OR ”Apache Hadoop” OR ”MapReduce” OR ”Map-Reduce”) AND TITLE-ABS-KEY(bioinformatics OR ”gene expression”)) AND PUBYEAR> 2008 Google Scholar query (GPU-based biclustering algorithms): ”biclustering” AND (”GPU” OR ”CUDA” OR ”GPGPU” OR ”OpenCL”) AND (bioinformatics OR ”gene expression”) -review -survey after:2008 As can be seen, each query was carefully crafted to match the syntax and indexing strategies of the corresponding platform. Advanced filters (e.g. publication year range, language, peer-reviewed journals) were also applied when available. 2.2 Eligibility criteria andscreening process Articles were managed using Zotero, which enabled deduplication and structured tagging. The selection was carried out in two stages: (1) title and abstract screening and (2) full-text evaluation. Separate inclusion/exclusion criteria were defined for each analysis. For survey analysis, studies were included if they were published reviews or surveys that analysed biclustering techniques applied specifically to biological data. Priority was given to work focused on bioinformatics or omics datasets, particularly those involving gene expression analysis. Studies were excluded if they did not relate Biclustering inbioinformatics using big data andHigh… Page 7 of 52 1123 to biological applications—such as reviews centred on marketing or social network data—or if they lacked peer-reviewed status (e.g. editorials or informal overviews). For the analysis of the HPC/Big Data biclustering algorithm analysis, articles were included if they proposed or applied biclustering techniques within parallel or distributed computing environments, including implementations using GPU or GPGPU acceleration. Works that described applications within Big Data frameworks such as Apache Spark or Hadoop were also eligible. The studies that were not included in this analysis were those that did not have any computational implementation, did not focus on scalability or performance, or only used simple toy datasets that were not relevant to processing large-scale biological data. The first analysis yielded 490 candidate review articles, of which 24 met all inclusion criteria. The second analysis retrieved 432 implementation articles, of which 24 were selected after full-text screening. 2.3 Data extraction andclassification For both analyses, metadata were extracted, including publication year, citation count (via Scopus/Google Scholar), journal quartile, target organism (if applicable), data type (e.g. microarray, RNA-Seq), computational platform (e.g. Spark, MPI, GPU), and evaluation metrics. Studies from the second analysis were further categorised by computational paradigm: • Parallel and distributed computing (e.g. MPI, OpenMP) • Big Data frameworks (e.g. Spark, Hadoop) • GPU-based approaches (e.g. CUDA, OpenCL) This classification enabled a comparative synthesis of algorithmic strategies and implementation patterns. 2.4 Review scope andtype This work combines both systematic and narrative review elements. The systematic components provide a reproducible structure and empirical foundation, while the narrative analysis enables critical interpretation of trends, bottlenecks, and methodological gaps. This hybrid approach allows for both evidence-based synthesis and forward-looking insight. 2.5 Evidence ofinternational interest andrelevance In both analyses, the articles included cover more than 20 countries, with strong representation from institutions in the USA, China, Germany, India, and Spain. Approximately 74% of the selected studies were published in journals ranked in Q1/Q2 in their category according to the Scimago Journal Rank (SJR). It is also interesting to note that of these publications ranked in Q1, the selected surveys exceed the 80% A.López-Fernández et al. 1123 Page 8 of 52 threshold, while for HPC biclustering algorithms they stand at 67%. Citation metrics also support the relevance of this field: the average citation count was 47.3, with a median of 31. These indicators reinforce the growing international interest in scalable biclustering techniques in biomedical research. 3 Related work Numerous biclustering reviews are available in the literature, with the most relevant ones detailed in Table2 of Appendix A. This table includes publication year, authors, principal application, and its scope: biclustering techniques (methods) and validation techniques (validation). Reviews can span multiple categories, as the groups are not mutually exclusive. Regarding biclustering strategies, S. Busygin etal. [20] offer an extensive analysis of the mathematical principles involved in the search for biclusters by means of various techniques that are classified based on their application domains. B. Pontes etal. [21] classify algorithms into metric and non-metric groups, employing diverse evaluation metrics such as iterative greedy search, meta-heuristics, clustering-based, probabilistic models, and linear algebra. Recent works have been published that provide a more comprehensive analysis, such as the study by A. José-García etal. [22], which specifically examines biclustering algorithms that use meta-heuristics as evaluation metrics. Concerning reviews of bicluster validation techniques, Tanay etal. [23] examine scoring model influence in the detection of significant biclusters, while Santamaria etal. [24] propose metrics to illustrate the consistency and efficacy of biclustering techniques to get biological information in the presence of noise exposure and overlap between biclusters. Although the aforementioned articles provide theoretical perspectives on biclustering approaches, they fail to address the issues posed by biological datasets and computational obstacles. Thus, additional evaluations seek to acquire knowledge by conducting comparative analyses of various methodologies. For example, Eren etal. [25] found that the computational efficiency of biclustering approaches declines with larger gene expression datasets and more biclusters. Additionally, not all methods can generate biologically enriched biclusters. Padilha etal.’s work [26] reiterated the findings of Eren etal. [25] and provided further details on the factors that influence computational efficiency, including noise in gene expression data, bicluster overlap, size, and quantity. Also, K. Nicholls etal. [27] compared various biclustering techniques, noting that adaptive algorithms demonstrate superior computational performance. They also observed a correlation between noise, dataset size, bicluster count, and decreased computational efficiency. More recently, Castanho etal. [28] have offered another study in which they introduce a unified taxonomy of biclustering concepts and detail a comprehensive analysis process. Although they identify three main implementation strategies for big data biclustering algorithms—parallel computing, GPU-based execution, and MapReduce—they limit the discussion to whether each method is a novel proposal or an adaptation of existing algorithms. Based on the highlights from the aforementioned reviews, it can be determined that biclustering efforts have long been trying to address the limitations and improve the quality of the results. However, refining meta-heuristics and addressing noise/ Biclustering inbioinformatics using big data andHigh… Page 9 of 52 1123 overlap challenges may not fully extract pertinent biological insights from vast datasets. Likewise, current bicluster validation methods may lack adequacy when applied in Big Data ecosystem. Regarding these concerns, Xie etal. [6] present a compelling study on the impact of biclustering techniques in current biomedical contexts. They highlight the lack of biclustering algorithms tailored for large gene expression datasets with high complexity. The authors advocate for the development of novel biclustering methods capable of managing biological and biomedical analyses within Big Data environments. Also, there are other works that are not reviews but that identify such problems. For example, in [29], the authors focus on Big Data biclustering approaches using HPC frameworks like Apache Hadoop and Spark, along with parallel systems with GPUs, to address scalability challenges. Traditional biclustering methods are deemed impractical for Big Data tasks due to their inability to analyse data efficiently. The authors outline key challenges, including leveraging parallelism, understanding hardware limitations, and addressing the increase in bicluster quantity and size with data volume. They recommend employing multiple validation measures when assessing biclustering performance in noisy environments or with similar biclusters. Additionally, they caution against over-interpreting statistical significance in gene enrichment and pathway analysis, advocating for additional filtering or comparison in Big Data biclustering approaches. As the analysis of current literature reviews on biclustering techniques shows, many proposals examine these techniques from various perspectives. However, none of these works addresses the challenges posed by the big data environment, such as processing enormous amounts of data and evaluating a large number of biclusters. This article aims to address this gap by providing a comprehensive study of biclustering techniques from the perspective of high-performance computing (HPC). 3.1 Critical analysis ofprevious work andidentified research gaps Although a significant number of challenges have contributed to the development of biclustering methods for bioinformatics and Big Data environments, several limitations and research gaps persist. First, many traditional biclustering algorithms exhibit limited scalability when applied to large-scale omics datasets. These methods are often implemented in sequential environments and lack efficient adaptations for distributed or parallel computing frameworks. Second, the interpretability of the resulting biclusters is frequently overlooked, which restricts their usability in biomedical contexts where the explainability of computational results is crucial for domain experts. Another recurrent issue is the absence of standardised evaluation methodologies; the diversity of metrics and validation strategies employed across studies hampers fair comparisons and reproducibility. Moreover, there is a general lack of integration of prior biological knowledge into the biclustering process. Gene ontology terms, known pathways, and regulatory interactions are often ignored, despite their potential to guide more biologically meaningful clustering. Furthermore, while modern machine learning paradigms such as deep learning and ensemble learning have shown success in other bioinformatic tasks, their application in biclustering remains limited. The reproducibility A.López-Fernández et al. 1123 Page 16 of 52 expression matrix into chunks corresponding to available processes. Each process is responsible for constructing local biclusters from its allocated data chunk, while the main process integrates these biclusters to produce the final output. As data transmission among processes is avoided during local bicluster construction, the authors find the communication cost to be negligible. However, they note that the method’s bottleneck arises from constructing biclusters with a single primary process, which can significantly affect computing performance, especially with exceptionally large datasets. To validate their approach, the authors employ a synthetic dataset measuring 12651 x 30 to generate 250 biclusters and claim similar results for a yeast gene expression dataset [76], though no supporting data is provided. Liu etal. [77] introduce P-Bicluster, a biclustering method designed for constructing shifting biclusters [78] from gene expression datasets containing both continuous and discrete values. The technique gathers shift patterns into a 2x2 matrix and utilises the deviation from each pattern to determine the shift trend. Biclusters are then generated by incrementally adding rows and columns to each pattern until their offset value falls below a predefined threshold. The authors evaluate their approach using real gene expression arrays [79] without noise control and compare results with other sequential methods, assessing bicluster quality based on deviation. However, validation lacks support from techniques ensuring the biological significance of biclusters. In terms of computational properties, P-Bicluster employs MPI for parallel computing on CPU clusters. The authors observe improved computational performance with several processes up to seven but note that performance declines beyond this threshold due to a lack of control over distributed memory architecture and process communications. Hence, the method is ineffective for continuous data transfer or large data volumes. Additionally, the study lacks details on parallelised tasks, CPU cluster utilisation, synchronisation, and data distribution, among other aspects. A. Nisar etal. [80] present a scalable and distributed algorithm derived from Ahmad et al.’s work [81], which reformulates the Bipartite Spectral Partitioning technique [82] using graph tracing to achieve optimal solutions. Their method involves positioning each node of the bipartite graph at the barycenter of its neighbours to identify optimal solutions. The biclustering process consists of two parallel tasks: crossover minimisation and bicluster identification. Initially, the gene expression dataset is horizontally split into homogeneous chunks, each assigned to an available process. Data normalisation is performed, followed by iterative row (local) and column (global) reordering. Each process conducts a local search for biclusters, and representative biclusters are globally distributed. The authors employ MPI/ C++ to accelerate computational performance and analyse algorithmic complexity and memory architecture management. They note that MPI-based biclustering algorithms incur communication costs as dataset size grows but claim independence from row count by using crossover minimisation. They demonstrate scalability by analysing synthetic datasets with up to 20 million rows and 64 fixed columns and using up to 256 processors. However, the study lacks analysis of real gene expression datasets and comparison with other methods or computational acceleration techniques. Biclustering inbioinformatics using big data andHigh… Page 17 of 52 1123 MFCM [83] is a biclustering technique derived from the C-Means fuzzy clustering algorithm [84] for generating biclusters from gene expression datasets. The method utilises MatlabMPI and an SPMD parallel computing model to distribute computational workload across multiple processors. The dataset is divided into chunks corresponding to MATLAB processes, followed by gene and sample clustering to construct biclusters. Gene centres are determined based on entropy rather than correlation coefficients, with iterations continuing until a threshold of rejected genes or a maximum iteration limit is reached. Real gene expression datasets are used, comparing results to a sequential version in MATLAB. Computational performance improves at all steps compared to the sequential version, but challenges may arise with increased processes or large datasets due to communication delays. The study lacks consideration of challenges in data distribution, memory architecture management, and data transfers across processes, and it only evaluates up to eight processes, leaving performance uncertainty with more. The Bioconductor runibic package [85] adapts the UniBic biclustering technique [86] for parallel computing to generate meaningful biclusters from gene expression data. The modification aims to accommodate high-throughput RNA-Seq, scRNASeq, and PCR datasets, improving runtimes for large datasets with hundreds of thousands of columns and thousands of rows. The authors provide detailed computational decisions for parallelising tasks, except for matrix discretisation. They reimplement the algorithm in C++11 with the OpenMP standard for enhanced performance. Restructuring the source code optimises features and addresses memory segmentation issues. OpenMP is utilised for tasks like computing the Longest Common Sequence (LCS) [87] between row pairs, distributed across CPU cores. However, evidence on the algorithm’s performance with huge datasets and the rationale behind OpenMP usage is lacking. Consequently, the scalability and performance of the OpenMP implementation depend on single-machine hardware. González-Dominguez et al. [88] developed ParBiBit, a parallel version of the BiBit algorithm [89], designed to extract biclusters from binary datasets. ParBiBit, implemented in C++11, uses MPI and POSIX threads to exploit both shared and distributed memory in a hybrid cluster environment. Tasks with high computational costs, such as pattern construction and row aggregation, are identified. MPI and POSIX threads distribute the workload across CPUs and nodes to enhance computational performance. However, data transfers between processes may saturate the network, impacting performance. López et al. [90] demonstrated ParBiBit’s degradation in performance with denser datasets and identified memory management concerns for large datasets. Fraguela etal. [91] addressed these issues with ScalaParBiBit, also using C++11, MPI, and POSIX Threads. Sequential optimisations and methodological tweaks were applied to enhance computational speeds and support larger datasets. The main difference lies in ScalaParBiBit’s approach to generating patterns; it aims to distribute pattern sets to optimise bicluster elaboration. However, noise sensitivity in datasets remains unaddressed across all versions. The COBRAC Convex biclustering algorithm [92] is capable of identifying potential biclusters along with the associated row and column clustering dendrograms. This algorithm operates by reducing the input dataset size and solving the convex weighted biclustering problem. The utilisation of convex clustering trees A.López-Fernández et al. 1123 Page 18 of 52 ensures robustness against minor input variations [93]. COBRAC is implemented in C/C++ for the algorithm itself and Python containers for broader accessibility. OpenMP is employed to accelerate the weighted convex biclustering task. While reducing the input dataset size improves execution time and enables handling larger biological datasets, further computational documentation is needed to elucidate their results and resource usage. Additionally, the rationale behind choosing OpenMP should be clarified, especially regarding its dependence on single-machine hardware. Although the authors prioritise improving usability, running simulations on a machine with 512 GB of RAM may not reflect typical computing environments. Further exploration with larger datasets is warranted to evaluate the algorithm’s scalability and performance impact. The ARBic algorithm developed by Ma etal. [94] is an accelerated rule-based biclustering algorithm designed to identify relevant gene condition patterns in biomedical datasets. Implemented in C++, ARBic employs OpenMP to parallelise the expansion phase of seed biclusters, which is computationally intensive due to the high dimensionality of omics data. OpenMP enables ARBic to exploit shared memory parallelism by distributing the processing of candidate seeds across multiple threads. This parallel strategy significantly reduces execution time during the exploration of the bicluster search space. However, the use of OpenMP confines the algorithm to shared memory architectures, limiting scalability for larger datasets that exceed the memory capacity of a single node. Furthermore, the insertion of results into shared structures requires critical sections, which may introduce synchronisation overhead and limit parallel efficiency. To enhance ARBic’s performance, future versions could adopt hybrid parallelism by integrating OpenMP with MPI or taskbased runtime systems, allowing distributed memory support and more effective workload balancing. Additionally, optimising memory access patterns and reducing contention in shared data structures could further improve scalability and parallel throughput. B. Baruah etal. [95] introduce EnsemBic, an assembly-based biclustering method that integrates several methods (Laplace Prior, iBBiG, and xMotif) to extract functionally relevant biclusters from gene expression data. The method identifies highquality biclusters using p-values derived from FuncAssociate 3.0 and employs an iterative approach to eliminate genetic similarities within biclusters. The authors emphasise that the design facilitates the execution of base algorithms in parallel, potentially decreasing computation time; however, the R implementation does not explicitly incorporate parallelisation techniques such as OpenMP, MPI, or native R libraries for parallel computing. This seeming flexibility indicates that parallelisation is either externally managed or subject to user implementation. EnsemBic has been assessed using four authentic data sets (two microarrays and two RNA-seq), demonstrating consistent enhancements over standalone methods in both topological metrics (internal density, modularity, ODF) and biological metrics (GO enrichment, pathway analysis, and validation by ChIP-seq). Moreover, the inclusion of the p-value as a quality criterion renders the approach sensitive to non-significant patterns and enhances its robustness against noise. The algorithm functions on real continuous data; however, its use with synthetic, discontinuous, or binary data has not been documented. Biclustering inbioinformatics using big data andHigh… Page 19 of 52 1123 There are additional biclustering techniques that are accelerated using conventional parallel models, but insufficient information is available due to the unavailability of software or a lack of description regarding the parallel model and computational resources. One such example is PBD-SPEA2 [96], an adaptation of the SPEA2 multi-objective evolutionary algorithm [97] for gene expression data. PBD-SPEA2 uses a fixed-length, integer-based encoding for evaluating bicluster sets. Objective functions include MSR [14], BSize [22], and rVAR [98], with the Cheng–Church [14] heuristic for producing new solutions. Parallel computing is applied to the dynamic coding scheme task, although specific parallelisation methods, standards, or models are not provided, nor is the algorithm’s source code available, hindering further exploration of computational features. In MPI-based biclustering algorithms, the size and diversity of input datasets play a crucial role in speeding up computing performance. Larger datasets often lead to memory management and data transfer challenges due to communication saturation among memory regions across processors in the cluster. ScalaParBiBit offers a solution to these limitations. On the other hand, biclustering algorithms using OpenMP face limitations inherent in this implementation, which primarily parallelises programs in a shared memory environment, avoiding communication overhead but restricting them to a single computer’s hardware resources. The COBRAC algorithm exemplifies these constraints by compressing dataset size to reduce execution time compared to its sequential version. Additionally, considerations such as task selection for parallelisation, architecture, and memory management must be addressed. While algorithms may face limitations due to hardware resources, such an approach ensures scalability without overloading or memory overflow. Furthermore, as dataset volume increases, the likelihood of obtaining more biclusters rises, requiring careful consideration during processing and memory management. 5.2 Parallel biclustering withbig data programming paradigms The emergence of Big Data programming paradigms addresses the limitations of traditional parallel and distributed models in handling large datasets efficiently. Consequently, programming paradigms like MapReduce, along with platforms such as Apache Hadoop, Apache Spark, and MATLAB, offer solutions for massively parallel or distributed biclustering. Two tables present the characteristics of the analysed biclustering methods for your reference. Table6 of Appendix A outlines the characteristics of biclustering techniques tailored for a Big Data environment. The first column lists the names of the biclustering algorithm or authors, if not specified. The second column indicates the Big Data platform used. The types of datasets used in experiments are detailed in the third column: 0 for synthetic datasets, 1 for real datasets, and 2 for a combination of both. The fourth column denotes whether comparisons were made with other studies: 0 for no comparison, 1 for comparisons with sequential methods, and 2 for comparisons with other Big Data-adjusted biclustering algorithms or deployment on different Big Data platforms. The fifth column notes the algorithm’s ability to handle noise. Data types used in experiments are A.López-Fernández et al. 1123 Page 20 of 52 listed in the sixth column, acknowledging that additional data types might also be compatible. Table7 of Appendix A delineates the computational characteristics of the discussed biclustering methods. The first column presents the algorithm’s name or authors’ names. The second column indicates the number of MR jobs required for generating results, where 0 signifies a small number and 1 indicates a higher number. The third column denotes the method’s capability for data splitting, crucial for balancing workload and data transfers [99]. The fourth column indicates whether methods have strategies to minimise I/O operations, which is critical for performance, especially in Apache Hadoop setups. The fifth column specifies if intermediate data is stored in main memory for in-memory computation, reducing disk I/O. The sixth column identifies whether MR jobs are asynchronous, impacting hardware and software resource utilisation. The seventh column evaluates the investigation of optimisation strategies, which are crucial for efficient parallelism. The final column indicates the type of scalability experiments conducted: 1 for increasing data volume, 2 for raising the number of cluster nodes, and 3 if both scalability tests were performed. BiTM-MR [100] is an early biclustering approach using the MapReduce paradigm, implemented with Apache Spark for extracting biclusters from massive datasets. The method transforms data matrices into a block structure organised as a topological map and employs two binary matrices to validate row–column relationships. Overcoming challenges like data loading, failure safety, and algorithm design, the authors use Apache Spark for fault correction, data distribution, and management. The method uses two MR jobs to iterate rows and columns and adjust parameters. The authors significantly improve computational efficiency by allocating independent MR jobs for rows and columns, minimising I/O operations, and enabling asynchronous processing. However, developing a data partitioning technique such that each MR job takes care of a group of homogeneous rows and columns is an additional computing factor that the authors have not accounted for and that could be useful for enhancing their algorithm design. Synthetic datasets with up to two million rows demonstrate near-ideal computational efficiency as dataset size and cluster cores increase. However, comparisons with other biclustering algorithms or performance acceleration resources would have added value. The authors note that Apache Spark initialisation impacts performance more on smaller datasets. Ruiqi etal. [101] adapted the non-negative matrix factorisation biclustering algorithm [102] using Apache Hadoop to handle 2D sparse matrices up to one million by one million. Their approach requires a non-negative matrix and a repetition count as input, using five MR jobs per iteration to update each output submatrix. However, this approach’s heavy reliance on MR jobs may impact computational performance on large datasets due to frequent disk access costs. Considerations such as reducing I/O operations, optimising data partitioning, and enabling asynchronous communication between map and reduce functions were not addressed. The study suggests room for significant improvement in both computational cost and scalability. An evaluation with synthetic datasets and a dataset exceeding one million rows and columns from the STRING database [103] was conducted, analysing computational efficiency concerning the number of nonzero interactions. However, there was no Biclustering inbioinformatics using big data andHigh… Page 21 of 52 1123 analysis of how well it scales with more Hadoop cluster nodes or how it compares to other biclustering methods or different resources for handling large datasets. MR-GABiT [104] is designed for genetic biclustering of time series datasets from microarray experiments. Despite the datasets being three-dimensional, the authors apply biclustering to extract local patterns at different time points, which are then compared to determine optimal patterns. Using the MATLAB MapReduce framework, they address dimensionality and scalability issues. Data partitioning by time points reduces MR job count, improving computational performance. Each MR task employs a map function to extract the best individual from data chunks, with reduce functions combining these best individuals. However, limitations include the inability to reduce disk I/O operations and reduce functions waiting for all map functions to complete. Evaluation with yeast cell cycle time series datasets [105] is conducted, but the scalability with increased cluster nodes is not tested, and there is no comparison with other genetic biclustering methods or Big Data resources. There are additional biclustering studies using the MapReduce paradigm where computational details are insufficient and the software is unavailable. One such work is MCC [106], an adaptation of the Cheng–Church (CC) biclustering technique [14]. MCC randomly generates numerous unique subarrays and executes more computations on them than the original CC algorithm, making it computationally intensive. While MapReduce is used, specifics on its adaptation, data partitioning, and distribution are not provided. Real gene expression datasets show that MCC works better than the sequential CC algorithm in terms of accuracy and speed. The authors acknowledge MCC’s inefficiency in memory management due to the need to store numerous intermediate arrays and the high number of I/O operations. As can be seen, this subsection highlights the scarcity of bioinformatics-related biclustering algorithms on Big Data platforms like Spark or Hadoop. Consequently, this subsection also includes biclustering algorithms that have been created in disciplines of science unrelated to bioinformatics but that are based on the MapReduce paradigm. Then, it is possible to create ais as to whether the lack of algorithms is due to the complexity of the biclustering technique’s definition or to scientific fieldspecific factors in bioinformatics. DisCo [107] rearranges submatrix rows and columns to meet a quality threshold, using key–value pairs to store rows. Bhatnagar etal. [108] developed a MapReduce biclustering technique using formal concept analysis (FCA) for binary datasets [109]. Their focus is on improving computational performance as dataset size increases, although they have not explored scalability with node numbers, and the BiBit sequential biclustering method achieves faster execution speeds in eight of the fourteen overall performance evaluations. Lin etal. [4] developed biclustering algorithms for the telecommunications field, with versions for Spark (SP-PLSS) and Hadoop (MR-PLSS). They aim to identify profitable customers with similar buying behaviours, resulting in better efficiency and scalability with Apache Spark for large datasets. This analysis demonstrates that, regardless of the field of study, biclustering techniques have not been widely implemented on Big Data platforms such as Apache Spark and Apache Hadoop. In addition, there is a significant gap between the present number of biclustering algorithms produced on these Big Data platforms and other types of computational acceleration applications. As a result, multiple hypotheses may emerge from diverse viewpoints. A.López-Fernández et al. 1123 Page 22 of 52 Many traditional biclustering methods need to constantly process one part of the dataset, which makes it difficult to design these methods [14, 89, 94, 110, 111]. Consequently, not all strategies easily adapt to the MapReduce paradigm due to this complexity. The algorithm design must consider factors like MR job count, dataset partitioning, minimising I/O operations, in-memory systems, and asynchronous communication between map and reduce functions. Processing large datasets can greatly enhance performance, but the differences in bioinformatics datasets might make it hard for these algorithms to work well, especially when there are only a few genes and experimental conditions [108]. Also, Big Data platform initialisation time impacts performance [49, 100], and issues like reduce functions waiting for map function completion and uncontrolled execution order can degrade performance. Biclustering algorithms on Apache Spark outperform those on Apache Hadoop due to intermediate I/O operations relying on memory rather than disk [4]. 5.3 Parallel biclustering onGPUs The GPGPU concept harnesses the processing resources of GPU devices to accelerate computational performance, making them viable alternatives or supplements to CPU processing for biclustering algorithms [112]. However, biclustering algorithms tailored for GPGPU must consider various factors outlined in subsection4.3 to optimise computational efficiency and support Big Data datasets. These factors are analysed for each biclustering algorithm adapted to GPGPU and presented in two tables. Table8 of Appendix A provides key features of biclustering techniques implemented on GPGPU. The first column lists algorithm names, while the second column indicates dataset types: 0 for synthetic, 1 for real, and 2 for both. The third column indicates if comparisons were made with other methods: 1 for sequential methods and 2 for other GPGPU-adapted algorithms. The fourth column notes if the approach handles dataset noise. The fifth column specifies the data types used, acknowledging that other types may also be applicable. The final columns detail supported data types for each algorithm. Table9 of Appendix A details the computational aspects of biclustering algorithms. The first column lists algorithm names. The second column specifies multiGPU support and the parallelism technique. Shared memory usage is noted in the third column. The fourth column indicates the use of coalescence. The memory allocation method is specified in the fifth column (0 for non-pinned, otherwise pinned). The data transfer method (global memory or unified memory) is described in the sixth column. Occupancy levels, reflecting resource utilisation, are in the seventh column. To better reflect the computational footprint of the surveyed tools, we classified resource occupancy into three categories based on estimated average usage of computational resources: low (<40%), medium (40–70%), and high (>70%). The final column indicates how synchronisation points are used. Arnedo-Fernández etal. [113] introduced an adaptation of the FLOC biclustering method [114], focusing on the computational intensity of calculating the MSR measure [14] for bicluster validation. Utilising CUDA for this task, they demonstrate improved computational performance compared to the sequential version, using up Biclustering inbioinformatics using big data andHigh… Page 23 of 52 1123 to 2000 x 2000 square synthetic datasets. However, they did not fully exploit the GPU’s capabilities, noting data transfer via the PCI-Express bus as a bottleneck and not considering multi-GPU workload splitting. The authors also discussed the impact of thread count per CUDA block on computational efficiency, observing diminishing returns with excessively high thread counts. Regarding memory management, the authors do not specify the various forms of memory present in a GPU device. Therefore, it is assumed that data is always used from the GPU’s global memory. Liu etal. [17] parallelised the task of processing column pairs in the Geometric Biclustering algorithm (GBC) [115]. They developed three parallel implementations: one using POSIX Threads, another using CUDA, and a third employing a field-programmable gate array (FPGA) [116]. In the CUDA version, columns are paired per CUDA block, with each thread responsible for combining a portion of each column using shared memory to ensure high coalescence and avoid memory conflicts. Real dataset experiments indicate that the GPU version achieves the highest speedup, while the FPGA version is the most energy-efficient. The authors discover that minimising data transfers between RAM and GPU global memory is crucial, especially for large datasets in the GPU version. In another work, Liu etal. [117] apply the GBC algorithm to identify neural processing patterns in microarray datasets, presenting three CUDA versions with distinct optimisations. The first version uses shared memory to store column pair chunks, ensuring no overhead regardless of chunk size. After loading the target column pair into shared memory, merging operations are conducted, and outcomes are stored in global memory. Moreover, the chunking of column pairs prevents the shared memory from overloading, necessitating a global synchronisation for every processed chunk. According to the authors, this global synchronisation incurs a cost by restoring inefficient column pairs in the global memory. The second version attempts operations on multiple column combinations to mitigate this, while the third version focuses on reducing index update time. In order to accomplish this, they examine the effect of continuous transfers between GPU devices and the CPU, a problem identified in prior work [17]. Real dataset experiments show the second version achieves the best speedup, emphasising the importance of data reuse, minimising global memory accesses, and distributing workloads between CPU and GPU for large datasets. Despite these optimisations, none of these versions implements a multi-GPU design, which could further enhance workload allocation and support for large datasets. Mejía-Roa etal. [118] introduced a multi-GPU adaptation of the non-negative matrix factorisation (NMF) [102] algorithm, using the MPI standard for multi-GPU synchronisation. The algorithm partitions the dataset into blocks and distributes it across GPUs, subsequently decomposing the data in global memory. A 1D block setup achieves coalescence by accessing consecutive memory addresses. Asynchronous CPU-to-GPU data transfers enhance computational performance. Real dataset experiments confirm that the GPU version outperforms the sequential CPU-based algorithm, with speedup increasing with dataset size. Scalability analysis reveals that factors like data transfers, synchronisation overheads, and dataset size are critical in multi-GPU designs [119]. For large datasets, the multi-GPU approach is preferable, while a single GPU suffices for smaller ones. A.López-Fernández et al. 1123 Page 24 of 52 The Condition-dependent Correlation Subgroups (CCS) [120] algorithm, developed in CUDA, fully utilises GPU devices for bicluster generation. Real and synthetic datasets demonstrate a 20x speedup compared to the sequential implementation. However, unlike previous works, the algorithm does not employ dataset chunking using CUDA block/thread decomposition, which limits its scalability to the GPU memory size. Shared memory expedites bicluster creation but may overload with excessive columns, capped at 200. Each CUDA block constructs a bicluster using a single thread, resulting in low occupancy. The method lacks support for multi-GPU architectures; thus, it does not explore workload distribution or scalability for large datasets. EBIC, developed by Orzechowski etal. [121], is a parallel evolutionary biclustering method designed to discover numerous relevant patterns with high precision. It uses a multi-GPU architecture to distribute dataset rows across GPUs, handling the entire bicluster generation process. The authors optimise workload distribution across GPUs, CUDA blocks, and threads by using shared memory for computing fitness functions and generating biclusters. Experimental results with synthetic and real datasets demonstrate EBIC’s computational efficiency, high coalescence, and occupancy. EBIC outperforms alternative techniques like CCS by up to 12 times for larger datasets. However, the algorithm’s scalability for Big Data is limited by the maximum support of 60000 rows per GPU device. González-Domínguez etal. [122] implemented CUBiBit, a multi-GPU version of the BiBit [89] biclustering algorithm. CUBiBit employs GPU devices to add rows to potential biclusters, with the remaining processing done on the CPU. In the case of a multi-core processor, it can exploit the multiple CPU cores while finding potential biclusters thanks to a parallelisation with POSIX threads. CuBiBit employs the GPU’s shared memory to store seeds for computational performance. However, this decision limits scalability for large datasets because this memory is tiny and the number of patterns increases as the dataset grows. Furthermore, the original method outperforms it for small datasets. gBiBit [90], another version of BiBit, uses a multiGPU architecture to process massive binary datasets efficiently. The methodology addresses data transfer, workload distribution, and resource utilisation by outperforming other adaptations and being the only modified version capable of handling big datasets. In another work, the authors created a Python package named bioScience that uses HPC to speed up various data mining techniques, such as the BiBit algorithm, using CPU and multi-GPU clusters [123]. The utilisation of GPU devices for accelerating biclustering algorithms has become a prevalent trend, with CUDA being the most used platform for this purpose. Various aspects need consideration when developing and optimising biclustering algorithms in a multi-GPU environment, as observed in the previous works. While GPGPU parallel algorithms generally outperform sequential ones in terms of computing performance, many struggle to handle massive datasets effectively. In the majority of the analysed algorithms, authors stress the importance of planning resource allocation for GPU utilisation. CCS, for instance, employs a CUDA block to parallelise a task, whereas CUBiBit, NMF, EBIC, and other algorithms use the threads of a CUDA block to do the parallelisation task. This decision is directly dependent on the algorithm’s occupancy rate and, consequently, on whether the Biclustering inbioinformatics using big data andHigh… Page 25 of 52 1123 implementation will use the full power provided by the grid of each GPU device. For algorithms supporting multi-GPU architectures, workload distribution among devices is carefully considered. These works use parallelisation techniques like POSIX Threads, MPI, and OpenMP to distribute tasks, but the most effective strategy remains unclear due to the variety of options available. Regarding memory-related elements, the authors of several works identify CPU–GPU data transfer as a performance bottleneck due to communication bus limitations, leading some to minimise transfers. Continuous global memory access can also impede efficiency, prompting the use of shared memory to boost speed. However, improper shared memory use, as seen in CuBiBit or CCS, can hinder performance on large datasets or smaller datasets. Thus, shared memory usage’s efficacy varies with dataset size and should be carefully assessed. Most algorithms employ fixed memory reservation and allocation strategies. Synchronisation between GPU devices is crucial in multi-GPU systems, but it varies based on algorithm technique. For example, NMF and gBiBit use asynchronous communications to minimise synchronisation and improve efficiency. 5.4 Performance considerations ingene expression data analysis Implementing biclustering algorithms using HPC technologies in big data environments has shown variable performance depending on the specific characteristics of the biological task. In the context of gene expression data analysis, which typically involves large-scale, high-dimensional, and sparse matrices, each computational model presents distinct advantages and limitations. Benchmark datasets are commonly utilised to assess the performance and scalability of HPC biclustering methods within a Big Data framework, ensuring both efficacy and biological significance. The Cancer Genome Atlas (TCGA) and the Genotype-Tissue Expression (GTEX) project [124] are the most frequently utilised resources. The TCGA offers extensive, high-dimensional gene expression data across many cancer types, serving as a valuable resource for assessing the scalability and precision of biclustering algorithms in real-world biomedical contexts. Conversely, GTEX provides comprehensive transcriptome data from several non-diseased human tissues, facilitating the identification of tissue-specific gene co-expression modules. The extensive utilisation of TCGA and GTEX stems from their public accessibility and comprehensiveness, as well as the chance they provide to evaluate algorithmic resilience in high-throughput settings, which is crucial for implementing biclustering approaches within HPC or Big Data frameworks. Synthetic datasets are utilised in various research to evaluate algorithm performance under controlled settings of noise, size, and structure. Table 1 provides a comparative summary of the main HPC paradigms discussed in this section, outlining their associated technologies, strengths, limitations, and typical applications in gene expression biclustering tasks. Traditional parallel and distributed models, such as MPI-based clusters, are proficient at managing large-scale gene expression datasets, especially in jobs characterised by significant data parallelism, such as gene or sample filtering. Nevertheless, their efficacy may diminish in iterative biclustering contexts due to inter-node A.López-Fernández et al. 1123 Page 32 of 52 Based on the research conducted in the preceding sections of this work, the primary challenges that biclustering must overcome in order to adapt to a Big Data ecosystem are outlined below. These challenges are categorised into four primary groups: datasets, biclustering approaches, biclustering validation approaches, and visualisation and interpretation of the results. 7.1 Data‑centric challenges Due to the exponential growth of biological and biomedical data alongside advancements in NGS [166] technologies, bioinformatics is confronted with significant challenges in storing and analysing vast datasets. The pace of data generation from sequencing outpaces the computational resources’ capacity to process such large volumes, leading to an expanding gap between them [167]. Numerous projects and repositories are emerging to store increasingly massive and intricate datasets. For instance, The Cancer Genome Atlas (TCGA) [18] amassed 2.5 petabytes of diverse biological data, including mRNA, miRNA, and protein expression data, along with histology slides and genetic variation data. The ENCODE project [168] focuses on annotating functional sequences in the human genome, expanding to include data from other organisms like mice, flies, and worms, totalling over five terabytes. The European Bioinformatics Institute (EBI) [169] stores more than 390 petabytes of raw biological data, making it one of the largest repositories globally. As Biclustering Pipeline in Big Data Bioinformatics with HPC Integration Data SourcesGenomic Data, Gene Expression Matrices (HDFS/Cloud Storage) PreprocessingNormalization, Feature Selection(Spark + GPU) Biclustering AlgorithmsDistributed FLOCK, Plaid(MPI + CUDA) VisualizationHeatmaps, D3.js, Tableau ApplicationsDrug Discovery, Multi-Omics Big Data HPC DATA SOURCES DATA PREPROCESING BICLUSTERING VISUALIZATION Validation & EvaluationGO Enrichment, Scalability Metrics VALIDATION APPLICATIONS Fig. 3 Bioinformatics Big Data biclustering pipeline with HPC integration (MPI/GPU), designed for scalability in applications such as drug discovery and multi-omics integration Biclustering inbioinformatics using big data andHigh… Page 33 of 52 1123 these repositories grow in volume, complexity, and diversity, extracting meaningful insights becomes increasingly challenging. The creation and preprocessing of datasets from NGS platforms can impact data precision and quality due to factors like low-quality reads, duplicate reads, or insertions/deletions [170]. Reads represent the sequenced base pairs (bp) from DNA fragments. Data transformation, such as from FASTQ to FASTA, is crucial for verifying data quality and eliminating noise [171]. However, the production and preprocessing of large datasets are time-consuming processes. To address these challenges, the scientific community is developing Big Data and HPC tools using distributed memory systems [172–174]. The continuous generation of biological and biomedical data presents challenges in storage and management. Cloud computing has emerged as an effective solution, as evidenced by various studies [175, 176]. Despite its benefits, efforts are underway to develop data compression techniques to reduce cloud computing costs [177]. Latency is another concern, as data retrieval from the cloud for scientific data analysis can be time-consuming. Shifting data processing to the cloud can address this issue, reducing latency and costs and accelerating result generation. Amazon AWS, for example, offers the Amazon AWS Genomics service, equipped with tools to process large volumes of genomic data efficiently [178, 179]. Cloud-based storage requires conducting data analysis computations in the cloud to mitigate latency in result generation. Scalable tools like SeqPig [180], BigBWA [181], GMQL [182], SeQuiLa-cov [183], and SeQual [184] ensure data integrity during quality control and preprocessing, especially when dealing with extensive raw sequencing data. 7.2 Algorithmic challenges Traditional biclustering methods have primarily focused on improving result quality, but they often struggle with handling large datasets, resulting in reduced usefulness. Challenges in producing biclusters from extensive datasets, influenced by various factors such as application domains, types, structures, or dimensions, contribute to this issue [6, 146]. Additionally, the characteristics of large biological datasets increase the likelihood of generating larger and more numerous biclusters, requiring the use of Big Data and HPC tools for expedited bicluster generation. To choose the appropriate Big Data or HPC application for designing a biclustering methodology, understanding the problem’s nature, parallelisable tasks, and dataset characteristics is crucial. For tasks requiring high processing capability with larger-than-usual but not massive datasets, parallel or distributed traditional models are recommended. Dataset size is pivotal, influencing memory management, data transfer, and communication latency, particularly in MPI-based methods when using distributed memory to capacity. This issue does not affect multi-threaded methods (e.g. OpenMP or POSIX Threads), which rely on a single computer’s hardware but preclude building CPU clusters for enhanced performance. To optimise multithreaded algorithm performance, it is recommended to limit parallel task complexity and compress datasets to control processing and storage costs. A.López-Fernández et al. 1123 Page 34 of 52 Platforms like Apache Hadoop or Apache Spark are recommended for handling large datasets requiring increased processing costs and computational resources, facilitating the development of biclustering solutions for such datasets. However, there exists a gap between the number of biclustering algorithms designed for these platforms and those for other applications aimed at improving computational performance. Adapting biclustering algorithms to these platforms to construct biclusters via dataset partitioning poses critical challenges such as the number of MapReduce jobs, intelligent dataset partitioning strategies, reduced I/O operations, and asynchronous communication between map and reduce functions. Furthermore, these biclustering techniques often exhibit poor performance on small datasets due to platform initialisation time, significantly reducing algorithm efficacy. Some comparative studies have shown that biclustering algorithms developed using Apache Spark offer superior computational performance compared to Apache Hadoop, attributed to reduced I/O operation costs and faster access speeds. Recently, GPU devices and the CUDA platform have become popular for accelerating biclustering algorithms, demonstrating optimal computational performance. However, not all GPU-based biclustering algorithms can handle huge datasets effectively, which poses a challenge in developing algorithms for such volumes. Maximising GPU power utilisation through resource planning is crucial, alongside minimising data transfers between CPU and GPU to mitigate bandwidth limitations. While shared memory offers speed, its use may lead to memory overflows in processing massive data volumes, requiring careful selection of its usage. Multi-GPU architectures can enhance performance, but workload distribution and synchronisation among GPU devices are essential. This distribution typically employs parallel or distributed methods like OpenMP, POSIX Threads, or MPI, while asynchronous communication across GPU devices has been shown to improve computational performance by omitting synchronisation. 7.3 Challenges invalidation Biclustering algorithms often generate a vast number of results over large datasets, which require subsequent validation [17, 29]. Typically, validation involves statistical or biological knowledge-based methods, often sourced from public databases, to ascertain the biological relevance of the generated outcomes [127]. Recent advancements in bicluster validation methods and tools have primarily focused on improving result accuracy and developing user-friendly interfaces. However, there is a notable absence of methodologies and tools capable of facilitating efficient comparative analysis of bicluster validation approaches. For instance, EBIC, a multi-GPU evolutionary biclustering method, limited its outcomes to 100 and validated them using the sequential software RGOStats [160], which operates solely with gene lists. There is a pressing need to develop validation techniques for large bicluster datasets. These techniques would enable researchers to validate biclusters derived Biclustering inbioinformatics using big data andHigh… Page 35 of 52 1123 from datasets managed by scientific initiatives or generated from next-generation sequencing. Additionally, it has been demonstrated that it is not possible to draw interesting biological knowledge from the huge number of biclusters generated by Big Data-adapted biclustering algorithms. 7.4 Visualisation andinterpretability issues After biclustering results are validated, researchers must focus on visualising and interpreting them in biological and biomedical contexts to derive reliable conclusions. Several visualisation techniques aid in interpreting biclustering algorithm outcomes. BicOverlapper [185] addresses the challenge of visualising bicluster overlap by using intersecting hulls, enabling the extraction of biological insights through interactive visualisation. BiVisu [186] employs Parallel Coordinate (PC) plots and objective metrics like MSR [14] and Average Correlation Value (ACV) [187] to assess bicluster homogeneity. BiDots [188] explores new visual and interactive methods for investigating weighted biclusters across domains. Other considerations in bicluster visualisation include the limitations of heatmaps and PC plots in demonstrating biological relevance [189]. The authors advocate for novel visualisation techniques and propose standards for biclustering algorithm outputs to facilitate visualisation tool integration. Additionally, network-based visualisation approaches simplify the interpretation of bicluster visualisation networks [159, 190]. These visualisation tools can help assess the coherence of biclusters formed by biclustering techniques, indicating their potential significance. However, it is still possible that some biclusters hold domain-relevant knowledge [146]. External sources or domain expert evaluation may be needed to confirm the authenticity of these biclusters. Despite efforts to standardise biclustering algorithm outputs and utilise visualisation and enrichment analysis tools for better biological interpretation, the sheer volume of biclusters generated in a Big Data setting can lead to a loss of functionality. Consequently, there is a need for visualisation tools tailored to enhance biological interpretation and manage large bicluster datasets. 8 Conclusions The growing volume of biological and biomedical data necessitates the adaptation of traditional biclustering methods to effectively manage large-scale datasets. To meet this challenge, recent advancements have focused on employing High Performance Computing (HPC) and Big Data frameworks—such as Apache Spark, GPU acceleration, and distributed/parallel architectures—to improve computational efficiency. Although these methodologies offer substantial benefits, they also introduce challenges related to communication overhead, memory management, and initialisation A.López-Fernández et al. 1123 Page 36 of 52 latency, particularly in MPI-based and MapReduce frameworks. GPU-based solutions, while powerful, require careful resource allocation, CUDA grid optimisation, and efficient data transfer between RAM and GPU memory to achieve peak performance. Beyond computational considerations, validation is essential for confirming the reliability and biological significance of biclustering outcomes. Contemporary validation strategies—based on statistical consistency or gene enrichment analysis—struggle to scale with the vast number of biclusters produced by modern algorithms. Consequently, many tools limit their assessments to a subset of results, highlighting the need for more scalable and automated validation techniques that leverage HPC infrastructure. This study has identified multiple avenues for improvement. Future research should focus on developing biclustering algorithms that are inherently scalable and optimised for distributed and heterogeneous computing environments. Improving the interpretability of biclustering results is also essential to facilitate adoption by biomedical researchers, which may involve the use of interpretable models or visualisation tools that clarify the biological relevance of the detected patterns. Standardised, open benchmarking frameworks that rely on real-world datasets are also needed to support reproducible and comparative assessments. One promising approach involves integrating prior biological knowledge— such as functional annotations or regulatory networks—into the biclustering process to guide the discovery of biologically meaningful structures. The field can also benefit from the incorporation of advanced machine learning methods, including deep learning, graph-based models, and emerging forms of artificial intelligence such as reinforcement learning and transformers. These technologies are expected to play a central role in the next generation of biclustering methods, enabling more adaptive, context-aware, and automated discovery of complex patterns in omics data. Expanding the application of biclustering to multi-omics integration and time series analysis could provide more comprehensive insights into dynamic biological systems. Furthermore, future algorithms should enhance robustness to noise through the use of robust statistical techniques and preprocessing strategies. Strengthening collaboration with experimental biologists is also key to validating computational findings invitro or invivo, thereby supporting their translational relevance. Looking ahead, anticipated mainstream trends include the integration of biclustering into automated machine learning (AutoML) pipelines, enabling non-expert users to conduct exploratory analyses in biomedical contexts. Additionally, the deployment of biclustering workflows within cloud-based HPC environments will enhance scalability and accessibility. Another important direction is the exploration of federated learning frameworks to enable secure Biclustering inbioinformatics using big data andHigh… Page 37 of 52 1123 and privacy-preserving biclustering across distributed biomedical datasets—a crucial consideration in multi-institutional and clinical settings. By addressing these limitations and opportunities, next-generation biclustering techniques can evolve into powerful, interpretable, and scalable tools for analysing high-dimensional biological data. The rising adoption of deep learning and AI-driven methodologies in bioinformatics represents a promising direction. Investigating how such models can be effectively adapted within HPC and Big Data ecosystems—particularly in alignment with the analytical strategies discussed in this paper—constitutes a significant area for future research. Appendix A: Detailed tables See Tables2, 3, 4, 5, 6, 7, 8 and 9 All the tables that are presented below are referenced in the main text. Table 2 A summary of the most significant biclustering reviews Year Authors Application Methods Validation 2004 Madeira etal. [19] Gene expression X 2005 Tanay etal. [23] Gene expression X 2007 Santamaría etal. [24] Gene expression X 2008 Busygin etal. [20] Biomedicine and text mining X 2010 Rastegar etal. [191] Gene expression X 2010 Verma etal. [192] Gene expression X 2013 Eren etal. [25] Gene expression X 2013 Orzechowski. [128] Gene expression X 2014 Oghabian etal. [193] Gene expression X 2014 Horta etal. [194] Gene expression X 2015 Pontes etal. [195] Gene expression X 2015 Pontes etal. [21] Gene expression X 2015 Mounir. [196] Gene expression X 2016 Biswal etal. [197] Gene expression X 2016 Anitha etal. [198] Gene expression X 2017 Padilha etal. [26] Gene expression X X 2018 Biswal etal. [199] Gene expression X 2018 Aouabed etal. [200] Biomedical X 2019 Xie etal. [6] Biological and biomedical X 2021 Nicholls etal. [27] Gene expression X 2021 Sozdinler, [189] Gene expression X 2022 Noronha etal. [146] Data mining and gene expression X 2022 José-García etal. [22] Gene expression X 2024 Castanho etal. [28] Biological and biomedical X X A.López-Fernández et al. 1123 Page 38 of 52 Table 3 Repositories where the source codes for the HPC biclustering algorithms are stored Year Method Source code URL 2007 RoBA [74] – 2008 DisCo [107] – 2008 P-Bicluster [77] – 2009 A.Nisar etal [80] – 2012 FLOC [113] – 2013 GBC [17] – 2014 GBC [117] – 2014 MCC [106] – 2014 BiTM-MR [100]https:// github. com/ Tugdu alSar azin/ sparkclust ering 2014 Cloudnmf [101]http:// admis. fudan. edu. cn/ proje cts/ Cloud NMF. html 2015 NMF [118]https:// github. com/ bioin focnb/ bionmfgpu 2015 Bhatnagar etal. [108] – 2015 MFCM [83] – 2016 MR-GABiT [104] – 2017 CCS [120]https:// github. com/ abhat ta3/ Condi tiondepen dentCorre lationSubgr oupsCCS 2017 PBD-SPEA2 [96] - 2018 EBIC [121]https:// github. com/ Epist asisL ab/ ebic 2018 Runibic [85]https:// www. bioco nduct or. org/ packa ges/ relea se/ bioc/ html/ runib ic. html 2018 ParBiBit [88]https:// sourc eforge. net/ proje cts/ parbi bit/ 2019 CuBiBit [122]https:// sourc eforge. net/ proje cts/ cubib it 2019 SP-PLSS [4] – 2021 ScalaParBiBit [91]https:// github. com/ fragu ela/ Scala ParBi Bit 2021 COBRAC [92]https:// github. com/ haidyi/ cvxbi clustr 2021 gBiBit [90]https:// github. com/ aurel iolfd ez/ gbibit Biclustering inbioinformatics using big data andHigh… Page 39 of 52 1123 Table 4 Main features of biclustering algorithms accelerated by traditional parallel and distributed models Data compatibility Year Method Datasets Comparative Noise Data Binary Discrete Continuous 2007 RoBA [74] 2 0 No Microarrays X 2008 P-Bicluster [77] 1 1 No Microarrays X X 2009 Nisar etal [80] 0 0 Yes Microarrays X 2015 MFCM [83] 1 2 No Microarrays X 2017 PBD-SPEA2 [96] 2 1 Yes Microarrays X 2018 Runibic [85] 1 1 Yes RNA-Seq X X 2018 ParBiBit [88] 2 1 No Microarrays X 2021 ScalaParBiBit [91] 0 2 No Microarrays X 2021 COBRAC [92] 1 0 Yes Microarrays X 2022 ARBic [94] 2 1 No Microarrays X 2023 EnsemBic [95] 1 1 Yes Microarrays & RNA-Seq X Table 5 Computational features of biclustering algorithms based on traditional parallel and distributed models Year Method Language Acceleration Optimisation Communications 2007 RoBA [74] MATLAB MPI No Yes 2008 P-Bicluster [77] ANSI C MPI No No 2009 A.Nisar etal [80] C/C++ MPI Yes Yes 2015 MFCM [83] MATLAB MPI No No 2017 PBD-SPEA2 [96] - - - - 2018 Runibic [85] C/C++ OpenMP No No 2018 ParBiBit [88] C/C++ MPI/POSIX No Yes 2021 ScalaParBiBit [91] C/C++ MPI/POSIX Yes Yes 2021 COBRAC [92] C/C++ OpenMP Yes No 2022 ARBic [94] C/C++ OpenMP No No 2023 EnsemBic [95] R - - - A.López-Fernández et al. 1123 Page 40 of 52 Table 6 Main features of biclustering algorithms supported by MapReduce platforms Data compatibility Year Method Platform Datasets Comparative Noise Data Binary Discrete Continuous 2008 DisCo [107] Hadoop 1 0 No - X 2014 MCC [106] - 1 1 No Microarrays X 2014 BiTM-MR [100] Spark 0 0 No - X 2014 Cloudnmf [101] Hadoop 2 0 No PPI X 2015 Bhatnagar etal. [108] Hadoop 1 2 No - X 2016 MR-GABiT [104] MATLAB 1 0 No Microarrays X X 2019 SP-PLSS [4] Spark 2 2 No – X X Biclustering inbioinformatics using big data andHigh… Page 41 of 52 1123 Table 7 Computational features of biclustering algorithms based on MapReduce platforms Year Method MR Jobs Partitioning I/O In-memory Asynchronous Speedup Scalability 2008 DisCo [107] 1 0 0 0 0 1 3 2014 MCC [106] 1 0 0 0 0 0 1 2014 BiTM-MR [100] 1 0 1 1 1 1 3 2014 Cloudnmf [101] 1 0 0 0 0 0 1 2015 Bhatnagar etal. [108] 1 0 0 0 0 0 3 2016 MR-GABiT [104] 0 1 0 0 0 1 1 2019 SP-PLSS [4] 1 1 1 1 0 1 3 Table 8 Main features of biclustering algorithms supported by GPGPU platforms Data compatibility Year Method Datasets Comparative Noise Data Binary Discrete Continuous 2012 FLOC [113] 0 1 No Microarrays X X 2013 GBC [17] 1 2 No Microarrays X 2014 GBC [117] 1 2 No Microarrays X 2015 NMF [118] 1 1 No Microarrays X 2017 CCS [120] 2 1 No Microarrays X X 2018 EBIC [121] 2 2 Yes Microarrays X X X 2019 CuBiBit [122] 0 1 No Microarrays X 2021 gBiBit [90] 2 2 No Microarrays & RNA-Seq X Table 9 Computational features of biclustering algorithms based on GPGPU platforms Year Method multi-GPU Shared mem. Coalescing Allocation Transfer Occupancy Sync. 2012 FLOC [113]No No No 0 Global Low No 2013 GBC [17]No Yes Yes 1 Global Low No 2014 GBC [117]No Yes Yes 1 Global High Yes 2015 NMF [118] Yes (MPI) Yes Yes 1 Global High Yes 2017 CCS [120]No Yes No 1 Global Low No 2018 EBIC [121] Yes (OpenMP) Yes Yes 1 Global High Yes 2019 CuBiBit [122] Yes (POSIX) Yes Yes 1 Global Low No 2021 gBiBit [90] Yes (POSIX) No Yes 1 Global High No A.López-Fernández et al. 1123 Page 48 of 52 115. Zhao H, Liew AW-C, Xie X, Yan H (2008) A new geometric biclustering algorithm based on the Hough transform for analysis of large-scale microarray data. J Theor Biol 251(2):264–274. https:// doi. org/ 10. 1016/j. jtbi. 2007. 11. 030 116. Gandhare S, Karthikeyan B (2019) Survey on fpga architecture and recent applications. In: 2019 International Conference on Vision Towards Emerging Trends in Communication and Networking (ViTECoN), pp. 1–4. https:// doi. org/ 10. 1109/ ViTEC oN. 2019. 88995 50 117. Liu B, Xin Y, Cheung RC, Yan H (2014) GPU-based biclustering for microarray data analysis in neurocomputing. Neurocomputing 134:239–246. https:// doi. org/ 10. 1016/j. neucom. 2013. 06. 049 118. Mejía-Roa E, Tabas-Madrid D, Setoain J, García C, Tirado F, Pascual-Montano A (2015) NMFmGPU: non-negative matrix factorization on multi-GPU systems. BMC Bioinformat 16(1):1– 12. https:// doi. org/ 10. 1186/ s128590150485-4 119. Mejía-Roa E, García C, Gómez JI, Prieto M, Tirado F, Nogales R, Pascual-Montano A (2011) Biclustering and classification analysis in gene expression using nonnegative matrix factorization on multi-GPU systems. In: 2011 11th International Conference on Intelligent Systems Design and Applications, pp. 882–887. https:// doi. org/ 10. 1109/ ISDA. 2011. 61217 69 120. Bhattacharya A, Cui Y (2017) A GPU-accelerated algorithm for biclustering analysis and detection of condition-dependent coexpression network modules. Sci Rep 7(1):1–9. https:// doi. org/ 10. 1038/ s4159801704070-4 121. Orzechowski P, Sipper M, Huang X, Moore JH (2018) EBIC: an evolutionary-based parallel biclustering algorithm for pattern discovery. Bioinformatics 34(21):3719–3726. https:// doi. org/ 10. 1093/ bioin forma tics/ bty401 122. González-Domínguez J, Expósito RR (2019) Accelerating binary biclustering on platforms with CUDA-enabled GPUs. Inf Sci 496:317–325. https:// doi. org/ 10. 1016/j. ins. 2018. 05. 025 123. López-Fernández A, Gómez-Vela FA, Gonzalez-Dominguez J, Bidare-Divakarachari P (2024) bioscience: a new python science library for high-performance computing bioinformatics analytics. SoftwareX 26:101666. https:// doi. org/ 10. 1016/j. softx. 2024. 101666 124. Ardlie KG, Deluca DS, Segrè AV, Sullivan TJ, Young TR, Gelfand ET, Trowbridge CA, Maller JB, Tukiainen T, Consortium G etal (2015) The genotype-tissue expression (GTEx) pilot analysis: multitissue gene regulation in humans. Science 348(6235):648–660. https:// doi. org/ 10. 1126/ scien ce. 12621 10 125. López-Fernández A, Gómez-Vela FA, Gonzalez-Dominguez J, Bidare-Divakarachari P (2024) bioscience: a new python science library for high-performance computing bioinformatics analytics. SoftwareX 26:101666. https:// doi. org/ 10. 1016/j. softx. 2024. 101666 126. Kerr G, Ruskin HJ, Crane M, Doolan P (2008) Techniques for clustering gene expression data. Comput Biol Med 38(3):283–293. https:// doi. org/ 10. 1016/j. compb iomed. 2007. 11. 001 127. Saber HB, Elloumi M (2015) A new study on biclustering tools, biclusters validation and evaluation functions. Int J Comput Sci Eng Surv 6(1):1. https:// doi. org/ 10. 5121/ ijcses. 2015. 6101 128. Orzechowski P (2013) Proximity measures and results validation in biclustering–a survey. In: International Conference on Artificial Intelligence and Soft Computing, pp. 206–217. https:// doi. org/ 10. 1007/ 978-364238610-7_ 20 129. Choi S-S, Cha S-H, Tappert CC (2010) A survey of binary similarity and distance measures. J Syst, Cybern Informat 8(1):43–48 130. Wang L, Zhang H, Chang H-W, Qin Q-M, Zhang B-R, Li X-Q, Zhao T-H, Zhang T-Y (2021) GAEBic: a novel biclustering analysis method for miRNA-targeted gene data based on graph autoencoder. J Comput Sci Technol 36(2):299–309. https:// doi. org/ 10. 1007/ s113900210804-3 131. Serrano-Rubio AA, Morales-Luna GB, Meneses-Viveros A (2021) Gene expression analysis through parallel non-negative matrix factorization. Computation 9(10):106. https:// doi. org/ 10. 3390/ compu tatio n9100 106 132. Belitser E, Nurushev N (2018) Local inference by penalization method for biclustering model. Math Methods Statist 27(3):163–183. https:// doi. org/ 10. 3103/ S1066 53071 80300 18 133. Wang H, Wang W, Yang J, Yu PS (2002) Clustering by pattern similarity in large data sets. In: Proceedings of the 2002 ACM SIGMOD International Conference on Management of Data, pp. 394–405. https:// doi. org/ 10. 1145/ 564691. 564737 134. Xiao F, Chen L, Sha C, Sun L, Wang R, Liu AX, Ahmed F (2018) Noise tolerant localization for sensor networks. IEEE/ACM Trans Netw 26(4):1701–1714. https:// doi. org/ 10. 1109/ TNET. 2018. 28527 54 Biclustering inbioinformatics using big data andHigh… Page 49 of 52 1123 135. Mao KZ, Tang W (2010) Recursive mahalanobis separability measure for gene subset selection. IEEE/ACM Trans Comput Biol Bioinf 8(1):266–272. https:// doi. org/ 10. 1109/ TCBB. 2010. 43 136. Najat N, Abdulazeez AM (2017) Gene clustering with partition around mediods algorithm based on weighted and normalized mahalanobis distance. In: 2017 International Conference on Intelligent Informatics and Biomedical Sciences (ICIIBMS), pp. 140–145. https:// doi. org/ 10. 1109/ ICIIB MS. 2017. 82797 07 137. Yip KY, Cheung DW, Ng MK (2004) Harp: a practical projected clustering algorithm. IEEE Trans Knowl Data Eng 16(11):1387–1397. https:// doi. org/ 10. 1109/ TKDE. 2004. 74 138. Aguilar-Ruiz JS (2005) Shifting and scaling patterns from gene expression data. Bioinformatics 21(20):3840–3845. https:// doi. org/ 10. 1093/ bioin forma tics/ bti641 139. Bozdağ D, Kumar AS, Catalyurek UV (2010) Comparative analysis of biclustering algorithms. In: Proceedings of the First ACM International Conference on Bioinformatics and Computational Biology, pp. 265–274. https:// doi. org/ 10. 1145/ 18547 76. 18548 14 140. Chen PY, Smithson M, Popovich PM, Y C, etal.: Correlation: Parametric and Nonparametric Measures vol. 139, (2002). Sage 141. Priness I, Maimon O, Ben-Gal I (2007) Evaluation of gene-expression clustering via mutual information distance measure. BMC Bioinformat 8(1):1–12. https:// doi. org/ 10. 1186/ 14712105-8111 142. Li C, Tang Z, Zhang W, Ye Z, Liu F (2021) GEPIA2021: integrating multiple deconvolution-based analysis into GEPIA. Nucleic Acids Res 49(W1):242–246. https:// doi. org/ 10. 1093/ nar/ gkab4 18 143. Roy S, Bhattacharyya D, Kalita JK (2012) Deterministic approach for biclustering of co-regulated genes from gene expression data. In: Advances in Knowledge-Based and Intelligent Information and Engineering Systems, pp. 490–499. https:// doi. org/ 10. 3233/ 978-161499105-2490 144. Ayadi W, Elloumi M, Hao J-K (2012) Pattern-driven neighborhood search for biclustering of microarray data. In: BMC Bioinformatics, vol. 13, pp. 1–11. https:// doi. org/ 10. 1186/ 1471210513S7S11 145. D’Urso P (2015) Fuzzy clustering. Handbook of cluster analysis, 545–574 146. Noronha MD, Henriques R, Madeira SC, Zárate LE (2022) Impact of metrics on biclustering solution and quality: a review. Pattern Recognit. https:// doi. org/ 10. 1016/j. patcog. 2022. 108612 147. Consortium GO (2019) The gene ontology resource: 20 years and still Going strong. Nucleic Acids Res 47(D1):330–338. https:// doi. org/ 10. 1093/ nar/ gky10 55 148. Kanehisa M, Goto S (2000) KEGG: Kyoto encyclopedia of genes and genomes. Nucleic Acids Res 28(1):27–30. https:// doi. org/ 10. 1093/ nar/ 28.1. 27 149. Zou D, Ma L, Yu J, Zhang Z (2015) Biological databases for human research. Genom, Proteomics & Bioinformat 13(1):55–63. https:// doi. org/ 10. 1016/j. gpb. 2015. 01. 006 150. Akoglu H (2018) User’s guide to correlation coefficients. Turkish J Emerg Med 18(3):91–93. https:// doi. org/ 10. 1016/j. tjem. 2018. 08. 001 151. Raudvere U, Kolberg L, Kuzmin I, Arak T, Adler P, Peterson H, Vilo J (2019) G: Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic Acids Res 47(W1):191–198. https:// doi. org/ 10. 1093/ nar/ gkz369 152. Kuleshov MV, Jones MR, Rouillard AD, Fernandez NF, Duan Q, Wang Z, Koplev S, Jenkins SL, Jagodnik KM, Lachmann A etal (2016) Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res 44(W1):90–97. https:// doi. org/ 10. 1093/ nar/ gkw377 153. Fan J, Fan D, Slowikowski K, Gehlenborg N, Kharchenko P (2017) UBiT2: a client-side webapplication for gene expression data analysis. bioRxiv, 118992 https:// doi. org/ 10. 1101/ 118992 154. Sun L, Zhu Y, Mahmood A, Tudor CO, Ren J, Vijay-Shanker K, Chen J, Schmidt CJ (2017) WebGIVI: a web-based gene enrichment analysis and visualization tool. BMC Bioinformat 18(1):1–10. https:// doi. org/ 10. 1186/ s128590171664-2 155. Liao Y, Wang J, Jaehnig EJ, Shi Z, Zhang B (2019) WebGestalt 2019: gene set analysis toolkit with revamped UIs and APIs. Nucleic Acids Res 47(W1):199–205. https:// doi. org/ 10. 1093/ nar/ gkz401 156. Tipney H, Hunter L (2010) An introduction to effective use of enrichment analysis software. Hum Genom 4(3):1–5. https:// doi. org/ 10. 1186/ 14797364-43202 157. Hung J-H, Yang T-H, Hu Z, Weng Z, DeLisi C (2012) Gene set enrichment analysis: performance evaluation and usage guidelines. Brief Bioinform 13(3):281–291. https:// doi. org/ 10. 1093/ bib/ bbr049 158. Hoadley KA, Yau C, Wolf DM, Cherniack AD, Tamborero D, Ng S, Leiserson MD, Niu B, McLellan MD, Uzunangelov V etal (2014) Multiplatform analysis of 12 cancer types reveals molecular classification within and across tissues of origin. Cell 158(4):929–944. https:// doi. org/ 10. 1016/j. cell. 2014. 06. 049 A.López-Fernández et al. 1123 Page 50 of 52 159. Lopez-Fernandez A, Rodriguez-Baena D, Gomez-Vela F, Diaz-Diaz N (2018) BIGO: A web application to analyse gene enrichment analysis results. Comput Biol Chem 76:169–178. https:// doi. org/ 10. 1016/j. compb iolch em. 2018. 06. 006 160. Falcon S, Gentleman R (2007) Using GOstats to test gene lists for GO term association. Bioinformatics 23(2):257–258. https:// doi. org/ 10. 1093/ bioin forma tics/ btl567 161. Gomez-Pulido JA, Cerrada-Barrios JL, Trinidad-Amado S, Lanza-Gutierrez JM, Fernandez-Diaz RA, Crawford B, Soto R (2016) Fine-grained parallelization of fitness functions in bioinformatics optimization problems: gene selection for cancer classification and biclustering of gene expression data. BMC Bioinformat 17(1):1–13. https:// doi. org/ 10. 1186/ s128590161200-9 162. Boutros A, Betz V (2021) FPGA architecture: principles and progression. IEEE Circuits Syst Mag 21(2):4–29. https:// doi. org/ 10. 1109/ MCAS. 2021. 30716 07 163. López-Fernández A, Rodríguez-Baena DS, Gómez-Vela F (2020) gMSR: a multi-GPU algorithm to accelerate a massive validation of biclusters. Electronics 9(11):1782. https:// doi. org/ 10. 3390/ elect ronic s9111 782 164. Greene CS, Tan J, Ung M, Moore JH, Cheng C (2014) Big data bioinformatics. J Cell Physiol 229(12):1896–1900. https:// doi. org/ 10. 1002/ jcp. 24662 165. DiMaggio PA, McAllister SR, Floudas CA, Feng X-J, Rabinowitz JD, Rabitz HA (2008) Biclustering via optimal re-ordering of data matrices in systems biology: rigorous methods and comparative studies. BMC Bioinformat 9(1):1–16. https:// doi. org/ 10. 1186/ 14712105-9458 166. Wordsworth S, Doble B, Payne K, Buchanan J, Marshall DA, McCabe C, Regier DA (2018) Using ‘big data’ in the cost-effectiveness analysis of next-generation sequencing technologies: challenges and potential solutions. Value Health 21(9):1048–1053. https:// doi. org/ 10. 1016/j. jval. 2018. 06. 016 167. Schatz MC, Langmead B, Salzberg SL (2010) Cloud computing and the DNA data race. Nat Biotechnol 28(7):691–693. https:// doi. org/ 10. 1038/ nbt07 10691 168. Luo Y, Hitz BC, Gabdank I, Hilton JA, Kagda MS, Lam B, Myers Z, Sud P, Jou J, Lin K etal (2020) New developments on the encyclopedia of DNA elements (ENCODE) data portal. Nucleic Acids Res 48(D1):882–889. https:// doi. org/ 10. 1093/ nar/ gkz10 62 169. Cantelli G, Bateman A, Brooksbank C, Petrov AI, Malik-Sheriff RS, Ide-Smith M, Hermjakob H, Flicek P, Apweiler R, Birney E etal (2022) The European Bioinformatics Institute (EMBL-EBI) in 2021. Nucleic Acids Res 50(D1):11–19. https:// doi. org/ 10. 1093/ nar/ gkab1 127 170. Bao R, Huang L, Andrade J, Tan W, Kibbe WA, Jiang H, Feng G (2014) Review of current methods, applications, and data management for the bioinformatics analysis of whole exome sequencing. Cancer Informat 13:13779. https:// doi. org/ 10. 4137/ CIN. S13779 171. Pabinger S, Dander A, Fischer M, Snajder R, Sperk M, Efremova M, Krabichler B, Speicher MR, Zschocke J, Trajanoski Z (2014) A survey of tools for variant analysis of next-generation genome sequencing data. Brief Bioinform 15(2):256–278. https:// doi. org/ 10. 1093/ bib/ bbs086 172. O’Driscoll A, Daugelaite J, Sleator RD (2013) ‘Big data’, Hadoop and cloud computing in genomics. J Biomed Inform 46(5):774–781. https:// doi. org/ 10. 1016/j. jbi. 2013. 07. 001 173. Luo J, Wu M, Gopukumar D, Zhao Y (2016) Big data application in biomedical research and health care: a literature review. Biomed Informat Insights 8:31559. https:// doi. org/ 10. 4137/ BII. S31559 174. Smowton C, Balla A, Antoniades D, Miller C, Pallis G, Dikaiakos MD, Xing W (2017) A costeffective approach to improving performance of big genomic data analyses in clouds. Futur Gener Comput Syst 67:368–381. https:// doi. org/ 10. 1016/j. future. 2015. 11. 011 175. Schadt EE, Linderman MD, Sorenson J, Lee L, Nolan GP (2011) Cloud and heterogeneous computing solutions exist today for the emerging big data problems in biology. Nat Rev Genet 12(3):224–224. https:// doi. org/ 10. 1038/ nrg28 57c2 176. Grossman RL, White KP (2012) A vision for a biomedical cloud. J Intern Med 271(2):122–130. https:// doi. org/ 10. 1111/j. 13652796. 2011. 02491.x 177. Kumar S (2021) Trends and advancements in genome data compression and processing algorithms. Soft Comput: Theor Appl. https:// doi. org/ 10. 1007/ 978981161696-9_ 15 178. Kaushik P, Rao AM, Singh DP, Vashisht S, Gupta S (2021) Cloud Computing and Comparison based on Service and Performance between Amazon AWS, Microsoft Azure, and Google Cloud. In: 2021 International Conference on Technological Advancements and Innovations (ICTAI), pp. 268–273. https:// doi. org/ 10. 1109/ ICTAI 53825. 2021. 96734 25 179. Jain S, Saxena A, Hesarur S, Bhadhadhara K, Bharti N, Kasibhatla SM, Sonavane U, Joshi R (2021) GenoVault: a cloud based genomics repository. BioData Mining 14(1):1–10. https:// doi. org/ 10. 1186/ s1304002100268-5 Biclustering inbioinformatics using big data andHigh… Page 51 of 52 1123 180. Schumacher A, Pireddu L, Niemenmaa M, Kallio A, Korpelainen E, Zanetti G, Heljanko K (2014) SeqPig: simple and scalable scripting for large sequencing data sets in Hadoop. Bioinformatics 30(1):119–120. https:// doi. org/ 10. 1093/ bioin forma tics/ btt601 181. Abuín JM, Pichel JC, Pena TF, Amigo J (2015) BigBWA: approaching the Burrows-Wheeler aligner to big data technologies. Bioinformatics 31(24):4003–4005. https:// doi. org/ 10. 1093/ bioin forma tics/ btv506 182. Masseroli M, Canakoglu A, Pinoli P, Kaitoua A, Gulino A, Horlova O, Nanni L, Bernasconi A, Perna S, Stamoulakatou E etal (2019) Processing of big heterogeneous genomic datasets for tertiary analysis of next generation sequencing data. Bioinformatics 35(5):729–736. https:// doi. org/ 10. 1093/ bioin forma tics/ bty688 183. Wiewiórka M, Szmurło A, Kuśmirek W, Gambin T (2019) SeQuiLa-cov: a fast and scalable library for depth of coverage calculations. Gigascience 8(8):094. https:// doi. org/ 10. 1093/ gigas cience/ giz094 184. Expósito RR, Galego-Torreiro R, González-Domínguez J (2020) Sequal: big data tool to perform quality control and data preprocessing of large NGS datasets. IEEE Access 8, 146075–146084 https:// doi. org/ 10. 1109/ ACCESS. 2020. 30150 16 185. Santamaría R, Therón R, Quintales L (2008) BicOverlapper: a tool for bicluster visualization. Bioinformatics 24(9):1212–1213. https:// doi. org/ 10. 1093/ bioin forma tics/ btn076 186. Cheng K-O, Law N-F, Siu W-C, Lau T (2007) BiVisu: software tool for bicluster detection and visualization. Bioinformatics 23(17):2342–2344. https:// doi. org/ 10. 1093/ bioin forma tics/ btm338 187. Teng L, Chan L-W (2006) Biclustering gene expression profiles by alternately sorting with weighted correlated coefficient. In: 2006 16th IEEE Signal Processing Society Workshop on Machine Learning for Signal Processing, pp. 289–294. https:// doi. org/ 10. 1109/ MLSP. 2006. 275563 188. Zhao J, Sun M, Chen F, Chiu P (2017) Bidots: Visual exploration of weighted biclusters. IEEE Trans Visual Comput Graphics 24(1):195–204. https:// doi. org/ 10. 1109/ TVCG. 2017. 27444 58 189. Sozdinler M (2021) A Review on Analysis and Visualization Methods for Biclustering. arXiv preprint 190. Merico D, Gfeller D, Bader GD (2009) How to visually interpret biological data using networks. Nat Biotechnol 27(10):921–924. https:// doi. org/ 10. 1038/ nbt. 1567 191. Rastegar-Mojarad M, Talatian-Azad S, Minaei-Bidgoli B (2010) A survey on biological data analysis by biclustering. In: 2010 International Conference on Educational and Information Technology, vol. 1, pp. 1–100. https:// doi. org/ 10. 1109/ ICEIT. 2010. 56077 92 192. Verma NK, Meena S, Bajpai S, Singh A, Nagrare A, Cui Y (2010) A comparison of biclustering algorithms. In: 2010 International Conference on Systems in Medicine and Biology, pp. 90–97. https:// doi. org/ 10. 1109/ ICSMB. 2010. 57353 51 193. Oghabian A, Kilpinen S, Hautaniemi S, Czeizler E (2014) Biclustering methods: biological relevance and application in gene expression analysis. PLoS ONE 9(3):90801. https:// doi. org/ 10. 1371/ journ al. pone. 00908 01 194. Horta D, Campello RJ (2014) Similarity measures for comparing biclusterings. IEEE/ACM Trans Comput Biol Bioinf 11(5):942–954. https:// doi. org/ 10. 1109/ TCBB. 2014. 23250 16 195. Pontes B, Girldez R, Aguilar-Ruiz JS (2015) Quality measures for gene expression biclusters. PLoS ONE 10(3):0115497. https:// doi. org/ 10. 1371/ journ al. pone. 01154 97 196. Mounir M, Hamdy M (2015) On biclustering of gene expression data. In: 2015 IEEE Seventh International Conference on Intelligent Computing and Information Systems (ICICIS), pp. 641– 648. https:// doi. org/ 10. 1109/ Intel CIS. 2015. 73972 90 197. Biswal BS, Mishra P, Mohapatra A, Vipsita S (2016) A survey on greedy based algorithms for biclustering of gene expression microarray data. In: 2016 International Conference on Information Technology (ICIT), pp. 124–128. https:// doi. org/ 10. 1109/ ICIT. 2016. 036 198. Anitha S, Chandran C (2016) Review on analysis of gene expression data using biclustering approaches. Bonfring Int J Data Mining 6(2):16–23. https:// doi. org/ 10. 9756/ BIJDM. 8135 199. Biswal BS, Mohapatra A, Vipsita S (2018) A review on biclustering of gene expression microarray data: algorithms, effective measures and validations. Int J Data Min Bioinform 21(3):230–268. https:// doi. org/ 10. 1504/ IJDMB. 2018. 097683 200. Aouabed H, Santamaria R, Elloumi M (2018) Biclustering impact in biomedical sciences via literature mining. Int J Biomed Data Min 7(134):2. https:// doi. org/ 10. 4172/ 20904924. 10001 34 Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. A.López-Fernández et al. 1123 Page 52 of 52 Authors and Affiliations AurelioLópez‑Fernández4· FranciscoA.Gomez‑Vela1· DomingoS.Rodriguez‑Baena1· FernandoM.Delgado‑Chaves2· JorgeGonzalez‑Dominguez3 * Aurelio López-Fernández [email protected] Francisco A. Gomez-Vela [email protected] Domingo S. Rodriguez-Baena [email protected] Fernando M. Delgado-Chaves [email protected] Jorge Gonzalez-Dominguez jorg[email protected] 1 Intelligent Data Analysis Group (DATAi), Universidad Pablo de Olavide, Ctra. Utrera, km. 1, ES-41013Seville, Spain 2 Institute forComputational Systems Biology, University ofHamburg, Notkestrasse 9, 22607Hamburg, Germany 3 Computer Architecture Group, Universidade da Coruña, Campus de Elviña, 15071ACoruña, Spain 4 Dpto. Lenguajes y Sistemas Informáticos, Universidad de Sevilla, Seville, Spain