scieee AI-readable full text Open interactive document viewer

GenAI-Augmented Polyglot Persistence for Microservices: Toward Self-Adaptive and Explainable Data Fabrics

Khule, Akshay

Abstract

Polyglot persistence lets microservices choose the best database for each workload, but today it is mostly static, manually configured, and hard to manage at scale. This paper presents GenAI-Augmented Polyglot Persistence (GAPP) — an architecture that uses generative AI and large language models (LLMs) inside the data orchestration layer. GAPP automates workload-to-database selection, supports natural-language query federation, adapts orchestration based on runtime patterns, and provides explainable data lineage. Our prototype shows 94% query accuracy, complete lineage coverage, and an 85% reduction in development effort compared to traditional pipelines. The results highlight how AI can enable self-governing data fabrics that unify DevOps and DataOps. We conclude with key research challenges and a roadmap for building AI-native database management systems.

Full text

GenAI-Augmented Polyglot Persistence for Microservices: Toward Self-Adaptive and Explainable Data Fabrics Akshay Khule Unisys Corporation Application Architect, Cloud & Data Engineering Email: [email protected] Abstract—Polyglot persistence enables microservices to select optimal database technologies for diverse workloads but remains largely static, manually configured, and difficult to govern at scale. This paper introduces GenAI-Augmented Polyglot Persistence (GAPP), an architecture that integrates generative AI and large language models (LLMs) into the data orchestration layer. GAPP enables automated workload-to-database mapping, natural language query federation, adaptive orchestration, and explainable data lineage. We present a prototype implementation that achieves 94% query accuracy, full lineage coverage, and an 85% reduction in development effort compared to traditional pipelines. Experimental results demonstrate the feasibility of AI-assisted, self-governing data fabrics bridging DevOps and DataOps. The paper concludes with open research challenges and a roadmap toward AI-native database management systems. Index Terms—Polyglot persistence, Generative AI, Data governance, Microservices, Federated query, Lineage, Cloud computing I. INTRODUCTION Polyglot persistence has become a foundational principle of microservice and cloud-native architecture design [1]–[3]. By allowing each service to select a database best suited to its data model and query pattern, organizations achieve modularity, scalability, and performance. However, maintaining and governing such heterogeneity introduces new challenges in consistency, compliance, and interoperability. Generative AI (GenAI) technologies, powered by large language models (LLMs) such as GPT-4 [4], now offer an opportunity to rethink data management. Instead of static database configurations, an intelligent orchestration layer can dynamically interpret workloads, select optimal persistence models, and even generate federated query plans. In this paper, we propose GenAI-Augmented Polyglot Persistence (GAPP) — an architectural paradigm where LLMs assist in workload analysis, federated query generation, governance automation, and dynamic orchestration. Our work builds upon principles from distributed systems and data governance frameworks [5], [6] and introduces a proof-of-concept (POC) implementation that connects relational, document, and graph databases through a reasoning-based orchestration layer. Our main contributions are as follows: •We introduce the GAPP architecture, which integrates LLM-based reasoning for adaptive polyglot persistence. •We design and implement a POC that translates natural language queries into federated multi-database execution plans. •We evaluate the system using a multi-store benchmark and compare it against traditional and template-based baselines. •We discuss implications for explainability, governance, and future self-adaptive data fabrics [7], [8]. II. BACKGROUND AND RELATED WORK Polyglot persistence is an established concept in distributed data management, emphasizing the use of multiple database paradigms (relational, document, key-value, and graph) within a single application ecosystem [1]. Industry leaders such as Uber and Netflix have adopted this pattern to achieve domainaligned data design and service-level isolation [3], [9]. A. Federated Query Engines Research in federated query engines, such as Presto and Trino [10], has focused on unifying heterogeneous data access through distributed SQL planners. These systems provide a powerful foundation but still require manual schema mapping and configuration, limiting their agility for evolving workloads. B. Data Governance and Lineage Modern governance frameworks such as OpenLineage and DataHub [5], [6] have standardized how metadata and lineage are collected and propagated, forming the backbone of modern data governance. Building on these capabilities, our approach introduces an AI-assisted orchestration layer that dynamically infers lineage and validates policy compliance, extending traditional rule-based extraction with context-aware reasoning. C. Generative AI for Data Systems Recent advances in GenAI demonstrate how LLMs can translate natural language into structured SQL queries [11], and even support reasoning across multiple sources [8]. Prior research has also explored AI-assisted data governance [7], though integration into distributed, multi-database orchestration remains largely unexplored. GAPP extends this line of work by embedding reasoning and governance awareness directly into the persistence layer. III. GENAI-AUGMENTED POLYGLOT PERSISTENCE ARCHITECTURE The GAPP architecture extends the traditional microservice data stack by introducing an intelligent reasoning layer that can interpret user intent, schema context, and workload characteristics to orchestrate data access dynamically. Figure 1 provides an overview of the proposed system. Fig. 1. GAPP Architecture: LLM-based planner generates validated, explainable multi-database query plans and logs lineage. A. Core Components 1. Workload Profiler: Captures service-level access patterns and SLAs for dynamic planning. 2. GenAI NL2Planner: Translates natural language into executable query plans leveraging LLM reasoning [4], [8]. 3. Query Validator: Enforces schema and policy compliance using metadata extracted via OpenLineage [5]. 4. Executors: Interface modules for PostgreSQL, MongoDB, and Neo4j [1]. 5. Lineage Recorder: Stores field-level transformations in a graph structure for auditability [12]. 6. Governance Layer: Integrates compliance rules and cost-performance optimizations [7]. The architecture emphasizes transparency, allowing data engineers and auditors to trace lineage across multiple persistence layers while benefiting from AI-driven reasoning. IV. PROOF OF CONCEPT (POC) DESIGN AND IMPLEMENTATION We implemented a POC to validate GAPP using three heterogeneous stores: PostgreSQL, MongoDB, and Neo4j. This setup demonstrates AI-assisted query planning, validation, and lineage generation across diverse persistence layers. A. Design Goals The POC aimed to prove the feasibility of: 1) Translating natural language queries into federated execution plans. 2) Dynamically adapting to schema changes and governance constraints. 3) Capturing field-level lineage automatically during execution. B. Execution Flow The GenAI NL2Planner analyzes a natural language query and schema metadata, producing a structured JSON plan containing SQL, aggregation, and Cypher fragments [11]. The Validator enforces constraints such as PII masking and tenant isolation. Executors run validated subqueries in parallel, and the results are merged and logged in the lineage store. C. Decision Flow Figure 2 illustrates the decision logic for selecting appropriate databases based on workload characteristics and schema type. Fig. 2. AI-assisted database selection in a polyglot architecture. D. Implementation Details The proof-of-concept (POC) implementation operationalizes the proposed GenAI-Augmented Polyglot Persistence (GAPP) architecture in a microservices environment and demonstrates how federated queries can be executed, validated, and explained across heterogeneous stores. The system was deployed using a hybrid model combining containerized microservices (for the orchestrator, AI reasoning, and lineage services) and PaaS-managed databases where available, including PostgreSQL, MongoDB, Neo4j, an Apache Spark orchestration service, and a GenAI reasoning service powered by GPT-4. To support self-adaptivity, the Workload Profiler continuously harvested service-level metadata—including schemas, access patterns, cost/latency statistics, and governance annotations—and registered them in OpenLineage and DataHub. These metadata inputs enabled the GenAI NL2Planner to dynamically translate natural-language queries into federated execution plans that adapt to schema changes, workload shifts, and governance policies. The Query Validator ensured plan correctness by crosschecking GPT-4-generated subqueries against authoritative schema and compliance metadata. Violations such as missing fields, deprecated attributes, or PII policy breaches triggered real-time corrective prompts back to the reasoning layer. After validation, the Executor modules dispatched store-specific subqueries concurrently through Spark, merged partial results, and monitored runtime characteristics for feedback into the profiler. To support explainability, the Lineage Recorder extracted field-level transformation paths directly from Spark’s QueryExecution logical plans. These were normalized into nodeand edge-based structures and persisted in Neo4j, enabling analysts to inspect provenance, understand transformations, and trace data dependencies across store boundaries. The Governance Layer integrated lineage, profiling, and policy metadata to generate human-readable explanations of planning decisions, storage selection, and compliance outcomes, consistent with the goal of building an explainable data fabric. This modular implementation validates the design principles of GAPP: (1) GenAI-based reasoning for federated query planning, (2) Self-adaptive orchestration driven by metadata and policy feedback loops, and (3) Explainable lineage-aware execution across polyglot microservices. V. EXPERIMENTAL RESULTS AND COMPARATIVE ANALYSIS We compared GAPP against manual polyglot pipelines and template-based natural-language translators. The benchmark involved 30 queries spanning analytical, relational, and graph operations. A. Evaluation Metrics Metrics included correctness, latency, hallucination rate, explainability, and governance completeness [7]. Table I summarizes the performance. TABLE I QUANTITATIVE COMPARISON BETWEEN BASELINES AND GAPP Metric Manual Template GAPP Query Success Rate 100% 62% 94% Latency (s) 2.3 1.6 2.1 Explainability (1–5) 2.0 3.1 4.6 Dev Effort (hr/q) 3.2 0.8 0.4 Governance Coverage Manual None Full B. Discussion GAPP’s reasoning capability yields near-human query accuracy while automating lineage capture. Although latency increases slightly due to LLM inference [8], the trade-off is justified by productivity and compliance gains. VI. CONCLUSION AND FUTURE RESEARCH ROADMAP This paper presented GenAI-Augmented Polyglot Persistence (GAPP), a framework integrating LLM reasoning within polyglot data orchestration. The proposed approach improves adaptability, transparency, and governance automation compared to traditional microservice data stacks. Our results confirm that integrating AI reasoning layers with metadata and lineage frameworks [5], [6] can reduce development time while enhancing explainability. Future work will explore AI-native databases [7], trust-aware model governance [7], and edge-aware federated persistence [13], [14]. Ultimately, GAPP moves data ecosystems toward selfgoverning, context-aware architectures capable of optimizing themselves through continuous AI feedback loops. REFERENCES [1] F. H. Anila Nuhiji, Diellza Mustafai Veliu, “Polyglot persistence in microservices: Managing data diversity in distributed systems,” arXiv preprint arXiv:2509.08014, 2025, explores database heterogeneity, comparative frameworks, and operational patterns in microservice ecosystems. [2] Uber Engineering, “Designing schemaless: Uber engineering’s scalable datastore using mysql,” https://www.uber.com/blog/ schemaless-part-one-mysql-datastore/, 2016, part one of a threepart blog series on Uber’s Schemaless datastore. [3] Netflix Technology Blog, “Evolution of the netflix data pipeline,” https: //techblog.netflix.com/2016/02/evolution-of-netflix-data-pipeline.html, 2016, describes the evolution of Netflix’s data platform and pipelines. [4] OpenAI, “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774, 2024, technical report describing architecture, multimodal capabilities, and evaluation of GPT-4. [5] O. Project, “Why an open standard for lineage metadata?” OpenLineage Blog, 2023, blog post explaining why OpenLineage free, open-standard specification enables more effective metadata lineage collection. [6] L. Engineering, “Datahub: A generalized metadata search discovery tool,” LinkedIn Engineering Blog, 2019, blog post introducing DataHub’s architecture and data catalog capabilities at LinkedIn. [7] L. W. Dietz, A. Wider, and S. Harrer, “Data governance automation through generative ai: Challenges and opportunities,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES). AAAI Press, 2024, pp. 1301–1309, explores how generative AI can automate policy validation, lineage tracking, and governance workflows. [8] Z. Qin, X. Li, Y. Chen et al., “Relational database augmented large language model,” arXiv preprint arXiv:2407.15071, 2024, proposes an LLM-agnostic memory architecture that leverages relational databases as an external memory for enhanced reasoning. [Online]. Available: https://arxiv.org/abs/2407.15071 [9] A. Jacob and U. Engineering, “Designing microservice architectures at uber scale,” Uber Engineering Blog, 2019, describes Uber’s domainoriented microservice and data architecture enabling independent scaling and polyglot persistence. [Online]. Available: https://www.uber. com/en-IN/blog/microservice-architecture/ [10] M. Traverso, D. Sundstrom, D. Phillips, and E. Hwang, “Presto: Sql on everything,” in Proceedings of the IEEE International Conference on Data Engineering (ICDE) Workshops. IEEE, 2019, pp. 180–188, describes Presto’s distributed SQL architecture and execution engine for federated query processing. [11] R. G. Ali Mohammadjafari, Anthony S. Maida, “From natural language to sql: Review of llm-based text-to-sql systems,” arXiv preprint arXiv:2410.01066, 2024, comprehensive survey of Text-to-SQL systems in the LLM era; useful background for NL→federated-query planning. [Online]. Available: https://arxiv.org/html/2410.01066v1 [12] A. Schoenenwald, S. Kern, J. Viehhauser, and J. Schildgen, “Collecting and visualizing data lineage of spark jobs,” Datenbank-Spektrum, vol. 21, pp. 179–189, 2021, describes how Spark execution plans can be used to extract fineand coarse-grained data lineage and visualize it as a graph. [13] P. P. Khine, “A review of polyglot persistence in the big data world,” Information, vol. 10, no. 4, p. 141, 2019, survey of polyglot persistence approaches and challenges in big-data systems (relational, document, key-value, columnar, graph). [14] H. G. Abreha et al., “Federated learning in edge computing: A systematic survey,” IEEE Access / Sensors (reviewed survey), 2022, systematic survey of federated learning and edge computing architectures, communication and privacy challenges — relevant to federated edge analytics and persistence. [Online]. Available: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8780479/