scieee AI-readable full text Open interactive document viewer

Towards Adaptive Pedagogical Policies: A Hybrid Reinforcement Learning and Large Language Model Framework for Intelligent Tutoring Systems under Indonesia's Curriculum

Wicaksono, Adit.Praditya; Andono, Pulung Nurtantio; Pujiono, Pujiono

Abstract

This vision paper proposes a hybrid architecture integrating Q-learning as a pedagogical decision engine and LLMs as a communication engine, orchestrated within an Adaptive Learning System (ALS) layer. The framework aims to support Indonesia’s Merdeka Curriculum by enabling adaptive pedagogical policies that are both strategic and conversational. Contributions include a principled integration of RL and LLM, localization to Indonesian secondary education, and a four-phase research roadmap. The paper avoids implementation details, focusing instead on conceptual clarity, curricular relevance, and dual-axis evaluation.

Full text

Towards Adaptive Pedagogical Policies: A Hybrid Reinforcement Learning and Large Language Model Framework for Intelligent Tutoring Systems under Indonesia’s Curriculum Praditya Wicaksono Master of Informatics Universitas Dian Nuswantoro Semarang, Indonesia [email protected]us.ac.id Pulung Nurtantio Andono Master of Informatics Universitas Dian Nuswantoro Semarang, Indonesia [email protected]us.ac.id Pujiono Master of Informatics Universitas Dian Nuswantoro Semarang, Indonesia [email protected]us.ac.id Abstract—Personalized learning is increasingly central to technology-enhanced education, yet practical deployments often oscillate between systems that are pedagogically strategic but rigid (classical Intelligent Tutoring Systems, ITS) and systems that are conversationally fluent but pedagogically myopic (Large Language Models, LLMs). This vision paper proposes a hybrid architecture that integrates Q-learning as a pedagogical decision engine and LLMs as a communication engine, orchestrated within an Adaptive Learning System (ALS) layer. The goal is to realize adaptive pedagogical policies that align with Indonesia’s Kurikulum Merdeka (Merdeka Curriculum), which emphasizes flexibility, learner agency, and differentiated pathways. We position the research gap at the intersection of strategy (RL/ITS), communication (LLM), and orchestration (ALS), and articulate contributions in terms of (i) a principled integration that preserves pedagogical intent while delivering human-like dialogue; (ii) a localization agenda grounded in the curricular and infrastructural realities of Indonesian secondary education; and (iii) a research roadmap covering conceptual design, limited-scope prototyping, technical and pedagogical evaluation, and scaling. The paper intentionally avoids implementation details, parameters, or code, focusing instead on novelty, scope, and actionable implications for researchers and practitioners. Index Terms—Intelligent Tutoring Systems, Reinforcement Learning, Q-learning, Large Language Models, Adaptive Learning Systems, Merdeka Curriculum, Indonesia, Educational Technology. I. Introduction The global shift toward personalization has exposed limitations in one-size-fits-all instruction. Learners differ in prior knowledge, motivation, pace, and preferred modalities, demanding systems that adapt to evolving needs rather than enforcing static sequences. Classical Intelligent Tutoring Systems (ITS)—particularly those employing Reinforcement Learning (RL)—have demonstrated promise in optimizing pedagogical strategies such as difficulty adjustment and content sequencing [1]. Yet many ITS deployments remain rigid, as strategy parameters and response forms are often predetermined and slow to react to real-time learner dynamics. In parallel, Large Language Models (LLMs) have shown exceptional capability in natural dialogue, scaffolding explanations, and responding to open-ended queries [4]. However, LLMs are typically reactive content generators without stable long-horizon pedagogical policies; conversation quality can be high while instructional coherence over time is not guaranteed. Indonesia’s Merdeka Curriculum (Kurikulum Merdeka) accentuates personalization, flexibility, and contextualization across diverse school settings. The curriculum’s ambition meets implementation constraints: variability in teacher readiness, resource limitations, and geographical disparities. A principled human–AI partnership is needed to scale differentiation without diluting pedagogical integrity. This paper proposes a hybrid ITS architecture where Q-learning governs what pedagogical action to take and when, while an LLM determines how to communicate the action in a human-like, student-sensitive manner. An ALS layer orchestrates data flows, profiles, and objectives, ensuring that micro-level decisions cohere with macro-level learning paths and curricular targets. II. Positioning and Novelty A. Research Gap Evidence suggests RL can learn adaptive teaching strategies and improve engagement [1]–[3], yet classical setups struggle with real-time responsiveness and rich feedback channels. Conversely, LLMs enable flexible, contextual dialogue but lack policy memory and principled instructional control [4]. ALS provides the macro-framework for differentiation and progress monitoring [5], but most integrations remain partial (e.g., scheduling without conversational scaffolding, or chat without policy coherence). The gap lies in a unified model that: (i) preserves pedagogical intent via a learned policy, (ii) expresses that intent via natural dialogue, and (iii) coordinates both within a curriculum-aware orchestration layer. B. Claims of Novelty (1) Policy-first integration of RL and LLM within ALS. We articulate a division of labor: Q-learning as a decision engine that selects pedagogical actions and LLM as a communication engine that realizes those actions conversationally. ALS binds the two, aligning policy with learner models, goals, and constraints. (2) Localization to the Merdeka Curriculum. We explicitly target the needs of Indonesian secondary schools: differentiated pathways, varied readiness, and teacher workload. The vision prioritizes curricular coherence, teacher oversight, and scalability under resource constraints. (3) Evaluation across technical and pedagogical axes. Beyond system performance, we emphasize effects on understanding, motivation, and engagement—critical for policy adoption and real-world impact. III. Conceptual Architecture A. Three-Layer View Decision Engine (Q-learning). The policy space comprises pedagogical actions (e.g., explain, prompt metacognition, give practice, offer challenge, review), chosen with respect to learner states (engaged, struggling, passive, curious, etc.). The focus is on adaptive policies that respond to recent evidence while remaining faithful to curricular targets. We deliberately avoid algorithmic detail; the essential point is that decisions are policy-driven and data-informed. Communication Engine (LLM). The LLM translates selected actions into dialogic turns that are age-appropriate, linguistically and culturally sensitive, and aligned with learner profiles (reading level, prior misconceptions, affect). The LLM is constrained by the decision engine: it does not “freelance” instruction, but expresses policy coherently and consistently. ALS Layer. ALS orchestrates long-horizon differentiation: learner profiling, path planning, prerequisite structures, assessment checkpoints, and alignment with Merdeka Curriculum outcomes. It mediates data exchanges and guards coherence between micro-decisions and macro-goals. B. Human-in-the-Loop Oversight Teacher dashboards surface analytics (e.g., progression indicators, engagement traces, misconception clusters) and provide veto/override mechanisms. This ensures accountability, supports formative assessment, and preserves teacher agency. The system aims to complement— not replace—teachers, particularly in contexts with high student–teacher ratios. C. Safety, Reliability, and Ethics We anticipate risks such as hallucinations, biased feedback, and over-automation. The architecture therefore incorporates: (i) grounded prompts and controlled generation for the LLM, (ii) content validation against curriculum-aligned repositories, (iii) transparent logging for auditability, and (iv) student privacy safeguards consistent with local regulations and institutional policies. IV. Research Roadmap A. Phase I: Conceptual Design •Formalize the policy–communication separation of concerns. •Define learner state taxonomies and action ontologies aligned with Merdeka Curriculum competencies. •Conduct expert reviews (educational technologists, curriculum specialists, teachers) to validate constructs and usability expectations. B. Phase II: Limited-Scope Prototyping •Implement a minimal vertical slice for one high-school subject (e.g., Physics or Informatics), focusing on a few well-defined competencies. •Integrate a basic dashboard for teacher oversight and feedback capture. •Pilot in a controlled environment (small class, afterschool program) to test feasibility and workflow. C. Phase III: Technical and Pedagogical Evaluation Technical. Evaluate responsiveness of the decision engine and stability of policy–dialogue alignment. Track convergence signals and robustness to noisy inputs (without exposing parameters). Pedagogical. Use pre/post measures of conceptual understanding, motivation, and engagement; include usability measures and qualitative feedback from teachers and students. Compare against conventional instruction or chat-only baselines to isolate the value of policy-driven dialogue. D. Phase IV: Expansion and Dissemination •Extend to additional subjects and competencies; stress-test in schools with varying resource profiles. •Document design principles and deployment playbooks for local edtech stakeholders. •Share findings via open preprints and conference/journal publications to foster collaboration. V. Implications A. For Research This agenda situates ITS research at the intersection of strategy learning, conversational pedagogy, and system orchestration. By cleanly separating policy from expression, we open avenues for studying how different action ontologies, feedback designs, and curriculum constraints shape learning trajectories. B. For Practice Teachers gain a policy-aware assistant that scales differentiation and reduces routine burdens, while preserving professional judgment. School leaders obtain analytics for targeted interventions. For learners, the system promises guidance that is both personable and purposefully sequenced. C. For Technology Vendors and developers can adopt the architecture as a modular blueprint: interchangeable policy learners, swappable LLM backends, and pluggable ALS components. This modularity supports responsible upgrades without destabilizing classroom practice. VI. Conclusion We have outlined a hybrid ITS vision that marries the pedagogical intentionality of RL with the communicative power of LLMs, orchestrated by an ALS layer tuned to Indonesia’s Merdeka Curriculum. The contribution is not algorithmic novelty per se but architectural clarity: a policy-first integration that keeps dialogue in service of instructional goals, localized to a national reform agenda. By focusing on conceptual soundness, localized relevance, and dual-axis evaluation (technical and pedagogical), we aim to catalyze deployable innovation rather than laboratorybound prototypes. Acknowledgment The authors thank collaborating teachers, curriculum experts, and school partners for formative feedback during the concept development. References [1] A. Iglesias, P. Martínez, R. Aler, and F. Fernández, “Learning teaching strategies in an adaptive and intelligent educational system,” Applied Intelligence, 2009. [2] Y. Zhang and I. Arroyo, “Q-learning in Intelligent Tutoring Systems,” Proc. AI-ED, 2001. [3] F. Dorça et al., “Adaptive learning systems: A review,” Computers & Education, 2013. [4] E. Kasneci et al., “ChatGPT for education,” Computers and Education: AI, 2023. [5] K. Sharma et al., “Conversational AI for personalized learning,” Proc. ACM L@S, 2023. [6] X. Chen et al., “Predictive analytics in adaptive learning,” IEEE Trans. Learning Technologies, 2022. [7] J. Bassen et al., “Adaptive scheduling for practice,” Proc. EDM, 2020.