Full text
Automated Assessment Environment For Programming Assignments Ishan Pandita Department of Computer Engineering Sri Sivasubramaniya Nadar College of Engineering Chennai, India 3122225001042 Sakaya Milton R Department of Computer Engineering Sri Sivasubramaniya Nadar College of Engineering Chennai, India Professor Sarthak Darshan Department of Computer Engineering Sri Sivasubramaniya Nadar College of Engineering Chennai, India 3122225001122 Abstract—This review examines automated assessment systems for programming education, focusing on the development and implementation of comprehensive evaluation environments. The review analyzes existing approaches to automated grading, machine learning integration in code assessment, and plagiarism detection mechanisms. Key findings indicate that modern automated assessment systems must address multiple evaluation criteria including correctness validation, performance analysis, code quality assessment, and academic integrity verification. The review synthesizes research on containerization-based execution environments, CI/CD pipeline integration, and AI-powered feedback generation. Current implementations demonstrate successful multi-language support with strong accuracy in complexity analysis, automated pipeline deployment using Kubernetes orchestration, and comprehensive database integration for scalable student management. Future directions emphasize enhanced AST-based performance analysis, expanded AI integration for nuanced evaluation, and extended programming language support. This review contributes to understanding the evolution of automated programming assessment and identifies critical gaps requiring further research in areas of advanced pattern recognition, real-time code quality evaluation, and sophisticated plagiarism detection mechanisms. Index Terms—automated assessment, programming education, code quality evaluation, plagiarism detection, continuous integration, containerization, machine learning, complexity analysis I. INTRODUCTION The exponential growth in computer science education has created unprecedented challenges in programming assignment evaluation. Traditional manual assessment methods struggle with scalability, consistency, and comprehensiveness as enrollment in programming courses increases worldwide [1]. This review examines the state-of-the-art in automated assessment environments, analyzing how modern systems address the multifaceted requirements of programming education. Programming assignment evaluation extends beyond simple correctness verification to encompass performance analysis, code quality assessment, and academic integrity verification. The emergence of artificial intelligence tools, particularly large language models capable of generating code, has introduced new dimensions to the academic integrity challenge [8]. Simultaneously, the COVID-19 pandemic accelerated adoption of remote learning, creating demand for robust automated assessment systems capable of offline operation while maintaining evaluation rigor [2]. This review systematically examines existing automated assessment approaches, synthesizes findings from recent research in machine learning-based code evaluation, and analyzes plagiarism detection mechanisms specifically designed for programming assignments. The review critically evaluates the strengths and limitations of current implementations, highlighting areas where existing solutions demonstrate effectiveness and identifying gaps requiring further research. The scope of this review encompasses automated testing frameworks, containerization technologies for secure code execution using Docker and Kubernetes, continuous integration and deployment (CI/CD) pipelines for assessment automation, database systems for managing evaluation data, AI-powered feedback generation mechanisms, and multi-criteria evaluation frameworks integrating correctness, performance, quality, and originality assessments [5]. The rationale for this review stems from the urgent need to understand how automated assessment systems can effectively support programming education at scale while maintaining educational value and academic integrity. As artificial intelligence becomes increasingly capable of generating functional code, educational institutions require sophisticated tools that can distinguish human-written code from machine-generated submissions while providing meaningful feedback that supports student learning [11]. II. LITERATURE SURVEY A. Early Approaches and Output Comparison Systems The earliest automated programming assessment systems focused primarily on output comparison, executing student
code against predefined test cases and comparing program outputs with expected results. These systems provided basic correctness validation but lacked sophistication in evaluation criteria and feedback quality [4]. While effective for introductory programming courses with clearly defined input-output specifications, output comparison systems proved inadequate for complex assignments requiring subjective evaluation of code quality, design patterns, and algorithmic efficiency. The limitations of early systems motivated research into more comprehensive evaluation frameworks that could assess multiple dimensions of programming competence simultaneously. This evolution reflected growing recognition that programming proficiency encompasses not only the ability to produce correct outputs but also the capacity to write efficient, maintainable, and well-documented code adhering to software engineering best practices. B. Multi-Criteria Evaluation Frameworks Novak and Kermek (2024) provide comprehensive analysis of assessment automation for complex student programming assignments, documenting the evolution from simple output comparison to sophisticated multi-criteria evaluation frameworks [9]. Their research identifies key trends including integration of static code analysis tools for quality assessment, dynamic performance profiling for efficiency evaluation, automated testing frameworks supporting diverse programming languages, and plagiarism detection mechanisms adapted for code submissions. The transition to multi-criteria evaluation frameworks represents significant advancement in automated assessment capabilities. Modern systems evaluate correctness through comprehensive test suites covering edge cases and boundary conditions, assess performance through algorithmic complexity analysis and runtime profiling, evaluate code quality through metrics including readability, documentation, and adherence to coding standards, and verify originality through plagiarism detection and AI-generated code identification [3]. C. Systematic Review of Current Tools Messer et al. (2024) conducted systematic review of automated grading and feedback tools for programming education, analyzing trends in tool adoption and identifying patterns in implementation approaches [10]. Their analysis reveals growing emphasis on providing meaningful feedback beyond simple correctness indicators, integration of machine learning techniques for understanding code semantics, adoption of containerization technologies for secure code execution, and implementation of scalable architectures supporting large student populations. III. METHODOLOGY A. Large Language Models for Programming Evaluation Recent developments in artificial intelligence have fundamentally transformed automated assessment methodologies. Large language models demonstrate remarkable capabilities in understanding code semantics, identifying logical errors, and generating natural language explanations of programming concepts [6]. This advancement enables automated assessment systems to provide feedback quality approaching human instructors. Mohammad et al. (2025) introduced StepGrade, a novel approach utilizing context-aware large language models for grading programming assignments [11]. Their research demonstrates LLMs’ potential to understand nuanced aspects of code quality traditionally requiring human expertise. StepGrade evaluates submissions through multi-stage analysis including syntactic correctness verification, semantic analysis of algorithmic approach, evaluation of code efficiency and optimization, assessment of readability and documentation quality, and generation of personalized feedback addressing specific weaknesses. B. AI-Powered Error Analysis and Feedback Traditional automated assessment systems struggle with submissions containing compilation errors, runtime exceptions, or incomplete implementations. AI-powered error analysis addresses this challenge through sophisticated understanding of code structure and common programming mistakes [7]. The application of large language models to error analysis enables systems to identify specific error causes, explain why errors occur in accessible language, suggest concrete fixes addressing root causes, and provide guidance for completing partial implementations. IV. PROPOSED SYSTEM ARCHITECTURE A. Assessment Pipeline Architecture The comprehensive assessment pipeline integrates multiple evaluation modules to provide holistic code analysis. Figure 1 illustrates the complete workflow from submission to final grading. The pipeline architecture orchestrates multiple evaluation stages: Submission Processing: Students and instructors submit code through dedicated interfaces, which route submissions to the backend server. GitLab server integration enables version control and collaborative development tracking. Compilation and Validation: The compiler stage performs initial syntax validation and compilation for compiled languages. Successful compilation proceeds to correctness evaluation, while errors trigger the Error and Incomplete Evaluator using LLM analysis for constructive feedback generation. Correctness Evaluation: Executes comprehensive test suites comparing program outputs against expected results. Test cases cover normal operation, edge cases, and boundary conditions to ensure robust correctness validation. Custom Evaluation: Enables instructors to define specialized evaluation criteria beyond standard correctness testing, supporting diverse assignment requirements and learning objectives. Performance Analysis: Evaluates algorithmic complexity using AST-based analysis and pattern matching. Determines time and space complexity with confidence scoring, providing insights into solution efficiency.
Fig. 1. Comprehensive assessment pipeline showing evaluation flow through compiler validation, correctness testing, custom evaluation, performance analysis, plagiarism detection, and code quality assessment, with LLM integration for error analysis and quality evaluation. AI Plagiarism Detection: Employs sophisticated detection mechanisms to identify code similarities between submissions and detect AI-generated code patterns, maintaining academic integrity. Code Quality Evaluation: LLM-based analysis assesses readability, documentation, coding standards adherence, and software engineering best practices. Generates detailed feedback with improvement suggestions. Grade Synthesis: The grader component aggregates results from all evaluation modules, applying configurable weighting schemes to compute final scores. Results persist in the database for record keeping and progress tracking. B. Performance Analysis Architecture The performance analysis system employs a sophisticated multi-layered architecture designed for accurate complexity evaluation across multiple programming languages. Figure 2 illustrates the comprehensive architecture integrating language-specific analyzers with core services. The architecture comprises four primary layers: Interface Layer: Provides multiple access points including CLI (Command Line Interface) for direct invocation, HTTP API supporting Flask/Gunicorn for scalable web access, and container runtime enabling isolated execution environments. This flexibility ensures accessibility across different deployment scenarios from development environments to production systems. Core Services Layer: The orchestrator component manages the analysis workflow, routing requests to appropriate analyzers based on language detection. The analyzer bridge serves as an adapter layer interfacing with native language analyzers, handling process management and result aggregation. Configuration management maintains analyzer settings and execution parameters. Language Analyzers Layer: Contains specialized analyzers for each supported language. The Python Analyzer operates inprocess using Abstract Syntax Tree (AST) parsing for direct code analysis. The Java Analyzer executes as an external JVM process, utilizing JavaParser for AST construction and complexity calculation. The C Analyzer runs as a native binary compiled from C source, employing libclang for sophisticated AST traversal. The JavaScript Analyzer operates through Node.js, leveraging Babel parser for modern JavaScript syntax support. Support Services Layer: Provides foundational capabilities including pattern and rule repositories containing algorithmic patterns for searching, sorting, and common data structures; complexity calculation engines implementing time and space complexity analysis algorithms; data models defining result structures and representations; and utilities for file detection, validation, and formatting. C. Deployment and Execution Architecture Figure 3 presents the deployment architecture utilizing Kubernetes for container orchestration, providing scalability and reliability for production environments. The deployment architecture implements several critical components: Frontend Services: Users interact through Monaco Editorbased interfaces, providing rich code editing capabilities with syntax highlighting and autocompletion. The ingress controller manages external access and load balancing across service replicas. API Endpoints: RESTful API services expose functionality for code submission, analysis requests, and result retrieval. Built on Flask with Gunicorn workers, the API layer handles concurrent requests efficiently while maintaining responsive performance. Caching and Queue Management: Redis provides highperformance caching for frequently accessed data and implements queue management for asynchronous analysis tasks. This architecture enables horizontal scaling by distributing work across multiple executor nodes. Database Persistence: PostgreSQL database maintains persistent storage for user data, submission history, analysis results, and system metrics. Connection pooling optimizes database access, while volume mounting ensures data durability across pod restarts. Executor Nodes: Isolated executor pods perform actual code analysis in secure, resource-limited containers. This separation
Fig. 2. Performance Analysis System Architecture showing integration of language-specific analyzers (Python, Java, C, JavaScript) with core services including orchestrator, analyzer bridge, and support services for pattern recognition and complexity calculation. Fig. 3. Kubernetes-based deployment architecture showing user interaction flow through frontend interfaces to backend API endpoints, with Redis caching, PostgreSQL database persistence, and isolated executor nodes for secure code execution. ensures that potentially malicious code cannot compromise system infrastructure while enabling parallel processing of multiple submissions. D. Backend Implementation Details The backend system implements several critical architectural patterns for reliability and scalability: Database Layer: PostgreSQL handles submission history storage, user session management, language configurations, and execution metadata. Connection pooling optimizes database access under concurrent load, while volume mounting ensures data persistence across system restarts. Caching Layer: Redis manages execution queues, provides fast access to frequently used data, and handles temporary storage for ongoing executions. This reduces database load and improves response times for repeated requests. API Server: Exposes RESTful endpoints for code submission with authentication and request validation. Manages communication between components and provides status updates. Environment variable configuration enables flexible deployment across different environments. Worker Nodes: Execute code in isolated containers with security measures and resource limits. Process queued submissions supporting multiple programming languages while managing execution timeouts and reporting results. Security Features: Containerized execution provides isolation, implements resource limitations, and enforces network restrictions. Secret key-based authentication secures API endpoints, while environment variable configuration maintains secure credential management. V. PERFORMANCE EVALUATION AND RESULTS A. Complexity Analysis Accuracy Comprehensive testing of the performance analysis module across 45 code samples demonstrates strong accuracy in algorithmic complexity detection. Table I summarizes languagespecific performance.
TABLE I LANGUAGE-SPECIFIC COMPLEXITY ANALYSIS ACCURACY Language Accuracy Sample Size Python 100% 13 files Java 88.89% 9 files C 84.62% 13 files JavaScript 70.00% 10 files Overall 86.67% 45 files Python achieves perfect accuracy due to robust AST parsing capabilities and comprehensive pattern matching. Java demonstrates high accuracy with occasional misclassification of complex recursive patterns. C analysis succeeds in most cases but struggles with sophisticated linearithmic patterns involving nested logarithmic operations. JavaScript presents the greatest challenge, with accuracy limitations stemming from dynamic language features and varied coding styles. B. System Performance Characteristics Performance testing reveals several critical system characteristics: Analysis Latency: Average complexity analysis completes within 0.8 seconds for typical code submissions under 500 lines. Analysis time scales sub-linearly with code size due to efficient AST traversal algorithms. Concurrent Processing: The distributed architecture supports up to 50 concurrent analysis requests with Redis queue management. Worker node isolation prevents resource contention between analyses. Database Operations: Query execution times remain under 0.12 seconds for typical retrieval operations, supporting responsive user interfaces. Connection pooling maintains performance under concurrent load. Container Provisioning: Docker container startup averages 2.3 seconds, enabling rapid scaling to handle submission spikes during assignment deadlines. C. Database Integration Fig. 4. Database entries showing pass/fail results for multiple testcases across 34 submissions. D. CI/CD pipeline Fig. 5. Actual GitLab CI/CD pipeline execution for validation, testing, and grading stages. Figure 4 demonstrates the database integration capabilities of the automated assessment system. The PostgreSQL database maintains comprehensive records of submission evaluations, including detailed test case results across multiple student submissions. Each database entry captures submission metadata, individual test case outcomes (pass/fail status), execution timestamps, and performance metrics. This persistent storage enables instructors to track student progress over time, identify common problem areas across the class, and provide targeted feedback based on historical submission patterns. The database architecture supports efficient querying for generating analytics dashboards and progress reports. E. Limitations and Areas for Improvement Despite strong overall performance, several limitations require attention: Linearithmic Pattern Recognition: The 75% accuracy for O(nlog n)complexity indicates need for enhanced pattern matching. Many linearithmic algorithms employ sophisticated divide-and-conquer strategies that superficially resemble simpler patterns. Dynamic Language Challenges: JavaScript’s 70% accuracy reflects difficulties in analyzing dynamic language features. Dynamic typing, functional programming constructs, and asynchronous operations complicate static analysis. Recursive Algorithm Analysis: While simple recursive patterns achieve high accuracy, complex recursive algorithms with multiple base cases or interleaved recursive calls present analysis challenges. Confidence Scoring Calibration: Current confidence scores require refinement to better reflect actual accuracy. Some high-confidence predictions prove incorrect, while conservative scoring underestimates accuracy for clearly identifiable patterns. F. Comparative Analysis Of Existing Systems Table II compares major automated assessment systems. TABLE II COMPARISON OF EXISTING SYSTEMS System Correctness Plagiarism Performance Evaluation Detection Analysis CodeRunner Yes No No MOSS No Yes No HackerRank Yes Partial Yes StepGrade Yes No No Proposed AAE Yes Yes (AI) Yes
G. Test Coverage Fig. 6. PyTest execution output and coverage report for a sample student submission. H. Grading Coverage Fig. 7. Fully automated grading result summary for a student submission. VI. PLAGIARISM DETECTION AND ACADEMIC INTEGRITY A. AI-Generated Code Detection The integration of DetectCodeGPT provides sophisticated capabilities for identifying machine-generated code [8], [13]. The tool analyzes multiple dimensions of code characteristics: Stylistic Consistency: AI-generated code often exhibits unusually consistent style throughout, lacking the natural variations present in human-written code. Variable naming conventions, indentation patterns, and comment styles maintain uniformity exceeding typical human variation. Comment Patterns: Machine-generated comments frequently reflect training data characteristics, using specific phrasings and structures common in documentation but rare in student code. Comment density and positioning follow patterns distinguishable from human practices. Algorithmic Approach: AI models tend toward textbook implementations of algorithms, lacking the creative variations, optimizations, and minor inefficiencies characteristic of human problem-solving. Error Patterns: Human code contains characteristic error types reflecting learning progression and common misconceptions. AI-generated code, when containing errors, exhibits different error patterns resulting from model limitations rather than conceptual misunderstandings. B. Traditional Similarity Detection Code similarity analysis employs multiple complementary techniques: Token-Based Comparison: Normalizes code by removing whitespace, comments, and formatting, then compares token sequences. This approach detects superficial modifications like variable renaming and comment changes. Abstract Syntax Tree Comparison: Analyzes structural similarity by comparing AST representations. This technique identifies algorithmically equivalent implementations despite syntactic differences. Control Flow Analysis: Examines program logic through control flow graphs, detecting similarities in algorithmic approach even when implementation details differ. VII. CONCLUSION This review has examined the state-of-the-art in automated assessment environments for programming education, synthesizing research on evaluation methodologies, machine learning integration, and plagiarism detection mechanisms. The analysis reveals significant progress in developing comprehensive systems addressing multiple evaluation criteria while maintaining scalability and reliability. Key contributions of current research include demonstration of effective multi-criteria evaluation frameworks achieving strong accuracy in complexity analysis across Python, Java, C, and JavaScript; successful application of containerization technologies with Kubernetes orchestration for secure, scalable code execution; implementation of sophisticated CI/CD pipelines enabling reproducible evaluation workflows; integration of large language models for error analysis and code quality assessment; and development of specialized tools for detecting AI-generated code patterns. The current state of knowledge indicates that modern automated assessment systems have achieved substantial functionality for comprehensive programming evaluation. Successful implementations demonstrate reliable correctness validation across multiple languages, effective database integration supporting persistent data management, functional orchestrated deployment architectures, and initial AI integration for sophisticated feedback generation. Critical gaps requiring further research include enhanced linearithmic pattern recognition improving accuracy rates,
JavaScript analysis improvements addressing dynamic language feature challenges, advanced recursive algorithm analysis for complex call patterns, confidence scoring calibration better reflecting actual prediction accuracy, and expanded plagiarism detection addressing evolving AI code generation capabilities. VIII. FUTURE DIRECTIONS A. Unified Evaluation Pipeline A unified evaluation pipeline should integrate compiler validation, AI plagiarism detection, correctness testing, and LLMbased evaluators under a single orchestrator. This centralized architecture would streamline the assessment workflow, reduce processing overhead, and enable coordinated evaluation across all quality dimensions. B. LLM-Based Intelligent Feedback Systems Future implementations will leverage large language models to provide sophisticated error analysis and code quality assessment. An intelligent error and feedback system will interpret compilation and runtime errors, generating natural language explanations with step-by-step correction guidance. The system will distinguish between logical, syntax, and conceptual errors for precise feedback. Additionally, an LLM-based code quality evaluator will assess readability, modularity, naming conventions, and efficiency while providing automated refactoring suggestions and style compliance checks. C. ML-Driven Plagiarism Detection Advanced plagiarism detection will employ deep learning to understand code logic and semantics beyond simple text similarity. The system will perform semantic equivalence analysis to detect plagiarism even when code is restructured or rewritten, support cross-language comparison to identify functional equivalence across different programming languages, and provide cluster-based visualization of code similarity groups for instructors. D. Grade Synthesizer Module A comprehensive grade synthesizer will aggregate all evaluator outputs into final assessments. The module will implement weighted scoring combining correctness, performance, quality, and plagiarism metrics, generate LLM-assisted explanations of grading decisions, support adaptive grading models adjusting weights based on difficulty and instructor configuration, and integrate with database systems for structured result storage and dashboard display. E. Backend Architecture Enhancements System scalability and reliability will improve through architectural modernization. Plans include adopting microservices architecture to separate compiler, evaluator, and AI modules for independent scaling, implementing event-driven workflows using message queues for asynchronous processing, deploying caching and load balancing with Redis and Nginx, establishing secure API gateways with token-based routing and rate limiting, and developing instructor configuration portals for custom rubric definitions. F. Multi-File and SQL Support The system will extend capabilities to handle complex assignments requiring multiple source files and database interactions. Multi-file project support will enable evaluation of modular programming assignments with proper dependency management and inter-file relationship analysis. SQL support will facilitate assessment of database design, query optimization, and data manipulation tasks. The evaluation framework will verify query correctness, assess performance efficiency, and analyze database schema design quality. REFERENCES [1] C. S., S., C., B., & J., S. K. (2023). A Secure and Scalable System for Online Code Execution and Evaluation using Containerization and Kubernetes. Journal of Emerging Technologies and Innovative Research (JETIR),10(2), c197-c202. [2] Raut, S. S., Bamane, P. D., et al. (2022). A Novel Approach for Detecting Malicious JavaScript Code Using Machine Learning. In 2022 6th International Conference on Computing, Communication and Control (ICCUBEA) (pp. 1-6). IEEE. [3] Cardoso, M., de Castro, A. V., Rocha, ´ A., Silva, E., & Mendonc¸a, J. (2020). Use of Automatic Code Assessment Tools in the Programming Teaching Process. In 1st International Computer Programming Education Conference (ICPEC 2020) (Vol. 81). Schloss Dagstuhl-LeibnizZentrum f¨ ur Informatik. [4] Patil, P. B., et al. (2023). Malicious JavaScript Detection using Machine Learning. In 2023 3rd International Conference on Pervasive Computing and Social Networking (ICPCSN) (pp. 1126-1131). IEEE. [5] Fang, Y., et al. (2019). A Machine Learning Approach to Detect Malicious JavaScript. In 2019 IEEE International Conference on Big Data (Big Data) (pp. 4334-4339). IEEE. [6] Rozi` ere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X., ... & Synnaeve, G. (2023). Code Llama: Open Foundation Models for Code. arXiv preprint arXiv:2308.12950. [7] Mitchell, E., Zhou, Y., Khazatsky, A., Singh, M., Piech, C., & Finn, C. (2023). DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature. arXiv preprint arXiv:2301.11305. [8] Su, J., Zhuo, T. Y., Wang, D., & Nakov, P. (2023). DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of MachineGenerated Text. arXiv preprint arXiv:2306.05540. [9] M. Novak and D. Kermek, “Assessment Automation of Complex Student Programming Assignments,” Education Sciences, 2024. [10] M. Messer, N. C. C. Brown, M. K¨ olling, and M. Shi, “Automated Grading and Feedback Tools for Programming Education: A Systematic Review,” ACM Transactions on Computing Education (TOCE), 2024. [11] A. Mohammad, K. Zamiri Azar, and H. Kamali, “StepGrade: Grading Programming Assignments with Context-Aware LLMs,” arXiv preprint, 2025. [12] S. Bama, N. Mala, P. Ranjani, M. Saswati, and A. T. Jeena, “Automated Computer Program Evaluation and Projects – Our Experiences,” 2024. [13] Y. Shi, H. Zhang, C. Wan, and X. Gu, “Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers (DetectCodeGPT),” in Proc. 47th Int. Conf. on Software Engineering (ICSE), 2025.