Abstrakt
Purpose: The study evaluated the effectiveness of AI tools in automatically assessing student database projects in content management systems (CMS). It focused on evaluation criteria covering data structure, performance, integrity, scalability, and business context.
Need for the Study: Assessing CMS database projects is complex, combining technical and business considerations. With increased reliance on data, efficient and objective evaluation methods are needed. AI offers potential for automation, but its reliability and alignment with expert judgment were unclear, motivating the study.
Methodology: Several AI models (e.g., ChatGPT, Google Gemini, Microsoft Copilot, DeepSeek, and Grok3) were used to grade a set of student CMS database projects. Their scores on the five criteria were compared to expert reference evaluations to gauge accuracy and consistency.
Findings: DeepSeek and Microsoft Copilot showed the smallest deviations from expert grades but struggled to distinguish the highest- and lowest-quality projects (indicating score smoothing). Grok3 was the most balanced, closely aligning with experts while preserving some variability. In contrast, ChatGPT and Google Gemini had larger deviations and tended to misjudge project quality (over- and under-scoring respectively).
Practical Implications: AI tools show promise for streamlining database project evaluation in educational and business contexts by speeding up grading and providing initial feedback. However, limitations like averaged-out scoring and difficulty capturing context mean they cannot replace human experts. Further refinement—such as better handling of outlier cases and integration with analytical tools—is recommended to improve reliability. With improvements, AI evaluators could become valuable support tools in academia and industry.