Abstract
This study aims to examine the challenges, interventions and AI literacy outcomes of an artificial intelligence (AI)-enabled project-based learning (PBL) curriculum that adopts five sequential phases in a primary school classroom in Taicang, Jiangsu, China. The teacher-researcher used the methodology of observation with self-study to document 6 sessions from Weeks 10–15 of the academic year 2025–2026 using structured observation diaries that were guided by the LOTPBL framework (Ukobizaba et al., 2025) and the AI-assisted POPBL model (Habibah et al., 2025). A qualitative content analysis and reflective thematic analysis (Braun & Clarke, 2021) were used to identify the challenge patterns within each phase and intervention strategies. The results indicate that technical operation problems and insufficient prompt quality were the main issues during the prototype development; and difficulty in filtering AI generated content was a critical issue in the second prototype session and after that. Based on the theories of constructivist scaffolding and Chiu's (2021) holistic K–12 AI curriculum design model, targeted interventions—including pre-set prompt templates, teacher modelling, case sharing, and structured group discussion—produced measurable improvements in students' AI use behaviours. The study contributes practitioner evidence to underresearched primary-level AI-PBL implementation.
Keywords: AI literacy, project-based learning, primary education, observational self-study, scaffolding, AI-enabled pedagogy
Introduction / Background
The swift integration of generative Artificial Intelligence (AI) tools into daily educational practices has raised unprecedented expectations for teachers and students, especially in primary school where students' developmental readiness meets the aspirational requirements of AI literacy education. National policies stated in the New Generation Artificial Intelligence Development Plan (State Council of the People's Republic of China, 2017) and later in the Ministry of Education have prioritized AI literacy as a key competency for all levels of education, and this has brought pressures to the institutions, which forces primary school teachers to add AI-related learning activities into their teaching, but without any special preparation or evaluated and tested teaching models for younger students. This institutional momentum highlights a global trend: The extent to which generative AI tools are already integrated into professional, civic, and creative life has shifted the question of introducing AI literacy education with structure to young learners from speculative to urgent pedagogical. However, for primary school teachers in China, this challenge is exacerbated by the fact that there is currently a lack of widely accepted and empirical validations of curriculum models specifically tailored for upper primary learners, who are also a particularly significant group to intervene in the early stages of AI literacy due to their developmental stage's capacity to think flexibly and emergent metacognitive abilities.
This dissertation shares a practitioner-led investigation of the sustained efforts of one teacher-researcher to fill this gap by designing and implementing a curriculum of five phases of AI project-based learning (PBL) at Taicang GaoXinqu No.3 Primary School in Jiangsu Province from December to January 2025. The study builds on the new, underresearched literature on integrating AI in PBL in primary school settings (Yim & Su, 2025; Crompton et al., 2024; Alfarwan, 2025), and provides an insider, longitudinal perspective on how problems emerge in distinct phases of a PBL implementation cycle, intensify, and respond to intervention. The use of PBL as a curricular vehicle for AI literacy is intentional and theoretically sound, because the requirement of authentic problem solving, iterative design, and collaborative inquiry in PBL provides opportunities to apply AI tools as genuine cognitive partners, instead of passive answer machines, to produce the evaluative and creative approaches necessary for higher order AI literacy. Concurrently, the sequential, phase-structured nature of PBL is particularly well suited for longitudinal observations as it provides a natural structure for observing and tracking emergence and resolution of different types of challenges over time, with each phase posing qualitatively different demands on students' AI interaction competencies.
The research problem is of an intellectual nature and is at a structural level. While there are existing frameworks for K–12 AI literacy, such as the four dimensions developed by Ng et al. (2021) (knowing, using, evaluating, creating,and ethics with AI), which set aspirational goals, the developmental abilities of primary school students cannot reach these goals without carefully planned scaffolding and active teacher mediation. Meanwhile, the AI-PBL integration literature has tended to represent challenge categories as static and cross-sectional processes, as opposed to dynamic and phase-dependent processes that change in response to pedagogical intervention (Crompton et al., 2024; Zha et al., 2025). The findings reveal a continuing disconnect between the "what" of the frameworks and the "how" that is available to classroom practitioners on how to achieve this goal, especially for young learners. This gap is practical, as teachers will see numerous examples of what can go wrong when implementing AI-PBL at the primary level in the literature that already exists, but relatively little of what is likely to go wrong, at which point in the sequence of AI-PBL, and what can be done about it. The current study was developed to meet this need directly by developing observationally-based evidence in support of a real classroom implementation of the current study, which is phase-anchored.
The present study tries to fill this void on two levels. Empirically, it produces phase specific, observationally based data on the emergence, rise, and decline of four types of challenge: technical operation barriers (N1.1); insufficient prompt quality (N1.2); difficulty in critically filtering AI output (N1.3); and limitations in self-directed learning (N1.4) over six sessions distributed across five phases of a PBL implementation. A unique aspect of the contribution of the study is the granularity of this documentation: challenge types are not simply co-occurring properties of AI-PBL contexts in general, but are sequentially structured – a challenge of one type provides the conditions for the emergence of the next. The study is particularly important in its documentation of how successful responses to the quality challenges in the prompts of Session 1 created the conditions for the content challenges to be apparent in Session 2, which is not possible in the cross-sectional or aggregated data. At the pedagogical level, it details the intervention strategies implemented (pre-set prompt templates, teacher modelling and guidance, structured group discussion, peer-to-peer learning, and successful case sharing), and explains how the type and frequency of teacher intervention changed over time as the students gradually internalized the skills introduced by the scaffold.
The study is based on two related theoretical frameworks. The theoretical account of how primary learners can become more AI literate and develop higher-order skills is provided by constructivist learning theory, specifically Vygotsky's (1978) zone of proximal development (ZPD) and its operationalisation in the classroom as scaffolding. In this context, the teacher's role is not only to correct students' errors but to function within the zone of proximal development (ZPD) between what can be done independently and what can be done with assistance, which is observable in the present study at the level of the session. Chiu's (2021) holistic K–12 AI curriculum design model offers a systemic explanation of the conditions (such as student relevance and teacher-student communication) under which those scaffolds are more likely to lead to enduring gains in competency. These two theories, when combined, make clear predictions about the evolution of the three variables – teacher guidance frequency, student AI usage behaviour, and emergence of higher-order AI literacy indicators – across a well-designed sequence of phases, which the phase-structured data in the study is set to assess.
The research questions are: The first is, what difficulties did they face when it came to implementing an AI-enabled PBL curriculum in the five phases? What solutions and support strategies were put in place and noticed to impact those challenges and how did student AI usage behaviour change? The questions are organized to be answerable from the self-observational data yielded from this particular implementation context and are related to and extend challenge categories and intervention logics found in the literature on AI-PBL as a whole. The formulation of their questions after the literature search is intentional – the exact wording of each question (especially the notion of 'phase specificity' in RQ1, and 'observable behaviour change' in RQ2) is directly related to the gaps that were identified in the literature search, and to theoretical predictions based on the conceptual and theoretical frameworks used in the study.
Methodology
Research Design: Observational Self-Study
This study adopts an observational self-study methodology, in which the teacher-researcher serves simultaneously as practitioner and analyst (Creswell & Poth, 2016). Data collection was carried out by systematically documenting my own teaching decisions, instructional strategies, and professional reflections across six sessions. This self-study methodology produces an insider perspective on my own practice as a teacher implementing an AI-PBL curriculum. I maintained a reflective diary after each session, focusing on my observations of general classroom dynamics (without recording any identifiable student data).
Data Collection
The data of this study comprise the teacher-researcher's observation diaries and reflective notes, prepared systematically after each class implementation through five phases of implementation. These records detail my professional observations of general classroom activities, including common patterns in how students interacted with AI tools (without linking any observation to specific individuals), the issues I noticed in group work, and the interventions I implemented using the LOTPBL framework (Ukobizaba et al., 2025) as an observation instrument with an established interrater reliability and adapting the phase structure from the POPBL-AI model (Habibah et al., 2025) to represent the sequence of five phases of implementation.
Observation data were recorded across six sessions spanning five key phases:
Phase 1 — Prototype Creation, Session 1 (Week 10, 8 December 2025): first AI interaction session; primary challenges documented
Phase 1 — Prototype Creation, Session 2 (Week 11, 15 December 2025): preset prompt templates introduced; continued prototype development
Phase 2 — Testing & Revision (Week 12, 22 December 2025): prototype modification using AI; teacher presented successful cases
Phase 3 — Prototype Optimisation (Week 13, 29 December 2025): poster refinement and AI-assisted speech script generation
Phase 4 — Final Presentation (Week 14, 5 January 2026): group presentations, peer evaluation, and voting
Phase 5 — Reflection (Week 15, 12 January 2026): individual and group reflection forms
The observation diary was organized around the nine LOTPBL dimensions, and included extra interpretive free-form sections to the diary that recorded thoughts that arose during the observation. Teacher guidance interventions were documented in session-anchored records which then allowed for frequency count disaggregation by session for descriptive purposes. The written reflection records, in conjunction with the longitudinal observation diaries, therefore, form the full data set for this study.
Figure 1. AI-enabled PBL classroom implementation.
Data Analysis
The data from the observation records has been qualitatively content analysed using reflective thematic analysis (Braun & Clarke, 2021). The analysis was conducted by repeated and in-depth reading of the diary texts, to identify and distil the core themes and patterns following the procedures described by Creswell and Poth (2016). The coding framework was created in an open-ended manner based on the identified types of challenges and intervention responses, ultimately resulting in 28 codes that were grouped into five thematic nodes: N1 Challenges, N2 Intervention Strategies, N3 PBL Implementation Elements, N4 AI Literacy Development, and N5 Student Growth and Reflection (see Appendix A for full codebook).Systematic organisation and retrieval of coded segments were achieved by using NVivo 15 (QSR International) qualitative data analysis software for coding all the data.
Research Ethics and Research Quality
This study was approved by the Academy of Future Education, Xi'an Jiaotong-Liverpool University (Student ID: 2468117; Supervisor: Na Li). The sole data source was my personal observation diary and reflective notes, which contain no identifiable information about any individual. The school's headteacher granted administrative permission for me to conduct self-study activities within my classroom. As no personal data was collected, no informed consent from parents or students was required. All records are stored securely on a password-protected and encrypted device, accessible only to me.
Results / Findings
The results are arranged according to the five phases of the implementation of PBL, which are also aligned with the observation design based on the phases. For each phase, findings are provided from two perspectives: the problems observed (in response to RQ1) and the actions taken and the observed impact on student AI usage behaviours (in response to RQ2). A cross-phase summary (Section 5.6) draws together the direction of change of challenges and interventions throughout the entire implementation process.
It is noted that there were two sessions during Phase 1 (Prototype Creation) in the actual classroom implementation, which happened in Weeks 10 and 11 (8 December and 15 December 2025), indicating that the prototype creation took more than one session. There are observation records for all six sessions.
Phase 1: Prototype Creation (Weeks 10–11, Sessions 1–2)
Session 1 (8 December 2025) — Challenges Encountered
The first prototype production session had the largest density and range of challenge cases (over the whole implementation). Two challenge types were recorded at the same time: technical barriers in operating the prompt (N1.1) and lack of prompt quality (N1.2).
Technical operation barriers (N1.1) was mainly the slow typing speed, which greatly affected the students' typing efficiency of inputting AI prompts. Some students had to repeat the process of inputting prototype design requirements into the AI platforms (Doubao AI and KIMI) because they did not have a high level of proficiency in typing the requirements. One group attempted to enter prototype design requirements three times before they were able to get a usable interaction. Of particular interest is that the session was marked as "Not Occurred" for the student-centred learning dimension (N3.1) as students had not yet been able to fully control the AI-assisted inquiry process.
The second main difficulty was lack of prompt quality (N1.2) and was noticed in several groups. Some students were unable to write follow-up questions if the AI response was unclear or if they didn't receive a sufficient response and multiple groups were unable to write prompts to match their prototype content. The observation record reveals the following patterns for group-level interactions: one group provided the AI with non-specific design ideas and lacked the ability to elaborate on the interaction; one group provided two sets of prompts, but neither one was acceptable to them; two groups had low overall efficiency while interacting with the AI. This session was marked as "Vague" in the AI Usage Record. Importantly, on this first session, all six dimensions of critical thinking observation (N4.1-4.6) and all six dimensions of communication skills observation were rated as "Not occurred," meaning that the students were not working at the level of critical thinking or communication with the AI-generated output.
Session 1 (8 December 2025) — Interventions and Observed Effects
During the session, the teacher offered some tips on typing shortcuts and typing efficiently to overcome the typing speed barrier. As part of the response to the quality deficits, two complementary interventions were provided: the teacher modelled how to use follow-up questions (N2.2 — Teacher Modelling & Guidance) and inter-group peer tutoring was organised, between higher and lower efficiency groups (N2.4 — Peer-to-Peer Learning). Also, the teacher broke down the design criteria into specific prompt parts at the component level to lessen the number of prompt parts that are complex. This session had the highest number of instances of teacher guidance in the implementation, 8.
Figure 2. Pre-set AI prompt template used during prototype creation.
Session 2 (15 December 2025) — Challenges Encountered
There was shown to be measurable improvement in Session 2 compared to Session 1. The teachers had prepared the preset prompt templates in advance (N2.1 — Pre-set Prompt Templates) and added them to the class presentation, and each group could use the relevant prompt templates to specify the prototype they were working on, using Doubao AI to write the content, KIMI to conduct research tasks, and Wenduoduo PPT AI to create slides. The increase of usage frequency of AI to "many times" (多次) and the change of question quality dimension from "Vague" to "Specific and Clear" in Session 1 was a direct reflection of the prompt formulation in the preset template. In AI output handling, it was noted as "Critically Modified", which is a further qualitative improvement as compared to Session 1.
Another challenge has come up though as a new one, which is students' inability to accurately filter the AI-generated content (N1.3). Some groups noted problems in discerning what information was relevant in what the AI had generated, and students needed guidance to concentrate on the main functionality of the prototype and not to add everything the AI had suggested. The challenge of critical filtering of AI output was not seen in Session 1 (where output quality was the main problem), but was clearly seen in Session 2 because good prompting led to the generation of more and longer outputs from which to choose. The reasoning dimension of critical thinking (N4.2 — evaluation of reliability of information sources) was rated as "Not occurred", indicating that students lacked systematic strategies to evaluate AI-generated information.
Session 2 (15 December 2025) — Interventions and Observed Effects
The teacher used two more interventions in response to the challenge of content filtering in addition to the preset prompt templates that are already there. First, a structured group discussion session was introduced (N2.3), where groups evaluated together the relevance of the AI-generated content for the main functions of their prototype. Second, the teacher gave explicit instructions to groups who had trouble selecting content to help them think about the central functions of the prototype they were creating, and used that as a standard to assess the AI's suggestions (N2.2). Successful cases would be shared during future lessons in the observation record, to expand student thinking regarding content selection strategies. The number of teacher guidance for this session was 5. Students Independent thinking time 20 minutes.
Phase 2: Testing & Revision (Week 12, 22 December 2025)
Prototype revision from peer and teacher feedback was the major activity of Phase 2. The students edited the colour of the AI-generated posters, edited the design elements, and modified the reward and punishment mechanisms in their prototypes that were made with PPT. The purpose of AI use was "Content Generation" and "Optimization", meaning that students have started using AI as a refinement tool instead of their primary source of content. The frequency of teacher guidance was reduced to 3, while the time for students to think independently was increased to 25 minutes (the highest in implementation until now).
Interventions and Observed Effects
The observation record highlights that the most impactful intervention in Phase 2 was the sharing of successful cases from previous groups at the beginning of the session (N2.5 — Successful Case Sharing), which "effectively addressed the students' previous confusion in selecting the AI content, helping students find the optimal entry point quickly". This inter-session intervention was planned at the end of Session 2 of Phase 1 and had an immediate and observable impact: students were able to make independent, critical changes to the AI-generated content in a way that was personalised, whereas in the previous sessions this was not the case.
Groups applied AI to support content optimisation, also making independent and personalised content optimisations according to the specific features of their project, thus reflecting the rational and critical use of AI tools. Groups presented their optimised prototypes and received peer feedback (N2.3) in the last 15 minutes, where they collaboratively discussed the prototypes to further clarify the direction for revisions. Improper copying of AI-generated content was noted to have been replaced by students making critical revisions based on their needs, as reflected in the overall evaluation.
Figure 3. Group collaboration during PBL project work.
Phase 3: Prototype Optimisation (Week 13, 29 December 2025)
AI use changed in this stage, with students mainly relying on AI to create and tweak their speech scripts, but not to write content. In group discussions, the PPT group edited the AI-generated scripts paragraph by paragraph, a more deliberate human/AI collaboration than the students had engaged in during Phase 1 when they had examined AI-generated outputs at the level of the full responses. This revision on sentence level is a tangible evidence of the N4.3 dimension (Critical Handling of AI Output). Both groups were able to finish their work within the session, one based on working with physical prototype refinement and the other on AI-assisted script creation.
Phase 4: Final Presentation (Week 14, 5 January 2026)
Phase 4 was designed as the session to present the results. There are no challenge incidents recorded on the observation record. AI was only sparingly used in this class, and then only in preparation for the class, and the value of its use was in the preparatory stage: “Students use AI to create and optimize the content of posters and PPTs, which is the foundation for this high-quality presentation. AI usage frequency during the class was "少量" (minimal/rare), and there was no additional guidance to be given by the teacher on the use of AI, with just one occasion when the teacher used AI.
Phase 5: Reflection (Week 15, 12 January 2026)
Structure and Observations
Phase 5 was designed as a whole class review and reflection. No teacher AI guidance was needed, and there was no use of AI during this session. The independent thinking time of students was 30 minutes, the longest of the whole implementation process, which was consistent with the dominant form of individualization and reflection of the tasks.
Groups then delved into deeper discussions and filled out group reflection forms in the third segment (16:55 – 17:10) on the project process, AI used and teamwork. In their group reflections, students commonly noted that AI-generated content (PPTs, pictures) required revision to fit their project needs, and that AI should be viewed as a helper rather than something to rely on entirely.
AI Literacy Reflections
The reflection session yielded high-quality evidence clarifying the students' AI literacy awareness.Some students identified subtler differences in how AI can be used in their learning, noting that interactions should be kept simple and contextualized for their age and school project level, indicating an understanding of the need to frame AI interactions appropriately.
Another reflection pointed out that the colour scheme of the poster created with AI appears to look bad, and that the colour of the poster was changed to become more appropriate to the campus, which is done by the aesthetic sense of the creators. Questions about AI's factual accuracy emerged from group discussions, indicating emerging critical awareness– showing initial critical consciousness towards the limitations of AI in terms of facts and truth. The group consensus that AI should be viewed as a helper rather than something to rely on entirely was the most explicit expression of the AI literacy orientation achieved by the intervention.The overall assessment indicated that the reflection session "contributed to having a rational understanding of AI tools" by the students.
Cross-Phase Summary
The full set of session by session AI usage indicators is presented in Table 1, and is used as the quantitative evidential basis for the cross-phase analysis that follows in Sections 5.6.1 – 5.6.3.
Table 1.
AI Usage Indicators Across Six Observation Sessions. All data recorded verbatim from Section C of each observation diary. Shading: Vague prompt quality highlighted in Session 1 (8 Dec. 2025); zero teacher AI guidance highlighted in Session 6 (12 Jan. 2026).
Table 2.
Critical Thinking (4.1–4.6) and Communication Skills (5.1–5.6): Session-by-Session Observation Results. ✓ = Occurred; ✗ = Not Occurred. Data recorded verbatim from Sections 4 and 5 of each observation diary.
Table 3.
Student Agency (N3.1) and AI Literacy Development (N4.1–N4.4): Session-by-Session Observation Results. Data derived from Sections B, C, and E of each observation diary. ✓ = Observed; ✗ = Not Observed; Emerging = partially demonstrated.
Intervention Trajectory and Teacher Guidance Frequency
There was a decrease in the number of times teachers gave guidance throughout the implementation: 8 during Session 1 (Phase 1), 5 during Session 2 (Phase 1), 3 during Phase 2, 4 during Phase 3 and 1 during Phase 4. In Phase 5 no teacher AI assistance was needed. This non-linear but overall downward trend confirms the constructivist expectation that scaffolded learning will lead to more and more autonomous use of AI and Chiu's (2021) suggestion that the communication between the teacher and students should be less and less directive as the scaffolding is internalised. The nature of the intervention also changed from providing direct technical assistance and prompt decomposition (Session 1) to effective case sharing and structured group discussion (Phases 2-3) to minimal facilitative presence (Phases 4-5).
AI Literacy Development Trajectory
Differentiated trajectories were identified for the four AI literacy development dimensions (N4.1–N4.4). AI Usage Purpose Awareness (N4.1) was present starting from Phase 2 as students were able to use AI only for particular optimisation tasks instead of seeing it as a generic content source. Evolution of Prompt Quality (N4.2) demonstrated the most obvious and swift progression from "Vague" in Session 1 to "Specific and Clear" by Session 2 and stayed at this level throughout Phase 4. The most contested dimension was Critical Handling of AI Output (N4.3) where it was noted to be absent in Session 1, emerging in modified form in Session 2 (students had the option to choose from the AI output but had some difficulty), independent and critical in Phase 2, and most richly articulated in Phase 5 reflections. Understanding of AI's role (N4.4) was most evident in phase 5 reflections where students were able to share principled and contextualised understandings of the capabilities and limitations of AI in the context of their project.
It is important to note that the N4.1–N4.4 trajectory documented here does not represent evidence of full achievement of AI literacy as conceptualised by Ng et al. (2021). The data shows evidence of emerging evaluative engagement, that is, students' critical modification of the contents generated by AI, identification of factual errors produced by AI and articulations of principled limitations of AI, in line with the operational definition of the “evaluating” dimension. This is best described as a trajectory towards the upper levels of Ng et al.'s (2021) framework, not necessarily the attainment of these levels, which is developmentally realistic for primary aged learners (Su & Zhong, 2022) and consistent with Wu and Zhang's (2025) finding that AI integration yields measurable gains in students' digital literacy when pedagogically structured.
Implications and Recommendations
Implications for Primary AI-PBL Curriculum Design
The most concrete implication of this research for the application of AI-PBL curriculum design in primary school is to expect the sequential appearance of various types of challenges and design the sequential intervention. The current research shows the following trend: Technical and prompt quality challenges are most pressing in the first exposure session; Content filtering challenges are evident in the second (when improved prompting results in outputs that need to be evaluated), and in the third session (with appropriate scaffolding in place), the students are able to use AI critically and independently. It is important that curriculum designers expect to have at least two prototype creation sessions (not one), and that the second session focus on content evaluation needs that better prompting will result in.
The preset prompt template scaffold should be used as a starting point for any primary implementation of AI-PBL, specifically as a first session intervention. The effectiveness in this study, of moving the question from Vague to Specific and Clear within one session, indicates that this is a true ZPD boundary for primary learners for the first time encountering AI tools. Most importantly, however, the template can be thought of as the first of the 2-session Phase 1 design, not as a one-session solution: the template allows for the prompt quality issue to be resolved, but the condition for the content filtering challenge to occur will be set and this will be addressed in the structured group discussion and successful interventions in case sharing that will follow in the second session.
Teacher preparation for AI-PBL needs to develop practitioners' ability to monitor the changing needs for guidance from session to session, and to expect the "success creates new challenge" phenomenon found in this study. It is not obvious that the reduced guidance in phases 5 and 6 is a good thing and is a consequence of the scaffolding design working as intended, as practitioners would have been used to a consistent level of guidance provision in session 1.
Implications for AI Platform Selection and Multi-Platform Design
Unlike other studies that only use one AI platform for their research-based tasks, this study utilized multiple AI platforms to generate content and create slides: Doubao AI for generating general content, KIMI for research-based tasks, and Wenduoduo PPT AI for slide production. The observation records indicate that the multi-platform strategy was more suitable to the prototype creation task's various output requirements than using only a single platform, and introduced a particular technical problem: tool switching in Session 2 resulted in delay for the PPT production group. Future curriculum designs should be multi-platform flexible and give advance information on tool adaptation to reduce the switching costs.
Discussion and Conclusion
Addressing RQ1: The Phase-Specific and Session-Specific Structure of AI-PBL Challenges
Findings answer RQ1 in detail and longitudinally, confirm and extend challenge categories in existing literature. The most theoretically important empirical finding is on the sequential emergence of challenge types during the implementation. The most common issues in Session 1 of Phase 1 were technical (slow typing speed) and prompt (formulating effective queries), which aligns with Crompton et al.'s (2024) description of insufficient technical skills as a structurally recurring issue for K–12 AI education. The observation record, however, showed the aspects of the technical challenge that could not be seen in aggregated, cross-sectional studies: The technical challenge in this context was not one of navigating a platform or encountering login problems, but typing speed, and this was effectively addressed in one session through the provision of targeted typing support and peer support.
The finding of content filtering difficulty (N1.3) in Session 2 and not in Session 1 is one that can only be understood in the context of the moment and cannot be inferred from cross-sectional data. In Session 1, students' use of AI was constrained by imprecise prompting, resulting in very basic responses that could be accepted or rejected as a whole. Preset templates were better in terms of prompting quality and provided more extensive output that was now deserving of evaluation and selection, by the end of Session 2. N1.3 challenge, in other words, was not a result of the failure of the Session 1 intervention, but a result of that intervention as it showed himself to be a more advanced challenge. This "success creates new challenge" is a hallmark of scaffolded learning progressions that are well designed and is a major contribution of phase-anchored observational research to the literature on AI-PBL.
Addressing RQ2: Intervention Effectiveness and the Scaffolding Logic
These findings include strong and specific evidence relating to the five intervention strategies coded in the codebook. The most immediately apparent and quantifiable outcome of the pre-set prompt template (N2.1) was the change from "Vague" (Session 1) to "Specific and Clear" (Session 2) to this level within one session after the intervention and for the duration of the implementation. This was the best fit in this data set with the ZPD scaffolding logic, because the template made the prompt formulation process a cognitive structure that could not be produced by the students on their own, thus allowing them to perform at the top edge of their capabilities (Vygotsky, 1978).
The progressive decrease of teacher guidance frequency across sessions (8, 5, 3, 4, 1, 0) is in line with the constructivist idea of a gradual scaffolding of teacher guidance. The non-monotonic element was a slight increase in the levels 3-4 between Phases 2 and 3, because students needed a little extra scaffolding to learn about the new task demand in Phase 3 (AI-assisted speech script generation), before applying their existing skills with AI to the new format. This contextual sensitivity in guidance frequency is generally in line with Chiu's (2021) model, in which communication between teachers and students should vary according to the level of task complexity, with more instruction provided when the task complexity is high, even if AI literacy general skills have been acquired.
The Phase 5 reflection data offers the best summative evidence on how effective the intervention is, but with the proper caveat as the data is self-reported. Their spontaneous comments about the limitations of AI (plant names that are not correct, pictures that "don't match what we wanted," PPT content that "needs to be revised") coupled with their awareness of the need to contextualise their interaction with AI for their own student level, and their principled group consensus that "AI is a helper, but it cannot be relied on entirely" taken together suggest that the higher-order AI literacy dimensions identified by Ng et al. (2021) were meaningfully engaged with, if not fully internalised, at the end of the implementation.
References
Alfarwan, A. (2025). Generative AI use in K-12 education: A systematic review. Frontiers in Education, 10: 1647573.
https://doi.org/10.3389/feduc.2025.1647573 Braun, V., & Clarke, V. (2021). Thematic analysis: A practical guide. Sage Publications.
Chiu, T. K. (2021). A holistic approach to the design of artificial intelligence (AI) education for K-12 schools. TechTrends, 65(5), 796-807.
https://doi.org/10.1007/s11528-021-00637-1 Creswell, J. W., & Poth, C. N. (2016). Qualitative inquiry and research design: Choosing among five approaches. Sage Publications.
Crompton, H., Jones, M. V., & Burke, D. (2024). Affordances and challenges of artificial intelligence in K-12 education: A systematic review. Journal of Research on Technology in Education, 56(3), 248-268.
https://doi.org/10.1080/15391523.2022.2121344 Habibah, L. B., Ibrohim, I., & Susilo, H. (2025). The effect of AI-assisted problem-oriented project-based learning on students' critical thinking and communication skills. JPBI (Jurnal Pendidikan Biologi Indonesia), 11(2), 656-668.
https://doi.org/10.22219/jpbi.v11i2.40667 Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, 100041.
https://doi.org/10.1016/j.caeai.2021.100041 State Council of the People's Republic of China. (2017). New generation artificial intelligence development plan.
http://www.gov.cn/zhengce/content/2017-07/20/content_5211996.htm Su, J., & Zhong, Y. (2022). Artificial intelligence (AI) in early childhood education: Curriculum design and future directions. Computers and Education: Artificial Intelligence, 3, 100072.
https://doi.org/10.1016/j.caeai.2022.100072 Ukobizaba, F., Maniraho, J. F., & Uworwabayeho, A. (2025). Lesson observation tool for project-based learning: a useful tool for learner-centered pedagogy enhancement. Frontiers in Education, 10, 1623269.
https://doi.org/10.3389/feduc.2025.1623269 Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes (Vol. 86). Harvard University Press.
Wu, D., & Zhang, J. (2025). Generative artificial intelligence in secondary education: Applications and effects on students' innovation skills and digital literacy. PLoS One, 20(5), e0323349.
https://doi.org/10.1371/journal.pone.0323349 Yim, I. H. Y., & Su, J. (2025). Artificial intelligence literacy education in primary schools: a review. International Journal of Technology and Design Education, 35(5), 2175-2204.
https://doi.org/10.1007/s10798-025-09979-w Zha, S., Qiao, Y., Hu, Q., Li, Z., Gong, J., & Xu, Y. (2025). Designing child-centric AI learning environments: Insights from an LLM-powered creative project-based learning study. International Journal of Human-Computer Studies, 204, 103602.
https://doi.org/10.1016/j.ijhcs.2025.103602