An exploratory pilot study of artificial intelligence
Within the context of rapid technological advancement and the accelerating global digital transformation, higher education institutions have increasingly adopted artificial intelligence–based applications as part of their instructional practices. In line with this direction, King Faisal University has sought to implement contemporary educational programs that enhance students’ cognitive skills, particularly those related to higher-order thinking and innovation. The present study was designed on the premise that the university learning environment provides a suitable context for examining the effectiveness of AI-based instructional interventions in developing productive thinking skills.
Productive thinking, which integrates both creative and critical thinking processes to achieve meaningful and practical outcomes, represents a key competency for university students navigating the complex demands of the 21st-century labor market. This construct encompasses multiple interrelated dimensions, including fluency, flexibility, originality, inference, interpretation, expansion, and imagination, which together form a comprehensive framework that can be systematically nurtured through AI-enhanced teaching strategies.
Accordingly, the central hypothesis of the study was formulated as follows:
There is no statistically significant difference at the significance level (α ≥ 0.05) between the mean scores of the experimental group students on the pre-test and post-test of the Productive Thinking Skills Scale.
This study employed a quasi-experimental design, specifically a single-group pre-test–post-test design, to investigate the effectiveness of an AI-based instructional program in enhancing productive thinking skills among female students at King Faisal University. Given the exploratory nature of implementing AI-based instruction in this specific educational context and the availability of participants, the study was conducted with nine female students as an initial investigation. This design allowed for a preliminary assessment of the program’s associative outcomes and feasibility.
The absence of a control group was a deliberate methodological decision aligned with the exploratory and pilot nature of the study. Institutional constraints, limited course enrollment, and ethical considerations related to equal access to innovative instructional practices prevented the formation of a comparison group. As such, the study was designed to generate preliminary evidence regarding feasibility and potential effectiveness rather than causal inference.
The study population consisted of all female students who enrolled in the educational technology course during the first semester of the 2024–2025 academic year at the College of Sharia, King Faisal University, in Al-Ahsa Governorate. The sample consisted of nine female students (n = 9) who were purposively selected from the study population to match the intervention’s objectives. Participation in the experimental group was limited to individuals who received the instructional program. Their ages ranged between 19 and 20 years, with a mean age of 19.4 years and a standard deviation of 0.5 years. All participants were female undergraduate students from the same academic program, reflecting a homogeneous sample in terms of educational background and cultural context. The internal and external validity limitations inherent to this small, single-group cohort are detailed comprehensively in the Limitations section.
The decision to utilize a sample of nine students was necessitated by institutional enrollment limits within the specialized Educational Technology course and the desire to maintain high instructional fidelity during this pilot phase. While this small n limits broad generalizability, the design serves as a foundational case study to explore the associative relationship between AI integration and cognitive development in a real-world higher education setting.
This sample size carries direct statistical consequences that should be made explicit. With n = 9, the design has adequate power to detect only large within-group effects and is underpowered for small or moderate ones; correspondingly, the 95% confidence intervals for the pre–post differences reported in Table 3 are wide, reflecting substantial sampling uncertainty around each point estimate. Readers should therefore treat the effect-size magnitudes as indicative of this exploratory pilot sample rather than as precise, generalizable parameters.
The study followed a systematic implementation process as illustrated in Figure 1. Before data collection, necessary approvals were obtained from the relevant institutional authorities, and informed consent was secured from all participants after explaining the study’s objectives, procedures, and ethical considerations.
Study procedure and data collection timeline.
As shown in Figure 1, the study was conducted over 16 weeks during the first semester of the 2024–2025 academic year. The pre-test was administered in Week 1 using the Productive Thinking Skills Questionnaire to establish baseline performance across seven dimensions. The AI-based instructional program was implemented over 13 weeks (Weeks 2–14), integrating AI-enhanced learning activities within the educational technology course. The post-test was administered in Week 15 using the same instrument to measure skill development. A reflective consolidation phase (Weeks 15–16) incorporated reflective journaling, peer discussions, and digital storytelling to foster metacognitive awareness and facilitate the transfer of skills to real-world contexts.
All data were anonymized, coded, and processed for quantitative analysis in alignment with the ethical standards for educational research. Throughout the process, participants were provided with opportunities for feedback, reflection, and clarification to ensure the integrity and reliability of the intervention outcomes.
The present study employed a comprehensive psychometric instrument to assess participants’ productive thinking skills. The assessment tool, developed and validated by Salah (2024), was explicitly designed to measure the multidimensional construct of productive thinking through seven empirically derived dimensions: originality, fluency, flexibility, inference, interpretation, expansion, and imagination. The instrument comprises 20 items strategically distributed across these dimensions to ensure comprehensive coverage of the theoretical framework: originality (3), fluency (4), flexibility (2), inference (4), interpretation (2), expansion (3), and imagination (2) (Appendix).
All items presented respondents with five-point Likert-type options (1 = needs development, 2 = developing, 3 = satisfactory, 4 = strong, and 5 = excellent), enabling a nuanced evaluation of participants demonstrated productive thinking abilities. The cumulative score across all dimensions provides a composite index of productive thinking capability, with higher aggregate scores indicating more advanced productive thinking skills. The psychometric properties of the instrument were rigorously evaluated. Internal consistency reliability was assessed using Cronbach’s alpha coefficient. The overall reliability of the scale was high (α = 0.84), indicating strong internal consistency. Furthermore, all sub-dimensions demonstrated acceptable reliability levels, with alpha coefficients exceeding the recommended threshold of 0.70, confirming the adequacy of the instrument for research purposes. Table 1 presents the Cronbach’s alpha reliability coefficients for the total scale and each sub-dimension of the Productive Thinking Skills Questionnaire.
Reliability coefficients of the sub-dimensions.
While the reliability evidence in the original submission focused solely on internal consistency (Total Scale α = 0.84), additional validation procedures were undertaken to strengthen the psychometric basis of the instrument. Face and content validity were established prior to the intervention by a panel of five expert judges in educational psychology and curriculum design, who evaluated each item for theoretical alignment, clarity, and cultural appropriateness; items falling below an 80% consensus threshold were revised accordingly. Construct validity was further examined through item-total correlation analysis conducted with a separate pilot group (N = 15) drawn from the same student population but not included in the main study. Pearson correlations between individual items and their respective sub-scale totals ranged from 0.58 to 0.79 (p < 0.01), indicating satisfactory item relevance and structural coherence. It should nonetheless be noted that this evidence, while stronger than internal consistency alone, still falls short of an independent, performance-based validation against an external criterion measure of productive thinking; consequently, post-test scores are best interpreted as self-perceived productive thinking rather than directly observed task performance. This distinction is revisited in the Limitations section.
The instrument was administered at two temporal points within the research design framework, once as a pre-intervention baseline measure and subsequently as a post-intervention assessment, to facilitate precise quantification of developmental changes in participants’ productive thinking abilities attributable to the AI-based instructional intervention program. This pre- and post-intervention approach allowed a robust comparative analysis of cognitive development outcomes across the experimental timeline.
The instructional intervention utilized a set of widely accessible generative and organizational artificial intelligence tools, including ChatGPT for idea generation, reasoning support, and feedback simulation, as well as Notion AI for structuring learning artifacts, facilitating reflective writing, and enabling collaborative documentation. These tools were selected due to their availability, ease of use, and relevance to higher education learning tasks. No custom-built AI systems were developed for this study; instead, commercially available platforms were integrated into instructional activities to reflect realistic classroom implementation.
To address the level of procedural detail requested during review, this subsection specifies the AI interaction protocol underlying the six core sessions summarized above. Across the 13-week instructional window, students engaged with ChatGPT and Notion AI in guided, in-class blocks of approximately 45–60 min per session (a 10-min instructor modeling phase, 25–35 min of guided student prompting and AI interaction, and a 10–15 min debrief/reflection phase), supplemented by an expected minimum of two independent prompting cycles per week logged by students between sessions. The instructor reviewed AI-generated outputs with students before they were incorporated into deliverables, using a standing rubric addressing relevance, originality, and logical coherence rather than accepting AI outputs at face value. Table 2 summarizes the core activity and a representative (not exhaustive) prompt structure for each targeted dimension; the complete prompt bank is available from the corresponding author upon request.
Illustrative AI interaction protocol by instructional session.
To enhance methodological transparency, an example of a typical AI-supported interaction is provided. During the brainstorming session, students entered the following prompt into ChatGPT:
“Generate five innovative instructional strategies that could improve collaborative learning among first-year university students, explaining the strengths and limitations of each strategy.”
A typical AI response included suggestions such as peer-led problem-solving workshops, gamified collaborative challenges, rotating discussion leadership, project-based learning teams, and digital collaborative portfolios. Rather than accepting these responses uncritically, students were instructed to evaluate each suggestion, identify weaknesses, modify the proposed strategies where appropriate, and justify their final selections. The instructor facilitated reflective discussions by asking participants to compare AI-generated suggestions with their own ideas and explain the reasoning behind revisions. This process emphasized that AI functioned as a cognitive support tool rather than a source of authoritative answers.
Implementation fidelity was monitored through two mechanisms: (1) the instructor’s session-by-session log, which recorded attendance, the AI tool used, the activity completed, and any deviation from the planned protocol; and (2) a review of students’ independent interaction logs against the minimum threshold of two prompting cycles per week. No sessions required substantive protocol deviation. These fidelity records are treated as supplementary process documentation for this exploratory pilot rather than as a formal fidelity index; future, larger-scale replications would benefit from a standardized fidelity checklist scored by an independent observer.
The experimental group was exposed to an instructional program specifically designed to answer the study’s first research question: “What is the effectiveness of an AI-based instructional program in developing productive thinking skills among female King Faisal University students?” To address this question, the researcher reviewed a wide range of relevant literature and previous Arabic and international studies focusing on AI and productive thinking skills. Based on this foundation, the instructional program was constructed with the following components:
Title of the instructional program: An AI-based instructional program for developing productive thinking skills among female students at King Faisal University.
Philosophical foundation: Within this framework, the instructor primarily functioned as a facilitator, guiding students in formulating effective prompts, evaluating AI-generated outputs, and reflecting critically on their learning outcomes. Learning tasks included AI-assisted brainstorming, adaptive quizzes, collaborative concept mapping, and project-based assignments. Assessment focused on formative evaluation through continuous feedback, reflective tasks, and performance on the Productive Thinking Skills Questionnaire rather than summative grading.
The program’s structure is based on interconnected stages that involve the careful selection of objectives, content, strategies, activities, instructional tools, and assessment methods. These are all designed to foster productive thinking skills that help students generate innovative ideas and solve problems more effectively.
Foundational principles: The program was built upon several key pedagogical and technological principles:
Utilization of strategies that encourage interaction, dialogue, and non-linear thinking, such as brainstorming, divergent thinking, mind maps, and problem-solving.
Incorporation of engaging activities and evaluation techniques that promote the development of productive thinking.
Emphasis on self-directed and continuous learning.
Use of varied AI tools and open, rich learning environments.
Encouragement of expressing multiple perspectives and scaffolding tasks from simple to complex.
Continuous and holistic assessment focusing on higher-order thinking, with diverse tools and immediate feedback.
Clearly defined roles: Instructors provide support and guidance, while students are active, independent, and engaged participants.
Immediate training and support for students to move beyond rote application toward autonomous problem solving and decision making.
Provision of a supportive atmosphere that respects freedom of expression, encourages curiosity, and allows time for reflection and exploration.
General objective: The program aims to develop productive thinking skills among female students at King Faisal University.
Specific Objectives: By the end of the program, students are expected to:
Understand and articulate the concept of productive thinking.
Identify AI applications relevant to productive thinking.
Show interest and engagement in productive thinking practices.
Use AI tools (e.g., ChatGPT, Notion AI) to generate new ideas and solve academic problems.
Create mind maps illustrating productive thinking processes.
Differentiate and apply various creative thinking skills, including fluency, originality, flexibility, interpretation, and inference.
Analyze AI-generated content in terms of fluency, originality, and logical interpretation.
Evaluate and compare AI- and human-generated outputs.
Design innovative projects using AI tools.
Demonstrate flexibility in adapting their thinking to evolving AI outputs.
Apply reasoning and inference using AI-assisted analysis.
Present AI-supported solutions to real-world problems.
Summarize and reflect on the types and stages of productive thinking.
Assess their own development in terms of productive thinking skills.
Demonstrate enthusiasm for applying these skills in new contexts that utilize AI technologies.
Instructional content: The program consisted of a series of structured sessions, each targeting a specific productive thinking skill:
Each instructional session was explicitly aligned with one or more dimensions of productive thinking. For example, the fluency session emphasized generating multiple responses to AI-generated problem scenarios, while the originality session focused on producing uncommon solutions evaluated against peer and AI-generated alternatives. Flexibility was addressed through tasks requiring students to reformulate solutions based on adaptive AI feedback. Interpretation and inference sessions emphasized meaning-making, pattern recognition, and drawing logical conclusions from AI-supported datasets and texts. Expansion and imagination were integrated through project-based tasks that required students to elaborate on initial ideas into feasible applications and envision alternative future scenarios. This structured alignment ensured systematic coverage of all seven productive thinking dimensions throughout the intervention.
The statistical analyses were conducted using the R programming language (version 4.3.2) and its relevant statistical packages. Descriptive statistics, including means and standard deviations, were computed to summarize participants’ responses across the study variables.
Although preliminary descriptive statistics suggested no substantial departures from symmetry, the very small sample size (n = 9) precluded reliable assessment of the normality assumption required for parametric procedures. Consequently, the Wilcoxon signed-rank test was selected because it provides a robust non-parametric alternative for paired observations without requiring the assumption of normally distributed difference scores. In addition to significance testing, effect sizes were calculated to determine the magnitude of the intervention’s impact. All analyses adhered to appropriate statistical assumptions and best practices for non-parametric data analysis.
The R packages used in the analysis included “psych” for computing reliability, “stats” for Wilcoxon tests, and “rstatix” for calculating effect sizes. The significance level was set at p < 0.05, and all statistical decisions were made based on two-tailed tests.
Effect sizes were calculated as r = Z/√N, following Rosenthal’s (1994) formula for rank-based tests, and interpreted using Cohen’s (1988) benchmarks (small ≈ 0.10, medium ≈ 0.30, large ≈ 0.50). Ninety-five percent confidence intervals for each dimension’s median pre–post difference (reported in Table 2) were computed using the exact distribution method implemented in the R stats and rstatix packages, which is appropriate for small-sample analyses and does not rely on large-sample asymptotic approximations.
Related Stories
AI News
New Zealand’s prime minister proposes banning children from using social media
21 minutes ago
AI News
KONGSBERG Launches Aegir Subsea Situational Awareness Sonars
54 minutes ago
AI News
Tonight's Operation Education: Local school district embracing Artificial Intelligence
54 minutes ago
AI News
AI Agents Become the API Economy’s Biggest New Customers
54 minutes ago
AI News
What can federal data collection tell policymakers and researchers about artificial intelligence in the U.S. labor market?
54 minutes ago
AI News
AUTOMA+ 2026 Highlights AI and Data Intelligence as Key Drivers of Clinical Trial Efficiency - healthcare
55 minutes ago
AI News
IBM Unveils Next Generation Dual
1 hour ago
AI News
Thought for the Week: The increasing rise of Artificial Intelligence and God
1 hour ago