Research Article

Disaggregating Speaking Gains: A Three-Arm Randomised Longitudinal Trial of AI-Only, Teacher-Led and Blended Speaking Instruction with Iraqi University EFL Learners

Authors

  • Muhannd Ali ABDULLAH Head of International Relations Section, Planning and Follow up- up Department, Scientific Research Commission Baghdad, Iraq,

Abstract

Artificial-intelligence (AI) applications for second-language speaking practice are rapidly expanding in higher education, including Iraq and the Kurdistan Region, where local evidence is dominated by perception surveys and small single-group designs. Although meta-analyses of dialogue-based computer-assisted language learning report moderate positive effects on L2 speaking, three instructional questions remain unresolved: whether AI practice produces broad speaking development or gains concentrated in temporal fluency; whether gains persist after practice ceases; and whether AI-supported instruction differs from time-equated teacher-led instruction rather than comparison conditions with unequal speaking time. This registered report proposes a three-arm, stratified cluster-randomised, longitudinal mixed-methods trial comparing AI-only, teacher-led, and blended AI–teacher speaking instruction under equivalent instructional time. Speaking outcomes are disaggregated into utterance fluency, pronunciation, grammatical accuracy, syntactic complexity, lexical diversity, and interactional competence, with a delayed post-test ten weeks after the intervention. Approximately 260 first- and second-year undergraduate EFL learners at CEFR A2–B1 level, nested within fifteen intact class clusters across three universities in Iraq and the Kurdistan Region, will be allocated to three arms. All arms will receive identical task content and approximately three hours of instruction weekly for fourteen weeks. A four-task speaking battery will be administered at pre-test, mid-point, immediate post-test, and delayed post-test. Temporal fluency will be extracted acoustically; pronunciation, comprehensibility, and interactional competence will be double-rated by trained raters with many-facet Rasch adjustment; and accuracy, complexity, and lexical measures will be derived from verified transcripts. Speaking anxiety, self-efficacy, willingness to communicate, self-regulated learning, and engagement will be measured at each occasion. Word error rate will be computed for each learner and modelled as a moderator of instructional effects. Linear mixed-effects models with learner and cluster random effects will estimate condition-by-time effects. ANCOVA will provide a confirmatory check, multilevel path-analytic mediation will use bootstrap intervals, and pre-specified false-discovery-rate control will be applied across a restricted primary outcome set. The design will provide a causal, sub-skill-level account of what each instructional mode produces, the first durability evidence for AI-mediated speaking gains following cessation, and evidence on whether automated-feedback validity functions as an individual-difference variable with instructional consequences.

Article information

Journal

International Journal of Linguistics, Literature and Translation

Volume (Issue)

9 (9)

Pages

07-25

Published

2026-08-28

How to Cite

ABDULLAH, M. A. (2026). Disaggregating Speaking Gains: A Three-Arm Randomised Longitudinal Trial of AI-Only, Teacher-Led and Blended Speaking Instruction with Iraqi University EFL Learners. International Journal of Linguistics, Literature and Translation, 9(9), 07-25. https://doi.org/10.32996/ijllt.2026.9.9.2

Downloads

Views

10

Downloads

4

Keywords:

artificial intelligence; second-language speaking; utterance fluency; blended instruction; retention; automatic speech recognition; registered report; Iraq.