Skip to content
PaperFren

Does practising with ChatGPT improve EFL students' writing?

Open paper intelligence

Students who practised academic writing with ChatGPT feedback scored clearly higher on a later writing test and reported more writing motivation than students taught with teacher feedback alone.

Source

Enhancing academic writing skills and motivation: assessing the efficacy of ChatGPT in AI-assisted language learning for EFL students

Song C, Song Y · Frontiers in psychology · 2023

doi.org/10.3389/fpsyg.2023.1260843Read the full paper ↗52 citationscc by

Study at a glance

Design
RCT — Students randomly assigned to 12 weeks of ChatGPT-assisted writing practice or conventional teacher-feedback instruction; pre/post IELTS writing tasks and motivation scale, plus interviews.
N
N=50 · 50 Chinese EFL undergraduates split between experimental and control groups (per-group sizes not stated); nine experimental students interviewed.
Population
Intermediate to upper-intermediate EFL undergraduates at one Chinese national university
Outcome
IELTS-rubric academic writing scores (overall, content, organisation, language use) and a seven-item writing motivation scale

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

Fifty Chinese undergraduates learning English were randomly split into two groups taught by the same teacher with the same materials for 12 weeks. One group wrote with ChatGPT giving real-time feedback on grammar, vocabulary and organisation; the other got conventional in-class teacher feedback. Both groups wrote IELTS academic tasks before and after, scored by two raters, and filled in a writing-motivation questionnaire; nine ChatGPT users were also interviewed.

What they found

After adjusting for pre-test scores, the ChatGPT group had higher overall writing scores (post-test means 59.12 versus 45.18, Cohen's d of 0.76) and higher scores on content, organisation and language use. Writing motivation also rose more in the ChatGPT group (d = 0.52). Interviewees described the tool as a personal tutor but worried about over-reliance, suggestions that did not fit their style, and whether gains would last without it.

The limits

What it doesn't show

The sample was small and from one university, and group sizes are not reported. One of the two raters was the second author, who also taught the course, so scoring was not blind and expectations could have leaked into grades. The groups differed in more than ChatGPT (e.g. anytime access and a web platform versus in-class-only feedback), and the authors acknowledge possible contamination between groups. There was no delayed follow-up, so it cannot show whether students write better once the tool is removed.

Key terms

EFL
English as a Foreign Language: learning English in a country where it is not the main community language.
ANCOVA
Analysis of covariance: compares group outcomes while statistically adjusting for a baseline score such as a pre-test.
Cohen's d
A standardised effect size: the difference between group means divided by their pooled standard deviation.
Zone of Proximal Development
Vygotsky's idea of the gap between what a learner can do alone and what they can do with support from a more capable partner.
Inter-rater reliability
How closely two independent scorers agree when grading the same work.

Flashcards

1 / 11

0 of 11 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 5

What was the control group's writing instruction?

Common questions

Does this prove ChatGPT makes students better writers?

Not on its own. It is one small randomised study where the ChatGPT group also got more flexible access and a non-blind rater helped score essays, so the size of the benefit could be overstated.

Were students just letting ChatGPT write for them?

The researchers coached students to write original text and use the AI only for feedback, and the post-tests were writing tasks scored by raters, but the paper does not describe how test conditions prevented AI use.

Did students have any complaints?

Yes. Some said suggestions were not always accurate for their context and that they became too dependent on the tool, and several doubted the improvement would last without it.

More on Feedback