Research method
Pre-Test/Post-Test Design
A pre-test/post-test design measures the same outcome before and after an instructional episode. The change score (or a covariance-adjusted post-test) is the claim. Designs in this set range from single-group modules with no control, through class-assigned contrasts, to randomised sessions. A rise from time 1 to time 2 is not, by itself, a randomised treatment effect: practice, maturation, concurrent teaching and who sits both tests can all move the score.
Education papers use pre/post when they want to show that essays, fraction tests, mentoring ratings, faculty attitudes or course grades moved across a window. They answer 'did the measured score change?' Their main limitation is the missing counterfactual: without a comparable no-treatment group, the post-test cannot isolate the module from everything else that happened in the course.
Evidence
What the evidence shows
Drawn from 10 studies in this library. Each finding starts with a plain-language takeaway, then the denser detail. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope with a short note on each study’s contribution. Challenged positions are labeled — they are not findings.
At Wageningen, 330 students from five domain courses joined an online supported peer-feedback module and 284 completed it. Argumentative-essay performance improved from pre-test to post-test with a large effect (Wilks’ λ = 0.65, F(7, 269) = 20.56, Partial η² = 0.35), including recognised argument elements, across bachelor and master levels. There was no no-feedback control, so practice and instructor effects are not isolated.
After exclusions, 66 Taiwanese sixth-graders (35 experimental, 31 control; from 89 enrolled) were randomly assigned to a dialogue ITS or traditional remedial teaching for fraction multiplication and division. Tests had 24 items (96 points); pretest reliability was high (α = 0.879). The experimental group outperformed controls on the post-test after adjusting for prior scores (F = 5.52). The session was short remedial tutoring, not a semester-long curriculum rewrite.
At a U.S. very-high-research-activity university, 1,731 of 19,752 invited students completed a voluntary survey with retrospective pretest–posttest items on Universal Design for Learning after the spring 2020 shift online (153 respondents were eligible for disability services). Students reported that UDL Engagement and Action & Expression were significantly poorer after moving online; Representation improved for some, including disability-eligible students. These are recalled perceptions of instructional quality, not objective learning gains.
A single-group faculty programme at Universidad Francisco de Vitoria enrolled 695 participants (mean age 45.46, SD 9.74). Attitudes toward innovation and the LMS improved, and professors reported greater knowledge of how to implement innovative teaching elements, with broadly similar patterns for full-time and part-time faculty. Without a control group, maturation and concurrent campus change remain competing explanations; student learning was not measured.
Classrooms were randomly assigned so that 505 biology students in 15 rooms at 13 universities received start- and end-of-term mentoring surveys; intervention rooms got emails highlighting three personal similarities with the instructor. There was no general effect on perceived similarity or mentoring quality; students who began feeling very different from the instructor did increase perceived similarity. Intervention classrooms were larger at baseline.
Sixty Malaysian second-year students in two intact classes of 30 were assigned by class to chatbot-supported project-based learning or a traditional control sharing content, duration and instructor. The chatbot class scored higher on learning performance (µ = 42.500 vs 39.933; d = 0.978). Need for cognition, motivational beliefs, creative self-efficacy and perception of learning did not differ. Class-level assignment is not individual randomisation.
Open questions
Tensions and limits
Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes — limits on how far one study travels — not a forced fight between papers.
Pre/post windows are not interchangeable designs. Peer feedback and faculty training are single-group rises. The ITS paper randomises after exclusions and adjusts for prior scores. Chatbots contrast two intact classes. Mentoring randomises classrooms and finds a null average with a subgroup gain. COVID UDL is a retrospective perception contrast, not a scored artefact. Calling all of them 'pre-test/post-test evidence that teaching improved' hides those differences.
- Online peer feedback that improves argumentative essays
- Faculty training that shifts teaching innovation attitudes
- Can a dialogue ITS fix fraction multiplication?
- Do educational chatbots raise project-course grades?
- Can showing shared interests improve student-teacher mentoring?
- Did remote COVID teaching feel less engaging to students?
Study Role Design N Population Outcome Online peer feedback that improves argumentative essays Supports OtherPre–post evaluation of a structure-focused online peer-feedback essay module N=284 · 284 of 330 students across five courses completed the module Bachelor and master students at Wageningen University across course domains Argumentative essay quality before vs after guided peer feedback Faculty training that shifts teaching innovation attitudes Supports OtherPre–post faculty development program evaluating attitudes toward teaching innovation N=695 · 695 university professors participated University faculty (full-time and part-time) in a faculty training program Attitudes toward transformative/innovative teaching after the development experience Can a dialogue ITS fix fraction multiplication? Supports OtherQuasi-experiment of dialogue-based fraction ITS vs traditional remedial instruction N=66 · 66 sixth graders analysed (35 ITS, 31 control) after exclusions from 89 enrolled Sixth-grade students in central Taiwan learning fraction operations Fraction multiplication/division achievement with ITS vs traditional remediation Do educational chatbots raise project-course grades? Supports OtherQuasi-experimental intact-class comparison of chatbot-supported vs conventional project-based learning N=60 · Two classes of 30 students each (chatbot vs conventional) Malaysian students in a team design course Achievement (and secondary motivation/creativity measures) with educational chatbot support Can showing shared interests improve student-teacher mentoring? Supports OtherCluster-randomized classroom trial emailing shared instructor–student similarities N=505 · 505 biology students in 15 classrooms across 13 US universities College biology students in semester-long research courses Perceived similarity and mentoring relationship quality, especially for initially dissimilar students Did remote COVID teaching feel less engaging to students? Supports Cross-sectionalRetrospective pretest–posttest survey of spring 2020 remote instruction quality N=1731 · 1,731 of 19,752 invited students at a U.S. R1 university University students experiencing the COVID shift to remote instruction Perceived UDL engagement/action/representation supports before vs after the shift What moves is not always engagement. Chatbot students gained on course performance (d ≈ 0.98) without significant affective/motivation differences. Forty Iranian EFL learners randomised to text-based versus multimodal CMC for fourteen sessions split engagement by dimension (text-based stronger on autonomous, behavioural and cognitive; multimodal on emotional and social) with similar satisfaction. A pre/post 'engagement improved' headline would mis-state both papers.
- Do educational chatbots raise project-course grades?
- Text or multimodal CMC: which engages EFL writers?
Study Role Design N Population Outcome Do educational chatbots raise project-course grades? Supports OtherQuasi-experimental intact-class comparison of chatbot-supported vs conventional project-based learning N=60 · Two classes of 30 students each (chatbot vs conventional) Malaysian students in a team design course Achievement (and secondary motivation/creativity measures) with educational chatbot support Text or multimodal CMC: which engages EFL writers? Supports Human experimentRandom assignment to text-based vs multimodal CMC for 14 treatment sessions N=40 · 40 intermediate Iranian EFL learners (20 per condition) Intermediate EFL learners in CMC writing groups Autonomy, multi-dimensional engagement, satisfaction, and writing performance by modality
Common misconceptions
If the post-test is higher, the intervention caused the gain.
Wageningen’s Partial η² = 0.35 and the faculty attitude rise are single-group changes. Practice, other coursework, self-selection and campus events are still in the score. A control or randomised contrast is what would isolate the module.
A retrospective pretest is the same measurement as a test taken before the change.
The COVID UDL items asked 1,731 volunteers to recall quality before versus after remote shift. Recall and who chose to answer can bias the contrast; the study does not score objective learning from a controlled redesign.
A large achievement d means the tool also raised engagement.
The chatbot class’s d = 0.978 is on learning performance. Motivational beliefs, need for cognition, creative self-efficacy and perception of learning did not differ from the traditional class.
Exam-style questions
Short-answer questions that ask you to explain or compare, not recall.
Wageningen reports Partial η² = 0.35 from pre-test to post-test on argumentative essays for 284 completers. What additional design piece would you need before attributing that gain to peer feedback rather than practice?
A no-feedback or alternative-feedback control group measured on the same essay elements. Without it, revision practice, instructor effects and other course activities remain in the change score.
The ITS paper randomises 66 sixth-graders after exclusions and adjusts the post-test for prior scores (F = 5.52). Why is that still not evidence to replace a year-long fraction curriculum?
The contrast is a short remedial session in one regional cohort. Durable classroom transfer was not tested, and traditional-control quality may vary. Covariance adjustment does not turn a brief sitting into a curriculum trial.
How does a retrospective UDL pretest–posttest on 1,731 of 19,752 invited students differ from the mentoring study’s start- and end-of-term surveys in 15 randomised classrooms?
UDL items reconstruct a campus-wide emergency shift among volunteers and measure perceived instructional quality. Mentoring surveys bookend an assigned email intervention (505 students) and still found no general mentoring-quality effect, with larger intervention classes at baseline. One is recalled pandemic perception; the other is a classroom-randomised contrast with a null average.
Two classes of 30 share an instructor; the chatbot class outperforms on achievement (d = 0.978) but not on motivation measures. What causal limit remains, and what outcome claim is already unsupported?
Assignment was by intact class, so class differences confound the chatbot. The nulls already block treating the chatbot as a general engagement or motivation booster; the supported claim is a short-window performance difference in one design course.
The studies
10 studies in this library bear on Pre-Test/Post-Test Design, ordered by citations. The first 8 are shown.
- Do educational chatbots raise project-course grades?
In a Malaysian team design course, students using an educational chatbot outperformed a traditional class on learning achievement, without clear gains on motivation or creative self-efficacy.
- Did remote COVID teaching feel less engaging to students?
After a sudden shift online, students reported poorer engagement and action/expression supports, while some representation access improved.
- Text or multimodal CMC: which engages EFL writers?
Text-based CMC boosted autonomy and cognitive/behavioral engagement, while multimodal CMC boosted emotional and social engagement.
- Does fantasy keep online gamified classes engaging?
Adding a fantasy narrative to a gamified online course helped sustain engagement after novelty wore off.
- Online peer feedback that improves argumentative essays
A structure-focused online peer feedback module improved university students’ argumentative essay quality across courses and degree levels.
- Can AR projects build teachers' digital competence?
Creating collaborative AR language projects improved pre-service teachers' awareness that pedagogy-plus-technology skills need more training.
- Can a dialogue ITS fix fraction multiplication?
Sixth graders using a dialogue-based math ITS outperformed peers who received traditional remedial instruction on fraction operations.
- Can showing shared interests improve student-teacher mentoring?
Sharing common interests only helps students who initially feel very different from their instructors feel closer to them, which indirectly improves their mentoring relationship.
Show 2 more studiesShow fewer studies
- Faculty training that shifts teaching innovation attitudes
A university faculty development program improved professors’ attitudes toward teaching innovation and LMS use, plus knowledge of innovative teaching elements.
- How did math PD leaders adapt during COVID?
Mathematics coordinators kept using effective PD feature clusters when shifting school-team PD online during COVID-19.
Learn alongside