Research method
Effect Size
An effect size is a number that says how large a contrast is, not merely whether p dropped below 0.05 — for example Cohen’s d, partial η², or a near-zero programme-level estimate. Size is still not importance: a large pre/post η² without a control can be practice, a tiny exam d can be the policy-relevant finding, and a 'significant' poorer UDL rating is a perception shift, not a ranked educational value. Effect size also does not settle which engagement dimension moved.
Engagement and learning papers report effect sizes when they want to say how much a format, module or group differed. They answer 'how big was this contrast in these units?' Their main limitation is comparability: Partial η² = 0.35 on essays, a FLEX estimate indistinguishable from zero across 133 courses, and high self-reported simulation engagement are not the same kind of 'size'.
Evidence
What the evidence shows
Drawn from 7 studies in this library. Each finding starts with a plain-language takeaway, then the denser detail. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope with a short note on each study’s contribution. Challenged positions are labeled — they are not findings.
A Swiss FLEX Bachelor’s format cut face-to-face time by about 51% versus a part-time programme. Researchers compared 2,822 FLEX exams with 11,638 part-time exams in 133 courses (2015–2019; FLEX n students = 278, part-time = 1,068). The overall effect size was close to zero and not significantly different from zero; exam differences were significant in only 24 of 133 courses, sometimes favouring FLEX and sometimes part-time. Equivalence on exams is not proof of identical soft-skill or career outcomes.
Wageningen’s online peer-feedback module (330 enrolled, 284 completers across five domain courses) improved argumentative-essay scores from pre-test to post-test with Partial η² = 0.35 (Wilks’ λ = 0.65, F(7, 269) = 20.56). That is a large within-sample change without a no-feedback control, so it is not an isolated treatment effect of peer feedback versus practice.
Among 1,731 of 19,752 invited U.S. research-university students, retrospective UDL items found Engagement and Action & Expression significantly poorer after the spring 2020 remote shift (153 respondents eligible for disability services; Representation improved for some of them). Significance here flags a perception contrast, not a Cohen’s d on learning and not a controlled redesign.
Forty intermediate Iranian EFL learners (20 text-based, 20 multimodal; fourteen sessions) split engagement: text-based stronger on autonomous, behavioural and cognitive dimensions; multimodal on emotional and social; satisfaction similar. A single 'engagement effect size' would have hidden that split. The window is fourteen sessions in one EFL context.
Polish primary students (survey n = 333, ages about 8–14) reported widespread home Internet and more screen time in the first COVID wave. Experts rated a randomly selected 60-student subsample (18%) on five national informatics criteria (instrument α = 0.76). Ratings were generally low; the group with school plus out-of-school distance informatics scored better on several criteria, especially programming/troubleshooting. This is not a randomised pedagogy trial.
Seventy-three Finnish 16–17-year-olds ranked four fictitious sugar-health texts by credibility more easily than they justified those ranks: mean justification score 20.21 out of 48. Time on task was the best predictor of justification quality. Ranking success with weak justifications is a mismatch of outcome sizes, not a classroom engagement d.
Open questions
Tensions and limits
Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes — limits on how far one study travels — not a forced fight between papers.
Near-zero and large are both in this set because the contrasts differ. FLEX’s programme-level exam effect is ~0 despite 51% less classroom time (only 24/133 courses significant). Peer feedback’s Partial η² = 0.35 is a single-group pre/post on essays. Science-simulation undergraduates (n = 1,034; 536 male, 498 female) reported high engagement and satisfaction with no control-group d at all. You cannot rank those as competing estimates of 'whether blended learning works'.
- Does half the classroom time hurt flexible learners?
- Online peer feedback that improves argumentative essays
- Science simulations, engagement, and learning styles
Study Role Design N Population Outcome Does half the classroom time hurt flexible learners? Supports OtherComparison of FLEX blended vs part-time exam results across many courses N=14460 · 2,822 FLEX and 11,638 part-time exams across 133 courses (2015–2019) Business Bachelor's students in FLEX blended vs part-time programmes Exam performance equivalence despite ~51% less face-to-face time Online peer feedback that improves argumentative essays Supports OtherPre–post evaluation of a structure-focused online peer-feedback essay module N=284 · 284 of 330 students across five courses completed the module Bachelor and master students at Wageningen University across course domains Argumentative essay quality before vs after guided peer feedback Science simulations, engagement, and learning styles Supports Cross-sectionalSurvey of engagement/satisfaction after semester-long science simulation teaching N=1034 · 1,034 undergraduates (536 male, 498 female) Undergraduates taught with simulations across science subjects at a gulf-country university Engagement and satisfaction with simulation teaching by subject, gender, and learning style Statistical significance and educational size part company. COVID UDL ratings were significantly poorer on two principles among volunteers. FLEX found significant course-level exam differences in a minority of 133 courses while the pooled effect sat at zero. 'Significant' is not a shared importance ranking.
- Did remote COVID teaching feel less engaging to students?
- Does half the classroom time hurt flexible learners?
Study Role Design N Population Outcome Did remote COVID teaching feel less engaging to students? Supports Cross-sectionalRetrospective pretest–posttest survey of spring 2020 remote instruction quality N=1731 · 1,731 of 19,752 invited students at a U.S. R1 university University students experiencing the COVID shift to remote instruction Perceived UDL engagement/action/representation supports before vs after the shift Does half the classroom time hurt flexible learners? Supports OtherComparison of FLEX blended vs part-time exam results across many courses N=14460 · 2,822 FLEX and 11,638 part-time exams across 133 courses (2015–2019) Business Bachelor's students in FLEX blended vs part-time programmes Exam performance equivalence despite ~51% less face-to-face time
Common misconceptions
A larger effect size means the finding matters more for policy.
FLEX’s near-zero exam effect after cutting classroom time by half is the policy-relevant magnitude: lots of time moved, little exam shift. A large uncontrolled η² on essays is a smaller warrant for requiring peer-feedback modules everywhere.
If p is significant, the effect was large.
UDL Engagement ratings were significantly poorer after remote shift among 1,731 volunteers; that is a detectable perception change, not a demonstrated large learning loss. FLEX’s few significant course contrasts sit inside an overall null.
High mean engagement on a survey is an effect size for simulations versus other teaching.
The 1,034-student simulation paper is a cross-sectional self-report after one semester in one teacher context. Learning style predicted engagement and satisfaction; there is no control-group contrast to convert into d.
Exam-style questions
Short-answer questions that ask you to explain or compare, not recall.
FLEX reduced on-site time by about 51% and compared 2,822 versus 11,638 exams in 133 courses. The overall effect size was near zero; 24 courses differed. What should a dean conclude about cutting classroom time?
In this Swiss business programme, exams were not systematically worse (or better) under FLEX, but a minority of courses did differ in both directions. That is not proof of identical soft skills or careers, and students were not randomly assigned to FLEX.
Why is Partial η² = 0.35 on Wageningen essays a different object from a between-format Cohen’s d?
It is a pre/post multivariate effect inside one module with no no-feedback control. A between-format d would compare peer feedback with an alternative. The η² can be large because almost everyone practised and revised.
EFL CMC splits engagement dimensions while satisfaction stays similar. How would reporting one engagement d mislead?
Text-based won autonomy/behavioural/cognitive; multimodal won emotional/social. A pooled engagement d could be near zero or pick a winner that does not exist on every subscale.
Students scored 20.21/48 on credibility justifications while ranking sources more easily. Which 'effect' is the instructional target if the goal is evaluating online health texts?
Justification quality, not ranking accuracy. Most of these 73 students could order the four texts yet gave weak or confused reasons (for example judging knowledge from design). Time on task predicted justification scores; ranking success would overstate critical evaluation.
The studies
7 studies in this library bear on Effect Size, ordered by citations.
- Does half the classroom time hurt flexible learners?
A flexible blended business programme with about half the on-site time showed essentially equivalent exam performance to part-time peers.
- Did remote COVID teaching feel less engaging to students?
After a sudden shift online, students reported poorer engagement and action/expression supports, while some representation access improved.
- Text or multimodal CMC: which engages EFL writers?
Text-based CMC boosted autonomy and cognitive/behavioral engagement, while multimodal CMC boosted emotional and social engagement.
- Polish primary informatics under COVID remote schooling
Polish primary students mostly had home ICT access during lockdown, but expert-rated informatics outcomes stayed low unless school distance lessons were paired with out-of-school courses.
- Online peer feedback that improves argumentative essays
A structure-focused online peer feedback module improved university students’ argumentative essay quality across courses and degree levels.
- Science simulations, engagement, and learning styles
Undergraduates reported high engagement and satisfaction with science simulations; engagement/satisfaction differed by subject and gender, and learning style predicted both.
- How do teens judge if online health articles are trustworthy?
High school students can easily identify which online health texts are untrustworthy, but they struggle to explain the reasons behind their judgments.
Learn alongside