Skip to content
PaperFren

Research method

Effect Size

An effect size is a number that says how large a contrast is, not merely whether p dropped below 0.05 — for example Cohen’s d, partial η², or a near-zero programme-level estimate. Size is still not importance: a large pre/post η² without a control can be practice, a tiny exam d can be the policy-relevant finding, and a 'significant' poorer UDL rating is a perception shift, not a ranked educational value. Effect size also does not settle which engagement dimension moved.

Engagement and learning papers report effect sizes when they want to say how much a format, module or group differed. They answer 'how big was this contrast in these units?' Their main limitation is comparability: Partial η² = 0.35 on essays, a FLEX estimate indistinguishable from zero across 133 courses, and high self-reported simulation engagement are not the same kind of 'size'.

Evidence

What the evidence shows

Drawn from 7 studies in this library. Each finding starts with a plain-language takeaway, then the denser detail. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope with a short note on each study’s contribution. Challenged positions are labeled — they are not findings.

  • A Swiss FLEX Bachelor’s format cut face-to-face time by about 51% versus a part-time programme. Researchers compared 2,822 FLEX exams with 11,638 part-time exams in 133 courses (2015–2019; FLEX n students = 278, part-time = 1,068). The overall effect size was close to zero and not significantly different from zero; exam differences were significant in only 24 of 133 courses, sometimes favouring FLEX and sometimes part-time. Equivalence on exams is not proof of identical soft-skill or career outcomes.

    1 study
    1. 1Does half the classroom time hurt flexible learners?
  • Wageningen’s online peer-feedback module (330 enrolled, 284 completers across five domain courses) improved argumentative-essay scores from pre-test to post-test with Partial η² = 0.35 (Wilks’ λ = 0.65, F(7, 269) = 20.56). That is a large within-sample change without a no-feedback control, so it is not an isolated treatment effect of peer feedback versus practice.

    1 study
    1. 1Online peer feedback that improves argumentative essays
  • Among 1,731 of 19,752 invited U.S. research-university students, retrospective UDL items found Engagement and Action & Expression significantly poorer after the spring 2020 remote shift (153 respondents eligible for disability services; Representation improved for some of them). Significance here flags a perception contrast, not a Cohen’s d on learning and not a controlled redesign.

    1 study
    1. 1Did remote COVID teaching feel less engaging to students?
  • Forty intermediate Iranian EFL learners (20 text-based, 20 multimodal; fourteen sessions) split engagement: text-based stronger on autonomous, behavioural and cognitive dimensions; multimodal on emotional and social; satisfaction similar. A single 'engagement effect size' would have hidden that split. The window is fourteen sessions in one EFL context.

    1 study
    1. 1Text or multimodal CMC: which engages EFL writers?
  • Polish primary students (survey n = 333, ages about 8–14) reported widespread home Internet and more screen time in the first COVID wave. Experts rated a randomly selected 60-student subsample (18%) on five national informatics criteria (instrument α = 0.76). Ratings were generally low; the group with school plus out-of-school distance informatics scored better on several criteria, especially programming/troubleshooting. This is not a randomised pedagogy trial.

    1 study
    1. 1Polish primary informatics under COVID remote schooling
  • Seventy-three Finnish 16–17-year-olds ranked four fictitious sugar-health texts by credibility more easily than they justified those ranks: mean justification score 20.21 out of 48. Time on task was the best predictor of justification quality. Ranking success with weak justifications is a mismatch of outcome sizes, not a classroom engagement d.

    1 study
    1. 1How do teens judge if online health articles are trustworthy?

Open questions

Tensions and limits

Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes — limits on how far one study travels — not a forced fight between papers.

  • Scope / different questions

    Near-zero and large are both in this set because the contrasts differ. FLEX’s programme-level exam effect is ~0 despite 51% less classroom time (only 24/133 courses significant). Peer feedback’s Partial η² = 0.35 is a single-group pre/post on essays. Science-simulation undergraduates (n = 1,034; 536 male, 498 female) reported high engagement and satisfaction with no control-group d at all. You cannot rank those as competing estimates of 'whether blended learning works'.

    3 studies
    1. 1Does half the classroom time hurt flexible learners?
    2. 2Online peer feedback that improves argumentative essays
    3. 3Science simulations, engagement, and learning styles

    Study comparison

    StudyRoleDesignNPopulationOutcome
    Does half the classroom time hurt flexible learners?2023SupportsOtherComparison of FLEX blended vs part-time exam results across many coursesN=14460 · 2,822 FLEX and 11,638 part-time exams across 133 courses (2015–2019)Business Bachelor's students in FLEX blended vs part-time programmesExam performance equivalence despite ~51% less face-to-face time
    Online peer feedback that improves argumentative essays2023SupportsOtherPre–post evaluation of a structure-focused online peer-feedback essay moduleN=284 · 284 of 330 students across five courses completed the moduleBachelor and master students at Wageningen University across course domainsArgumentative essay quality before vs after guided peer feedback
    Science simulations, engagement, and learning styles2022SupportsCross-sectionalSurvey of engagement/satisfaction after semester-long science simulation teachingN=1034 · 1,034 undergraduates (536 male, 498 female)Undergraduates taught with simulations across science subjects at a gulf-country universityEngagement and satisfaction with simulation teaching by subject, gender, and learning style
  • Scope / different questions

    Statistical significance and educational size part company. COVID UDL ratings were significantly poorer on two principles among volunteers. FLEX found significant course-level exam differences in a minority of 133 courses while the pooled effect sat at zero. 'Significant' is not a shared importance ranking.

    2 studies
    1. 1Did remote COVID teaching feel less engaging to students?
    2. 2Does half the classroom time hurt flexible learners?

    Study comparison

    StudyRoleDesignNPopulationOutcome
    Did remote COVID teaching feel less engaging to students?2021SupportsCross-sectionalRetrospective pretest–posttest survey of spring 2020 remote instruction qualityN=1731 · 1,731 of 19,752 invited students at a U.S. R1 universityUniversity students experiencing the COVID shift to remote instructionPerceived UDL engagement/action/representation supports before vs after the shift
    Does half the classroom time hurt flexible learners?2023SupportsOtherComparison of FLEX blended vs part-time exam results across many coursesN=14460 · 2,822 FLEX and 11,638 part-time exams across 133 courses (2015–2019)Business Bachelor's students in FLEX blended vs part-time programmesExam performance equivalence despite ~51% less face-to-face time

Common misconceptions

Exam-style questions

Short-answer questions that ask you to explain or compare, not recall.

FLEX reduced on-site time by about 51% and compared 2,822 versus 11,638 exams in 133 courses. The overall effect size was near zero; 24 courses differed. What should a dean conclude about cutting classroom time?

In this Swiss business programme, exams were not systematically worse (or better) under FLEX, but a minority of courses did differ in both directions. That is not proof of identical soft skills or careers, and students were not randomly assigned to FLEX.

Why is Partial η² = 0.35 on Wageningen essays a different object from a between-format Cohen’s d?

It is a pre/post multivariate effect inside one module with no no-feedback control. A between-format d would compare peer feedback with an alternative. The η² can be large because almost everyone practised and revised.

EFL CMC splits engagement dimensions while satisfaction stays similar. How would reporting one engagement d mislead?

Text-based won autonomy/behavioural/cognitive; multimodal won emotional/social. A pooled engagement d could be near zero or pick a winner that does not exist on every subscale.

Students scored 20.21/48 on credibility justifications while ranking sources more easily. Which 'effect' is the instructional target if the goal is evaluating online health texts?

Justification quality, not ranking accuracy. Most of these 73 students could order the four texts yet gave weak or confused reasons (for example judging knowledge from design). Time on task predicted justification scores; ranking success would overstate critical evaluation.

The studies

7 studies in this library bear on Effect Size, ordered by citations.

  • Does half the classroom time hurt flexible learners?

    A flexible blended business programme with about half the on-site time showed essentially equivalent exam performance to part-time peers.

    International journal of educational technology in higher education · 2023 · 14 citations

  • Did remote COVID teaching feel less engaging to students?

    After a sudden shift online, students reported poorer engagement and action/expression supports, while some representation access improved.

    International journal of educational technology in higher education · 2021 · 9 citations

  • Text or multimodal CMC: which engages EFL writers?

    Text-based CMC boosted autonomy and cognitive/behavioral engagement, while multimodal CMC boosted emotional and social engagement.

    Education and information technologies · 2023 · 6 citations

  • Polish primary informatics under COVID remote schooling

    Polish primary students mostly had home ICT access during lockdown, but expert-rated informatics outcomes stayed low unless school distance lessons were paired with out-of-school courses.

    Education and information technologies · 2022 · 5 citations

  • Online peer feedback that improves argumentative essays

    A structure-focused online peer feedback module improved university students’ argumentative essay quality across courses and degree levels.

    Education and information technologies · 2023 · 4 citations

  • Science simulations, engagement, and learning styles

    Undergraduates reported high engagement and satisfaction with science simulations; engagement/satisfaction differed by subject and gender, and learning style predicted both.

    Education and information technologies · 2022 · 4 citations

  • How do teens judge if online health articles are trustworthy?

    High school students can easily identify which online health texts are untrustworthy, but they struggle to explain the reasons behind their judgments.

    Education and information technologies · 2022 · 3 citations

Learn alongside