
The importance of robust research evidence
- What constitutes robust research
- What do effect sizes mean?
- What do p-values mean?
- The limitations of average scores
- Why reading and spelling ages should not be used
What constitutes robust research evidence
Rigorous research studies that evaluate educational interventions should meet at least six of the eight indicators listed below to be considered methodologically sound or rigorous:
1. the use of a comparison control group(s) or treatment group(s);
2. the random assignment of participants to treatment or control groups
3. the establishment of baseline equivalence between groups;
4. the inclusion of sufficient information regarding the participant sample for study replication;
5. the use of standardized scores (age equivalent scores such as reading and spellings ages should not be used as they are unreliable);
6. enough information to calculate effect sizes;
7. the report of the fidelity of treatment; and
8. the use of blind evaluators.
Read More
Six or more is the threshold for rigour, however, some indicators like random assignment or blind evaluators might not be possible given the research design and are the most commonly missing indicators. In rare cases, another indicator may not apply but there must be a strong methodological rationale and the threshold of 6 must be met. Having a comparison/control group, using standardised scores and reporting effect sizes are considered essential for rigour.
No indicator should be omitted if it is possible within the research design. Whilst the use of blind evaluators is one of the two criteria most commonly omitted, it is particularly important to ensure there is no bias. For example, in my PhD research on spelling, the research assistants marking and cross marking the standardised tests did not know whether the tests papers they were marking were from control or experimental schools. The assessment of spelling in independent writing samples, and the evaluation of the quality of independent writing samples, was completed by experienced teachers and inter-rater reliability was established. These teachers did not know whether the samples of writing were from the control or experimental schools. To ensure the validity of the data, and protect me as the author of the intervention, this was a rigorously controlled study. I did not administer the tests, collect any of the data or mark or evaluate any of the tests and samples of writing. This was to ensure the research met the criteria of blind evaluation. The data was entered into SPSS by staff in the School of Education at the university.

What do Effect Sizes mean?
Effect sizes show the magnitude of the difference between groups. The effect size reflects the strength of the interventionโs impact. It shows how meaningful or substantial the effect is in practical terms.
0.05 very small (almost no difference between groups)
0.2 small effect (modest improvement )
0.5 medium effect ( a meaningful, clearly detectable difference between groups)
0.8 large effect (a large, strong difference between groups)
1.0+ very large effect ( a substantial and easily noticeable difference between groups)
What do p-values mean?
P-values provide a measure of the extent to which the result is due to chance.
| p-value | Meaning | Interpretation | Strength of evidence |
| p < .05 | Less than a 5% chance the result is random | Statistically significant | Moderate evidence |
| p < .01 | Less than a 1% chance the result is random | Highly significant | Strong evidence |
| p < .001 | Less than a 0.1% chance the result is random | Very highly significant | Very strong evidence |
| p < .0001 | Less than a 0.01% chance the result is random | Extremely significant | Extremely strong evidence |

The limitations of average scores. How average scores can mask low scores
Research evidence to support intervention programmes often focuses on group averages to make claims about effectiveness.
Where effect size and statistical significance are moderate or low (small or very small) then it is clear that a group of children are not progressing at an expected rate for their age, or may even be regressing.
The extent of difficulties experienced by children who underachieve is hidden in the average score.
Failure to highlight this group results in a lack of transparency and expectations that the intervention will benefit everyone when this is not the case.
Why reading ages and spelling ages should not be used
Reading and spelling ages (also known as age-equivalent scores) are considered unreliable because they are overly simplistic. They are simply the median raw score for a particular age. Standardized scores show how a child performs relative to their peers, rather than comparing them to an arbitrary average. Due to the psychometric problems associated with reading and spelling ages that seriously limit their reliability and validity, these scores should not be used. (Bracken, 1988; Reynolds, 1981).
Standardised scores are a more accurate representation of an childโs ability because they are based not only on the mean at a given age level but also on the distribution of scores. Reading ages and spelling ages are not ratio or interval scales of measurement. They should not be added, subtracted, or averaged. For this reason they should not be used for statistical analysis.
Standardized scores can be arithmetically compared and provide a measure of how a child performs relative to their peers, rather than comparing them to an arbitrary average.
Bracken, B.A. (1988). Ten psychometric reasons why similar tests produce dissimilar results. Journal of Psychology, 26, 155-166.
Reynolds, C.R. (1981). The fallacy of “two years below grade level for age” as a diagnostic for reading disorders. Journal of School Psychology, 19 (4), 350-358.
