The methodology that coerces what it claims to measure
Standard self-report measures of student engagement — the Student Engagement Instrument, the Student Engagement Questionnaire, and the Student Course Engagement Questionnaire — share a format. Students are presented with a statement and asked to indicate the strength of their agreement on a scale. Selecting nothing isn’t an option. The respondent must position themselves relative to each statement, even when none of the available positions matches their actual experience.
For neurotypical students this is mild epistemic friction. For autistic students it is a methodological coercion. Camouflage — what Laura Hull and colleagues established with the development and validation of the CAT-Q in 2019 — is the effortful adaptation to expectations and neurotypical norms. Forced-response scales reward camouflage. Selecting the closest-fitting agreement statement, even when the underlying experience is something the statement doesn’t name, is a small act of compliance. The compliance is then recorded as data.
When the measurement instrument requires camouflage to complete, the instrument is not catching what it claims to catch. It is catching the mask the autistic respondent had to put on to interact with the instrument at all. This is the problem Julie Bailey and Sara Baker at the Faculty of Education at Cambridge set out to address with the Measure of Engagement with Learning Activities — MELA — validated across three phases in their 2026 paper.
What Bailey & Baker validated — three phases, 213 participants, 18 items
Phase 1 tested face and content validity. Thirty-three autistic undergraduates and five student support experts were presented with a long list of 32 candidate items across four engagement constructs — behavioural, emotional, cognitive, and social. Items were rated for clarity, meaningfulness, and relevance. Seven autistic students then participated in semi-structured interviews about their learning experiences to ensure the items reflected what was actually salient. The interview transcripts surfaced which subconstructs students brought up unprompted — a different signal from what they rated highly on a survey.
Phase 2 tested convergent validity with 175 students (53 autistic). MELA item selection was compared against established measures of the underlying processes: the Adult ADHD Self-Report Scale for attentional difficulties, the Difficulties in Emotion Regulation Scale, the Meta-Cognitions Questionnaire, the Generalised Anxiety Disorder-7 scale. The hypothesised relationships replicated. Concentrating on MELA correlated with lower attentional difficulties; distracted correlated with higher ones. Frustrated correlated with emotion regulation difficulties. Overwhelmed and superficial correlated with metacognition scores. Isolated correlated with anxiety; collaborating correlated inversely.
The non-finding is worth naming. Enthusiastic was not associated with emotion regulation. Bailey and Baker suggest this likely reflects enthusiasm being driven more by the content of the learning activity itself than by the student’s regulation capacity. Engagement is a function of the interaction between student and context, not a property the student carries from one context to the next.
Phase 3 tested reliability with 38 participants returning the following term. For comparable learning contexts, no significant differences were found between test and retest scores across any of the 18 items. The measure is stable across time when the context is held constant. Average completion time: 15 to 30 seconds per learning activity, brief enough to capture the same student across multiple contexts in a single survey session.
The final 18-item measure is a menu, not a scale. Behavioural: focused, distracted, exhausted, concentrating. Emotional: enjoyed, frustrated, enthusiastic, satisfied, comfortable, bored. Cognitive: overwhelmed, purposeful, superficial, thorough. Social: misunderstood, collaborating, isolated, independent. Students select the items that reflect their experience in a specific learning context. The option to select nothing — to opt out of any item that doesn’t reflect the experience — is methodologically central. It is what counteracts camouflage.
The vocabulary autistic students actually use — and the scales that exclude it
The face validity data from Phase 1 produced a finding the conceptual antecedent paper had only sketched. Words like isolated, misunderstood, withdrawn, overwhelmed, exhausted, frustrated, purposeful, and superficial were rated as meaningful by over two-thirds of autistic undergraduates. Isolated was rated meaningful by 81.8% of respondents. Misunderstood by 75.8%. Overwhelmed by 78.8%. Anxious by 85.7%. Exhausted by 84.2%.
None of these words appear on the standard engagement scales Bailey and Baker examined in developing MELA. The Student Engagement Instrument (Appleton et al., 2006), the Student Engagement Questionnaire (Coates, 2011), and the Student Course Engagement Questionnaire (Handelsman et al., 2005) all use statements with agree/disagree scales — typically 30 to 150 items, intended to be used once. The constructs underlying those statements were derived from research with predominantly neurotypical student populations. The vocabulary of the apparatus measuring student engagement has been built around behaviours and dispositions that pattern-match neurotypical learner experience.
Autistic students describing their learning use a different vocabulary. The words above did not have to be invented to fit the measure. They surfaced because autistic students used them when asked what mattered. The apparatus catching student engagement has been operating with a vocabulary that excludes the words the population it claims to measure actually uses. When the apparatus is recalibrated to include those words, what counts as “engagement” includes phenomena standard measures simply could not detect.
This is the structural argument operating one layer up from the diagnostic apparatus. The corpus has been arguing that diagnostic categories like ADHD and autism catch mechanism profiles at threshold-crossing intensity against an externally imposed demand structure. The research apparatus that studies the populations diagnostic categories produce is itself built on the same demand structure. The vocabulary of standard engagement scales is calibrated for neurotypical learner experience. Asking autistic students to report their engagement through that vocabulary is asking them to translate their actual experience into a register the instrument can hear.
The translation process is camouflage by another name. The MELA design’s opt-out architecture — the ability not to select an item that doesn’t fit — refuses the translation requirement. The student reports in their own register, or doesn’t report at all.
The cognitive engagement that doesn't show up on standard surveys
A subtle finding from Phase 1 sits inside the paper that warrants drawing out. Cognitive engagement items — purposeful, superficial, reflecting, thorough, strategic, overwhelmed, pressured, aimless — scored lower on the quantitative face validity survey than behavioural and emotional items. Students rated cognitive items as less clear, less meaningful, and less relevant in the survey responses.
Yet in the seven follow-up interviews, cognitive engagement subconstructs were mentioned more often than behavioural or emotional ones. Purposeful was mentioned 19 times. Collaborating — a social construct — 16 times. Superficial 10 times. Misunderstood 8 times. By contrast, concentrating and effort were not mentioned at all by any of the seven interviewees describing what mattered to them about their learning. The behavioural items that scored highest on the survey were the ones students did not talk about when asked to describe their experience in depth.
The reading is that autistic students engage cognitively with their learning in ways standard surveys haven’t been catching. The depth, style, and quality of the reasoning matter to the student. The survey format flattens this into items the student can clearly recognise on a checklist but doesn’t necessarily think to surface as central. The interview format surfaces what is actually being engaged with. Standard engagement measures rely entirely on the survey format. The cognitive dimension of autistic student engagement has been systematically under-measured because the instrument’s format selects against it.
The apparatus pathologises by design
The architectural reading the MELA validation paper supports — and that Bailey and Baker name explicitly in their discussion of implications for research methods — is that the existing research apparatus measuring neurodivergent students has been pathologising them in the design of its own instruments. The forced-response scale produces compliance behaviour. The vocabulary excludes the words students actually use. The unit of analysis locates engagement in the student rather than in the student-context interaction. Each of these design choices reflects an assumption about what kind of learner the instrument is calibrated for.
The neurodivergent student then encounters the instrument, performs the camouflage the instrument coerces, has their compliance recorded as engagement data, and is subsequently studied as a population with poorer engagement outcomes. The apparatus produces the deficit it then records. The student is pathologised by the design choices made before the measurement event began.
This is the demand-structure-as-variable pattern the corpus has been tracing in workplaces, in NHS care pathways, in homelessness services, in criminal justice arcs — operating now inside research methodology itself. The substrate is consistent. The demand structure shapes what gets caught. Recalibrate the demand structure — different vocabulary, different response format, different unit of analysis — and the substrate is suddenly visible in a register the previous measurement could not reach.
Bailey and Baker frame this carefully as a methodological contribution. The broader implication they don’t quite spell out, but that the validation data demands: research findings on neurodivergent student engagement produced by standard measures are not findings about neurodivergent student engagement. They are findings about how neurodivergent students perform compliance with measurement instruments designed for someone else. The body of literature built on those findings rests on a methodological premise the MELA validation has now demonstrated is unsound.
What this requires from education research going forward is straightforward and demanding. Existing engagement-related findings about neurodivergent populations should be treated as findings about the measurement-instrument-interaction, not about the students. New work on neurodivergent learner experience needs instruments built with the populations the work claims to study. The apparatus needs to be rebuilt before it gets to keep claiming what it has been claiming.
Citations
Bailey, J. & Baker, S. T. (2026) — Neurodiversity and learning in context: Development and validation of a self-report measure of engagement with learning activities (MELA)
Bailey, J. & Baker, S. T. (2022) — Rethinking engagement with learning for neurodiverse students
Hull, L., Mandy, W., Lai, M.-C. et al. (2019) — Development and Validation of the Camouflaging Autistic Traits Questionnaire (CAT-Q)
Milton, D. E. M. (2012) — On the ontological status of autism: The ‘double empathy problem’
Bottema-Beutel, K., Kapp, S. K., Lester, J. N. et al. (2021) — Avoiding Ableist Language: Suggestions for Autism Researchers
