Scale Opens Doors. Assessment Validity Walks Through Them

Scale Opens Doors. Assessment Validity Walks Through Them

Education systems have spent the past decade achieving something genuinely impressive. National testing programmes now run across hundreds of thousands of students in a single sitting, often spanning multiple delivery formats and testing windows within the same cycle. The ambition underpinning most of this expansion is sound: more test takers and more administrations produce a larger, richer dataset, and a larger dataset opens up more meaningful comparison across regions, year groups and demographics.

Realising that ambition fully, though, requires looking closely at what scale introduces alongside its benefits.

What Scale Reveals About Comparability

When an assessment scales, it rarely scales as a single, uniform instrument. It scales across a range of conditions: some students sit the test on paper, others on screen; some schools run the exam in one window, others a fortnight later; some regions have reliable broadband and devices, others do not. Each of these variations introduces what psychometricians call mode effects: differences in performance that stem not from what a student knows, but from how the test was delivered.

A recent study of New York’s transition to computer based testing for fifth and eighth graders, published in Kappan during 2025, found that inconsistent scores across paper and digital formats can affect data validity and create interpretive challenges for the decision makers who rely on those results for accountability and funding decisions. The researchers were not arguing against computer based testing. They were pointing out something instructive: when a system shifts delivery modes at scale, the resulting scores may look like a single continuous dataset while actually representing two different measurements stitched together.

Germany’s nationwide VERA assessment, which tests around 700,000 eighth graders annually across mathematics, German and foreign languages, has navigated the same challenge. A 2025 analysis of mode effects in this large scale assessment found that technology based versions of the test were often measurably harder than their paper equivalents, even when both were designed to assess the same underlying competencies. This is not a failure of programme design; it is a known psychometric phenomenon that becomes more visible, and more worth addressing, as programmes grow.

The volume of data an institution collects says nothing about whether that data is measuring the same thing across every student who contributes to it. A national exam producing ten million data points spread across three delivery formats and two testing windows is not automatically more informative than one producing a hundred thousand data points from a single, controlled administration. Recognising that distinction is what allows assessment professionals to get the most out of the scale they have worked hard to build.

Turning a Design Challenge Into a Structural Advantage

For institutions running mixed delivery exams across large and varied populations, comparability is first and foremost a design opportunity. Maintaining equivalence across paper and digital cohorts, different testing windows and varying levels of device access is entirely achievable, and the profession already has well developed tools for doing it.

The institutions managing this well share a common approach: they treat delivery mode as a variable to be actively controlled rather than a background condition to be accepted. That means building item banks that account for mode sensitivity at the authoring stage, running parallel pilot programmes that generate mode effect data before full rollout, and applying equating procedures that allow results from different delivery conditions to be placed on a common scale with confidence.

An enterprise level platform built to manage this complexity from the outset gives assessment teams the operational leverage to do exactly that. When comparability is engineered into the infrastructure rather than added later, institutions can expand their reach without compromising the validity of what they are measuring.

What the Score Actually Represents

There is a second consideration layered on top of mode effects, and it has moved quickly from an emerging question to a practical one. As generative AI tools have become integrated into browsers and operating systems, the boundary between what a student knows and what a student can retrieve in real time has become more important to establish, particularly in unsupervised or lightly supervised online assessments.

A 2025 study examining a postgraduate data science course offers a useful data point. Students who scored an average of around 92 out of 100 on take home assessments where AI use was permitted scored an average of around 63 on proctored exams where it was not: a gap of nearly 30 percentage points with a large effect size. The researchers described this as an AI inflation effect, a measurable difference between what an assessment records and what a student can demonstrate unaided. It is not an indictment of any institution; it is a signal that assessment design needs to keep pace with the tools students now have access to.

The implication for scale is worth sitting with. An institution can run an assessment across a million students and produce a complete, well structured dataset, while the construct the test was designed to measure has gradually shifted. Completeness of data and validity of data are not the same property. Recognising that distinction early is what gives assessment professionals the clearest path to acting on it.

Assessment Design as the Decisive Variable

Understanding where AI inflation is most likely to emerge gives educators a clear and actionable basis for programme review. The conditions that tend to concentrate risk are identifiable: unsupervised take home tasks weighted heavily in the final grade, open ended written or coding assignments submitted without proctoring, and formats that predate the wide availability of generative AI tools. Mapping which assessments sit in that category is a straightforward diagnostic starting point.

From there, the practical options are well established. Supervised components can be reweighted to carry more of the grade. Task design can shift toward approaches that are harder to complete by retrieval alone: applied scenarios, constrained problems, or responses that require students to explain their reasoning step by step. Where take home tasks remain appropriate, they can be complemented by shorter proctored checks that allow educators to verify whether performance is consistent with what the student produces unaided.

None of this requires dismantling existing assessment programmes. It requires reading them with fresh eyes and making targeted adjustments where the evidence suggests the current format would benefit from revision.

Building Assessment Infrastructure That Keeps Pace

The institutions getting the most out of scale are the ones that treat expansion as an opportunity to sharpen what their assessments are actually measuring, alongside how many people they can measure it for.

That process is more tractable than it has ever been. Modern assessment platforms generate item level performance data that makes it possible to identify which tasks show mode sensitivity, which show score distributions consistent with AI inflation, and which are holding up well under changed conditions. Educators who engage actively with that data are well positioned to make targeted, evidence based improvements rather than relying on intuition alone.

Large scale testing programmes, properly designed and actively managed, remain one of the most powerful tools education systems have for comparing performance across regions, tracking progress over time, and identifying where additional support is needed. The profession has built something valuable. The next step is making sure the infrastructure around it is doing full justice to that ambition.