Counted, Not Understood: How Britain's Research Metrics Are Quietly Reshaping What Scientists Dare to Ask
There is a particular kind of frustration that settles over a scientist who has spent years pursuing a question nobody else is asking. It is not the frustration of failure — failure, at least, produces data. It is the frustration of being structurally discouraged from asking the question in the first place. Across British universities, that frustration is becoming harder to distinguish from ordinary professional life.
The Research Excellence Framework, or REF, was conceived as a mechanism for accountability — a means by which public investment in science could be directed toward institutions demonstrably producing valuable work. The logic was sound, and the intention was genuinely reformist. What it has produced, however, is a research culture in which the definition of "valuable" has contracted around what can be measured, submitted, and scored within a fixed cycle.
The Architecture of Caution
The REF operates on a roughly six-year cycle. During each cycle, universities submit selected outputs — journal articles, monographs, datasets — for assessment by expert panels. Institutions are ranked, and funding allocations are influenced accordingly. The incentive structure this creates is not subtle: departments under pressure to perform submit work that is legible to assessors, work that fits established disciplinary categories, and work that can be completed and published within the relevant window.
Long-range, speculative, or genuinely interdisciplinary research does not fit neatly into this architecture. A project that requires a decade of data collection before any publishable conclusions are possible is, from a departmental planning perspective, a liability. A researcher who spends three years developing a novel methodology that ultimately fails to yield positive results has, within the REF's accounting logic, produced nothing. The system does not reward intellectual courage; it rewards legible productivity.
Dr Priya Subramaniam, a computational biologist at the University of Edinburgh, describes the effect in terms that will be familiar to many of her colleagues. "There is a version of my research programme that I would pursue if I were thinking purely about the science," she says. "And there is a version I actually pursue, because I have to think about what I can show for myself in four years' time. Those two versions are not the same research programme."
What the Numbers Cannot See
The problem is not that metrics are inherently corrupting. Accountability frameworks serve a legitimate function, and the alternative — distributing research funding on the basis of reputation and personal networks — carries its own well-documented distortions. The difficulty is more specific: the metrics in current use are poorly calibrated to capture the kind of scientific work that tends, historically, to be most consequential.
Citation counts, for instance, reward work that is immediately legible to existing research communities. A paper that synthesises well-understood findings in a novel way will often be cited more rapidly than one that opens an entirely new line of inquiry. The latter may take years to accumulate citations, not because it is less valuable, but because its intellectual neighbourhood is, by definition, sparsely populated at the outset.
Journal impact factors present a related problem. High-impact journals have established editorial preferences, and those preferences reflect the current consensus about what constitutes important science. Research that challenges foundational assumptions, or that works at the margins of recognised disciplines, frequently struggles to find a home in prestigious venues — not because it is poor work, but because it does not conform to the genre expectations of journals optimised for a broad, citation-generating readership.
Professor Alistair Drummond, a historian of science at University College London, draws a pointed historical comparison. "If you apply REF-style metrics retrospectively to the careers of scientists we now regard as transformative, a significant number of them look like moderate performers for long stretches of their working lives," he observes. "Darwin spent years producing what looked, to his contemporaries, like natural history of limited scope. Continental drift was professionally dangerous territory for decades. The metrics would not have helped either case."
The Early-Career Trap
The consequences fall most heavily on those least able to absorb them. Early-career researchers — postdoctoral fellows, lecturers on probation, those working toward permanent positions — operate in a labour market of considerable precarity. Their career prospects depend substantially on publication records, and their publication records are assembled during precisely the period when intellectual risk-taking would, in a more forgiving system, be most appropriate.
The effect is a kind of professional sorting. Researchers who are temperamentally inclined toward speculative or long-horizon work either learn to suppress those inclinations or find themselves at a disadvantage relative to peers who have concentrated on producing a steady stream of competent, publishable work in established fields. Over time, the population of researchers advancing through British academic science skews toward the latter type — not because the system has selected for intelligence or creativity, but because it has selected for a particular relationship with risk.
Dr James Okafor, a climate scientist at the University of Leeds who has recently moved into a permanent post after several years on fixed-term contracts, is candid about the calculations involved. "I did not pursue the research questions I found most interesting during my postdoc years. I pursued research questions I thought I could finish. I am not proud of that, but it was entirely rational given my circumstances, and I suspect most of my peers made similar decisions."
Reforming the Instrument
The Research Excellence Framework is currently under periodic review, and there is no shortage of proposals for its revision. Some advocate for greater weighting of research process and infrastructure, rather than outputs alone. Others argue for the inclusion of explicit provisions for high-risk, high-reward research, following models piloted by funding bodies such as the Advanced Research and Invention Agency (ARIA), which was established in part to circumvent the conservatism of conventional grant assessment.
UKRI has, in recent years, made increasing reference to supporting "transformative" and "ambitious" research within its strategy documents. Whether this rhetorical shift is accompanied by structural change in how institutions are evaluated and funded remains, in the view of many working academics, an open question.
What is clear is that the problem cannot be resolved by exhortation alone. Telling researchers to take more intellectual risks while leaving in place the incentive structures that penalise risk is not a coherent policy position. If British science is to remain capable of genuine discovery — rather than the efficient production of incremental knowledge — the instruments by which it is measured will need to be rebuilt with that ambition in mind.
The audit, in other words, requires an audit of its own. And the question it needs to answer is not merely whether research is being produced, but whether the research being produced is worth the asking.