How Science Works: Method, Peer Review, and Publication
Research reaches the public as headlines, but it is produced by a long chain of design, review, funding and measurement. Understanding that chain is the most durable way to judge a scientific claim.
Why the process matters more than the finding
Most people meet science in its finished form: a headline, a press release, a confident sentence on a news channel. That form hides almost everything that determines whether the claim deserves belief. A finding is the last link in a long chain that begins with someone deciding which question is worth asking and ends with an editor deciding which manuscript is worth printing. Every link in that chain can strengthen a result or quietly weaken it, and none of them are visible in the headline itself.
The useful habit, then, is not to memorise conclusions but to ask how a conclusion was produced. Who measured what, under which conditions, compared against what? Was the comparison fair? Could the same result have appeared by chance, or through some path the researchers did not consider? These questions do not require technical training. They require only the recognition that scientific knowledge is a process with known failure modes, and that those failure modes are reasonably easy to learn.
The method is a discipline, not a recipe
School textbooks present the scientific method as a numbered sequence: observe, hypothesise, experiment, conclude. Working science rarely proceeds so tidily. A geologist cannot run an experiment on a mountain range. An epidemiologist cannot randomly assign people to breathe polluted air. Astronomers observe but cannot intervene at all. What these fields share with laboratory science is not a procedure but a commitment: claims must be exposed to the possibility of being shown wrong, and the exposure must be genuine rather than ceremonial.
That commitment shows up in specific practices. A researcher states in advance what would count as evidence against the idea being tested. Measurements are made so that someone else could repeat them. Alternative explanations are named and, where possible, ruled out. Controls carry much of the weight here, because a comparison group separates the thing being studied from everything else happening at the same time. Without one, improvement over weeks might reflect a treatment, the natural course of an illness, or the season.
From an interesting idea to a testable claim
A hypothesis becomes scientifically useful only when it is specific enough to fail. The claim that a diet is healthy cannot be tested; the claim that a particular dietary change alters a particular measurable outcome in a defined population over a defined period can be. Specificity forces the researcher to commit in advance to what will be measured, in whom, and against which comparison. It also makes the eventual result interpretable, because readers can see what question was actually asked rather than the broader question the title suggests.
This is why pre-registration has become common. Researchers deposit their planned methods and analysis in a public registry before collecting data. The purpose is to close off a subtle route to false findings, sometimes described as the many paths a single dataset offers. Analysts face dozens of defensible choices about which participants to include, how to group them, which outcome to treat as primary, and which adjustments to apply. Making those choices after seeing the data, even in good faith, tends to produce results that look stronger than they are.
Inside the peer review process
When a manuscript reaches a journal, an editor first decides whether it is worth sending out at all; many are declined at this stage for scope or apparent weakness. Those that survive go to independent researchers in the same field, who read the manuscript and write assessments recommending acceptance, revision or rejection. Authors typically respond to criticisms, revise, and resubmit, sometimes across several rounds. The editor, not the reviewers, makes the final call, weighing the reports against the journal's judgement of interest and fit.
It is important to understand what reviewers usually do not do. They rarely see raw data. They almost never repeat the experiments. They cannot easily detect fabricated numbers or manipulated images unless something is visibly odd. Reviewers work unpaid, in time carved out of their own research, and often outside their exact speciality. Peer review is therefore best understood as a structured expert critique that catches many errors of reasoning, method and presentation, not as a certificate that a finding is true.
The mechanics also vary. Single-blind review lets reviewers see the authors but not the reverse, allowing reputation to colour judgement. Double-blind review conceals both, though guessing is easy in small fields. Open review publishes reviewer names and reports, trading candour for accountability. Each model manages a different weakness, and none removes them all.
Preprints: fast, public, and unfinished
A preprint is a manuscript posted publicly by its authors before, or instead of, formal peer review. Preprint servers exist across physics, biology, medicine, economics and other fields, and they solve a real problem: journal review can take many months, during which findings that other researchers need remain locked away. In fast-moving situations, particularly outbreaks, that delay has genuine costs. Preprints let work circulate immediately and let criticism arrive from the whole field rather than from two or three selected reviewers.
The corresponding risk is that a preprint carries no external check whatsoever. Anything from careful work by an established laboratory to a badly flawed analysis can appear on the same server, looking identical. The reasonable stance is neither to treat preprints as authoritative nor to dismiss them as worthless, but to read them as what they are: a claim made in public, awaiting scrutiny. Where a preprint drives a news story, it is worth asking whether independent researchers have commented, and whether the work has since been published or quietly abandoned.
Sharing data, code, and materials
Open science describes a cluster of practices aimed at making research inspectable: depositing datasets in public repositories, publishing analysis code, sharing experimental materials and protocols, and making papers freely readable. The motivation is straightforward. If the underlying data and code are available, other researchers can check whether the stated analysis actually produces the stated numbers, try alternative reasonable analyses to see how fragile the conclusion is, and reuse the material for questions the original team never considered.
The obstacles are real rather than merely cultural. Health data about identifiable individuals cannot simply be posted online. Preparing a dataset so that a stranger can understand it takes considerable effort that career systems rarely reward. Researchers who assembled a hard-won dataset may reasonably want time to work with it first. Registered reports offer a partial answer for a different problem: journals review the design before results exist and commit to publishing whatever the study finds, removing the incentive to chase attractive outcomes.
Money shapes the questions that get asked
Research requires salaries, equipment, consumables and time, which means someone decides what gets funded. Government agencies, philanthropic foundations, universities and companies each fund according to their own priorities, and those priorities determine which questions receive sustained attention. Diseases that mainly affect poorer populations have historically attracted less commercial investment than conditions common in wealthy markets, not because the science is less tractable but because the funding logic differs. The shape of the evidence base reflects the shape of the funding base.
Grant cycles also shape behaviour within funded work. Short funding periods favour projects that can produce publishable output quickly, which disadvantages long-term observation, careful replication and negative results that confirm something does not work. Reviewers of grant applications tend to reward novelty, so proposals to repeat someone else's study compete poorly against proposals to try something new. None of this implies dishonesty. It means the literature systematically over-represents the kinds of studies the system finds easy to fund.
Declaring interests is a start, not a settlement
Conflict-of-interest declarations require authors to state financial relationships that a reader might reasonably consider relevant: funding for the study, consulting fees, patents, shareholdings, paid advisory roles. The point is not that funded researchers lie. It is that interest operates through subtler routes, including which comparison is chosen, which outcome is treated as primary, how borderline results are framed, and whether an unfavourable study is written up at all. Disclosure lets readers apply their own discount rather than being kept unaware.
Non-financial interests matter too and are almost never declared. A researcher who has built a career defending a particular theory has a stake in its survival. Institutions gain from prestigious findings; journals gain from attention-grabbing ones. The practical response is not cynicism but calibration. A single industry-funded study should be weighted differently from a body of consistent evidence produced by independent groups using different methods, and that weighting should be conscious rather than reflexive.
What the replication crisis revealed
Over the past two decades, coordinated efforts to repeat well-known published studies found that a substantial share of results did not hold up when the experiments were run again with larger samples and pre-specified analyses. The pattern appeared most visibly in parts of psychology and biomedicine, but concerns have been raised across many fields. The episode was uncomfortable precisely because the original studies had passed peer review, appeared in respected journals, and had been cited for years as established knowledge.
The diagnosed causes were structural rather than scandalous. Studies with few participants produce unstable estimates, and the ones that happen to look striking are the ones that get published. Flexible analysis choices made after seeing data inflate apparent effects. Journals have historically preferred positive, surprising results to null findings, so the published record is a biased sample of the research actually conducted. Add pressure to publish frequently, and a system emerges that reliably generates some proportion of results that will not survive re-testing.
A failure to replicate is not by itself proof that the original was wrong or dishonest, since the repeat study may differ in population or procedure, and it too can be underpowered. What the crisis established is a change of default: a single striking study is a reason for interest, not a reason for belief, and confidence should scale with independent repetition by other groups.
Calibration and the quiet work of measurement
Every instrument drifts. Balances, spectrometers, thermometers, pressure sensors and air-quality monitors all change their response over time with temperature, humidity, vibration, contamination and ordinary wear. Calibration is the routine of comparing an instrument against a reference of known value and correcting or documenting the difference. In well-run laboratories this comparison is traceable through a chain of successively better references up to internationally maintained standards, so that a measurement made in one city is meaningfully comparable to one made elsewhere.
This unglamorous work underwrites everything above it. A carefully designed trial analysed with impeccable statistics still yields nothing if the assay drifted midway through. Systematic measurement error is particularly dangerous because it does not look like noise; it shifts results consistently in one direction and can survive every statistical check. The same logic applies to public data. When comparing air-quality readings from low-cost sensors against reference-grade monitoring stations, differences in calibration and siting can account for much of the apparent disagreement.
Questions worth carrying into the next headline
Faced with a new claim, a few questions do most of the work. What kind of study was this, and what can that design establish? How many participants or observations, and over how long? Was there a comparison group, and was it fair? Is the reported outcome the one that matters to people, or a laboratory marker standing in for it? Is this the first study to report the effect, or one of many? Who paid for it, and does the headline match what the paper concluded?
None of this requires distrust of science. The point is closer to the opposite. Science earns its authority through mechanisms that assume human beings are fallible and self-interested, and that build in comparison, criticism, disclosure and repetition to compensate. Readers who understand those mechanisms can tell a robust finding from a fragile one, hold journalism to a higher standard, and remain appropriately unmoved by the next confident announcement that overturns everything.
Sources & References
Editorial Team
Editorial
In-house writers and editors producing original explainers, guides, and analysis. Articles cite authoritative public sources where helpful.