How to tell whether something is true before you do it to somebody
This is the subject students most reliably dread and most reliably dismiss, and it is the one that decides whether the rest of their career consists of doing what they were shown or doing what works. Nursing has a long history of practices that everybody knew were correct and that turned out to harm people, and each of them was overturned not by a strong opinion but by somebody collecting evidence properly. The purpose of this manual is not to turn nurses into researchers. It is to make a nurse able to read a claim — in a journal, in a policy, in a training session, on a drug company's leaflet, from a confident senior colleague — and work out whether there is any reason to believe it.
Nursing and medicine have a long record of practices that everyone was certain about and that turned out to cause harm — positions for infants that increased deaths, routine procedures that were never necessary, wound treatments that delayed healing, bed rest prescribed for conditions it worsened. In each case the practice was taught, examined and defended by experienced people. What changed it was not a better opinion but somebody measuring properly.
Evidence-based practice is not a rule that research beats everything. It is the combination of the best available evidence, the clinician's own experience and judgement, and the values and preferences of the patient. A study showing that a treatment works on average says nothing about whether this person wants it, and an intervention nobody will accept has an effectiveness of zero however strong the trial.
Very few nurses will design a study. Almost every nurse will be told that something is best practice, will be handed a new protocol, will be sold a product, or will be asked why they do something a particular way. The skill this subject provides is the ability to ask where that came from and to understand the answer.
Taught as terminology, this subject is genuinely tedious. Taught as a set of questions — how would you know, how could you be fooled, what would change your mind — it is one of the more interesting things in nursing, because it is about how anybody knows anything at all.
Does mouth care help is not answerable. Whether twice-daily toothbrushing in ventilated adults reduces pneumonia compared with sponge swabs is answerable. The discipline of narrowing a question until it can be answered is most of the work, and a question that cannot be stated precisely usually indicates that the thinking behind it is not finished.
A clinical question generally names the population, the intervention or exposure, the comparison, and the outcome. Who are we talking about, what are we doing, compared with what, and what result are we measuring. Leaving out the comparison is the commonest omission, and it is the one that makes results uninterpretable, because better is meaningless without better than what.
A question about whether a treatment works needs a comparison between groups. A question about how common something is needs a survey. A question about what it is like to live with a condition needs interviews. Choosing a design because it is prestigious rather than because it fits the question is a standard error, and trials are the most frequently misapplied.
The most useful nursing research questions usually come from noticing that something on the ward is done in two different ways by two different people, or that a policy exists that nobody can explain. Those are exactly the points where evidence is likely to be thin and where finding out would change something.
Case reports describe one patient, case series a handful. They cannot establish that anything caused anything, because there is nothing to compare against, and they are frequently the first sign of something important. Their value is in generating questions and in flagging rare events, and their limitation is that they can only ever do that.
A cross-sectional study measures a population at one point in time and is the right design for asking how common something is. It cannot establish sequence, because exposure and outcome are measured together — which is why a survey finding that people with back pain sit more cannot tell you whether sitting caused the pain or the pain caused the sitting.
A cohort study identifies a group, records their exposures, and follows them to see what happens. Because exposure is recorded before the outcome occurs, sequence is established, which is a real advantage. The weakness is that the groups differ in many ways besides the exposure, and some of those differences also affect the outcome.
A case-control study starts with people who have the outcome and people who do not, and looks back at their exposures. It is efficient for rare outcomes and for long delays between exposure and disease. It is also vulnerable to people remembering their past differently depending on whether they became ill, which is a real and well-documented effect.
In a randomised trial, participants are allocated to groups by chance. This is the only design that reliably makes the groups similar in everything, including things nobody thought to measure, which is precisely what allows a difference in outcome to be attributed to the intervention. It is not always possible, not always ethical, and not always necessary.
Qualitative designs use interviews, observation and analysis of what people say to understand experience, meaning and process. They answer questions numbers cannot — why a protocol is not followed, what it is like to be told a diagnosis. Judging them by whether they had a large sample or a control group is a category error and a common one.
A result applies, strictly, to people like those who were studied. A trial conducted in men aged under sixty-five in a specialist centre tells you much less about an eighty-year-old woman with three other conditions in a district hospital. Checking who was actually included is the first and most neglected step in reading a paper.
A random sample gives everyone in the population a known chance of selection and allows the findings to be generalised. A convenient sample — whoever was available, whoever volunteered — does not, and volunteers differ systematically from those who do not volunteer. Neither is automatically wrong; describing a convenient sample as representative is.
Too small a study cannot detect a real effect, and reports no difference when a difference exists. Too large a study detects differences so small they do not matter to anybody. The important question about a negative result is therefore always whether the study was capable of finding the effect it was looking for.
People who leave a study part-way through are rarely a random subset — they are often those who did worst or tolerated the intervention least. Losing a substantial proportion therefore biases the result in a predictable direction, which is why the number lost, and any statement about who they were, matters as much as the headline finding.
Bias in research means a systematic error that pushes the result consistently in one direction, and it usually happens to careful, honest people who did not intend it. Understanding it as a structural problem rather than a moral one is what allows somebody to look for it in their own work, which is the only place they can do anything about it.
Selection bias occurs when the people in the study differ systematically from those they are meant to represent, or when the comparison groups differ from each other at the start. It is the reason randomisation exists, and the reason a study comparing patients who chose an intervention with those who did not is so hard to interpret.
If the people assessing outcomes know who received the intervention, their judgement shifts, particularly for outcomes involving opinion — wound appearance, pain scores, improvement. Blinding the assessor addresses this. Where blinding is impossible, as with most nursing interventions, an objective outcome measured by someone uninvolved is the next best defence.
A confounder is a third factor related both to the exposure and to the outcome, which makes them appear connected when they are not. Studies showing that coffee drinkers had more lung disease were describing smoking. Recognising this pattern is one of the most useful habits this manual can give, because it applies to headlines as much as to papers.
Studies finding an effect are more likely to be written up, submitted and published than studies finding nothing. The consequence is that the published literature systematically overstates effects, and that a review of published studies inherits the distortion. This is why trial registration before results are known matters so much.
Whole populations are systematically under-represented in health research — women in cardiac trials, people who do not speak the dominant language, those without stable addresses, and almost everyone in the poorest countries. The consequence is that the evidence base itself is skewed, not merely individual studies within it, and the groups with the least evidence are frequently those with the greatest need.
Two things occurring together does not establish that one produced the other. They may be connected through a third factor, the sequence may be the reverse of what is assumed, or it may be chance. Nearly every misleading health claim in public circulation is an association reported as a cause, and recognising the move is the single most useful skill in this subject.
A causal claim is stronger when the exposure clearly preceded the outcome, when more exposure produces more effect, when the association is large, when it is found repeatedly in different populations by different methods, when there is a plausible mechanism, and when removing the exposure removes the effect. No single one of these proves anything; together they build a case.
Random allocation makes the groups comparable in everything including unmeasured factors, so a difference afterwards can be attributed to the intervention. This is the whole reason trials sit where they do, and it is also why a badly conducted trial with unequal groups or heavy dropout loses precisely the advantage it was built for.
Some questions cannot ethically or practically be randomised — nobody will allocate people to smoke, or to be poor. For these, causation is established by accumulating consistent observational evidence with the features above. The absence of a trial is therefore not the absence of evidence, which is an argument frequently made in bad faith.
A treatment that reduces risk from two in a thousand to one in a thousand halves the risk, which sounds substantial, and prevents one event per thousand people treated, which sounds modest. Both statements are true. Reporting only the relative figure is the commonest way of making a small effect look large, and it is used constantly in promotional material.
The number of people who must receive a treatment for one of them to benefit is a plain-language way of expressing the same information, and it is far harder to dress up. A number of ten is a strong treatment; a number of several hundred may still be worthwhile for a serious outcome but changes the conversation about side effects and cost.
A confidence interval expresses the range within which the true value plausibly lies. A narrow interval indicates a precise estimate; a wide one indicates that the study cannot say much. An interval that includes the possibility of no effect means the study has not established one, whatever the headline figure inside it appears to show.
A statistically significant result means the finding is unlikely to have arisen by chance alone if there were truly no effect. It does not mean the effect is large, important, or relevant to your patients. A significant but tiny difference and a non-significant but substantial one are both common, and both are routinely misreported.
If a study measures twenty outcomes, one will typically appear significant by chance alone. This is why the main outcome should be stated before the study begins, and why a finding presented from a subgroup that was not specified in advance should be treated as a suggestion for future work rather than as a result.
Single studies disagree with each other routinely, for reasons of chance, population and method. A systematic review searches for all the studies on a question using a stated method, appraises them, and summarises what they collectively show. The systematic part is what distinguishes it from an article in which somebody cites the studies they happened to like.
Where studies are similar enough, their results can be combined statistically to produce a more precise estimate. Where they are not similar enough — different populations, different interventions, different outcomes — combining them produces a precise answer to no real question. A good review says which situation it is in.
A guideline turns evidence into advice, and that step involves judgements about benefit, harm, cost, feasibility and values that the evidence alone does not settle. This is why good guidelines state both the strength of the recommendation and the certainty of the evidence separately, and why two honest groups can read the same evidence and advise differently.
A guideline describes what is usually right for a typical patient. Departing from it for a particular patient is legitimate and sometimes necessary, provided the reason is thought through and recorded. Following a guideline that plainly does not fit the person in front of you, because it is the guideline, is not good practice and is not a defence.
Two systematic reviews of the same question sometimes reach different conclusions, usually because they searched differently, included different studies, or judged quality differently. This is not a scandal; it is visible disagreement about method, and the useful response is to read what each included and excluded rather than to pick the one that agrees with you.
Qualitative research asks what something means, how it is experienced, and why things happen as they do. It is the right tool for understanding why staff do not follow a protocol, what patients actually fear about a procedure, or how a diagnosis changes somebody's sense of themselves. Numbers cannot reach these, and nursing questions are frequently of this kind.
Data come from interviews, group discussions, observation and documents, and are analysed by systematically identifying and organising themes rather than by counting. Good qualitative work shows its method, includes data that complicate the argument, and quotes participants at enough length that the reader can judge the interpretation.
Qualitative research is not judged by sample size or by generalisability. It is judged by whether the method suited the question, whether the analysis is transparent, whether the researcher's own influence is acknowledged, and whether the conclusions are traceable to the data. Criticising it for having a small sample is a category error.
Many practical questions need both: a trial to find out whether something works, and qualitative work to find out why it was not used or how it was experienced. An intervention that works in principle and is never implemented has achieved nothing, and the reasons are almost always qualitative.
The rules governing research on people were written in response to research that harmed people, including studies conducted on prisoners, on patients who were not told, and on communities selected because they could not refuse. This history is the reason the requirements feel heavy, and knowing it makes them read as protections rather than as paperwork.
Participants must understand what is proposed, what the risks are, that participation is voluntary, and that they can withdraw at any time without their care being affected. That last clause matters most where the researcher is also the clinician, because a patient may reasonably fear that refusing will change how they are treated.
Children, people lacking capacity, prisoners, employees and patients dependent on the researcher are all in positions where consent may not be free. Research involving them is not prohibited, because excluding them would leave them without evidence to guide their care, but it requires additional safeguards and independent scrutiny.
Participants are entitled to have their identity protected, their data stored securely, and their information used only for what they agreed to. Small studies in small communities present a particular difficulty, because describing a case in enough detail to be useful may identify the person to anyone local.
A study can be approved and still be conducted unethically — by pressuring people to participate, by continuing when harm becomes apparent, by not reporting findings the sponsor dislikes, or by abandoning a community once data collection ends. Approval is a checkpoint, not a guarantee, and the obligations continue after it.
Outbreaks and disasters create pressure to act without evidence and simultaneously offer the only opportunity to generate it. Doing research during an emergency is both necessary and more dangerous ethically, because people are frightened, alternatives are scarce and consent is harder to make genuinely free. The response is not to suspend the safeguards but to plan the studies in advance, so that when the emergency arrives the protocols and the approvals already exist.
Research asks what should be done. Audit asks whether what should be done is actually being done here. The two are frequently confused, and the distinction matters practically because audit usually does not require the same ethical approval and because audit answers the question most wards actually have.
Audit sets a standard, measures current practice against it, identifies the gap, changes something, and then measures again. The final step is the one most often omitted, and an audit that stops after measuring has demonstrated a problem and fixed nothing. Most audit in most institutions stops there.
Quality improvement uses small, rapid cycles of change with measurement after each, rather than one large redesign. The advantage is that a change that does not work is discovered quickly and cheaply, and that people affected by the change are involved in shaping it, which is usually why changes succeed or fail.
Improvement efforts fail for consistent reasons: nobody asked the staff who do the work, the new way is slower than the old, the equipment is not where it needs to be, there is no feedback so people cannot see any effect, and the person driving it left. Anticipating these is more useful than any technique.
Audit data about individuals rapidly becomes performance management, and the moment it does, recording becomes defensive and the data stops being true. Measuring the system rather than the person, and feeding results back to the team rather than upwards, is what keeps the numbers honest enough to be worth collecting.
Start with sources that have already done the appraisal — systematic reviews, guidelines from recognised bodies, summaries of evidence — before searching for individual studies. Most clinical questions have been reviewed by somebody, and finding that review is faster and more reliable than assembling an answer from primary papers.
A useful search uses the concepts in the question rather than a sentence, combines synonyms, and narrows by design where appropriate. Searching once with a single phrase and concluding that nothing exists is the most common failure, and it usually means the vocabulary of the literature differs from the vocabulary of the ward.
Much research sits behind payment that most nurses cannot afford personally. Open-access journals, freely available guidelines from health bodies, institutional access where employment provides it, and requesting a copy directly from the author — which is normal and frequently successful — are the routes that exist. Pretending the barrier does not exist helps nobody.
Before reading in detail, check who funded it, who conducted it, whether it was registered before it started, whether it is peer reviewed, and how old it is. None of these decide the matter alone, and together they tell you how much attention the paper deserves.
Librarians in health institutions are trained in exactly this and are consistently the most under-used resource available to a nurse with a question. Where no librarian exists, a colleague who has done a degree recently, a professional body's enquiry service, or an author contacted directly will usually help. The habit of asking a person, rather than concluding that the answer is unavailable, finds most of what is findable.
There is a well-documented gap of years between evidence being established and practice changing, and the gap is caused by the difficulty of changing behaviour rather than by ignorance. Simply telling people the evidence is among the least effective ways of changing what they do, which is inconvenient and well established.
Changes are adopted more readily when the new way is easier than the old, when respected colleagues do it, when people receive feedback on their own performance, when reminders are built into the work rather than added to it, and when those affected helped design it. Most successful change uses several of these together.
A great deal of evidence assumes equipment and staffing that are not available, and the useful question becomes which part of the intervention carries the effect and whether that part can be delivered. Abandoning an intervention because the full version is impossible is a common and costly response.
Removing practices that do not work is as valuable as adopting ones that do, and considerably harder, because nobody is promoting removal and because the practice usually feels caring. Routine tasks performed because they have always been performed are worth asking about, and the question of what would happen if we stopped is a legitimate one.
Improvements driven by one enthusiastic person reliably decay when that person leaves. What survives is change built into the equipment, the form, the default option, the order of a checklist or the layout of a room — because those keep working when nobody is paying attention. Designing for the day after you leave is the difference between a project and an improvement.
Read the aim, then the conclusion, and ask whether the conclusion answers the aim. A surprising number of papers conclude something adjacent to what they set out to examine, and noticing that gap early saves reading the rest carefully.
Look at who was included and excluded, how many there were, and how many finished. These three tell you whether the study could answer its question and whether its answer applies to your patients, and they are printed in a table that takes thirty seconds to read.
Find the one outcome the study was designed around and look at the size of the effect, the confidence interval, and whether it is presented in absolute terms. Ignore the abstract's adjectives. If the effect is presented only as a percentage reduction, find the actual numbers before forming a view.
Check funding, whether the study was registered in advance, whether the outcome reported matches the outcome planned, and whether harms are reported at all. A paper that reports benefits in detail and harms in a sentence is telling you something about itself.
The mean is the arithmetic average and is pulled by extreme values; the median is the middle value and is not. Where a distribution is skewed — length of stay, income, waiting time — the median describes a typical person far better, and a paper reporting a mean for such data is either careless or making something look better than it is. The spread matters as much as the centre, because two groups with the same average can be entirely different.
A statement that eighty per cent of patients improved means nothing until you know how many patients there were. Eighty per cent of five is four people. Percentages calculated on small numbers move dramatically with one extra case, which is why good reporting gives both the percentage and the actual count, and why a percentage presented alone should prompt the question.
A test that is highly sensitive rarely misses people who have the condition, so a negative result is reassuring. A test that is highly specific rarely flags people who do not, so a positive result is meaningful. Screening tests are usually built to be sensitive, which is why they produce many false positives and why a positive screen is the beginning of investigation rather than a diagnosis.
When a condition is rare, even a good test produces more false positives than true ones, simply because there are so many more people without the condition than with it. This is genuinely counterintuitive, it is the reason mass screening for rare conditions causes harm as well as benefit, and it is one of the few statistical ideas that changes what a nurse says to a frightened patient.
Most published health research is conducted in a small number of wealthy countries, in well-staffed institutions, on populations that do not resemble most of the world. A finding is evidence about the setting in which it was produced, and transferring it elsewhere is a judgement about whether the mechanism travels even when the resources do not. That judgement should be made explicitly rather than by default.
Trials routinely exclude the old, the pregnant, the very unwell, those with several conditions, and those who do not speak the local language — which between them describe a large proportion of real patients. The result is strongest evidence for the patients who need it least. Noticing the exclusions is part of reading a paper honestly.
Much of nursing is complex, delivered by people, adapted to the individual, and impossible to blind. A trial of a communication approach cannot conceal which arm a patient is in. This does not mean such interventions cannot be evaluated; it means the design must fit, and that dismissing nursing evidence for failing to look like a drug trial is a misunderstanding of both.
For a great many practical nursing questions there is no adequate evidence at all, and the honest answer is that we do not know. What follows is reasoning from mechanism, from experience and from the patient's own preference, made explicitly and revisited. Claiming evidence that does not exist is worse than admitting the gap, because it cannot be corrected later.
Studies funded by an organisation with a commercial interest in the result are, on average, more likely to produce findings favourable to that interest. This is a measured effect across many fields rather than an accusation about any individual study, and it operates through design choices, comparator selection and which results get published rather than through fabrication.
A new intervention can be made to look good by comparing it with nothing, with an inadequate version of the alternative, or with a dose or method chosen to underperform. When reading a comparison, the question is not only whether the new thing won but whether the thing it beat was given a fair run.
Measuring a substitute outcome — a blood value, a score, a scan appearance — rather than something that matters to a patient makes a positive result much easier to obtain. A treatment that improves a number without improving how long or how well somebody lives may still be useful, but the claim being made should be stated in those terms.
Product information, sponsored study days and materials supplied with equipment are marketing, however educational they look, and they are effective precisely because they are useful. Using them is reasonable; treating them as an unbiased summary of the evidence is not, and the simplest defence is to notice who produced the page before reading it.
An audit completed and filed, or a project presented once and forgotten, has cost real time and changed nothing beyond the person who did it. Sharing does not require a journal article — a ward presentation, a poster, a written summary for the team, or a note in the handover file all count, and all are more than most improvement work receives.
Almost every research report follows the same structure: why the question matters, what was already known, what was done, what was found, and what it means. Following that order makes writing far easier than it appears, because each section answers one question and nothing else belongs in it.
State what you set out to measure, including anything that did not work. Give the numbers, including the denominators. Say what the limitations were before a reader finds them. A report that acknowledges its own weaknesses is more persuasive than one that does not, because the reader can tell that somebody was thinking.
Everyone who contributed substantially should be named, and being senior is not by itself a contribution. Data collected by students and healthcare assistants is frequently the whole basis of a project, and omitting those who gathered it is both common and wrong. Agreeing authorship at the start prevents most of the disputes.
Most clinical uncertainty arrives in a form no paper addresses directly: this patient, with these three conditions, refusing this part of the plan, in this ward with this staffing. The realistic use of evidence is not to find the answer but to know which general findings apply, how strongly, and where the judgement begins. A nurse who expects the literature to decide will be disappointed; one who expects it to inform will use it constantly.
In practice the choice is between a quick look at a reliable summary and no evidence at all. Knowing two or three trustworthy sources well, and being able to search them quickly, is worth far more than a thorough appraisal technique that is never used because it takes an hour. Speed here is not a compromise; it is what makes the habit survive a shift.
Patients are frequently told either that something definitely works or nothing at all, when the truth is usually that it helps some people somewhat. Saying plainly that the evidence is limited, what it does suggest, and what the alternatives are, respects the person and is also more durable — because when the advice later changes, they were not told a certainty that turned out false.
Practice changes, which means much of what you are taught now will be revised during your career, and some of it will be revised because it was harming people. Holding current practice as the best available account rather than as truth is what allows a professional to change without feeling attacked, and it is the attitude this entire subject is trying to build.
When somebody says this is how we do it here, ask where that came from. Asked as a genuine question rather than a challenge, it is almost always answered generously, and it occasionally uncovers that nobody knows. Those are the moments where evidence-based practice actually begins, and they require no journal access, no statistics and no permission.
The reliable topics are matching a design to a question, the difference between association and causation, informed consent and voluntariness, the distinction between audit and research, and the interpretation of a result. Calculation is rarely required; interpretation almost always is.
Treating a large relative reduction as a large benefit; treating a non-significant result as proof of no effect; treating a qualitative study as weak because the sample was small; treating ethics approval as making a study ethical. Each of these appears repeatedly because each is genuinely tempting.
Take any health claim you encounter — a headline, a leaflet, a policy — and work out what design would be needed to support it, what could have biased it, and whether the claim is about association or cause. Ten minutes on a real claim teaches more than an hour of terminology.
When told that something is best practice, ask where it comes from, and ask it as a genuine question rather than a challenge. When a routine cannot be explained by anybody, that is worth noticing. These two habits are the whole of this subject applied, and they cost nothing.