Reviews, databases, design,
and basic biostats: the filters
“Let’s do a systematic review — no IRB, easy publication.” “I have database access — let’s study sepsis.” Both can be real paths; both are graveyards of resident projects. The preventable killers are usually design failures, not effort failures: a question that isn’t a question, a question already answered by better data, or a question your resources cannot answer. One hour of design literacy — plus a biostatistician’s intake link saved to every phone — prevents the most preventable tragedy in resident research: six months of chart review that can’t answer anything.
Why this block
Beyond case reports, the resident-accessible research world is systematic reviews, retrospective database studies, surveys, education research, and quality improvement — and each kills intern projects in a characteristic way, always at the design stage, always preventably. The single highest-yield guest in this whole curriculum belongs to this block: a biostatistician or data scientist, for thirty minutes, whose entire assignment is one message delivered memorably — come see us before you collect data. An intern who remembers that sentence has absorbed the block.
What interns leave able to do
- Distinguish narrative, scoping, rapid, and systematic reviews — and say when a meta-analysis is not appropriate.
- Give the honest pitch and the honest warning on systematic reviews as resident projects.
- Apply the three design filters before any project starts: real question? already answered? answerable with our resources?
- Write a question in PICO form and screen it with the FINER criteria.
- Explain the IRB’s purpose, the separate gates beside it — privacy, data-use agreements, security — and the rule: the institution’s documented determination before data, never self-issued, renewed when the project changes.
- Name the classic observational biases — selection, misclassification, confounding, immortal time — and why “adjust for everything” is not an analysis plan.
- Practice clean data work: defined variables, a data dictionary, piloted abstraction, pristine raw data.
- Interpret power, p-values, confidence intervals, and effect measures at concept level — and know when to call the biostatistician.
- Name the overlooked resident-sized lanes: surveys, education research, policy-change impacts, and publishable QI.
The three pitches
A PGY-2 corners you: “Case reports are small-time. Let’s do a systematic review and meta-analysis — no IRB, no data collection, just papers. My friend did one and it’s basically a free publication.”
What do you ask her before you say yes?
Your institution has a clinical data warehouse and a structured data-capture platform; you cannot run drug trials or prospective interventional studies as an intern. Three interns pitch you:
- “Sepsis outcomes.” “I want to look at sepsis outcomes in our hospital’s data. Huge problem, tons of patients. We’ll find something.”
- “PPIs and C. diff.” “I want to study whether proton-pump-inhibitor use is associated with C. difficile infection in our inpatients.”
- “The biomarker.” “I want to study whether a novel inflammatory biomarker’s trajectory predicts decompensation in cirrhosis. The warehouse doesn’t capture that lab — but maybe we could get it added to orders for a year and see.”
For each pitch: is it a real question? Is it already answered? Can our resources answer it?
After the mixer, a co-intern deflates: “I don’t have a database or a biomarker. I admitted patients, I sat through the new duty-hour rules rollout, I helped run the intern bootcamp evaluation, and half my clinic patients never show. None of that is research. Research isn’t for me.”
Is she right? What is she standing next to without seeing it?
Reviews, honestly — starting with the taxonomy
The words are not interchangeable, so define them first: a narrative review synthesizes expert-selected literature; a scoping review maps what exists on a broad question; a rapid review trades comprehensiveness for speed and says so; a systematic review answers a focused question with a protocol, an exhaustive documented search, explicit screening, and risk-of-bias assessment — and a meta-analysis is a statistical synthesis inside a systematic review, appropriate only when the studies are similar enough to pool. Heterogeneous studies get a systematic review without pooling; forcing the forest plot is a methods error, not a bonus. Choosing the right type for the question and the team — and labeling the product honestly — is half the craft.
The honest pitch: a systematic review can be an excellent resident project — usually no IRB, no patient data, a learnable methodology, and it teaches critical appraisal better than almost anything. The honest warning: the workload is real — a protocol, registered where eligible; a search strategy built with a medical librarian and saved in full: database, platform, date, exact strategy, filters, counts, because the search is the method section; dual screening of hundreds to thousands of abstracts with a named conflict-resolution process; extraction on a piloted form; risk-of-bias assessment with a design-appropriate tool; and reporting to the PRISMA standard1 — commonly six to twelve months of steady work at resident pace, scaled to scope, and never a free publication. The question test applies with full force: the world does not need the fifth meta-analysis of a settled question, and the flood of low-quality reviews has made editors openly skeptical. The verdict for interns: a full systematic review is legitimate with a genuinely unanswered question, a mentor who has published one, a librarian, and someone who knows the analytic methods. Missing one of the four — shrink honestly instead of quitting: a well-executed scoping or narrative review with transparent methods can be the right resident-sized product, wearing its true label.
The three design killers
Pitch 1 — the inexact question. “Sepsis outcomes, we’ll find something” is a topic, not a question: no exposure, no comparison, no outcome, no hypothesis — and “we’ll find something” is a fishing expedition whose catch is junk, for reasons the biostats primer makes exact. The fix is PICO — population, intervention or exposure, comparison, outcome — followed by the FINER screen: feasible, interesting, novel, ethical, relevant. A question that passes both is a project; a topic with enthusiasm is a future regret.
Pitch 2 — already answered, with better data. The PPI–C. difficile association has been studied in enormous cohorts and meta-analyzed for years; a four-hundred-chart single-center retrospective cannot overturn, refine, or add to that — only add noise. The filter: could your data plausibly change what the best existing evidence says? If bigger and better studies have answered it, your smaller, confounded dataset is not the tool. This is block 1’s search-first lesson at study scale: a half-day in the literature precedes any project.
Pitch 3 — unanswerable with the resources present. The warehouse does not capture the biomarker, and “get it added to orders for a year” converts a chart review into a prospective — possibly interventional — study: full review, funding, consent, years. The same killer wears other costumes: the database lacks the exposure timing, the outcome, the confounders, or the sample size. And the boundary stated precisely: independently launching a drug trial or prospective interventional study is not a resident-feasible project — but joining one is. Established trials run on team members who screen, recruit, collect data, and keep protocol adherence, and that role teaches research consent, safety reporting, and data quality from the inside — with abstracts and secondary analyses as real outputs. Different role, same room. The rule: the feasibility audit happens at week zero — does the data source contain your exposure, your outcome, your confounders, with valid timing, at adequate n? Answered with the person who runs the data, before anything else.
Then model the save: pitch 1, redesigned aloud — a defined sepsis population, a timestamped early-intervention exposure versus later, a concrete outcome by a fixed day — with time zero declared (the moment follow-up starts, chosen so nobody accrues survival before their exposure: the immortal-time trap), the confounders named in advance, and the analysis plan written and versioned before anyone looks at outcomes — because “adjust for everything available” is not a plan, it is fishing wearing statistics. Checked against the literature, checked against the warehouse’s actual fields, screened with FINER. Along the way, name the classic observational traps at recognition level: selection bias (who got into your data and why), misclassification (the code that isn’t the disease), confounding (the third variable driving both), immortal time (the survival you accidentally required). One hour of design turns a topic into a question and a question into a plan. That transformation is the whole block.
The IRB, and clean data practice
The IRB’s purpose comes first, because it reframes everything: it exists to protect human subjects — born from real abuses, codified in the Belmont Report’s three principles: respect for persons, beneficence, and justice.2 When the IRB seems slow, that is what it is slow for. It reviews risk, consent or its waiver, and the privacy of data — and its labels are pathways, not a simple risk ladder: “exempt” means exempt from specified federal requirements, not from ethics, privacy, institutional policy, or review — some exempt categories themselves require limited IRB review — and “expedited” names a review procedure, not a promise of speed. Beside the IRB sit separate gates answering different questions: the privacy office, data-use agreements, information security, and data governance. The single most important rule survives every nuance: you never self-determine. Whether a project is research, human-subjects research, exempt, QI, or preparatory-to-research is your institution’s documented call, made through its required route3 — obtained before you access or use data for the project, and renewed when the project changes: new data elements, identifiers, linkages, populations, or aims can invalidate the original determination. Timelines run weeks; spend them on the protocol, the dictionary, and the literature — design work you needed anyway.
Clean data practice, boarded: define every variable before extracting anything — which definition, which baseline, timed from which timestamp; build a data dictionary as the contract between everyone who touches the data; use your institution’s structured capture platform, never free-text spreadsheet wandering; pilot the abstraction on a handful of charts and revise the form before scaling; if several people abstract, test agreement on the same charts first — or your dataset is three datasets stitched together; decide the missing-data plan up front; and keep raw data pristine, analyzing copies, versioning everything — code and outputs included, deviations documented, null results reported honestly, provenance treated as part of the science. If it cannot be reproduced, it is not science.
What “database” means — and the questions each kind demands
“I have database access” can mean eight different things, each with its own population, blind spots, and gatekeepers:
| The kind | What to ask before promising anything |
|---|---|
| Your EHR / clinical data warehouse | Which fields are actually structured vs. buried in notes; the request path, analyst support, and lead time |
| Disease or procedure registry | Who enters the data and how completely; publication and authorship rules attached to it |
| Claims / administrative data | Codes are billing artifacts, not diagnoses — validation, coding changes over time, lag |
| National public datasets | Unit of observation, survey weights and design, linkage limits, citation and use rules |
| Trial datasets & repositories | Access agreements, approved secondary-use questions, timelines |
| Research networks | Governance, site agreements, who runs queries and who authors what |
Common to all of them: missingness, coding drift, access agreements, security environments, costs, and the institutional reviews above. The exercise that beats the table: walk your own institution’s request path once — the data dictionary, the intake form, the analyst — before you need it.
The team — and the week-zero rule
You are the domain expert and the worker bee: the clinical question, the literature, the abstraction, the interpretation, the writing — not the methodologist. The data scientist knows the warehouse’s data model, extracts the cohort, and builds the pipeline that makes your dataset verifiable. The biostatistician co-develops the design, the power or sample-size justification, and the analysis plan with you — methodological leadership sits with expertise, and the plan belongs to the team — and is consulted before data collection, with your question and the data source’s contents, never with a finished spreadsheet. Many institutions offer free trainee consults; the intake link goes on every phone before this block ends. Give the guest the closing line and let them deliver it: “The saddest meetings I take are with residents who spent six months collecting data I can’t use. The happiest are the ones who come at week zero with a question. Be the second one.”
Biostats at concept level — no math, all meaning
- The frame: assume no effect (the null), then ask whether the data are surprising under that assumption. Everything else is detail on that one idea.
- Power is the probability of detecting an effect of a given size if it is really there — driven by sample size, effect size, and variability. Underpowered studies are the quiet scandal of resident research: real effects dissolve into noise, months are wasted, and subjects’ data served no end. Power is calculated before collection, with the biostatistician — and where a project is descriptive or hypothesis-generating, the honest requirement is a sample-size justification, not a ritual power calculation.
- The p-value is the probability of results this extreme if the null were true — not the probability your hypothesis is correct — and the cliff at 0.05 is a convention, not a law of nature. Always read it with the effect size and its interval.
- The confidence interval, read carefully: the textbook meaning is that under repeated sampling, 95% of intervals built this way would contain the true effect — the usable shorthand is the range of effect sizes reasonably compatible with your data, under the model. Narrow means precise; an interval spanning no-effect means uncertainty; and big effect, wide interval means promising-but-underpowered — the two most misread words in resident abstracts.
- One 2×2, several stories: the same data yield a relative risk, an absolute risk reduction, and a number needed to treat — and abstracts love the relative number because it sounds bigger. Odds ratios are their own measure: they approximate relative risk only when the outcome is rare, and when the outcome is common, reading an OR as if it were a relative risk overstates the effect; know which you are reading before repeating it at journal club.
- Statistical versus clinical significance: enormous samples make trivial differences “significant,” and small samples can leave huge effects uncertain. A p-value speaks to how surprising the data would be under the null — never to importance, and never to the probability your hypothesis is true; a finding worth acting on needs precision and a size that matters.
- Models have assumptions. Missing data, clustering and repeated measures, overfitting, unvalidated models — each can silently invalidate a result, which is one more reason the analysis plan is co-developed in advance rather than improvised over a spreadsheet.
- Multiplicity — why fishing fails: test twenty hypotheses at the conventional threshold and one “significant” result is expected by chance alone. This is why the primary outcome is pre-specified, why exploratory findings are labeled exploratory, and why pitch 1’s “we’ll find something” produces publishable-looking garbage.
- The retrospective caveat, tattooed: observational data show association, not causation — confounding is the rule, not the exception. Adjust for what you measured, confess what you didn’t, and write the limitations section like you mean it.
The overlooked lanes — the deflated co-intern’s gold mine
- Surveys. “Half my clinic patients no-show” is a survey study of barriers waiting to be designed — with real methodology: validated instruments where they exist, a defined sampling frame, a pilot for clarity, a defensible response rate, and honest attention to who didn’t answer, because nonresponse is the survey’s confounding. And the rule that surprises interns: surveying humans — patients or co-residents — is human-subjects research and needs an IRB determination, per the always-rule: the IRB decides, not you.
- Educational innovations. “I helped run the bootcamp evaluation” is education research: a needs assessment, a curriculum intervention, pre-and-post measures, behavior follow-up — publishable with proper design, in a field where residency programs are full of natural subjects. The catch: the evaluation must be designed in advance, not stapled on afterward.
- Requirement and policy-change impacts. Living through a duty-hour rollout is standing inside a natural experiment: before-and-after comparisons around mandated changes — handoff requirements, documentation rules, schedule redesigns — are real health-services research, and institutions change requirements constantly while almost never studying them. Designed properly: enough observations on both sides of the change, attention to secular trends and whatever else changed at the same time — a single before-and-after snapshot is the weakest version, and the interrupted-time-series shape is the strong one.
- Quality improvement, the publishable version. The bootcamp session called QI first-class scholarship; the publishable form has a defined problem with baseline measurement, an intervention, repeated measures over time, balancing measures — what did the intervention break? — and reporting to the SQUIRE standard.4 QI usually needs an IRB determination even when deemed non-research — same rule as always.
The meta-point, landed: residents overlook these lanes because they don’t look like “real research.” They are real research — often more publishable than a shaky database study, squarely within resident resources, directly useful to your own program — and fellowship and job committees count them. Two more doors, while the list is open: the team-member role on an existing trial (pitch 3’s lesson, turned positive), and the case-series, diagnostic-accuracy, and secondary-analysis lanes your mentors already work.
AI in the research workflow — the governance rules
The same three-tier rules from the hub, applied to research tasks specifically:
- Searching: term expansion and method explanations can help; generated citations are not a search — validate everything against the librarian-built strategy, and open every reference.
- Screening and extraction: only under a protocol, on institutionally approved systems, with validation samples, human adjudication, and the model, version, prompts, and settings recorded — a tool never silently replaces an independent reviewer.
- Code and statistics: develop on synthetic or approved data, inspect every line, test against known cases, preserve code and outputs — and the statistical review still happens.
- Writing and figures: the publication rules apply — disclosure of tool and purpose, no authorship, full human accountability5 — and figures come from verified code and data, never generation.
And the distinction that catches teams: an institutionally approved tool is approved for privacy — validity for your scientific task is a separate question you still owe.
Pocket card
- Topic ≠ question. PICO it, then FINER it.
- Search first: answered by better data? Different question.
- Feasibility audit at week zero: exposure · outcome · confounders · n — in YOUR data source.
- The institution decides, not you — documented determination before data, renewed when the project changes. Exempt ≠ unreviewed.
- Biostatistician before collection. Power — or an honest sample-size justification — before data.
- A p-value is not certainty, importance, or truth. Effect size + interval, always. Retrospective = association.
- Time zero declared; analysis plan versioned before outcomes are seen. “Adjust for everything” is fishing.
- Surveys · education · policy-change · QI · trial-team roles: the overlooked, publishable, resident-sized lanes.
- AI works inside protocols and approved tools — it never decides, and you disclose.
- Leave with: a one-page concept — question, design, time zero, variables, biases, data source, approvals, team.
Notes
Every pitch and every intern in this block is a fictional composite. The final block is the mentorship long game.
Sources
- Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://pubmed.ncbi.nlm.nih.gov/33782057/ ↩
- National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. (1979). The Belmont report: Ethical principles and guidelines for the protection of human subjects of research. U.S. Department of Health and Human Services. https://www.hhs.gov/ohrp/regulations-and-policy/belmont-report/index.html ↩
- Office for Human Research Protections. (n.d.). Exempt research determination FAQs. U.S. Department of Health and Human Services. https://www.hhs.gov/ohrp/regulations-and-policy/guidance/faq/exempt-research-determination/index.html Canonical federal guidance on exemption determinations; the linked page blocks automated access, and institutional processes govern in practice — verify through your own HRPP. ↩
- Ogrinc, G., Davies, L., Goodman, D., Batalden, P., Davidoff, F., & Stevens, D. (2016). SQUIRE 2.0 (Standards for QUality Improvement Reporting Excellence): Revised publication guidelines from a detailed consensus process. BMJ Quality & Safety, 25(12), 986–992. https://pubmed.ncbi.nlm.nih.gov/26369893/ ↩
- International Committee of Medical Journal Editors. (n.d.). Use of artificial intelligence in publishing. In Recommendations for the conduct, reporting, editing, and publication of scholarly work in medical journals. https://www.icmje.org/recommendations/browse/artificial-intelligence/ ↩
This page is a teaching framework for facilitated small-group education, at concept level by design: it states no statistical thresholds for real analyses and no regulatory determinations for real projects. Study design, review levels, and analysis plans belong to your IRB, your methodologists, and your mentors. Last reviewed July 2026.