Why Prompting Matters
An AI’s answer can only be as good as the question behind it. “Garbage in, garbage out” has always been a data problem, but with large language models (LLMs), the garbage is just as likely to be in the prompt.
In evidence synthesis, this isn’t an abstract concern. A poorly worded question doesn’t just produce a mediocre answer, it can introduce inconsistency, bias, or missed studies at scale, repeated across thousands of citations. And because the same question is applied to every record, a small ambiguity doesn’t cost you one bad decision. It costs you the same bad decision made thousands of times over.
That’s why prompting can no longer be treated as a “nice to have” skill for a handful of power users. It’s becoming a core literacy for anyone using AI tools in their workflow, much the way search literacy became a baseline professional skill in the 2000s. Knowing how to ask a good question is quickly becoming as important as knowing how to read the answer.
Why DistillerSR Is Positioned to Lead This Conversation
DistillerSR has been building and refining AI-enabled workflows since 2016, well before AI was mainstream in systematic review tools.
DistillerSR’s earlier AI features were deterministic: they learned from reviewer decisions and gave the same output for the same input. Our newest capabilities, Smart Screening and Smart Evidence Extraction, are built on large language models, and that changes what matters. With LLMs, the quality of the prompt directly affects the quality of the result: how a screening or extraction question is structured determines how accurately the model makes decisions and pulls data. A clearer question gives a more accurate, more defensible output.
Our Professional Services, Customer Success, and Training teams have led many customers through full implementations, so we’ve seen the full range of what works and what doesn’t. For us, this isn’t theory. It’s operational knowledge from real customers, real edge cases, all with regulatory implications. This is the practical insight gained only by directly helping review teams resolve complex challenges.
What We’ve Learned: Best Practices for Well-Formed Questions
The following guidance is the result of implementations and training review teams across academic, public sector, pharmaceutical, and medical device organizations.
- Give the model context, not fragments. A question like “Include?” gives the AI nothing to work with. The question itself needs to carry the full context, what’s being evaluated, and against what criteria.
- Write complete, well-structured questions. Full sentences outperform fragments. The model responds to structure the same way a human reviewer would.
- Define your terms. Spell out what “device,” “population,” or “outcome” means for this review rather than assuming shared understanding, the model has no institutional memory to fall back on.
- Use plain language. Avoid unexplained jargon or acronyms that force the model to guess at intent.
- One question, one ask. Split compound questions, like “What is the preferred treatment? Include?” into separate, single-purpose questions.
- State the default. Tell the model explicitly what to select when nothing clearly applies, rather than leaving that judgment call implicit.
- Remind the model of its inputs. If it’s working from title and abstract only, not full text or workflow metadata, the prompt should reflect that constraint so the model doesn’t reason beyond what it actually has access to.
What We’ve Learned: Build Clear, Distinct Answer Options
Good questions are only half the equation. The answer options you build around them matter just as much.
- Clarity and distinctness. Every option should be unambiguous and jargon-free.
- One concept per option. Don’t merge overlapping reasons into a single answer choice, it forces the model (and human reviewers) to guess which part is actually driving the decision.
- No overlapping ranges. This is especially important in radio or categorical questions. Avoid ranges like 0–9, 9–18, 18+, which create ambiguity at the boundaries.
- Order objective before subjective. Put administrative criteria (language, species) ahead of judgment-based ones (population, outcome), so the easier calls are made first.
- Use validations on text questions. Numeric validation, for example, constrains the response format and reduces downstream cleanup.
- Keep option lists manageable. More than 20 options slows response time without adding meaningful precision.
- Pilot before scaling. Test a question on edge-case abstracts before trusting it at scale across the full citation set.
Bringing It Together
Here’s what these principles look like in practice, a weak prompt rebuilt into a strong one.
Weak: “Include?”
Strong: “Based only on the title and abstract, select the single most appropriate screening decision for this reference in a systematic review of heart/coronary stents (device Z560). Evaluate the exclusion criteria in the order listed; if none clearly apply based on the available information, select ‘Include for further review.'”
The difference isn’t stylistic, it’s structural. The strong version gives the model scope (title/abstract only), context (the specific device and review type), sequence (evaluate criteria in order), and a default (what to do when nothing clearly applies). Each of those elements removes a place where inconsistency could creep in.
Why This Matters Beyond the Output
In evidence-based domains like the public sector, academia, pharmaceutical, and medical devices, prompting carries extra weight. It’s not just about getting a better output. It’s about defensibility, auditability, and trust in the review process. Being able to show, after the fact, exactly what was asked, why, and how a human validated the result.
DistillerSR’s role isn’t just to build regulatory-grade AI in workflows, our expert managed services are designed to support our customers in getting the most out of it. Learn more about DistillerSR’s Expert Managed Services.







