Selecting and Using Large Language Models
Model selection, prompting, and bounded workflow design for public health large language model use. The material is maintained separately so each operational question has a stable, focused reference.
- Identify the evidence and controls relevant to this decision area
- Distinguish technical performance from operational and population impact
- Apply the included framework without extending claims beyond the cited evidence
Introduction
This focused reference is part of the broader Selecting and Using Large Language Models overview. It preserves the detailed methods, examples, and exercises while reducing page size and improving direct navigation.
Choosing the Right LLM: A Decision Framework
Landscape Overview (2025)
AI models evolve rapidly. This section describes the state of major LLMs as of December 2025 (GPT-5.2 released December 11, Claude Opus 4.5 released November 24, Gemini 3 Pro released November 18, Grok 4.1 released November 17). For current information, always check: - Model provider websites for latest versions - Benchmark comparisons (e.g., LMSys Chatbot Arena) - Independent reviews and comparisons - Release notes and announcements linked in each model description below
The principles for choosing LLMs remain stable even as specific versions change.
Major LLM Options
OpenAI GPT-5 family (via ChatGPT, API) - Current versions: GPT-5.2 (released December 11, 2025) with Instant, Thinking, and Pro tiers; GPT-5.1-Codex-Max for agentic coding - Strengths: 400K token context window, 38% fewer errors than GPT-5.1, ~30% fewer hallucinations, leading reasoning, PhD-level expertise across domains - Weaknesses: Higher API costs (1.4x GPT-5.1 pricing at $1.75/$14 per million tokens), rate limits on free tier - Best for: Complex reasoning, code generation, multimodal analysis, research assistance, general use - Access: Free (GPT-5.2 with limits), ChatGPT Plus ($20/month - higher limits), ChatGPT Pro (unlimited GPT-5.2 + GPT-5.2 Pro access), Enterprise ($60+/user/month) - Context window: 400K tokens (~1,000 pages) - Learn more: OpenAI Platform | GPT-5.2 announcement
Anthropic Claude family (via Claude.ai, API) - Current versions: Claude Opus 4.5 (released November 24, 2025) - most intelligent model, Claude Sonnet 4.5 (released September 29, 2025), Claude Haiku 4.5 (released October 15, 2025) - Strengths: Leading agentic coding (Opus 4.5 outperforms Gemini 3 Pro and GPT-5.2 on SWE-bench), 30-hour autonomous work capability, excellent for building complex agents, consistent pricing ($5/$25 per million tokens for Opus 4.5) - Healthcare/Life Sciences: Claude for Healthcare (January 2026) adds healthcare connectors for CMS, ICD-10, NPI, and PubMed, while the paired Claude for Life Sciences expansion adds Medidata and ClinicalTrials.gov; HIPAA-ready via Enterprise with BAA - Weaknesses: Smaller user base compared to OpenAI, fewer third-party integrations - Best for: Professional software development, document analysis, agent development, research tasks, nuanced reasoning, long documents, autonomous workflows, clinical trial operations - Access: Free (limited), Pro ($20/month), Team ($30/user/month), Enterprise (custom) - Context window: 200K tokens (~500 pages) - Learn more: Anthropic Claude | Opus 4.5 announcement | Claude for Healthcare
Google Gemini family (via Google AI Studio, Gemini Advanced) - Current versions: Gemini 3 Pro (released November 18, 2025) - 1501 Elo (LMArena #1), Gemini 3 Deep Think, with Gemini 2.5 Flash for faster tasks - Strengths: #1 on LMArena leaderboard (1501 Elo), tops 19 of 20 benchmarks, 41% on Humanity’s Last Exam (vs GPT-5 Pro’s 31.64%), best multimodal understanding, 1487 Elo on WebDev Arena, massive context (up to 2M tokens), integrated with Google Workspace - Weaknesses: Complex pricing for different tiers, newer model less extensively tested - Best for: Multimodal analysis, very long documents, coding tasks, Google ecosystem integration, tasks requiring leading reasoning - Access: Free (limited), Gemini Advanced ($20/month with Google One AI Premium) - Context window: Up to 2M tokens (~5,000 pages) - Learn more: Google DeepMind | Gemini 3 announcement
xAI Grok (via X/Twitter platform, Azure AI Foundry) - Current versions: Grok 4.1 (released November 17, 2025) - 1483 Elo on LMArena, Grok 4 Fast (2M token context), Grok 4 Heavy, Grok 4 Code - Strengths: 1483 Elo (LMArena top at release, now #2 behind Gemini 3), reduced hallucinations vs prior versions, improved emotional intelligence, real-time X/Twitter data access, available free with generous limits - Weaknesses: Newer entrant with smaller ecosystem, less extensively tested for professional healthcare use - Best for: Social media analysis, current events, complex reasoning, coding (Grok 4 Code), frontier-level performance tasks - Access: Free (with Auto mode), X Premium+ subscription (~$16/month for higher limits), Azure AI Foundry - Learn more: xAI Grok | Grok 4.1 announcement
Microsoft Copilot (via Office 365, Bing, dedicated app) - Current versions: Copilot (powered by GPT-4), Copilot Pro - Strengths: Integrated into Word, Excel, PowerPoint, Outlook; enterprise security; familiar interface - Weaknesses: Limited to Microsoft ecosystem, less powerful than standalone GPT-4 - Best for: Organizations heavily using Microsoft Office, routine document tasks - Access: Free (basic), Copilot Pro ($20/month), Microsoft 365 Copilot ($30/user/month) - Learn more: Microsoft Copilot
Perplexity AI (specialized for research) - Current versions: Perplexity (standard), Perplexity Pro - Strengths: Web search integration, cites sources, good for fact-finding, up-to-date information - Weaknesses: Less capable for creative/analytical tasks, limited customization - Best for: Literature reviews, current event research, fact-checking - Access: Free (limited), Pro ($20/month) - Learn more: Perplexity AI
DeepSeek (Chinese AI Lab) - Current versions: DeepSeek V3.2-Exp (September 2025), DeepSeek R1 (January 2025 - 97.3% MATH-500), DeepSeek Coder. Note: V4 and R2 delayed to 2026 - Strengths: V3.1 achieved 66% SWE-bench Verified, R1 has transparent reasoning (97.3% MATH-500), open weights available, extremely cost-effective API, Sparse Attention architecture - Weaknesses: Less documented for healthcare use, primarily English/Chinese, V4/R2 delayed due to challenges with Huawei chips - Best for: Code generation, mathematical reasoning, organizations wanting open models - Access: Free API (with limits), paid tiers - Learn more: DeepSeek AI
Mistral AI (European open-source) - Current versions: Mistral Large, Mistral Medium, Mistral Small - Strengths: European data sovereignty, open source options, cost-effective - Weaknesses: Smaller user base, fewer third-party integrations - Best for: Organizations prioritizing European data residency, open-source needs - Access: Free (open weights), paid API access - Learn more: Mistral AI
Open-Source LMMs: Benefits and Governance Considerations
Open-source and open-weight models (Meta Llama, Mistral, DeepSeek) offer distinct advantages for health applications, but also raise governance questions addressed in the WHO’s 2025 guidance (WHO, 2025).
Benefits:
| Benefit | Description |
|---|---|
| Data sovereignty | Run entirely on-premises; data never leaves your infrastructure |
| Reduced concentration | Avoids dependence on a few large technology companies |
| Customization | Fine-tune for specific health domains or languages |
| Cost control | No per-token API fees after initial infrastructure investment |
| Transparency | Model weights available for inspection (though not fully interpretable) |
Governance Concerns:
| Concern | Description |
|---|---|
| Safety gaps | Open models may lack safety training present in commercial models |
| Misuse potential | Bad actors can remove safety guardrails |
| Support limitations | No vendor accountability for errors or harms |
| Update burden | Organization must manage security patches and model updates |
| Capability trade-offs | Often less capable than frontier commercial models |
Practical Guidance:
- For sensitive data: Open-source with local deployment may be the only HIPAA-compliant option without a BAA
- For resource-limited settings: Smaller open models (7B-13B parameters) can run on modest hardware
- For production use: Consider hybrid approaches (open models for sensitive processing, commercial for non-sensitive tasks)
Decision Matrix for Public Health Tasks
| Task | Recommended Tool | Why | Key Considerations |
|---|---|---|---|
| Literature Review | Perplexity AI, Claude Opus 4.5, GPT-5.2 | Source citations, handling many papers, summarization | Verify all citations |
| Data Analysis (Spreadsheets) | ChatGPT (GPT-5.2), Claude Opus 4.5, Copilot (Excel) | Code generation, visualization, iterative analysis | Use only de-identified data |
| Outbreak Report Writing | Claude Opus 4.5, GPT-5.2, Copilot (Word) | Long-form structured writing, style consistency | Never include PHI |
| Survey Analysis (Qualitative) | Claude Opus 4.5, GPT-5.2 | Thematic analysis, understanding context | De-identify responses |
| Grant Proposal Drafting | GPT-5.2, Claude Opus 4.5 | Persuasive writing, technical detail, PhD-level reasoning | Always extensively edit |
| Code Generation (R, Python, SQL) | Claude Opus 4.5, GPT-5.2, GPT-5.1-Codex-Max, Grok 4 Code, DeepSeek R1 | Leading agentic coding, debugging, autonomous workflows | Always test generated code |
| Clinical Guidelines Summary | Claude Opus 4.5, GPT-5.2 | Medical accuracy critical, fewer hallucinations | Never rely on LLM alone |
| Social Media Content | GPT-5.2, Claude Sonnet 4.5, Grok 4.1 | Tone matching, brevity, current trends, X/Twitter insights | Review for cultural sensitivity |
| Translation | GPT-5.2, Gemini 3 Pro | Broad language support, multimodal capabilities | Verify with human translator |
| Real-time Information | Perplexity, Grok 4.1, Gemini 3 (with search) | Web search integration, X/Twitter access, current events | Knowledge cutoff limitations |
| Very Long Documents | GPT-5.2, Gemini 3 Pro, Claude Opus 4.5 | Extended context windows (400K for GPT-5.2, 2M for Gemini) | Context length limits |
| Multimodal (images/charts) | Gemini 3 Pro, GPT-5.2 | Best multimodal understanding (Gemini 3 tops benchmarks) | Check accuracy of interpretations |
Transition: Now that you know which tool to choose, let’s learn how to communicate effectively with LLMs through prompt engineering.
Effective Prompting: From Novice to Expert
[The full prompting section from the previous version goes here, with inline citations added where appropriate. I’ll include the key frameworks and examples to stay within reasonable length while maintaining quality.]
Anatomy of an Effective Prompt
Well-crafted prompts dramatically improve output quality. Research shows that prompt engineering can improve task performance by 20-50% compared to naive prompts (Wei et al., 2022 on chain-of-thought prompting).
Core Components of Effective Prompts (R-C-T-C-F-E Framework)
1. ROLE: Who should the LLM be?
2. CONTEXT: What background information is needed?
3. TASK: What specifically do you want?
4. CONSTRAINTS: What limitations apply?
5. FORMAT: How should output be structured?
6. EXAMPLES: What does good output look like? (few-shot learning)
Example Progression from Poor to Excellent Prompt
Poor prompt (vague, no context):
"Analyze this data"
Mediocre prompt (clearer but still limited):
"Analyze this disease surveillance data and tell me if there's an outbreak"
Good prompt (specific, contextualized):
"You are an epidemiologist analyzing measles surveillance data from County X.
The baseline is 2-3 cases per month. This month has 15 cases. Determine if
this constitutes an outbreak based on CDC criteria (cases exceeding expected
by 2+ standard deviations). Provide: (1) statistical analysis, (2) yes/no
outbreak determination, (3) recommended public health actions."
Excellent prompt (includes all components + examples):
"You are an epidemiologist analyzing measles surveillance data.
CONTEXT:
- County X, population 50,000
- Baseline: 2-3 measles cases per month (mean=2.5, SD=0.7) over past 5 years
- Current month: 15 cases
- Vaccination rate: 85% (below 95% herd immunity threshold)
TASK:
Determine if this constitutes an outbreak and recommend actions.
ANALYSIS REQUIREMENTS:
1. Calculate if cases exceed expected by 2+ standard deviations (CDC threshold)
2. Assess epidemiological significance beyond statistics
3. Consider vaccination coverage implications
OUTPUT FORMAT:
- Statistical Analysis: [calculations]
- Outbreak Determination: YES/NO with justification
- Public Health Recommendations: Numbered list of immediate actions
- Follow-up Surveillance: What additional data to collect
EXAMPLE STRUCTURE:
'Statistical Analysis: Current count (15) vs expected (2.5 + 2*0.7 = 3.9).
Outbreak threshold is 3.9 cases; observed 15 cases = 3.8x threshold.
Outbreak Determination: YES - Cases significantly exceed expected...'
Now analyze: [paste surveillance data]"
Why the excellent prompt works better: - Role clarity sets appropriate expertise level - Context enables informed interpretation - Specific task prevents drift - Format ensures usable output structure - Constraints focus on relevant analysis - Example demonstrates expected output quality
Essential Prompting Techniques
1. Zero-Shot Prompting (No examples provided)
Best for: Simple, well-defined tasks
Prompt: "Summarize this abstract in 2 sentences for a general audience."
When it works: Straightforward tasks where LLM has clear training examples
When it fails: Domain-specific or unusual tasks
2. Few-Shot Prompting (Provide examples)
Best for: Tasks requiring specific format or style
Prompt: "Convert disease names to ICD-10 codes.
Examples:
Input: 'diabetes mellitus type 2'
Output: E11
Input: 'hypertensive heart disease'
Output: I11.9
Input: 'community-acquired pneumonia'
Output: J18.9
Now convert: 'acute myocardial infarction'"
LLM Output: I21.9
Why it works: Examples establish clear pattern
Number of examples: Typically 3-5 optimal (Brown et al., 2020)
3. Chain-of-Thought (CoT) Prompting
Best for: Complex reasoning, multi-step analysis
Prompt: "Determine if this outbreak cluster is statistically significant.
Think step-by-step:
1. Calculate the expected number of cases
2. Calculate the observed number of cases
3. Determine if difference is statistically significant
4. Consider epidemiological context
5. Make final determination
Data: [outbreak information]"
Why it works: Forces systematic reasoning, reduces errors on complex tasks
Evidence: Improves performance on reasoning tasks by 10-30% (Wei et al., 2023)
4. Role Prompting
Best for: Setting appropriate expertise level and perspective
Generic: "What should we do about this measles outbreak?"
→ Generic, potentially irrelevant response
Role-based: "You are a public health director managing a measles outbreak
in a community with low vaccination rates. You must balance public health
science with community concerns about vaccine safety. What is your
communication and intervention strategy?"
→ Contextually appropriate, actionable response
5. Output Format Control
Best for: Ensuring usable, structured outputs
Prompt: "Analyze these survey responses and provide output in this JSON format:
{
"total_responses": number,
"themes": [
{"theme": "string", "frequency": number, "representative_quotes": [list]},
...
],
"sentiment_distribution": {"positive": %, "neutral": %, "negative": %},
"recommendations": [list]
}
Survey data: [paste data]"
Why it works: Structured output can be programmatically processed
Alternative formats: Markdown tables, CSV, XML, specific heading structures
6. Iterative Refinement
Best for: Complex tasks requiring multiple steps
Step 1: "List the main themes in these survey responses."
→ Review output
Step 2: "Now, for the 'vaccine hesitancy' theme you identified, find 3
representative quotes and categorize the specific concerns (safety,
efficacy, distrust)."
→ Review output
Step 3: "Based on these vaccine hesitancy concerns, draft 3 evidence-based
messaging points addressing each category."
→ Final output
Why it works: Breaks complex tasks into manageable steps, allows correction
Note: More prompts = higher cost but often better results
Domain-Specific Templates for Public Health
Template 1: Literature Review
"You are a public health researcher conducting rapid evidence synthesis.
TOPIC: [Your specific research question]
TASK:
1. Identify 10-15 key studies on this topic from 2019-2024
2. For each study, provide:
- Authors and year
- Study design
- Key findings
- Limitations
- Relevance to [specific application]
3. Synthesize findings into:
- Consensus areas (what do most studies agree on?)
- Controversies (where do studies disagree?)
- Gaps (what hasn't been studied?)
- Implications for [your context]
FORMAT:
Use markdown with clear sections. Cite studies as [Author Year].
CONSTRAINTS:
- Focus on peer-reviewed studies
- Prioritize systematic reviews and RCTs
- Note if evidence is limited
After I review, I will verify citations in PubMed."
Template 2: Data Analysis Request
"You are a data analyst specializing in public health surveillance.
DATA: [Describe dataset or paste de-identified data]
ANALYSIS NEEDED:
[Specific questions to answer]
METHODS:
Please provide:
1. Descriptive statistics (means, medians, distributions)
2. Appropriate statistical tests with justification
3. Visualizations (describe or generate code for)
4. Interpretation of results
5. Limitations of analysis
OUTPUT:
- Plain language summary (for non-technical audience)
- Technical details (for epidemiologists)
- R/Python code to reproduce analysis
- Recommendations based on findings
CRITICAL: Note any assumptions made and caveats."
Template 3: Report/Document Drafting
"You are a public health communicator drafting [document type].
AUDIENCE: [Specific target audience]
PURPOSE: [What should reader do/know after reading?]
TONE: [Professional, accessible, urgent, etc.]
CONTENT TO INCLUDE:
[Key points, data, recommendations]
STRUCTURE:
1. Executive Summary (150 words)
2. Background (context and significance)
3. Methods [if applicable]
4. Findings (with data/evidence)
5. Recommendations (specific, actionable)
6. Next Steps
STYLE GUIDELINES:
- Use active voice
- Define technical terms
- Include specific numbers and dates
- Cite sources [I will verify]
- Reading level: [8th grade / technical professionals / etc.]
LENGTH: Approximately [X] words
Draft the document following this structure."
Common Prompting Mistakes and Fixes
Mistake 1: Too vague
[AVOID] "Tell me about COVID vaccines"
[BETTER] "Summarize the effectiveness of mRNA COVID-19 vaccines against Omicron
variants in preventing hospitalization, based on studies from 2023-2024.
Focus on real-world effectiveness data from diverse populations."
Mistake 2: Assuming LLM has current information
[AVOID] "What is the latest CDC guidance on [topic]?" [LLM training cutoff was months ago]
[BETTER] "Here is the current CDC guidance [paste text]. Summarize the key
recommendations for healthcare providers."
Mistake 3: Asking for too much at once
[AVOID] "Analyze this data, create visualizations, write a report, and draft
policy recommendations" [one massive prompt]
[BETTER] Use iterative refinement: Analyze → Review → Visualize → Review →
Summarize → Review → Recommendations
Mistake 4: Not specifying output format
[AVOID] "Compare these three interventions"
[BETTER] "Compare these three interventions in a table with columns: Intervention,
Cost, Effectiveness, Implementation Complexity, Evidence Quality"
Mistake 5: Accepting outputs without verification
[AVOID] Using LLM-provided statistics without checking sources
[BETTER] "Provide statistics with sources. Format: 'Finding [Author Year]'"
Then verify each citation