AI in Public Health Emergency Operations
Emergency operations converts incomplete, changing information into coordinated action. AI can reduce clerical load and help teams find patterns, but the command problem is not an information-retrieval problem alone. Authority, objectives, resource constraints, legal duties, field intelligence, public communication, and consequences must remain visible in every AI-supported workflow.
- Place AI inside public health emergency operations center and incident-management functions.
- Separate decision support from command authority.
- Design source-linked, time-stamped, reviewable emergency workflows.
- Evaluate operational value, false reassurance, automation bias, and failure recovery.
- Integrate model incidents and near misses into exercises and after-action improvement.
Introduction
WHO defines an emergency operations centre as a physical or virtual location that coordinates information and resources for incident-management activities. Its PHEOC framework emphasizes a concept of operations, plans and procedures, information management, trained staff, exercises, and unity of effort across response agencies (WHO, 2015). A 2026 WHO framework for national public health agencies places emergency management alongside surveillance, laboratories, risk communication, clinical guidance, countermeasure deployment, legal authority, financing, and evidence use (WHO, 2026).
Incident management supplies a role and authority structure. Public health agencies may lead, support, or coordinate with broader all-hazards structures. FEMA describes all-hazard incident management teams as operating within the Incident Command System, with assigned command and general staff roles and written incident action plans for extended incidents (U.S. Fire Administration, 2026). The exact legal structure varies by jurisdiction. AI implementation must fit the authorized structure in force.
Where AI Fits
| Emergency function | Bounded AI support | Required human control |
|---|---|---|
| Situation awareness | Classify incoming reports, extract events, detect duplicates, draft summaries | Validate source, time, location, significance, and unresolved contradictions |
| Common operating picture | Join approved feeds, generate maps, flag missing data | Confirm data currency, geography, denominator, and release audience |
| Planning | Retrieve plans, compare objectives with assigned actions, draft planning products | Incident command sets objectives and approves the incident action plan |
| Logistics | Match requests with inventory, identify shortages, support routing | Logistics staff validate quantities, priority, availability, and transport constraints |
| Public information | Translate or adapt approved messages, monitor questions and rumors | Authorized public information staff approve content and timing |
| Documentation | Draft situation reports, decision logs, shift handoffs, and after-action timelines | Named staff verify material facts, decisions, and open actions |
The safest initial use cases are assistive and reversible. Retrieval from a controlled corpus, structured extraction, duplicate detection, draft generation, and quality checks leave accountable staff in control. Systems that initiate alerts, change resource priority, publish warnings, or alter operational status have a higher evidence and governance burden.
The Emergency AI Control Loop
2. Establish the approved information boundary
Use a controlled set of plans, procedures, data feeds, maps, directories, and primary sources. Record version and acquisition time. Separate verified facts, credible but unverified reports, model estimates, assumptions, and recommendations.
3. Require traceable outputs
Situation-report drafts and briefing products should link each factual assertion to its source, display a data timestamp, identify missing periods or jurisdictions, and surface contradictions. Unsupported text is removed, not polished.
4. Review against the incident objective
Accuracy alone is insufficient. Reviewers ask whether the output changes the current incident objective, resource assignment, protective action, communication need, or information requirement. Low-value output should not consume command attention.
5. Record the decision and correction path
Log the output version, reviewer, decision, time, source set, edits, and downstream use. Establish how an error is corrected in every product that received it.
6. Monitor, suspend, and fall back
Set stop conditions for stale feeds, unexplained output changes, source-link failure, privacy breach, increasing error, overload, or loss of qualified review staff. Maintain a tested non-AI workflow.
What AI Should Not Control
AI should not independently:
- activate or deactivate an emergency response;
- set incident objectives or command structure;
- issue evacuation, isolation, quarantine, treatment, or other protective-action orders;
- allocate scarce life-safety resources without authorized review;
- publish warnings or health guidance;
- infer facts from absent data or conceal uncertainty;
- replace field verification, laboratory confirmation, epidemiologic investigation, or legal review;
- retain sensitive incident data outside approved systems.
Generative fluency is especially hazardous under time pressure because a coherent briefing can hide stale or unsupported claims. The correct response to missing evidence is a visible gap, not a completed narrative.
Evaluation for Emergency Use
Evaluate the task and the workflow, not only the model. Useful measures include:
- time from source arrival to reviewed product;
- recall of high-priority signals and false-priority rate;
- unsupported factual claims per reviewed product;
- source-link and timestamp completeness;
- correction time and downstream correction coverage;
- reviewer workload and disagreement;
- performance during surge volume, connectivity loss, and staff turnover;
- subgroup, language, geography, and jurisdictional error;
- time to safe manual fallback;
- effect on the named operational decision.
Exercises should include stale feeds, conflicting reports, missing jurisdictions, malicious or malformed inputs, incorrect translations, unavailable reviewers, and sudden policy changes. A system that performs only under clean data and normal staffing is not ready for an emergency.
Incident and Near-Miss Learning
AI failures belong in the same improvement system as other operational failures. Record false reassurance, missed signals, incorrect prioritization, privacy exposure, unsupported text, source confusion, automation bias, and delayed fallback as incidents or near misses according to consequence. Preserve enough evidence to reconstruct the event, assign corrective actions, test the fix, and update exercises and procedures.
After-action review should distinguish technical cause from organizational cause. A model defect, weak source control, unclear authority, poor interface, excessive workload, inadequate training, and failure to act on a known warning require different remedies. Closure requires evidence that the corrective action works in a realistic exercise.
Implementation Checklist
- Map the tool to a PHEOC or incident-management function and named owner.
- Define authority, approved data, intended decision, and prohibited uses.
- Require source-linked, time-stamped, uncertainty-aware outputs.
- Validate with realistic spatial, temporal, and surge conditions.
- Establish approval, correction, suspension, and fallback procedures.
- Train every operational period on tool limits and escalation.
- Monitor error, workload, drift, privacy, and subgroup effects.
- Exercise failure scenarios before activation and after material changes.
- Capture incidents and near misses in the after-action system.
- Retire the tool when benefit no longer exceeds operational burden or risk.