Most exercises I have sat through were built to be finished and check a box or get some training hours in rather than to teach anybody anything, and the tell is always the same: the scenario was written after somebody picked the date, the objectives were written after the scenario, and the after action report listed communications as a finding for the fourth year running with no name beside it. This piece is about the design decisions that separate an exercise that finds real problems from a well catered demonstration, starting with objectives and ending with a corrective action somebody actually owns (another challenging part: getting a person to take ownership).

The exercise designed to be passed, and how to spot one

An exercise designed to be passed has a recognizable shape. The scenario stays inside the capability the agency already has, so the hazard is a two alarm structure fire rather than a six hour hazmat release that outlasts the on duty shift, and the demand curve never crosses the line where somebody has to say no or ask for help. Players are briefed in advance on what is expected of them, sometimes down to the order in which units will be requested. Nothing is taken away during play, which means the radio system works, the generator starts, the bridge is open and the person who knows the password is present and awake.

The evaluation follows the same logic. The evaluator is often the person who wrote the plan being tested, or the shift supervisor of the people being observed, and both of those arrangements produce a document that says the plan worked. Senior officials attend as observers and end up being briefed by the same staff who are supposed to be playing, which quietly converts an exercise into a presentation. When the hotwash begins, the first speaker is the ranking person in the room, everyone else calibrates to that, and the useful complaints go home in the parking lot.

I want to be blunt about why this happens, because it is usually not laziness. Exercises cost overtime, they are visible to elected officials and to the press, and a grant deliverable is due. Nobody gets promoted for producing an after action report that says the county cannot sustain an EOC past twelve hours. The incentive runs toward a clean document, and unless the person building the exercise decides in advance that finding problems is the deliverable, the design will drift toward comfort on its own without anybody choosing it.

The stairwell drills Rick Rescorla ran again and again for Morgan Stanley in the World Trade Center were resented by traders and by management, and the reason they were resented is the reason they worked, because they interrupted real business at inconvenient times and demanded real movement down real stairs. If nobody in your organization minds your exercises, that is worth examining before the next one is scheduled.

Objectives first, and no more than a handful

The Homeland Security Exercise and Evaluation Program, which FEMA maintains and revises periodically and which grant conditions frequently require for federally funded exercises, is explicit that objectives drive the exercise and the scenario serves the objectives. That ordering is the single most violated rule in the field. Somebody decides on a train derailment because the railroad offered a tank car, and the objectives get reverse engineered from the scenario two weeks before play, which guarantees they will be vague enough to be met.

A usable objective names a function, a performer and an observable result. “Test communications” is not an objective because no outcome would falsify it, whereas establishing a common tactical channel between fire, EMS and the sheriff’s office and confirming that all three can hear traffic from the incident commander within fifteen minutes of arrival is an objective, because at the end of the exercise somebody can say yes or no and point to the time. HSEEP describes objectives that are specific, measurable, achievable, relevant and time bound, and the measurable part is where most drafts fail.

Keep the number small. Three to five objectives is a working range for a tabletop and for most functional exercises, and even a full scale exercise involving several agencies is better served by five sharp objectives than by fifteen that each get thirty seconds of attention. Every objective you add takes evaluator capacity away from the others, because somebody has to be positioned where that objective is being performed, with a data collection form in hand, for the whole period in which it might occur.

Objectives should also come from somewhere defensible rather than from the planning team’s imagination. The obvious sources are your hazard identification and risk assessment, the corrective actions still open from the last two exercises, findings from real incidents in the last year, and any capability your jurisdiction has claimed on paper but never demonstrated. If your threat and hazard identification process names ice storms as your highest probability disruptive event and you have run three active shooter exercises in a row, the gap in your program is already documented in your own files.

The test for an objective

Write the objective, then ask what observable thing would prove it was not met. If you cannot describe the failure, the objective is unmeasurable and you will not learn anything from it, because whatever the players do on the day will satisfy the wording. Objectives that survive that test tend to contain a time, a named function or a specific piece of equipment.

Discussion based and operations based: picking the form that fits the objective

HSEEP divides exercises into two families. Discussion based exercises include seminars, workshops, tabletop exercises and games, and they move at the speed of conversation with players talking through what they would do rather than doing it. Operations based exercises include drills, functional exercises and full scale exercises, and they involve real action in real time, whether that means a single company practicing one skill, an EOC and dispatch center working a simulated event against a clock, or units and equipment actually deploying to the field.

The choice between them should follow the objective and nothing else. If your question is whether two agencies interpret a mutual aid agreement the same way, whether the public information officers know who approves a joint statement, or whether the county attorney and the finance director agree on emergency procurement authority, a tabletop will surface that in three hours and a full scale exercise will not surface it at all, because in the field those people are not in the conversation. If your question is whether the portable radios work in the basement of the hospital, no amount of discussion will answer it and somebody has to walk down there with a radio.

A workshop is the right tool more often than people think, because it produces a product rather than a performance. If you need a revised evacuation annex, a decision matrix for EOC activation, or an agreed list of which agency owns which task, put the right eight people in a room for a day with a facilitator and a draft, and come out with a document. Calling that a tabletop and pretending it evaluated anything is a common bit of credit taking that muddies your exercise record.

Cost drives the sequencing whether anybody admits it or not. A tabletop needs a room and a facilitator for a morning, while a full scale exercise costs overtime for every player, apparatus out of service, controllers, evaluators, moulage, logistics and often a year of planning, and it will waste all of that if the basic coordination questions have never been talked through. Running the discussion based work first has nothing to do with maturity models and everything to do with not spending a year’s exercise budget to discover that two departments define staging differently.

Building the scenario backward from what you need to see

Once the objectives are fixed, the scenario is a delivery mechanism, and you build it by asking what conditions would force each objective to be exercised. If the objective is transfer of command between agencies, the scenario needs an incident that starts as one discipline’s problem and becomes another’s, which a hazmat release inside a structure fire or a rescue that turns into a crime scene will do. If the objective is sustained operations past one shift, the scenario has to run long enough or jump forward in time far enough that relief has to be arranged, and that is why a time jump written into the master scenario events list is often more valuable than another explosion.

Realism matters in the conditions rather than in the theatrics. Players will forgive a simulated hazard and simulated victims, and they will disengage immediately if the resource picture is dishonest. The most common dishonesty is availability, where the scenario assumes every unit is in quarters, every neighboring department can send help and the state can deliver a resource in two hours because the exercise would stall otherwise. The fix is to write the resource constraints into the scenario explicitly and to give the simulation cell standing instructions about what is unavailable and why.

Degrade one thing on purpose in every exercise, and pick something the plan quietly assumes. Take out the primary radio channel or one repeater site, close a bridge, make the EOC’s internet circuit fail, take the incident commander out with a simulated medical emergency, or make the building you always use for a shelter unavailable because it is inside the evacuation zone. These removals are where the plan’s unwritten assumptions become visible, and they cost nothing to write into the master scenario events list.

The master scenario events list itself is the planning team’s control document, listing each scheduled event, the time it is delivered, who delivers it, who receives it, the objective it supports and the response the planning team expects. That last column is the one people skip, and it is what lets a controller recognize at the time that play has diverged from the design, rather than discovering it three weeks later when the evaluator notes do not add up.

Injects that test something and injects that are just narration

An inject is a piece of information or a demand introduced during play, and most injects in most exercises fail because they only describe the scenario. A message saying a second alarm has been struck and additional units are responding tells players something and asks nothing of them. An inject earns its place when it requires a decision, forces coordination with somebody outside the room, imposes a constraint, or delivers information that contradicts what players already believe.

The injects that reliably produce findings fall into a few categories. Resource denial works, where the mutual aid engine you requested is committed elsewhere and the simulation cell says so, and the players have to decide what they will stop doing. Information conflict works, where the first report says the plant released chlorine and the plant manager calls twenty minutes later to say it was anhydrous ammonia, which tests whether anybody reconciles the two before the protective action decision goes out. External demand works, where a television station calls the EOC director directly for a statement or a school superintendent asks whether to hold students, because those calls are real, they arrive at the worst moment and they are almost never in the plan. A legal or financial inject works too, where somebody has to authorize a purchase or a curfew and the person who can sign is unreachable.

Injects must be delivered the way the information would actually arrive. If you hand a player a printed card describing a radio transmission, you have tested nothing about the radio system, the dispatcher or the player’s ability to hear traffic in a loud room. Phone calls should come in by phone, radio traffic should come over the radio on the channel it would really use, walk in reports should be delivered by a person walking in, and social media rumors should appear on a screen somebody has to be watching. This is the cheapest single upgrade available to most exercise designers, since it costs planning time rather than money.

Leave room for unscheduled injects. Controllers should carry two or three contingency injects for the case where players solve the problem faster than expected, and they should also be authorized to hold or delete a scheduled inject when play has moved past it. A rigid list delivered on the clock regardless of what players are doing turns the exercise into a script reading, and the evaluation data you collect afterward describes a situation the players were never really in.

The mistake I see most often

Planning teams brief the players on the scenario in detail before play begins, usually because somebody worried that people would look unprepared. That single decision removes the assessment, the information gathering and the early decision making from the exercise, which is where most real failures live. Brief players on safety, on the rules, on the artificialities and on how to reach a controller, and let them find the incident the way they would find a real one.

Controllers, evaluators, the simulation cell and safety

Controllers and evaluators do different jobs and should not be the same people. A controller runs the exercise, delivers injects, keeps play on schedule, answers player questions about what is real, and can start, pause or stop the exercise. An evaluator observes and records against a specific objective, does not intervene, does not coach and does not answer questions, and leaves with data rather than opinions. Combining the roles produces an evaluator who is too busy running things to watch, and a controller who unconsciously steers players toward the outcome the evaluation form wants.

Evaluators need to be told what to watch for and given a structured place to write it down. HSEEP provides exercise evaluation guides tied to core capabilities and their associated tasks, and whatever form you use, it should list the objective, the tasks that make it up, space for times, and space for a narrative description of what actually happened rather than a score. What you want on that page is an account with timestamps: the request was made at 0914, it reached the EOC at 0931, nobody acknowledged it, and the requesting officer asked again at 0952. Numeric ratings without a narrative are nearly useless when you sit down to write findings.

Bring evaluators from outside the organizational chain when you can. A neighboring county’s emergency manager, a retired chief, a hospital emergency preparedness coordinator or somebody from the state agency will write down things your own staff will not, both because they have nothing at stake and because they do not share the assumptions that hide the failure. Offer the same service back to them, since that trade costs nothing and improves both programs.

The simulation cell exists so players can interact with a world that is not there. Staff it with people who know the real agencies they are simulating, give each one a role card with what that agency can and cannot provide and how long it takes, and run it out of a separate room with its own phone lines. Safety is a separate function with a separate person, and for any operations based exercise there must be at least one safety officer with unambiguous authority to halt play, a briefed medical plan, and a plainly spoken phrase that everyone knows means a genuine emergency rather than part of the scenario. The common convention is the phrase “real world emergency” repeated, and whatever you choose has to appear in the player briefing, in the controller handbook and on the badge every participant is wearing.

The hotwash, and the after action report that follows it

The hotwash happens immediately after play ends, while people still remember the order of events, and it is a facilitated debrief of participants rather than a critique delivered by the exercise staff. The facilitation matters more than anything else in the room. Start with the players who were furthest from the decisions, hold the senior officials until last, and ask questions about specific moments rather than general impressions, because asking how it went produces nothing while asking somebody to walk you through what happened between the evacuation order and the first bus arriving produces the finding.

Keep the hotwash short, thirty to forty five minutes for most exercises, and take notes visibly so people can see their point being captured. Separate the debrief for players from the debrief for controllers and evaluators, which should happen the same day and covers what went wrong with the exercise itself, since a badly delivered inject or a simulation cell that gave away the answer will otherwise contaminate the findings. Collect written participant feedback forms as well, because the person who will not speak in front of a deputy chief will write two sentences on paper that are worth the whole session.

The after action report is where honesty either survives or dies. A useful report describes what happened against each objective with times, states plainly whether the objective was met, and separates three different kinds of problem that routinely get mixed together: the exercise artificiality, which is an artifact of how the exercise was built and is not a real gap; the individual performance issue, which belongs in training records rather than in a published report; and the systemic gap, which is a shortfall in plans, equipment, staffing, agreements or authorities that would recur on a real incident. Only the third kind belongs in the improvement plan.

Set your own deadline for the draft and hold it, because a report that circulates ninety days later arrives after everybody has moved on and after the budget cycle has closed. Send the draft to participating agencies for factual correction with a short return window, distinguish factual corrections from requests to soften a finding, and accept the first while declining the second in writing.

Corrective actions with a name, a date and a funding source

An improvement plan is a table of corrective actions, and each line needs four things to be real: a specific action described concretely enough that somebody could do it, the capability or objective it addresses, a named individual who owns it, and a completion date. A named individual means a person or at minimum a position, not an agency, because an action assigned to the fire department is assigned to nobody in particular. If the action costs money, the line should say where the money comes from, whether that is the operating budget, a grant cycle, or a request that has to go to the governing body with a date by which it must be submitted.

Split findings by what it would take to fix them. Some corrective actions are procedural and can be closed in a month by rewriting a page of the plan or changing a dispatch protocol, some are training actions that go into the next quarter’s schedule, and some require equipment or staffing and will take a budget cycle or a grant award. The last group needs an interim mitigation written beside it, because the gap is open in the meantime and the organization should know what it is doing about it, and a plan that treats a radio template change and a new frequency license as the same kind of item will leave both of them open.

Track them somewhere that a person other than the emergency manager can see. A simple list reviewed as a standing item at a meeting that already happens, whether that is the monthly department head meeting or a quarterly LEPC session, is worth more than an elaborate tracking system nobody opens. The discipline that matters is closure, which means the owner reports what was done and someone verifies it, and verification for most actions means testing the fix in the next exercise rather than accepting an email that says it is handled.

The hardest cases are the findings that recur. When interoperable communications or resource requesting appears in your after action reports three years running, the useful fact is no longer the finding itself but that the corrective action was never owned or never funded, and that belongs in the next report in those words and in front of whoever controls the budget. The report of the U.S. House Select Bipartisan Committee investigating the preparation for and response to Hurricane Katrina, published in 2006 under the title A Failure of Initiative, examined the 2004 Hurricane Pam exercise for southeast Louisiana and found that follow on planning work identified through that exercise remained incomplete when the storm arrived, which is the most expensive documented case available of exercise findings that nobody carried to completion.

Close the loop in the next exercise

Write at least one objective in each exercise that directly tests a corrective action from the previous one. This is the only verification method that does not depend on somebody’s assurance, it costs nothing to design, and it changes the internal politics of the improvement plan, because owners know their item will be exercised in front of the same people who saw it fail.

What to do at your agency

  • Pull the last two after action reports this week and list every corrective action that has no named owner or no completion date, then take that list to your next department head meeting and assign a person and a date to each line.
  • Before you pick a scenario for your next exercise, write three to five objectives on one page, and for each one write the observable evidence that would show it was not met, discarding any objective you cannot fail.
  • Ask the emergency manager in a neighboring county to supply two evaluators for your next exercise and offer the same in return, with an agreement in writing that their written observations go into your report unedited except for factual corrections.
  • Add one degradation to your next exercise design, specifically the loss of your primary radio channel, your EOC internet circuit or your usual shelter facility, and write it into the master scenario events list with a delivery time.
  • Convert every inject on your current draft list to the medium it would really arrive in, so that radio traffic goes over the radio, phone calls come by phone and nothing is handed to a player on a card describing something they should have heard.
  • Name a safety officer for any operations based exercise, put the real world emergency phrase in the player briefing and on the participant badge, and confirm that the safety officer has unqualified authority to stop play.
  • Add a standing five minute item called open corrective actions to a meeting that already exists on your calendar, and have the owner of each open item report status out loud rather than by email.

Takeaways

  • Objectives come before the scenario, and an objective that cannot be observed to fail will not teach you anything regardless of how well the exercise runs.
  • Discussion based exercises answer questions about agreements, authorities and interpretation, while operations based exercises answer questions about whether equipment, people and procedures actually work in the field, and the objective determines which one you need.
  • An inject that only describes the scenario is narration, and the injects that produce findings deny a resource, contradict earlier information, impose an external demand or require a decision somebody has to be found to make.
  • Controllers run the exercise and evaluators observe it, the roles should be held by different people, and evaluators from outside your chain of command will record things your own staff will not.
  • Deliver injects through the medium the information would really arrive in, and degrade at least one thing the plan quietly assumes will be available.
  • The hotwash belongs to the participants, works best when junior people speak first and senior officials speak last, and should be paired with written feedback forms and a separate controller and evaluator debrief the same day.
  • An after action report should separate exercise artificialities and individual performance issues from systemic gaps, and only the systemic gaps belong in the improvement plan.
  • A corrective action is real only when it has a named owner, a completion date, a funding source where money is involved, and verification in a later exercise, and a finding that recurs for three years is evidence that no one ever owned it.
Questions or a different view?

Reach me through the contact page. I read every message.