Representative interview topic

General Interview: How do you communicate during a live incident?

GeneralMedium
Offer.cc Editorial TeamPublished Updated

Question

A core service is failing and the technical team has not found the root cause. How would you communicate with internal teams, customers, and executives? What would you do when channels diverge?

Question and context

This question tests judgment, writing, and coordination during an incident, not a retrospective. Assume impact may still be growing, an incident manager, technical lead, and communications owner exist, and the root cause is unknown. You need an initial notice, a cadence of updates, mitigation and recovery messages, and a correction path.

It fits technical leads, SREs, platform engineers, customer engineers, and general roles that coordinate across teams. Separate the internal workflow from the public status page. Do not turn an unverified hypothesis into a conclusion or leave affected people without information while waiting for root cause.

What the interviewer is evaluating

A strong answer first sets severity and audiences from business impact, then assigns a single source of truth, a communications owner, and a cadence. It states what is known, unknown, in progress, and when the next update will arrive. It addresses security, data-loss, and compliance escalation, plus channel consistency, stale messages, recovery, and a later post-incident review.

Clarifying questions to ask

  • Which users, regions, functions, and data are affected? Is the scope full, partial, or unknown?
  • Could this involve security, privacy, data loss, or a regulatory duty? That changes approval and notice paths.
  • Who are the incident manager, technical lead, and communications owner? Who can approve external wording?
  • Which internal and external channels exist, and can customers access a status page or targeted notice?
  • What update cadence is expected? Should you still send a “still investigating” update without new evidence?

A 30-second answer framework

“I first establish impact, severity, and any security or data risk, then assign an incident manager, technical lead, and communications owner around one source of truth. I acknowledge the issue quickly with known impact, current investigation or mitigation, and the next update time. Internal messages include roles, the working channel, and escalation path; external messages contain only customer-relevant facts and actions. I keep the cadence even without a new conclusion. Every channel uses the same state and incident ID, and recovery messages include confirmation, impact summary, and the post-incident review path.”

Step-by-step answer

Step 1: Assess impact and communication boundaries

Establish who is affected, what they see, when it started, whether it is expanding, and whether security or data risk exists. Severity determines 24/7 response, executive escalation, legal review, or privacy involvement. Label unknowns as unknown; do not replace evidence with the most optimistic or alarming assumption.

Step 2: Set roles and one source of truth

The incident manager owns priorities and decisions, the technical lead owns hypotheses, mitigation, and evidence, and the communications owner owns internal and external messages. Everyone shares one incident ID, state document, and timeline; chat, status page, email, and tickets are distribution channels. Communications cannot change technical facts, and engineers should not publish guesses outside the approval path.

Step 3: Send the initial notice

Once the incident is reasonably confirmed, publish a short message with the affected product, current symptom, investigation state, temporary user action, and next update. Internal messages can include severity, on-call channel, owner, and escalation path; external wording should avoid internal jargon and unverified causes. If security or data impact is unknown, say it is being assessed and use the separate security-notification process when required.

Step 4: Set an update cadence and template

Choose a cadence that matches severity, such as every 30 minutes; update earlier when evidence changes. Keep each message to current state, user impact, actions underway, and next update time. Even without new evidence, say that impact remains under investigation or mitigation. Internal and external detail can differ, but state, time, and impact must agree.

Step 5: Resolve channel conflicts and corrections

When messages conflict, stop copying the old wording and return to the source of truth for time, impact, and state. The communications owner publishes a correction that names what changed instead of silently overwriting history. Use the same incident ID and state transitions everywhere; mark stale pages resolved or link them to the final summary.

Step 6: Communicate mitigation, recovery, and residual risk

Mitigation is not full recovery. Distinguish the mitigation action, users who may remain affected, data-consistency checks, and the next validation. After recovery, state the confirmation time, impact window, whether users should retry or sign in again, and whether a post-incident review will follow. If scope is only discovered later, send a targeted notice rather than presenting an estimate as final.

Step 7: Turn communication into an improvement loop

Review time to first notice, cadence adherence, channel consistency, support volume, and which information caused confusion. Create owned, testable actions for templates, status-page access, on-call roles, and escalation paths. Improve communication alongside the technical review, but do not write an in-progress public update as a post-incident root-cause conclusion.

Example of a strong answer

“I would first establish affected users, regions, functions, start time, and data risk, then set severity from business impact. The incident manager owns priorities, the technical lead maintains hypotheses and evidence, and the communications owner maintains messages; all three share one incident ID, state document, and timeline.

After confirming the issue, I would send an initial notice quickly: what is affected, whether we are investigating or mitigating, what users can do, and when the next update will arrive. Internal wording adds severity, working channel, and escalation contact; external wording contains customer-relevant facts and does not guess at root cause. I would keep the promised cadence even if there is no new conclusion.

If the status page, email, and support response diverge, I would use the source of truth and have the communications owner publish a correction with the incident ID. I would distinguish mitigation, recovery, and residual risk, then publish the impact window, retry guidance, and post-incident review path. Finally, I would review cadence, reach, consistency, and support load and assign owned improvements.”

Common mistakes

  • Waiting for root cause before notifying → users cannot judge risk → state known impact, action, and next update first.
  • Writing separately for every channel → state and time conflict → maintain one source of truth and incident ID.
  • Copying internal jargon externally → customers do not know what to do → adapt language by audience while preserving facts.
  • Saying only “mitigated” → users assume complete recovery → separate mitigation, recovery, validation, and residual risk.
  • Offering no cadence → silence looks like loss of control → commit to a cadence and update without new conclusions.
  • Silently editing a wrong message → trust and auditability decline → publish a timestamped correction and retain history.

Follow-up questions and answers

Follow-up 1: The cause is unknown and customers ask whether data is safe. What do you say?

State completed checks and ongoing assessment, such as “We have no evidence of data exposure so far, and the security team is still investigating.” Do not turn “not found yet” into “definitely none.” Use the security and privacy notice path if its threshold is met.

Follow-up 2: Should you update when there is no progress?

Yes. Follow the commitment and state whether impact changed, which hypothesis is being tested, what validation is pending, and when the next update will arrive. If the cadence changes, explain the new cadence and why.

Follow-up 3: A subset of customers is affected. Will a public status page create panic?

Choose based on scope and whether targeted channels can reach everyone quickly. If the affected set is unknown or normal channels are unavailable, a public page can supplement targeted notices. Describe affected functions and user-visible symptoms without unnecessary internal detail.

Follow-up 4: When should the post-incident review be published?

Publish recovery confirmation and the known impact window first. Publish the review when evidence and impact analysis are sufficient. If customers need earlier context, provide an uncertainty-labeled preliminary summary and add a validated timeline and actions later.

Public sources

Related questions