IT Management Antipatterns - Operations

Operations become ineffective when teams are trapped in reactive work and recurring problems are never addressed at their root. This section explores antipatterns that normalize constant firefighting, increase operational risk, and leave little time for automation, prevention, and continuous improvement.

Articles in this series

The Firefighter Factory

The Firefighter Factory is the organizational antipattern in which recurring incidents, outages, urgent requests, and operational crises consume so much attention that teams have insufficient time to eliminate their underlying causes. The organization has become very good at putting out fires while continuing to manufacture the conditions that create them.

How to recognize it

Typical symptoms include:

  • The same categories of incidents repeatedly occur.
  • Engineers spend a substantial proportion of their time responding to urgent operational problems.
  • A small number of experienced employees are repeatedly called to resolve crises.
  • Major incidents generate immediate attention, while preventive work struggles to obtain priority.
  • Post-incident reviews identify corrective actions that are never completed.
  • Temporary fixes remain in production for months or years.
  • Technical debt is acknowledged but continually postponed.
  • Monitoring detects failures but little effort is invested in eliminating their causes.
  • Teams regularly work outside normal hours to restore services.
  • Emergency changes are common.
  • Planned engineering work is frequently interrupted by incidents.
  • Reliability improvements compete with feature delivery and consistently lose.
  • The organization praises people who recover systems at 3 a.m. more visibly than people who prevent incidents from occurring.
  • Manual operational procedures continue despite repeated human errors.
  • Known capacity, resilience, security, or maintenance problems remain unresolved until they cause an outage.
  • Incident counts may decrease temporarily, but the same structural weaknesses remain.
  • "We'll fix it properly later" is a common conclusion.

Negative effects on the organization

Incidents are inevitable in complex IT systems. No mature organization should expect to eliminate operational incidents entirely. Strong incident-response capabilities are therefore essential. Teams need monitoring, alerting, escalation procedures, communication mechanisms, technical expertise, and the ability to restore service quickly. The antipattern is not firefighting. It is allowing firefighting to substitute for engineering.

The Firefighter Factory often creates organizational heroes. A critical system fails late at night. An experienced engineer receives the call. They immediately recognize the symptoms. They connect to the system, execute a sequence of commands that few other people understand, and restore service within minutes. The next morning, management thanks them for their dedication.

But there is a deeper question: why did the organization need a hero in the first place? Perhaps the recovery procedure could have been automated. Perhaps the knowledge could have been distributed across the team. Perhaps monitoring could have detected the problem earlier. Perhaps the architecture could tolerate the failure automatically. Perhaps the underlying defect has already caused five previous incidents.

If the organization celebrates recovery without investing in prevention, it unintentionally creates an incentive structure in which heroic intervention is more visible than quiet reliability. The engineer who spends three weeks redesigning the system so nobody ever needs to wake up at 3 a.m. may receive very little.

Prevention has an unusual management problem: success is difficult to see. The absence of failure can therefore make preventive engineering appear less valuable than reactive work, even though prevention may have saved far more money and disruption. Good management needs to recognize incidents that did not happen, not only successful responses to incidents that did.

The Firefighter Factory can become self-sustaining. The cycle looks like this:

Technical debt → incidents → emergency work → preventive work postponed → more technical debt → more incidents

The organization is consuming the capacity required to escape the problem simply to survive the consequences of the problem. Incident response correctly prioritizes restoring service over architectural elegance. During an outage, a temporary workaround can be exactly the right decision. The problem begins after the incident. A temporary workaround creates an implicit obligation: someone must return later and implement the durable solution. If that work is not prioritized, the temporary solution gradually becomes permanent. The organization has converted an emergency decision into technical debt without consciously deciding to do so.

Many organizations conduct incident reviews. This is useful, but the existence of a postmortem does not itself demonstrate organizational learning. The organization has documented the lesson without learning it. Google's Site Reliability Engineering guidance emphasizes that postmortems should produce actionable information and support a culture of learning rather than blame. A useful postmortem should therefore change something.

Another danger is treating incident investigation as a search for one defective component or one human mistake. Complex failures frequently require several conditions to align. This point also connects the Firefighter Factory with the Blame Machine antipattern: blame focuses attention on the individual who triggered the incident, while systemic analysis examines why the organization allowed an ordinary human error to produce significant consequences.

The direct cost of the Firefighter Factory is obvious, the indirect cost can be larger. Planned engineering becomes unpredictable because teams cannot know how much capacity will be consumed by emergencies. Project estimates lose reliability. Technical improvements are repeatedly postponed. Senior engineers become operational bottlenecks because they possess critical troubleshooting knowledge. The organization also becomes increasingly resistant to change. If every deployment carries significant operational risk, the natural reaction is to deploy less frequently and introduce more approval processes. Eventually the organization may respond to unreliable engineering by making engineering even harder.

Negative effects on people

Continuous firefighting has a substantial human cost. Occasional emergencies can be stimulating, permanent emergencies are exhausting. Evenings and weekends become unpredictable. Experienced employees become permanently reachable because too few people understand critical systems. People spend their time solving the same categories of problems instead of improving systems or developing new skills.

Eventually frustration appears: "We already told management this would happen."

When engineers repeatedly identify risks but the organization refuses to prioritize remediation until failure occurs, incidents stop being unexpected technical events. They become management decisions whose consequences were deferred.

A capable operations team can keep a badly managed environment functioning for surprisingly long periods. From outside the team, everything may appear functional. This creates a dangerous illusion. Management sees services continuing to operate and concludes that additional investment is unnecessary. But reliability is being provided through human effort rather than system design. When those people leave, the hidden fragility becomes visible.

The Healthy Pattern: Engineer the Fire Away

The healthy alternative is not to eliminate incident response. It is to use incidents as inputs into continuous engineering improvement. Every significant incident should create an opportunity to make the system easier to operate, harder to break, or faster to recover. Corrective actions should have explicit owners and priorities. Recurring incident categories should receive greater attention than isolated failures. Manual recovery procedures should be automated where the investment is justified. Critical operational knowledge should be distributed rather than concentrated in individual heroes. Monitoring should evolve from simply detecting failures toward detecting the conditions that precede them. Architectures should be designed so that ordinary component failures do not automatically become customer-facing incidents.

Teams also need protected capacity for reliability work. If every available engineering hour is allocated to feature delivery, preventive work will always lose to immediate business demand. Leadership must therefore recognize reliability, maintainability, automation, and technical-debt reduction as legitimate engineering outcomes rather than activities performed only when "there is spare time."

And incentives matter. Recognize the engineer who resolves a difficult production outage. But also recognize the engineer who eliminates the recurring outage permanently.

The principle is:

Reward people for putting out fires, but reward the organization for making sure the same fires do not return.

A mature IT organization is not one that never experiences incidents. It is one that becomes progressively harder to surprise with the same incident twice.

The Blame Machine

The Blame Machine is the organizational antipattern in which failures trigger a search for the person responsible rather than a search for the conditions that allowed the failure to occur.

A production outage happens. The first question becomes: "Who did this?"

Once a person has been identified, the organization feels that it has found the cause. The investigation effectively stops there. But identifying the person who triggered an event is not necessarily the same as identifying why the organization failed. The Blame Machine converts complex organizational and technical failures into simple stories about individual mistakes.

How to recognize it

Typical symptoms include:

  • Incident investigations quickly focus on identifying who made the mistake.
  • Managers ask "Who changed this?" before asking what happened.
  • Individual names feature prominently in discussions of failures.
  • Employees are publicly criticized for operational mistakes.
  • Postmortems describe human error as the root cause.
  • People become reluctant to admit mistakes.
  • Engineers avoid volunteering information that might associate them with an incident.
  • Employees quietly fix problems without reporting them.
  • People avoid making difficult decisions because a wrong decision could later be used against them.
  • Meetings about failures feel more like interrogations than investigations.
  • Email and chat histories are searched primarily to establish who authorized a decision.
  • Teams spend substantial effort demonstrating that a problem originated somewhere else.
  • Departments blame each other for incidents.
  • Managers protect their own teams by transferring responsibility to another team.
  • People copy managers on emails to establish evidence that they raised concerns.
  • Known problems remain hidden until they become impossible to conceal.
  • The same categories of failures recur despite people being reprimanded for previous incidents.

A particularly revealing question is: do people feel safer reporting a mistake immediately, or trying to hide it and hoping nobody notices?

Negative effects on the organization

Suppose an engineer accidentally deletes an important production resource. It is easy to conclude that the root cause is an engineer error. But this explains very little. Why could a single command delete a critical resource? Why did the engineer have that level of access? Why was the operation not protected by additional safeguards? Why was the resource being modified manually? Was the engineer working under time pressure? Was the procedure clear?

Each question moves the investigation away from the individual event and toward the system that transformed an ordinary human mistake into a significant failure. People will eventually make mistakes. A resilient organization assumes this and designs systems accordingly.

Calling a person the root cause creates the illusion that removing or correcting that person solves the problem. Replace the engineer and another engineer may eventually make the same mistake. If a predictable human mistake can produce catastrophic consequences, the organization should ask why the system provides so little protection against that mistake.

The more useful question is:

"What made this error possible, and what made its consequences so severe?"

This does not remove individual responsibility. It expands the analysis beyond it.

The Blame Machine creates one of the most dangerous effects in an IT organization: it reduces the flow of bad news. If the person who reports an issue becomes the focus of uncomfortable questioning, public criticism, or career consequences, employees adapt. A small operational issue that could have been resolved immediately can therefore grow into a major incident because employees hesitate to raise it.

The organization believes that blame creates accountability. In reality, it may create information latency. And information latency is dangerous.

Google's SRE guidance emphasizes blameless postmortems precisely because effective incident analysis depends on people being able to provide accurate information about what happened without the investigation becoming a search for punishment. DORA's work on generative organizational culture similarly emphasizes high cooperation, sharing risks, and treating failures as opportunities for inquiry rather than immediately assigning blame.

The effects extend beyond operational incidents. Suppose an engineer believes that a major project is in trouble. The architecture is not working as expected. The deadline appears unrealistic. A security risk has been discovered. In a healthy organization, raising that information early is valuable. Management receives more time to respond. In a blame-oriented culture, however, delivering bad news can be personally risky. Employees therefore have an incentive to soften negative information. By the time reality becomes impossible to hide, management has lost the time it needed to respond.

When people expect retrospective blame, they naturally begin protecting themselves. Decisions become documented primarily for defensive purposes. Emails contain long explanations of responsibility. Managers are copied unnecessarily. Teams insist on formal approvals before taking action. Employees become reluctant to make decisions outside narrowly defined responsibilities. This is enormously expensive. Energy that could have been spent solving technical and business problems is redirected toward organizational self-protection.

One objection to blameless cultures is that they supposedly allow poor performance to go unchallenged. That is a misunderstanding. A blameless incident review is not a promise that behaviour can never have consequences. It is a commitment to understand what happened before deciding what the consequences should be.

Before assigning individual responsibility, management should ask several questions. Was the expectation clear? Was the person appropriately trained? Did they have the information required to make the decision? Did they have sufficient time and resources? Was the procedure realistic? Did the system provide appropriate safeguards? Would another competent person in the same situation reasonably have made a similar decision?

These questions help distinguish individual failure from system failure. If a well-trained engineer following normal organizational practices can make one mistake and take down the entire production environment, the engineer may have triggered the incident, but the organization designed the conditions that made the incident possible.

Blame tends to flow downward. This can distort incident analysis. Suppose engineers previously identified a resilience weakness. They proposed a solution. Management decided that the improvement was too expensive or that feature delivery had higher priority. Months later, the predicted outage occurs. A poor investigation asks: "Why did operations allow the service to fail?" A better investigation also asks: "Why did the organization previously decide to accept this risk?"

Finding somebody to blame can make management feel that the problem has been solved. But if the underlying system remains unchanged, the risk remains. This creates false remediation. In reality, nothing may have changed about the technical or organizational conditions that produced the failure. Punishment can therefore become a substitute for engineering improvement.

Negative effects on people

The personal consequences of the Blame Machine can be significant.

People become afraid of mistakes. Experimentation declines. Employees avoid unfamiliar responsibilities. Junior engineers hesitate to act without senior approval. Senior engineers become conservative. This damages learning. Organizations that punish every unsuccessful outcome teach people to avoid taking risks.

Eventually the safest career strategy becomes: do exactly what you were told, make as few independent decisions as possible, and ensure somebody else approved everything. It does not produce a high-performing engineering organization.

The Healthy Pattern: Accountability Without Blame

The healthy alternative is a culture that separates learning from punishment while preserving genuine accountability. When a failure occurs, the first objective should be understanding. Establish the timeline. Understand what people knew at each point. Examine the decisions that appeared reasonable at the time rather than judging them only with hindsight. Identify technical conditions, organizational constraints, process weaknesses, communication failures, and management decisions that contributed to the outcome. When genuine individual performance or conduct issues exist, they should still be addressed. But they should be addressed through appropriate performance-management or disciplinary processes rather than turning every technical investigation into a trial.

This distinction protects both sides. Employees can participate honestly in incident analysis without assuming that every admission will be used against them. Management retains the ability to address genuine negligence, misconduct, or persistent performance problems when supported by evidence. Most importantly, leaders need to model the behaviour they expect.

The principle is:

Hold people accountable for their behaviour, but investigate failures as properties of the system.