Anyone who checks alerts every morning while on vacation is not necessarily demonstrating exceptional dedication. It may instead indicate that critical processes depend too heavily on a single person. That creates a risk for the entire organization.
Constant availability has two major side effects. First, IT leaders miss out on genuine downtime, leaving less room for strategic work during the regular workweek. Second, teams can become accustomed to waiting for top-level approval before making important decisions. If the CIO has to approve every major outage response, unnecessary coordination loops can delay recovery.
The good news is that this can be changed. An IT organization that runs reliably for two weeks without its leader is not a matter of luck. It is the result of good preparation. The following sections explain how to build that resilience.
Maturity Over Heroics
A useful framework is the Capability Maturity Model Integration, or CMMI. It defines five maturity levels: Initial, Managed, Defined, Quantitatively Managed, and Optimizing.
At lower maturity levels, processes may depend more heavily on individual people, project-specific procedures, and manual coordination. As process maturity increases, responsibilities, workflows, and standards become more clearly defined.
Level 3, Defined, is an important step. At this stage, processes are defined and documented across the organization. This can help reduce dependencies on individual people. Teams know which processes apply and how to respond in typical situations.
For an IT organization that needs to keep operating when its leaders are absent, this degree of standardization can already provide a solid foundation. Levels 4 and 5 go further by adding quantitative management and continuous optimization, among other elements.
Error Budgets: Let the Numbers Decide, Not the Boss
Site Reliability Engineering (SRE) provides a tool that can make discussions about availability and changes more objective: the error budget. It defines how much deviation from a specified service target is acceptable within a given period.
For an availability-based Service Level Objective (SLO), the permitted downtime can be calculated as follows:
Allowed downtime = time period × (1 − SLO)
With a target availability of 99.9 percent, that amounts to roughly 43 minutes over 30 days.
It is important to distinguish between a Service Level Objective (SLO) and a Service Level Agreement (SLA). An SLO defines a target for service quality. An SLA, by contrast, is a contractual agreement with customers and may specify consequences if guaranteed service levels are not met.
In practice, an internal SLO can deliberately be set more aggressively than the contractual SLA threshold. This allows an organization to respond before a contractual commitment is breached.
The key is the policy behind the number. As long as sufficient error budget remains, teams can deploy and make changes according to predefined rules. Once the budget is exhausted, a defined error-budget policy takes effect. For example, new features or nonessential changes may be paused while stability takes priority.
The CIO no longer has to personally weigh speed against reliability for every individual case.
Replace Instead of Repair
Another building block can be immutable infrastructure. With this approach, running systems are generally not modified manually after deployment. Infrastructure is typically defined as code and provisioned automatically.
Instead of painstakingly repairing a faulty server by hand, the server can, where the architecture allows it, be replaced with a newly created instance based on a defined configuration. This reduces manual changes to production systems to exceptional cases.
Depending on the environment, such a replacement can happen very quickly. At the same time, systems become more reproducible. New instances are created from the same defined configuration instead of gradually diverging over the years through individual manual changes.
However, immutable infrastructure is not a cure-all. A new server instance alone will not protect against application code defects, corrupted data, configuration errors, or an outage at an external service provider.
That still requires reliable backups, monitoring, tested recovery procedures, runbooks, and people who know how to use them.
When the Call Is Really Necessary
Letting go does not mean being unreachable when something goes wrong. It means defining in advance when escalation to IT leadership is actually necessary. In some organizations, incidents assigned the highest internal priority automatically trigger an escalation to IT leadership, even when technical redundancy has already done its job.
A better approach is to base escalation paths not only on technical priority but also on potential business impact.
For example, if a core switch fails and redundant infrastructure takes over as designed, the incident can initially remain with the responsible incident team.
Escalation to senior leadership may become appropriate or necessary if, for example, one of the following occurs:
- Data loss exceeds the accepted threshold: The Recovery Point Objective (RPO) of a critical system is exceeded.
- Major security incident: Under Germany’s BSIG, statutory reporting obligations may apply to significant security incidents involving particularly important and important entities. An initial notification generally must be submitted without undue delay and no later than 24 hours after becoming aware of the incident. A further notification with an initial assessment of the incident follows within 72 hours.
- Personal data breach: If personal data is affected by a data breach, the reporting obligation under Article 33 of the EU General Data Protection Regulation (GDPR) may also apply. The competent supervisory authority must generally be notified without undue delay and, where feasible, within 72 hours of becoming aware of the breach if it is likely to result in a risk to the rights and freedoms of individuals.
- Loss of a site: A fire, flood, or comparable event renders a site unavailable and the business continuity plan must be activated.
The call should therefore primarily come when legal, financial, or strategic decisions are actually required. Technical problems that remain within predefined boundaries, by contrast, can be handled independently by the responsible team.
Before and After
The difference between a person-dependent IT organization and a well-organized one becomes particularly clear in everyday operations.
Who makes the decisions during an incident?
In a highly person-dependent IT organization, important decisions are often escalated up the chain. In a process-oriented IT organization, teams can make decisions within clearly defined authority and rely on runbooks and established procedures.
What does the documentation look like?
In the first scenario, much of the knowledge resides in the heads of individual employees or leaders. In the second, key processes, responsibilities, infrastructure, and recovery procedures are documented in a way that others can follow.
When is IT leadership called?
In a person-dependent organization, a major technical disruption may trigger escalation on its own. In a well-prepared IT organization, leadership is primarily involved when defined business, legal, or security-related thresholds are crossed.
What happens when the boss is on vacation?
Without clear responsibilities, uncertainty and additional coordination loops can arise. When responsibilities and decision-making authority have been defined in advance, normal operations can continue even during an extended absence.
How quickly are problems resolved?
If decisions require approval from individual leaders, recovery can be delayed. Automation, runbooks, and clear responsibilities, by contrast, can help teams respond to disruptions more quickly.
What does IT leadership focus on?
In a highly person-dependent organization, leaders are regularly pulled into operational issues. An IT organization that can operate independently creates more room for strategy, architecture, and the long-term development of IT.
Five Steps to Take Before Vacation
None of this should be built during the week before departure. These structures should be established as part of normal operations and reviewed again before extended absences.
1. Assign a deputy with real authority. Appoint a backup and give that person the necessary decision-making and approval rights. A deputy without sufficient authority can become just as much of a bottleneck as an absent boss.
2. Review and strengthen runbooks. Identify the most important risk scenarios and document the required steps. A qualified team member should be able to carry out a recovery without relying on knowledge that exists only in the CIO’s head.
3. Grant admin rights only when needed. Privileged Access Management (PAM) can be used to grant elevated permissions based on need and for a limited period. Access can be logged and revoked afterward.
4. Rehearse the emergency. Before an extended absence, simulate a scenario in which IT leadership is unavailable for a defined period. This can expose problems with responsibilities, approvals, or documentation before a real incident occurs.
5. Make status information available asynchronously. Dashboards, ticketing systems, and written status updates reduce reliance on spontaneous meetings or direct questions to IT leadership.
Conclusion
Good IT should not depend on its leader being available around the clock. It needs resilient systems, documented processes, clear escalation rules, and a team empowered to make decisions within defined boundaries.
Automation, runbooks, error budgets, controlled access rights, and an empowered deputy do more than reduce dependence on individual people. They make the organization as a whole more resilient.
If IT leadership can take two weeks off without checking alerts by the pool every morning, that is not a sign of a lack of commitment. It is a sign that the organization has learned to function without relying on individual key people.