Network Enhancers - "Delivering Beyond Boundaries" Headline Animator

Showing posts with label ITIL. Show all posts
Showing posts with label ITIL. Show all posts

Tuesday, January 29, 2013

Cloud Service Assurance Process

 
We will discuss the service assurance aspect of the services that are provisioned through the cloud service fulfillment process. We will discuss the fundamental service assurance processes and corresponding best practices that are typically employed to successfully manage the services that are delivered to customers.
 
 
Service assurance is a combination of fault and performance management. Cloud service assurance requires fault and performance management of cloud infrastructure that is comprised of network, compute and storage in addition to the applications that run on services platform itself. Reporting an incident to the appropriate management system for technical remediation and further, to informational dashboards in near real time for end user notification is a fundamental requirement. It is also essential the cloud service fulfillment processes be closely coupled with cloud assurance processes, since cloud customers may only be using the infrastructure for a defined period of time. For example, customers may use a cloud service as a test/dev environment only for the period of time when they are developing/testing a service. Previously, customers stayed on the network for very long periods of time, hence the need to spin up and down services with the appropriate assurance monitoring and management did not happen as part of the service lifecycle.
 
Cloud service providers must resolve service-related problems quickly to minimize outages and revenue loss. Cloud assurance solutions give the customer and the operational team maximum visibility into service performance and cost-effective management of SLAs coupled with service impact analysis.
 

Cloud End-to-End Service Assurance Flow

Figure 1 below shows the typical end-to-end service assurance steps:
Picture1.jpg
Figure 1: Cloud Assurance Flow – data collection to display
Unlike the cloud service fulfillment cycle that starts with the customer and ends with a provisioning task into the service platform resources, service assurance events occur in the resources affecting services with variable levels of degradation. Eventually the customer is made aware of the service degradation via notification unless it is fixed pro-actively before it is perceived by the customer.
The steps illustrated in Figure 1 are explained below:
 
 
  1. The incident/fault and performance events are sent by the infrastructure devices comprised of network devices, compute/server platforms, storage devices, and applications.
  2. The domain managers collect these events from infrastructure devices through the CLI, Simple Network Management Protocol (SNMP) polling, or traps sent by the infrastructure devices. Typically separate domain managers are used for the network, compute, and storage devices.
  3. The domain managers receive the messages from the infrastructure devices and de-duplicate/filter the events.
  4. The device events received by the domain managers have device information only and do not contain any service information. The domain manager will enrich the events by looking into the service catalog. In other words, the events are mapped to the services to determine which services are affected because of the events in the infrastructure. It is also possible to map to customers and determine the customers who are impacted by the events.
  5. The service impact is assessed and the information is shown on a dashboard for the operations personnel. This will help the operations people to prioritize the remediation efforts.
  6. The service impact information is forwarded to other locations, including mobile devices.
  7. The service impact information is sent to the service desk (SD).
  8. The service impact information is sent to the service-level manager (SLM).
  9. The SD proactively manages customers, informs them of service impacts, and keeps them up to date on the remediation efforts. Also, the service impact is checked against the SLA to determine any SLA violations and business impacts.
  10. The events, SLA violations, trouble tickets, and so on are displayed on a portal for various consumers, such as customers, operations people, suppliers, and business managers.

Best Practices for Cloud Service Assurance using ITILv3 Principles

 
ITILv3 provides the IT life cycle processes: service strategy, service design, service transition, service operate and Continuous Service Improvement (CSI). Applying these processes is a good way to establish service assurance processes for data center virtualization and cloud management. Figure 2 shows cloud service assurance flow based on ITILv3:
Picture2.jpg
Figure 2: Cloud service provisioning flow based on ITIL V3 principles

Mapping ITILv3 phases to ‘Infrastructure as a Service’ requirements

 
 
Figure 2 also shows some of the items that need to be considered in each of the five phases of the cloud service life cycle for cloud assurance. More details are provided in the following sections along the lines of ITIL V3 phases.
1. Service Strategy: In the cloud strategy phase for the cloud assurance consider the following topics:
  • Architecture Assessment
  • Business Requirements
  • Demand Management
 
 
2. Service Design:
The following items should be considered, taking input from the service strategy phase:
  • Service Catalog Management
  • Service Level Management
  • Availability Management
  • Capacity Management
  • Incident Management
  • Problem Management
  • Supplier Management
  • Information Security Management
  • Service Continuity Management
 
 
3. Service Transition:
Consider the following items in this phase:
  • Change Management
  • Configuration and setting up all the Assurance Systems
  • Service asset management in the CMS (CMDB)
  • Migration from current state to target state (people, processes, products, and partners)
  • Staging and Validation of all systems and processes for assurance
 
 
4. Service Operate: Cloud Operate phase is where the service provider takes possession of the management of cloud operations from the equipment vendors, system integrators and partners, and monitors and audits the service using the monitoring systems (FCAPS) to ensure the SLAs are met. Consider the following items in the cloud Operate phase:
  • Service Desk (function)
  • Incident Management
  • Problem Management
  • Event Management
  • Other IT day-to-day activities
 
 
5. Cloud CSI Phase: Continuous Service Improvement (also referred as Optimization) for cloud assurance involves improving on the operations by adding best practices to the processes, tools and configurations.
  • Auditing the configurations against best practices and changing the configurations as appropriate
  • Fine tuning the tools and processes based on best practices
  • Adding new products and services, and performing assessments to ensure the new services can be incorporated into the current environment. If not, determine the changes required and go through the cloud life cycle, starting from cloud strategy.
 

Sunday, January 27, 2013

Cloud Service Fulfillment Process

Courtesy - Venkata Josyula, Malcolm Orr, Greg Page

This Chalk Talk from Greg Page and Venkata (Josh) Josyula discusses the fundamental processes and corresponding best practices that are typically employed to successfully fulfill a customer's "service request" for IT infrastructure in a real-time environment.


Cloud End-to-End Service Provisioning Flow



Figure 1 below shows the typical end-to-end provisioning steps for a customer interacting with a Cloud Service Provider (CSP) through a self-service portal and ordering a service. When appropriate, the customer would receive a confirmation from the CSP that the service request is fulfilled and is ready for use (typically via the portal in addition to a confirmation email).


Picture1a.jpg

Figure 1: Typical End-to-End IaaS provisioning steps

(click to view larger image)



The steps illustrated in Figure 1 are explained as follows:


  1. The customer logs on to the portal and is authenticated by the identity management.
  2. Based on the customer’s entitlement, the portal extracts a subset of services that the user can order from the service catalogue and constructs a ‘request catalogue’.
  3. The customer selects a service, e.g. a virtualized web server. Associated to this service is a set of technical requirements such as the amount of vRAM, vCPU etc. in addition to business requirements such as high availability or SLA requirements.
  4. The portal now raises a change request with the service desk which, when approved, will create a service instance and notify the portal. In most cases the approval process is automatic and happens quickly. The service request state is maintained in the service desk and can be queried by the customer through the self-service portal.
  5. The service desk raises a request with the IT process automation tool to fulfill the service. The orchestration tool extracts the technical service information from the service catalogue and decomposes the service into individual parts, such as compute resource configuration, network configuration, and so on.
  6. In the case of our virtualized web server running on Cisco UCS (Unified Computing System), we have three service parts: the server part, the network part, and the infrastructure part. The provisioning process is initiated.
  7. The virtual machine running on the blade or server is provisioned using the server/compute domain manager.

8&9. The network, including firewalls and load balancers, as well, the storage is provisioned by the network, network services and storage domain managers.

10-13.Charging is initiated for billing / chargeback and the Change management case is closed and the customer is notified accordingly.




Best Practices for Cloud Service Fulfilment using ITILv3 Principles.





ITILv3 provided the IT life cycle processes: service strategy, service design, service transition, service operate and Continuous Service Improvement (CSI). Applying these processes is a good way to establish service provision processes for data center virtualization and cloud provisioning. Figure 2 shows cloud service provisioning flow based on ITILv3.


Picture2a.jpg

Figure 2: Cloud service provisioning flow based on ITIL V3 principles

(click to view larger image)




Figure 2 shows some of the items that need to be considered in each of the five phases of the cloud service life cycle to provision a cloud. Data center virtualization and cloud-computing technologies can have a significant impact on IT service delivery, cost, and continuity of services; but as with any transformative technology, the adoption is greatly influenced by the up-front preparedness and strategy. IT governance can help the CIOs to become the agent of change and be an active partner in laying out the company's strategy.


The IT organization's success factors include the following:

  • Technology decisions driven by a business strategy (not the other way around)
  • Sustaining the IT activities as efficiently as possible
  • Speed to market
  • Technology architecture aligning with the business initiatives.

Mapping ITILv3 phases to ‘Infrastructure as a Service’ requirements


1. Service Strategy: With the preceding principles in mind, the following high-level tasks are done during the service strategy phase:
  • Cloud architecture assessment
  • Operations (people, processes, products, and partners [the 4Ps])
  • Demand management
  • Financial management or value creation (ROI)
  • Risk management
  • Service Design Phase

2. Service design: The following items should be considered, taking input from the service strategy phase:
  • Service catalogue management
  • Orchestration
  • Security design
  • Network configuration and change management (NCCM)
  • Service-level agreements (SLA)
  • Billing and chargeback

3. Service Transition: The following items are considered in this phase:
  • Change management
  • Service asset and configuration management: maintained.
  • Orchestration and integration
  • Migration, staging and validation




4. Service Operate: All ITILv3 phases are important, but this phase draws the most attention because 60%–70% of the I.T. budget is spent in dealing with day-to-day operations. In this phase, the service provider monitors and audits the service to ensure that the SLAs are met. The following items are considered in the service operate phase:

  • Service desk (function),
  • Incident management,
  • Problem management, Service fulfillment,
  • Event management
  • Access management



5. Cloud CSI Phase: This phase is, as its name implies, a proactive methodology to improve the IT organization activities using best practices, as opposed to reactive responses. CSI interacts with all the aforementioned phases to improve each of the phases through feedback loops. Specific activities that can be quickly adopted in this phase include:

  • Audit the configurations in all the infrastructure devices. The inventory and collection of data from all devices must be done to ensure the manageability of the cloud infrastructure.
  • Identify infrastructure that is end of life (EOL), end of service (EOS).
  • Fine-tune the management tools and processes based on best practices.




Finally, note that adding new products and services requires assessment to ensure that the new services can be incorporated into the current operating environment without sacrificing the quality of service to customers.
 

Thursday, July 26, 2012

Key ITIL Processes For Cloud Computing


With cloud computing, the only difference that is going to be in ITIL world would be that the ITIL processes could no longer be ignored. In my experience, today in many organizations ITIL processes exist in silos. Key processes like configuration management are ignored. In most of the processes, an option to bypass the process exists, i.e. the process adherence and its compliance with best practices are significantly low.

But with cloud computing ITIL processes can no longer be neglected. Service Asset and Configuration Management (SACM) process will become utmost important along with the IT security and demand management processes. To highlight the key processes in cloud computing environment, it is prefered to list 5 key processes from customer's and service providers’ perspective in the order of their importance.

Customer perspective:
1) Security Management
2) Service Continuity Management
3) Incident Management
4) Change Management
5) Release & Deployment Management
This perspective is important for a customer while finalising a vendor. These are the processes which concerns a customer the most.

Service provider perspective – External Facing
1) Service Level Management
2) Service Portfolio Management
3) Service Catalog Management
4) Financial Management
5) Supplier Management
These are the key processes which the service provider needs to focus on while approaching a customer.


Service provider perspective – Internal Facing
1) Service Asset & Configuration Management
2) Demand Management
3) Financial Management
4) Request Fulfillment
5) Capacity Management
These processes are important for a cloud service provider for their internal organization in order to provision customers’ requests.

Tuesday, July 24, 2012

4 Essential steps for Successful Incident Management

It never hurts to go back to basics. Recently, we were surprised at the confusion of some organizations about the process of incident management, so we thought – why not to put a quick incident management primer down on paper?

For successful incident management, first you need a process – repeatable sequence of steps and procedures. Such a process may include four broad categories of steps: detection, diagnosis, repair, and recovery.

1 – Detection

Identification Problem identification can be handled using different tools. For instance, infrastructure monitoring tools help identify specific resource utilization issues, such as disk space, memory, CPU, etc. End user experience tools can mimic user behavior and identify users’ POV problems such as response time and service availability. Last but not least, domain-specific tools enable detecting problems within specific environments or applications, such as a database or an ERP system.

On the other hand, users can help you detect unknown problems that are not reported by infrastructure or user behavior monitoring tools. The drawback with problem detection by users is that it usually happens late (the problem is already there), moreover the symptoms reported may lead you to point to the wrong direction.

So which method should you use? Depending on your environment, the usage of the combination of multiple methods and tools would be the best solution. Unfortunately, no single tool will enable detecting all problems.

Logging events will allow you to trace them at any point to improve your process. Properly logged incidents will help you investigate past trends and identify problems (repeating incidents from the same kind), as well as to investigate ownership taking and responsibility.

Classification of events lets you categorize data for reporting and analysis purposes, so you know whether an event relates to hardware, software, service, etc. It is recommended to have no more than 5 levels of classification; otherwise it can get very confusing. You can start the top level with something like Hardware / Software / Service, or Problem / Service request.

Prioritization lets you determine the order in which the events should be handled and how to assign your resources. Prioritization of events requires a longer discussion, but be aware that you need to consider impact, urgency, and risk. Consider the impact as critical when a large group of users are unable to use a specific service. Consider the urgency as high when the impacted service is of critical nature and any downtime is affecting the business itself. The third factor, the risk, should be considered when the incident has not yet occurred, but has a high potential to happen, for example, a scenario in which the data center’s temperature is quickly rising due to an air conditioning malfunction. The result of a crashing data center is countless services going down, so in this case the risk is enormous, and the incident should be handled at the highest priority.

2 – Diagnosis

Diagnosis is where you figure out the source of the problem and how it can be fixed. This stage includes investigation and escalation.

Investigation is probably one of the most difficult parts of the process. In fact, some argue that when resolving IT problems, 80% of the time is spent on root cause analysis vs. 20% that is spent on problem fixing. With more straightforward problems, Runbook procedures may be very helpful to accelerate an investigation, as they outline troubleshooting steps in a methodical way.

Runbook tip: The most crucial part of the runbook is the troubleshooting steps. They should be written by an expert, and be detailed enough so every team member can follow them quickly. Write all your runbooks using the same format, and insist on using the same terms in all of them. New team members who are not familiar yet with every system will be able to navigate through the troubleshooting steps much more easily.

Following the runbook can be very time consuming and lengthen the recovery time immensely. Instead, consider automating the diagnostic steps by using run book automation software. If you build the flow cleverly and weigh in all the steps that lead to a conclusion, automating the diagnostics process will give you quick answers, and help you decide what your next step is.

Escalation procedures are needed in cases when the incident needs to be resolved by a higher support level.

3 – Repair

The repair step, well… it fixes the problem. This may sometimes involve a gradual process, where a temporary fix or workaround is implemented primarily to bring back a service quickly. An incident repair may involve anything from a service restart, a hardware replacement, or even a complex software code change. Note that fixing the current incident does not mean that the issue won’t recur, but more on that issue in the next step.
In this case too, straightforward repairs such as a service restart ,a disk cleanup and others can be automated.

4 – Recovery

The recovery phase involves two parts: closure and prevention.

Closure means handling any notifications previously sent to users about the problem or escalation alerts, where you are now notified about the problem resolution. Moreover closure also entails the final closure of the problems in your logging system.

Prevention relates to the activities you take, if possible, to prevent a single incident from occurring again in the future and therefore becoming a problem. Implement two important tools to help you in this task:

RCA process (Root Cause Analysis) The purpose of the RCA process is to investigate what was the root cause that led to the service downtime. It is important to mention that the RCA process should be performed by the service owners, who are not necessarily the ones who solved the specific incident. This is an additional reason why incident logging is so important – the information in the ticket is crucial for this investigation process.

And finally, Incident reports – while this report will not prevent the problem from occurring again, it will allow you to continually learn and improve your incident management process.

Saturday, July 14, 2012

Network Operation Center (NOC) Best Practices – Part 3: Processes

This is the third part of our 3-part blog series discussing Network Operation Centers (NOC’s) best practices. The first post was dedicated to NOC tools. The second provided some useful tips regarding NOC knowledge and skills. In this last part, we’ll address processes.

What are the operational, structured processes that you should implement for effective and repeatable results? Here are our top ones.

Escalation

A table of escalation will ensure that all team members are clear on the proper protocol and channels for escalating issues. This table should also include all areas and skills covered by the NOC and the people who are trained to cover those areas.

For example, see the table below defining the escalation procedures for DB related problems.

Time FrameEscalate ToMethod
0+15minsDB on callSMS
0+30minsDB on callPhone
0+60minsDB Group LeaderPhone
0+90minsUNIX & DB Project ManagerSMS
0+120minsUNIX & DB DirectorSMS

A critical problem that was not solved within 30 minutes is escalated up the management ladder, until a response and/or ownership is taken. At every step of the process, it is recommended to involve all personnel up to the current level. So when an SMS is sent to the project manager, it is also sent the DB on call and Group Leader.

Prioritization

The process of prioritizing incidents is different in each NOC, and therefore should be clearly defined. Incidents should never be handled on a first come, first served basis. Instead, the shift manager should prioritize incidents and cases based on the importance and impact on the business. Issues that have a greater impact on the business should obviously be handled first.

Understanding the prioritization of incidents in terms of their business impact should be part of the NOC training. The entire team should be familiar with the NOC “Top 10” projects, and have an understanding of what signifies a critical incident. It could be the temperature rising in the data center, a major network cable breaking or a service going down.

Obviously, common sense is very useful. Clearly the shift leader should be able to determine that an incident that jeopardizes the entire data center has a higher priority than a request to verify why an individual server is down.

Incident handling

The process of handling incidents applies both to NOC operators and shift managers. Both roles should be familiar with the specific process of handling incidents with the greatest impact on users.

Incident handling process should cover issues such as:
  • Full technical solution, if available.
  • Escalation of issue to appropriate personnel.
  • Notification of other users who may be directly or indirectly affected by issues.
  • ‘Quick solution’ procedures or temporary workarounds for more complex problems that may take longer to completely resolve.
  • Incident reporting. An incident report, completed once the incident is resolved, helps improves the service when the next incident occur, or may also prevent the recurrence of the same incident.
Employing the proper tools, skills and processes in your NOC will allow you to run more efficient network operations and ensure smooth day to day operations as well as meeting the demands from the IT department.

Tuesday, July 10, 2012

Network Operation Center (NOC) Best Practices – Part 2: Knowledge & skills


This is the second of our 3 parts blog series discussing Network Operation Centers (NOC’s) best practices. The first post was dedicated to NOC tools. This part is dedicated to knowledge and skills. By ‘knowledge and skills’ we do not mean the obvious technical knowledge, network ‘know-how’ your team members must hold in order to run day-to-day operations, but rather –

How you can ensure your team’s skills are used to their best potential, and how to keep those skills up to date over time.

Clearly define roles

Definition of roles may vary between data centers and will depend on team size, the IT environment and tasks. Still, there should be a clear distinction between the roles and responsibilities of operators vs. shift supervisors in the NOC.

Why does it matter?

Mainly matters because of Decisions making. Without clearly defined roles and responsibilities, a disagreement between operators may lead to late decisions and actions, or to no decisions taken at all. This may affect customers, critical business services, and urgent requests during off hours.
It should be clearly defined, therefore, that a shift manager makes the final decisions.

Tasks division

Another potential problem caused by a lack of role definition is the division of tasks between operators and the shift leader.

A shift manager should be responsible for: prioritizing tasks, assigning work to operators based on their skills, verifying that tickets are opened properly and that relevant personnel are notified when required, escalating problems, communicating with management during important NOC events, sending notifications to the entire organization, preparing reports, and making critical decisions that impact many services, such as shutting down the data center in case of an emergency.

Operators, on the other hand, are responsible for handling the technical aspect of incidents – either independently or by escalating to another team member with the required skills. Operators are also responsible for following up and keep tickets up to date.

While it might sound as if operators lack independence and responsibility, this is not the case. When faced with technical challenges, operators’ input and skills are probably the most critical for resolution and smooth NOC operation. Operators provide additional insights into problems, and can provide creative solutions when the standard procedures fail to work.

Invest in orientation program for new NOC employees

How often have you started a new job, without receiving any orientation, mentoring or guided training?

Failing to provide proper training to new NOC operators always has consequences. A new NOC operator may not know where to find a procedure or how to execute it; be confused about who should handle a task – the NOC, service desk, or higher level of support; or in a more severe case, take a decision that causes equipment damage or results in downtime of critical business services.

Therefore, an extensive training program should be put in place for new NOC employees. This is definitely a challenge, considering the lack of resources, particularly in small NOCs. Ideally, such a program would consist of one week of classroom training followed by three weeks of hands-on training under the supervision of a designated trainer.

A new employee should only be trained by an experienced member of the NOC, preferably a shift leader. The trainer should be released from all duties during the entire period of the training – in order to ensure that the training does not gradually fade between all the urgent shift tasks.

The training program should be updated on an ongoing basis, and should include topics such as required users and permissions, technical knowledge, known problems, troubleshooting, teams and important contacts.

Communication and collaboration

Within your Organization

Establishing a solid communication flow between NOC members and other IT teams has many advantages. It propels professional growth, provides opportunities for advancement in the organization, and makes it easier to approach other teams when requiring assistance. But most importantly – it allows NOC personnel to see the larger picture. NOC members that are aware of projects, services and customers’ needs, simply provide better service.

A designated member of the NOC should attend weekly change management meetings. That person should communicate any issues or upcoming activities, such as planned downtime, to the rest of the team.

Define NOC members as focal points for important IT areas, such as NT, UNIX, Network, or a specific project is another good practice. These members should attend the meetings of the relevant teams, deliver new information and knowledge to the rest of the NOC, and handle specific professional challenges.

Within NOC Team

Another important form of communication is within the NOC team itself. There are clear advantages to having a strong connection and collaboration between NOC team members. Members are more willing to help each other, information is shared more easily, and the general atmosphere encourages collaboration when addressing problems, as opposed to an individualized approach.
Team communication is a challenge when the NOC team is geographically spread out or located in different countries. Because cultural and language differences can cause confusion and misunderstandings, spending efforts on building team communication and collaboration are even more critical.

Wednesday, July 4, 2012

Network Operation Center (NOC) Best Practices – Part 1: Tools

Today, Network Operation Centers (NOC’s) are under a great pressure to meet their IT organization’s demands. However, many NOCs struggle to meet these demands with insufficient tools, knowledge or skills.

In this 3 parts blog post series, we will provide tips on how to ensure you have the right tools, knowledge and processes in place to improve and manage your NOC’s performance and response time.

This first part of our ‘NOC best practices’ series is dedicated to tools, which are an essential element in NOC management and a key feature for improvement.

A ticketing system

A ticketing system will enable you to keep track of all open issues, according to severity, urgency and the person assigned to handle each task. Knowing all pending issues will help you to prioritize the shift’s tasks and provide the best service to your customers.

Knowledge-base system

Keep a one centralized source for all knowledge and documentation that is accessible to your entire team. This knowledge base should be a fluid information source to be continuously updated with experiences and lessons learned for future reference and improvements.

Reporting and measurements

Create reports on a daily and monthly basis. A daily report should include all major incidents of the past 24 hours and a root cause for every resolved incident. This report is useful and essential for the shift leaders and NOC managers. It also keeps the rest of the IT department informed about the NOC activities and of major incidents. Compiling the daily reports into a monthly report will help measure the team’s progress. It will also show areas where improvements can be made or indicate any positive or negative trends in performance.

Monitoring

There are two types of monitoring processes relevant to NOC:
(1) Monitoring infrastructure and (2) User experience.

A monitoring infrastructure can consist of the servers, the network or the data center environment. User experience monitoring involves the simulation of user behavior and activities in order to replicate problems and find the most effective solutions. Implementing a service tree model that connects the monitoring infrastructure with an affected service will allow your team to alert other areas that may be affected by the problems experienced.

IT Process Automation

Implementing IT Process Automation significantly reduces mean time to recovery (MTTR) and helps NOCs meet SLA’s by having a procedure in place to handle incident resolution and to consistently provide high quality response regardless of complexity of the process. IT Process Automation empowers a Level-one team to deal with tasks that otherwise might require a Level-two team. Some examples include password reset, disk space clean-up, reset services etc. IT Process Automation is also a major help with reducing the number of manual, routine IT tasks and free up time for more strategic projects. 

Wednesday, June 20, 2012

4 Essential steps for Successful Incident Management

It never hurts to go back to basics. Recently, we were surprised at the confusion of some organizations about the process of incident management, so we thought – why not to put a quick incident management primer down on paper?


For successful incident management, first you need a process – repeatable sequence of steps and procedures. Such a process may include four broad categories of steps: detection, diagnosis, repair, and recovery.

1 – Detection

Identification Problem identification can be handled using different tools. For instance, infrastructure monitoring tools help identify specific resource utilization issues, such as disk space, memory, CPU, etc. End user experience tools can mimic user behavior and identify users’ POV problems such as response time and service availability. Last but not least, domain-specific tools enable detecting problems within specific environments or applications, such as a database or an ERP system.

On the other hand, users can help you detect unknown problems that are not reported by infrastructure or user behavior monitoring tools. The drawback with problem detection by users is that it usually happens late (the problem is already there), moreover the symptoms reported may lead you to point to the wrong direction.

So which method should you use? Depending on your environment, the usage of the combination of multiple methods and tools would be the best solution. Unfortunately, no single tool will enable detecting all problems.

Logging events will allow you to trace them at any point to improve your process. Properly logged incidents will help you investigate past trends and identify problems (repeating incidents from the same kind), as well as to investigate ownership taking and responsibility.


Classification of events lets you categorize data for reporting and analysis purposes, so you know whether an event relates to hardware, software, service, etc. It is recommended to have no more than 5 levels of classification; otherwise it can get very confusing. You can start the top level with something like Hardware / Software / Service, or Problem / Service request.

Prioritization lets you determine the order in which the events should be handled and how to assign your resources. Prioritization of events requires a longer discussion, but be aware that you need to consider impact, urgency, and risk. Consider the impact as critical when a large group of users are unable to use a specific service. Consider the urgency as high when the impacted service is of critical nature and any downtime is affecting the business itself. The third factor, the risk, should be considered when the incident has not yet occurred, but has a high potential to happen, for example, a scenario in which the data center’s temperature is quickly rising due to an air conditioning malfunction. The result of a crashing data center is countless services going down, so in this case the risk is enormous, and the incident should be handled at the highest priority.

2 – Diagnosis

Diagnosis is where you figure out the source of the problem and how it can be fixed. This stage includes investigation and escalation.

Investigation is probably one of the most difficult parts of the process. In fact, some argue that when resolving IT problems, 80% of the time is spent on root cause analysis vs. 20% that is spent on problem fixing. With more straightforward problems, Runbook procedures may be very helpful to accelerate an investigation, as they outline troubleshooting steps in a methodical way.

Runbook tip: The most crucial part of the runbook is the troubleshooting steps. They should be written by an expert, and be detailed enough so every team member can follow them quickly. Write all your runbooks using the same format, and insist on using the same terms in all of them. New team members who are not familiar yet with every system will be able to navigate through the troubleshooting steps much more easily.

Following the runbook can be very time consuming and lengthen the recovery time immensely. Instead, consider automating the diagnostic steps by using run book automation software. If you build the flow cleverly and weigh in all the steps that lead to a conclusion, automating the diagnostics process will give you quick answers, and help you decide what your next step is.

Escalation procedures are needed in cases when the incident needs to be resolved by a higher support level.

3 – Repair

The repair step, well… it fixes the problem. This may sometimes involve a gradual process, where a temporary fix or workaround is implemented primarily to bring back a service quickly. An incident repair may involve anything from a service restart, a hardware replacement, or even a complex software code change. Note that fixing the current incident does not mean that the issue won’t recur, but more on that issue in the next step.

In this case too, straightforward repairs such as a service restart ,a disk cleanup and others can be automated.

4 – Recovery

The recovery phase involves two parts: closure and prevention.

Closure means handling any notifications previously sent to users about the problem or escalation alerts, where you are now notified about the problem resolution. Moreover closure also entails the final closure of the problems in your logging system.

Prevention relates to the activities you take, if possible, to prevent a single incident from occurring again in the future and therefore becoming a problem. Implement two important tools to help you in this task:

RCA process (Root Cause Analysis) The purpose of the RCA process is to investigate what was the root cause that led to the service downtime. It is important to mention that the RCA process should be performed by the service owners, who are not necessarily the ones who solved the specific incident. This is an additional reason why incident logging is so important – the information in the ticket is crucial for this investigation process.

And finally, Incident reports – while this report will not prevent the problem from occurring again, it will allow you to continually learn and improve your incident management process.

Wednesday, March 2, 2011

Cost per Incident

“how do you determine cost per incident?”

Cost per incident is a variation of cost per call or cost per contact, all of which are excellent ways to understand the impact of incidents, calls, or contacts on the business.

The calculation is fairly straightforward. Cost per incident is the total cost of operating your support organization divided by the total number of incidents for a given period (typically a month).

Cost per incident = total costs/total incidents

To accurately calculate cost per incident you must:
  • Log all incidents. You may also find it beneficial to distinguish between incidents (unplanned events) and service requests (planned events). Such a distinction will enable you to more accurately reflect business impact
  • Identify stakeholders. Determine all of the support people involved in the incident management process (e.g., service desk and level two/three technical and application management teams). 
  • Identify associated costs. Create a list of every cost associated with your stakeholders including salaries, benefits, facilities, and equipment. Many of these costs will be included in your annual budget or can be obtained from financial management.
As illustrated in the ROI calculator, tremendous cost savings can be realized by reducing the number of incidents. This can be accomplished by using trend and root cause analysis techniques via the problem management process, as well as through integration with processes such as change management, service asset and configuration management, and release and deployment management.

Savings can also be realized by reducing the number of incidents escalated from the service desk to other lines of support, or by increasing the number of incidents handled via self-help services such as a web-based knowledge management system. Some organizations report on cost per incident per channel (e.g., per call, e-mail, chat, walkup) to better understand these savings.

Sunday, February 20, 2011

Dealing with Major Incidents

We know that the primary goal of the incident management process is to restore normal service operations as quickly as possible and to minimize any adverse impact on business operations. This will insure the highest levels of service quality and availability are delivered to the user community, guaranteeing that the business is receiving value and facilitating the outcomes it wants to achieve.

The value this process produces for the business is in the ability to:
  • detect and resolve incidents quickly, resulting in higher availability of IT services.
  • align IT activities to real time business priorities and dynamically allocate resources as necessary.
  • identify potential improvements to services, through the analysis of incident trends.
So it sounds like we have everything covered as long as we handle all incidents in the same consistent and proceduralized manner. Well not so fast. What happens when we have an incident that affects a major business process and in turn creates a major impact to the business?

For these types of situations we need to have a separate procedure, with shorter escalation time scales and greater urgency in responding to “Major Incidents”. First we must agree on a definition of just what constitutes a major incident and how it will be integrated into the overall incident prioritization system.

Note: Many organizations that I have corresponded with confuse this separate process with problem management. A major incident may increase in impact to the business thus increasing in the priority it needs to be addressed by the ITSM processes but it still remains an incident and never becomes a problem.

Where necessary , the major incident procedure should include the formation of a separate and dynamic major incident team (under the leadership of the incident manager) to concentrate their efforts on the particular incident alone and insure that adequate resources are engaged and solely focused on providing a swift resolution to the impact at hand. Problem management can be involved if the underlying cause needs to be discovered at the same time, but the incident manager must ensure that restoration of services and root cause analysis are kept separate and that impact reduction is the priority.

Saturday, February 12, 2011

Problem Management Techniques

Perhaps one of the most underused yet powerful processes from ITIL is Problem Management. Many people recognize the importance of Problem Management, especially in relationship to Incident Management. Yet when I ask students if they have implemented a Problem Management process the response is often “We plan to...” or “We started but did not get too far…” or “Not yet.” So what is keeping companies and individuals from using Problem Management to its full effectiveness? I propose that some of the reason is fear, uncertainty and doubt about how to go about “doing” Problem Management.

By understanding that the Problem Management process has a number of techniques and tools available to help a service provider indentify root cause and recommend permanent resolution we may be able to remove some of the fear, uncertainty and doubt.

What are some of the techniques and how could we apply them? Let’s take a brief look and see what we can uncover.
  • Chronological Analysis: This time based approach lets us look at the order of events to help identify a chain of cause and effect. It is particularly helpful when looking at potential problems that have developed slowly over a period, yet have distinct symptoms.
  • Pain Value Analysis: This technique helps us narrow in on particular impacts that result from incidents and events that are troublesome for the business. It allows us to separate and prioritize items when we are facing a number of problems simultaneously.
  • Brainstorming: This collaborative method allows a number of subject matter experts and knowledgeable individuals to come together to share prospective ideas about the cause of a problem. This permits us to have many eyes and minds seeing a problem from different perspectives, thus broadening our problem solving capabilities.
  • Kepner-Tregoe Problem Solving Method: This well known method provides a formal approach and set of steps for identifying root cause, recommending a solution and implementing the solution. When formality and structure is needed—e.g. a major problem or large impact—this method would prove particularly useful.
  • Ishikawa Diagram (Fishbone): This visual technique allows the service provider to frame a question and develop possible causes using a structured, categorized approach. The technique is one of the most powerful and really is hard to do in an incorrect manner. This might be a good starting point for some of the other techniques.
  • Pareto Analysis: This prioritization helps the service provider focus in on the problems with the greatest impact. Based on the idea that 80% of incidents result from 20% of the problem causes, the method allows tackling problems in the order of greatest return. When coupled with the Ishikawa diagram it can be a very powerful tool. 
These Problem Management techniques provide us with a toolkit for identifying root cause and then determining an appropriate recommended permanent resolution. They seem difficult and complex to use but in reality they are each a straightforward method. The struggle may come when determining in which circumstance to apply which technique. Although I mentioned briefly some possible uses, these methods can be used in a variety of circumstances and with any number of problem types.

The success of Problem Management comes from taking the time to execute the process and use some or all of the techniques provided. Making time for Problem Management may detract from time spent resolving Incidents. The irony is that if you do not make the time for Problem Management you will only ever have time for Incident Management, which can consume large amounts of resources. Problem Management is not as resource hungry, but it requires special devotion and attention to see it succeed.

Monday, January 31, 2011

Why ITIL is still the heartbeat of IT services

Its perspective at the sharp end of ITIL projects means APM Group is well placed to see the benefits of implementing ITIL. Jessica Barry, APMG’s accreditor co-ordinator, explains why ITIL is still an extremely useful business tool and Version 3 is exactly what the industry needs to continue improving.

The issues surrounding ITIL and service management generate some heated debates.  On this website alone there are lots of examples of people who are pro-ITIL and evangelise on its benefits, but there are just as many who think it is removed from the working realities faced by service providers.
The APM Group has the role of accreditor for the ITIL scheme.  We accredit the examination institutes who work with training companies who in turn help end users adopt and embed ITIL.  We also manage the examination scheme, so we are well placed to understand the debate about ITIL, and answer some of the more burning questions that people are discussing.

Our view, naturally, is that ITIL is an extremely useful method, especially with the added dimensions offered by Version 3.  Our latest figures show that about 500,000 V3 certificates have been issued to candidates and every month there are more candidates taking the qualifications than the previous month.  The international appeal of ITIL V3 is also extending, and we have translated papers into 20 languages.

 Translations are triggered when there is demand – usually measured by a significant number of Foundation examinations in a particular country.  itSMF’s international chapters push for translations when they’ve got candidates asking for them – this in itself is a testament to ITIL’s value.

Demand for ITIL Masters

We are currently formulating the ITIL Master level of the certification scheme.  The ITIL Master is a high level qualification where candidates demonstrate practical application of their ITIL knowledge.  At Expert level you need to prove you know the guidance and the processes, but Masters goes further.  The IT service management industry told us they needed and wanted the Masters qualification because this level of maturity was missing in previous schemes.

We’ve got several ITIL experts devising the certification including: Sharon Taylor, Dave Cannon, Kevin Holland, Carol Hulm, Dave Wheeldon, and Vernon Lloyd.  They all have a wealth of experience in assessing individuals and defining what constitutes a robust and skilled ITIL knowledge and experience base.  We expect Masters to be widely available from mid 2011.

First hand ITIL experience

One of the candidates taking part in the Masters Pilots is Mathieu Mathieu Notéris, senior ITSM consultant at Sogeti Belux, a provider of IT services in Belgium and Luxembourg.  Sogeti helps clients develop, implement and manage practical IT solutions and Mathieu believes ITIL certifications are the evidence that you are able ‘to do IT’.  He got involved in the pilots because he wanted ‘to stay on top of the wave’ and believes the qualifications give service managers credibility.  Mathieu particularly likes the fact that the ITIL Masters level is based on the candidate’s real-life experience of service management.

“If you are a manager or high-level senior consultant this is a way to establish your own value based on real-life, proven experience and not just on an artificial business cases as was the case with the Service Manager certification from Version 2.  Results and evidence are there.  Lessons learned are real.  To argue, convince and therefore manage, you definitely need to be convinced yourself,” Mathieu says.

But he doesn’t think ITIL is just about the process and says we shouldn’t forget that behind the machines, the programmes, the business and the processes, there are human beings.  He says service management is as much about the discussions, factoring in different opinions, confronting issues, persuading and being persuaded, as it is about process.

He advises service managers to take all the opportunities open to them to enhance their learning and try to remember all lessons: both positive as well as negative.  “One remembered bad situation is a situation you’ll be able to detect, recognise and hopefully avoid in the future,” he says.

The Ten Golden Rules of Change Management

We all know that nothing stays the same – no matter how much we might wish it did. In business, there will be new services introduced and unpopular ones discontinued in response to business requirements.
Technology is no different and new applications will need to be incorporated and others dissolved. Equally, the workforce ebbs and flows with new employees to allow access, even internally bringing people in and out of project teams. Change happens, it’s a fact of life, but it needs to be handled competently if the desired benefits are to be realized.

When it comes to reflecting these changes in IT systems, complexity and a lack of process has trumped best practices. Organizations are ready to take action – but the enormity of the situation means many do not know where to start, or even how.

Where does it all go wrong?

We’ve all been there when we’ve arrived bright and breezy on a Monday morning, only to find that a change that has been implemented to the system over the weekend causes things not to work as they should. The result is IT staff, with their backs firmly up against the wall, scramble around the network, making random changes without planning or authorization, in a desperate effort to find and "fix" the fault. The reality is that all of these little tweaks can cause additional problems that may not come to light at first, instead rearing up at a later date when no connection to the original change is made, causing the system to fail.

A more worrying issue is when these changes aren’t identified. There have been numerous examples of how failing to control the process has led to breaches and compliance violations that can be traced back to misconfigured systems.

How can we make it right?

There are many steps that should be followed when defining a change management process. Our top tips for handling this are :

Step 1: Graphically build existing workflow processes within your organization. Ideally they should fit IT Infrastructure Library (ITIL) guidelines, or at the very least the organization’s pre-defined processes.

Step 2: Define how changes are requested and what supporting documentation is required. This could be as simple as an e-mail request to a triplicate form that is completed and officially submitted to a pre-determined person(s).

Step 3: Define the change management process so everyone knows what will happen and by when. This should include how changes are prioritized, the timeframes involved, how they can be tracked, how they’ll be implemented. As part of this step, you should also define the appeals process. It should cover how a person is informed that their request has been declined, the reason why it failed and what happens next.

Step 4: If you don’t already have one, establish a change advisory board (CAB). This should include a representative from each area of the business. This team will have responsibility for reviewing all requested changes, checking the change is complete, that it meets business needs, the impact on other areas of the organization both adversely or positively, and finally whether doing so would introduce risks. If declined, this needs to be communicated back to the originator with the reasons why. If approved, the team would then be responsible for communicating the change to those affected by it.

Step 5: Design the change. During this stage it is important that any conflicts or insecurities are identified and rectified to avoid expensive repercussions at a later point. It may be prudent for this role to be split in two with an administrator to verify that the change remains within corporate compliance. This is easier said than done for some changes because in making a change to meet one standard you could quite easily be in breach of another so this role can be key in determining what is and isn’t allowed. There is technology available that can help with this veritable minefield to automate, check and flag potential compliance conflicts.

Step 6: Implement and document the change. In an ideal world, the person who implements the change should be different from the person who designed it to avoid a conflict of interest.

Step 7: Verify that the change has been made, that it has been executed correctly and that only authorized changes have been implemented.

Step 8: Have a backup plan. If a change has been implemented that has had an adverse effect on the system, rather than blindly making changes, it should be reversed, reassessed and re-implemented once the point at which it failed has been identified and rectified.

Step 9: Audit the change process. It is important to check that approved improvements have been made – after all they’ve been identified as beneficial to the organization.
Step 10: Regularly reflect on the change management process to identify any sticking areas that can be ironed out.

Manual vs Automated

Many organizations still rely on manual documentation of their network configurations. This then means that access requests are also manually processed, which opens up a number of pitfalls:

Extended costs of manual changes: Manual requests, typically, can take anywhere up to 10 days, which is time that could be used elsewhere. A further issue is the time wasted checking, verifying and implementing unnecessary or duplicated changes that an automated process would have identified.

Risks of something breaking: As the workflow is all on "paper," it is virtually impossible to check all potential break points without automating the process. Sometimes the team will have spent a great deal of time engineering the change upfront for it to be denied by the risk team because approval is sought too late in the process.

Lack of audit and accountability: As everything is on paper, and often hurried, processes will go out the window with paperwork submitted after the change has been implemented, if at all. This can be as basic as having no idea who requested the change or why it was needed in the first place.

Impossible to define and enforce compliance: With just documentation to determine what the organization’s compliance requirements are, administrators are left to use their own personal judgment to determine whether a new rule introduces risks.

Quality of service: As it takes longer to submit and implement changes, the service is poor – both internally and often externally. This can lead to revenue loss – for example, failing to provide access to an online sales system or CRM that has a revenue potential means that every week implementation is delayed, more money is lost.

As this article demonstrates, change is happening frequently and organizations need to keep pace and manage the process completely if they’re to reap the rewards.

Whether you do so manually, or invest in technology to give you a winning edge, time and tide wait for no man. You need to be able to react to changes in your environment, competently, for your team to come out fighting.

My Blog List

Networking Domain Jobs