Tuesday, November 30, 2010

The Importance of Asset Criticality

The Importance of Asset Criticality

Here’s the scenario. You’re a recently hired Asset Manager/Reliability Engineer and you’ve been tasked with defining and implementing plant-wide Reliability initiatives. Your initial assessment reveals a significant number of asset breakdowns, overall asset health is degrading, and reactive/emergent maintenance is the norm. It’s a daunting task and you’re not entirely sure where to begin.

When determining which assets to address first, the asset criticality/priority matrix should serve as your guide. Initiatives such as RCM, FMEA, Bill of Material (BOM) development, PM/PdM application, etc., should be targeted toward the most critical assets first with an eventual progression towards the least critical. In your new role, you should first review the Master Equipment List (MEL) for completeness, accuracy and prioritization. Properly ranked/prioritized assets take into consideration all aspects of an organization and have been ranked using mathematical formulas or quantitative analysis, thus eliminating the “gut” feel and subjectivity from the ranking process. Criticality criterion regarding Maintenance, Production/Operations, Safety, Environmental and Quality should be developed and personnel from the aforementioned departments should be represented during the ranking process. An added benefit to having asset criticalities is greater accuracy when prioritizing work during Planning & Scheduling activities. So in conclusion, utilize those asset criticalities to “eat the elephant” one bite at a time and make an overwhelming task seem much more manageable and achievable.

Tuesday, November 23, 2010

Work Order Maintenance Tip

Work Order Maintenance Tip

When is a work order truly complete? There’s more to work order completion than simply performing the actual maintenance tasks and changing the WO status to “Complete” within the CMMS. Although tasks will vary depending on the type of work performed, consider the following activities to ensure a successful WO completion.

  • Perform general housekeeping activities and return the work area to an operating condition. Work area should be clean of rags, grease/oil, trash, etc. and all items have been properly disposed. Scaffolding, safety barrier tape, etc. is removed as required.

  • The Craft have notified Operations personnel that the equipment is ready for Post Maintenance Testing (PMT). Job related LOTO is removed and equipment PMT is satisfactorily performed.

  • All unused job material/parts are returned to stores.

  • All specialty tools and equipment are returned to their proper location.

  • All work permits are closed-out as required.

  • WO completion information is captured (hardcopy/electronically in CMMS)
    • Detailed description of work performed

    • Proper Failure Code information is documented (Failure/Cause/Remedy)

    • As Found/As Left conditions

    • Any materials not originally issued/purchased against the WO. Compare against the asset BOM and Job Plan to see if these materials should be added.

    • Labor hours for all craft

    • Start/Finish time

    • Job Plan feedback such as missing material, inaccurate procedures and improvements.

    • Recommendations for adjusting PM frequency

  • If follow-up work is required (additional repairs, modifications, etc.), a separate WO should be entered into the CMMS.

  • If the nature of the work met the requirements to trigger a Root Cause Failure Analysis (RCFA), all documentation, failed parts, etc. should be provided to individuals responsible for conducting the RCFA.

  • If a repairable spare was removed, ensure the spare is returned to the appropriate location for repairs and the “move” history of this spare is captured using the CMMS rotating item/asset functionality.

  • If a new asset was installed, ensure all related information is captured and updated in the CMMS including nameplate information, the asset BOM, Job Plan and PM/PdM information, etc.

  • New PdM baseline readings are taken as required.

  • Drawings and schematics are updated to reflect any changes.

  • All change control documentation is completed as required.


A properly completed work order will benefit many departments within an organization. For example, good housekeeping practices align with a facility’s safety and environmental directives. Storeroom & purchasing personnel will use this information to streamline their inventories and improve their services to the craft person. Detailed and accurate job plan feedback will improve the planning & scheduling process. Reliability engineering personnel will use this information to improve asset reliability. Incorporating the aforementioned work order closeout activities as a part of the work control process is crucial for a facility if they’re to achieve their overall asset management and reliability initiatives.

Tuesday, November 9, 2010

Don't Let Them Shut You Down: Part 3

Don't Let Them Shut You Down: Part 3

Also to consider is OSHA code 29 CFR Part 1910.119. The section listed below specifically states that equipment will be maintained and the history of maintenance will be documented. It goes one-step further by identifying safe operations as part of the requirement.

    1910.119(d)(3)(iii)
  • For existing equipment designed and constructed in accordance with codes, standards, or practices that are no longer in general use, the employer shall determine and document that the equipment is designed, maintained, inspected, tested, and operating in a safe manner.

  • 1910.119(j)(3)
  • Training for process maintenance activities. The employer shall train each employee involved in maintaining the on-going integrity of process equipment in an overview of that process and its hazards and in the procedures applicable to the employee's job tasks to assure that the employee can perform the job tasks in a safe manner.

  • 1910.119(j)(4)iv
  • The employer shall document each inspection and test that has been performed on process equipment. The documentation shall identify the date of the inspection or test, the name of the person who performed the inspection or test, the serial number or other identifier of the equipment on which the inspection or test was performed, a description of the inspection or test performed, and the results of the inspection or test.

  • 1910.119(j)(5)
  • Equipment deficiencies. The employer shall correct deficiencies in equipment that are outside acceptable limits (defined by the process safety information in paragraph (d) of this section) before further use or in a safe and timely manner when necessary means are taken to assure safe operation.


A significant factor in maintenance is the potential for change in a piece of capital equipment. In the Texas City BP case study, there were several pieces of instrumentation that had been changed without proper documentation. This affected several business processes downstream, specifically, startup procedures. Due to historical events like Three Mile Island and Chernobyl, the Nuclear Regulatory Commission (NRC) has created excellent guidelines for configuration & change management.
    10 CFR 50.65
  • (a)(1) Each holder of an operating license for a nuclear power plant under this part and each holder of a combined license under part 52 of this chapter after the Commission makes the finding under § 52.103(g) of this chapter, shall monitor the performance or condition of structures, systems, or components, against licensee-established goals, in a manner sufficient to provide reasonable assurance that these structures, systems, and components, as defined in paragraph (b) of this section, are capable of fulfilling their intended functions. …When the performance or condition of a structure, system, or component does not meet established goals, appropriate corrective action shall be taken….

  • (4) Before performing maintenance activities (including but not limited to surveillance, post-maintenance testing, and corrective and preventive maintenance), the licensee shall assess and manage the increase in risk that may result from the proposed maintenance activities. …

  • (b) The scope of the monitoring program specified in paragraph (a)(1) of this section shall include safety related and non-safety related structures, systems, and components, as follows:

  • Safety-related structures, systems and components that are relied upon to remain functional during and following design basis events to ensure the integrity of the reactor coolant pressure boundary, the capability to shut down the reactor and maintain it in a safe shutdown condition, or the capability to prevent or mitigate the consequences of accidents that could result in potential offsite exposure comparable to the guidelines in Sec. 50.34(a)(1), Sec. 50.67(b)(2), or Sec. 100.11 of this chapter, as applicable.

  • Non-safety related structures, systems, or components:

    • That are relied upon to mitigate accidents or transients or are used in plant emergency operating procedures (EOPs); or

    • Whose failure could prevent safety-related structures, systems, and components from fulfilling their safety-related function; or

    • Whose failure could cause a reactor scram or actuation of a safety-related system.


Some may suggest that in addressing CFR’s, all that is necessary is to draft a procedure or policy. That may be true until the production process has a failure that affects product quality (e.g. Pharmaceuticals) or people are injured or killed (e.g. Chemical Processing, Pharmaceuticals, Aviation, Manufacturing). This exposes the company to higher risk and has the resultant negative publicity. In regulated industries, federal marshals can walk in with a warrant and walk out with executives in handcuffs.

The most basic parts of a reliability program will address compliance requirements, mitigate or eliminate failures and reduce the cost to maintain assets that fall under various CFR’s. The reason for this is the components of a comprehensive reliability program go deep into the business processes across organizational units as well as the company as a whole. In addition, new regulations or stringent enforcements of existing regulations may not become necessary.

In order to have a successful reliability program there must be the following components:
  • Maintenance History Tracking

  • Standard Operating Procedures (SOP)

  • Materials Specification Requirements

  • Materials Stores Management (MRO)

  • Management of Change (MOC)

  • Maintenance Business Process

  • Design for Reliability


These components provide a foundation to build good data on the asset base, thus improving equipment health. The information gathered gives the entire business unit the power to make well-informed decisions. Everything from product quality to throughput capacity can be identified in the context of reliability.

Taking the road to reliability enables the organization to do the right thing for its company and the community in which the company resides. Eliminating failures and exceeding regulatory requirements will reduce government intervention and lead toward a more proactive organization/environment. Less government oversight will reduce operating expenses once thought necessary to achieve compliance.

Tuesday, November 2, 2010

Don't Let Them Shut You Down: Part 2

Don't Let Them Shut You Down: Part 2

Another significant event was the power outage in 2003, which blacked out parts of New York, Ohio, and Pennsylvania. The congressional investigation focused heavily on the system failures from the overloading of the grid. They talked about weak infrastructure and an aging power grid.

However, the root cause of the system wide failure was not discussed. According to the investigation, deferred maintenance was the initial cause.

Several trees scheduled for trimming contacted power lines in Ohio, tripping the first substation. Why this occurred is not discussed, but new regulations were introduced just as fast as Congress could write them. Among them was a new requirement for tree trimming to a specific distance from the power line.

In March 2005, the Texas City, TX BP refinery experienced an explosion that was felt for miles. This particular event was responsible for the deaths of 15 people and injuries to another 180 people.

The explosion occurred during the startup of a raffinate tower after a maintenance shutdown. There were many factors in this case which contributed to the resulting explosion. Among them, operator fatigue, outdated standard instructions, and malfunctioning instruments among others.

This event also had a more obvious management connection. BP’s drive toward cost cutting had reduced head counts and forced the plant to rely heavily on contract labor. There had also been a halt put on any work efforts to update the process equipment. Much of the Texas City equipment had outlived its planned life cycle. This series of decisions and equipment health created the potential for a large disaster, and the worst-case scenario was realized.

The Reliability Connection

Maintenance is not a passive player in the events listed above. It was an integral part of each incident and contributed significantly to the outcome. A holistic reliability approach to the maintenance programs in each of these cases could have prevented their occurrence and outcomes.

The link lies in understanding the Codes of Federal Regulation (CFR). There are several common themes in CFR’s that a comprehensive reliability program will address and/or exceed.

21 CFR Part 211.67; FDA
Equipment cleaning and maintenance.
  • (a) Equipment and utensils shall be cleaned, maintained, and sanitized at appropriate intervals to prevent malfunctions or contamination that would alter the safety, identity, strength, quality, or purity of the drug product beyond the official or other established requirements.

  • (b) Written procedures shall be established and followed for cleaning and maintenance of equipment, including utensils, used in the manufacture, processing, packing, or holding of a drug product. These procedures shall include, but are not necessarily limited to, the following:

  • (1) Assignment of responsibility for cleaning and maintaining equipment;

  • (2) Maintenance and cleaning schedules, including, where appropriate, sanitizing schedules;

  • (3) A description in sufficient detail of the methods, equipment, and materials used in cleaning and maintenance operations, and the methods of disassembling and reassembling equipment as necessary to assure proper cleaning and maintenance;

  • (4) Removal or obliteration of previous batch identification;

  • (5) Protection of clean equipment from contamination prior to use;

  • (6) Inspection of equipment for cleanliness immediately before use.

  • (c) Records shall be kept of maintenance, cleaning, sanitizing, and inspection as specified in 211.180 and 211.182


The example contained in the FDA code illustrates this reliability connection. Section (a) and (b) specifically instruct the company to identify responsibilities for maintenance and draft standard work procedures.

40 CFR Part 68.73; vessel mechanical integrity
  • Written Procedures

  • Training For Process Maintenance Activities

  • Inspection and Testing

  • Equipment Deficiencies (“Operator shall correct deficiencies…”)

  • Quality Assurance (“Correct Materials, Correct Design, Proper installation.)


This EPA Regulation is specifically for vessel integrity. This is an example where industry specific codes are addressed through a comprehensive reliability program. Note, there are statements in the code requiring written procedures, training and corrective actions for equipment deficiencies.

Tuesday, October 26, 2010

Don't Let Them Shut You Down: Part 1

Don't Let Them Shut You Down: Part 1

Maintenance and Regulatory Compliance

Many maintenance organizations do not realize how they are affected by regulatory compliance. However, much of what Maintenance does has a direct affect on compliance with federal regulations and can cost the company millions of dollars if their actions cause an out-of-compliance situation.

Understanding this from within maintenance organizations varies from complete ignorance to maintenance by decree. The latter is a maintenance strategy that is defined by fear of the regulatory bodies and causes paralysis when trying to adopt proactive maintenance strategies. On both ends of the spectrum, reactive maintenance is the prevailing strategy and changing to a true proactive reliability program is very difficult. Regardless of their understanding of the regulations which governs their business, the maintenance program is most likely not mitigating equipment failures. This, therefore, leaves the organization open to risk and regulatory oversight.

New Regulations

When the unhappy constituents of a congressional district or local government entity call their representatives, it is often the starting point of new regulations. Specific events can also drive the government to draft new regulations or increase enforcement of existing ones. This call for action becomes, particularly, loud when people are hurt or killed because of corporate neglect.

It is important to understand how and why new regulations start because this understanding is the key to preventing further creation of new ones or heavy-handed enforcement of existing ones. Both of these things can be avoided if we as industrial professionals take a proactive approach.

There are many recent events that have caused significant news coverage and bad publicity for all of industry. Some events have led to congressional investigations. It is a bad day for maintenance and engineering when a CEO has to testify on Capitol Hill for a catastrophic equipment failure that affects the public. Maintenance has moved into the limelight, and is receiving public attention, perhaps for the first time in the history of “wrench turning”.

Case Studies

Many examples of maintenance culpability have been documented and reported on in the last several years. All of the events mentioned within many of the case studies written could have been avoided if a well-developed maintenance plan and reliability program were in place.

On April 4 2008, the FAA grounded Southwest Airlines 737-300 aircraft in order to perform airframe inspections. While there was no impact to the public in terms of safety, this event became high profile because of news reporting.



According to the investigation reports, Southwest had deferred several airframe inspections. These inspections were to identify cracks in the fuselage of a certain size. Upon closer examination of this case, it turns out that the inspection had been developed by Southwest and far exceeded the minimum requirements of Boeing. Even the FAA recognized this as a non-critical inspection and as such, issued a non-mandatory airworthiness bulletin.

    “A progressive inspection for fuselage skin cracking was initially distributed to operators in the form of a "non-mandatory" Service Bulletin (SB) that provided "risk mitigation" actions that operators were encouraged to incorporate into their maintenance program. This Service Bulletin was based, in large part, on an inspection program developed by Southwest Airlines. …cracks in the fuselage skin on the Boeing 737 airplanes were identified and mitigated well before they could pose a safety of flight issue. …the FAA did not regard the skin cracking as an "immediate threat" to the safety of flight of the airplane.”


Even though the FAA deemed there to be NO Safety risk with these deferments, they fined Southwest Airlines $10MM and delayed thousands of passengers. The lesson to take away from this event is, “Do what you tell the regulators you are doing.”

Tuesday, October 19, 2010

Elevating Maintenance and Reliability Practices The Financial Business Case: Part 5

Elevating Maintenance and Reliability Practices The Financial Business Case: Part 5

Costs – Top performers begin with some form of gap analysis to understand the current state of relevant practices and to measure gaps that exist between current state and top performance. From that point, calculating costs to close gaps is objective and fairly accurate. Major investment categories typically include:

  • Development of Corporate Standards for work management, materials management, configuration change management and reliability excellence

  • Development of a Roll-out and Implementation Strategy taking advantage of work done at one plant as appropriate for other plants

  • Creation or Improvement of Foundational Information (Functional Location Hierarchy, Master Equipment List, Spares Materials Catalog, Bills of Material/Parts Lists)

  • Objective Criticality Ranking of Equipment

  • Methodical Analysis of Failure Modes, using combination of Reliability Centered Maintenance Analysis (RCM), Failure Modes and Effects Analysis (FMEA), and templating where appropriate, to determine the optimum PM and PdM activities that need to be deployed for your population of equipment

  • Based on methodical analysis, perform PM Optimization, eliminating unnecessary PMs, deploying recommended PdM, and creating the PM/PdM work orders in the CMMS system to automatically schedule these activities

  • Creation of Balanced Metrics Measurement system

  • Training and Awareness

  • Culture Change and Rewards System Alignment

  • Compliance Monitoring and Continuous Improvement


There is a lot of guidance that can be used to estimate the costs of closing gaps, but for purposes of this article, suffice to say that while these costs are not insignificant, in the context of the benefits and the financial business case, they are almost always easily justifiable, with typical Returns on Investment (ROI) from 8:1 to 16:1 and higher, and with Internal Rates of Return (IRR) from 50% to 250% or higher.

Summary:

Well, in summary, what do we know and what do we believe?

  • We know what good looks like, and a big part of that picture can be summed up with the phrase “More Predictive and Less Preventive”. We know that Predictive Maintenance is driving a large percentage of work on a daily basis at the top performing plants, and this, of course, is good news for the readers of this magazine. Our time has come!

  • We know that the top performers achieved their success using remarkably similar practices – regardless of their industry, so we shouldn’t spend a lot of time debating what good looks like.

  • We know that you can’t piecemeal your way to prosperity - the top performers attacked the opportunity holistically – weaving all of the aspects of a top-level practice carefully together to unlock the hidden benefits.

  • We know that even the top performers have been unable to uniformly elevate their maintenance and reliability practices across the entire enterprise, and we believe there are good business reasons for trying to do so, including reduced cost of implementation company-wide (vs. taking a plant-by-plant approach) and increased ROI.

  • We believe the size of the opportunity is three quarters of one Trillion dollars annually in the U.S. alone, and could exceed $2 Trillion world-wide!
  • We know the direct benefits will come from maintenance spend reduction, spare parts inventory reduction, reduced energy consumption, improved quality, reduced scrap and increased throughput/asset utilization.

  • We believe there is a correlation between success of any corporate improvement initiative – whatever it is - and improved reliability practices. The indirect benefits come from unlocking hidden benefits in other parts of the business previously thought to be unrelated to reliability, and they can be substantial.


Finally, we know that the financial business case for reliability – including predictive maintenance - is here, and the awareness in your executive suite is emerging. If you are involved in predictive maintenance, I urge you to be confident in what you are doing because the role you are playing is essential for your corporation to achieve success – and the executives in your company are figuring that out!

Robert S. DiStefano
Chairman and CEO
Management Resources Group, Inc.
Southbury, Connecticut

Tuesday, October 12, 2010

Elevating Maintenance and Reliability Practices The Financial Business Case: Part 4

Elevating Maintenance and Reliability Practices The Financial Business Case: Part 4

How Big Are The Benefits?

Recently, we studied statistics from the United States Department of Commerce, including their measurement of what they call “Net Stock of Private Fixed Assets” in various industries. This measurement is a close proxy of Replacement Asset Value (RAV). In 2003 (the latest statistics available from the USDOC), there were $4.9 Trillion of physical assets on the ground in United States industry. We applied our Four Quartile Benchmark Statistics of Maintenance Spend as a percentage of RAV, and we dollarized the value of elevating Fourth Quartile plants to the First Quartile in maintenance spend, moving the Third Quartile plants to the First Quartile, and moving the Second Quartile plants to the First Quartile. As you can see from the following chart, industry wastes approximately $183 Billion in excess maintenance spend annually in the United States alone!


* Calculated from Department of Commerce Current-Cost
Net Stock of Private Fixed Assets in 2003 (Total $4.9 Trillion)


Further, we can assume from numerous published case studies that three to seven times the maintenance spend reduction benefit is accomplished in operational benefits (including increased uptime, improved quality, more efficient production scheduling, reduced waste, reduced energy consumption, reduced inventories, etc.). Taking the conservative end of that statistic (three times maintenance spend reductions), you can see from the chart that another $553 Billion in “Productivity Losses” can be re-claimed through the maintenance and reliability improvements, making the financial business case in the United States alone $738 Billion in annual, recurring benefits. What is this number world-wide? Good question. We are currently trying to quantify that with good data, however our intuition is that, if the U.S. opportunity is conservatively estimated at three quarters of one Trillion dollars, the world-wide annual benefits could be $2 Trillion or more!

The following chart depicts the Reliability Adoption Life Cycle.


Assuming that 25% of plants have figured this all out (top quartile), the market is at the Early Adopter/Early Majority stage. 75% of plants have improvements to make and work to do. If we look for an example of a company that has uniformly elevated their practices fleet-wide, there are no examples, so we are still looking for the innovators. It should be pointed out though, that attacking the opportunity fleet-wide will ease the journey by reducing the level of effort necessary to implement the practices and make the changes. Attacking this fleet-wide should leverage work done once for reuse avoiding the re-inventing of reliability over and over again. The resultant lower investment should make it easier to justify the expenditures for foundational and culture change work, enhance the Return on Investment and speed the Rate of Return.

What Are the Benefits in Your Company?

Quantifying the potential benefits, as well as the likely costs to improve performance, in your corporation, is necessary. Here is some guidance.

Benefits - Here are some of the major benefit categories with some guidance on how to calculate the potential:

  • Maintenance Spend Reduction: Calculate your maintenance spend as a percentage of Replacement Asset Value (RAV), and dollarize the improvement to top quartile performance (approximately 2 – 4% of RAV or better). If you are currently spending 5 – 6% or more, this benefit could be significant. The benefit comes from eliminating unnecessary work, working more efficiently, reducing the need for abundant stocked spares, eliminating collateral damage thereby reducing use of spare parts, reducing use of contractors, reducing overtime.


  • Inventory Reductions: Calculate your stocked inventory value (include satellite spares, etc.) as a percent of RAV, and dollarize the improvement to top quartile performance (approximately 0.5% - 1.5% of RAV). The actual reduction will yield on average $0.20 cents on the dollar of reduction (some inventory will have to be scrapped). This is a one-time benefit. In addition, the recurring annual avoided carrying costs will be on average 25% of the full inventory reduction value – annually.


  • Energy Consumption Reduction: Published guidelines show us that smoother running rotating equipment and leak-free operation of water, steam and compressed gas handling equipment will consume from 3% to 14% less energy (electricity, fuel).


  • Increased Uptime: Increased Asset Utilization can have a variety of substantial financial benefits to a company, including selling more product on the existing capital assets (assuming the demand for the additional product is present), or reducing the cost of goods made on the capital assets through more stable operations (even if the demand for additional product is not present). Two downtime areas should be targeted: Unscheduled Maintenance-related Downtime and Scheduled Maintenance Downtime. Unscheduled Maintenance-related Downtime can eventually be almost eliminated. Scheduled Maintenance Downtime in a plant heavily dependent on time-based Preventive Maintenance strategies can be reduced by from 30% to as much as 60% (depending on the starting point). Dollarizing the value of this varies from business to business, however remember that these benefits can be as much as 3 to 7 times larger than the maintenance spend reduction!


  • Improved Quality: Typically, scrap material and rejected/returned off spec product is measured accurately in most corporations. Calculate the value of the scrap material and assume that between 5% and 16% of that value can be eliminated through sound reliability practices. In addition, calculate the value of the rejected/returned product and assume that between 1% and 5% of that value can be eliminated through sound reliability practices. These statistics will vary business to business.