Wednesday, December 31, 2008

The Exception-Tolerant Organization - Part 5

As explained in my introductory post on the Exception-Tolerant Organization, ETOs do not tolerate waste, stifling of innovation, or inflexibility. These are the corrosive agents that prevent an organization from smoothly handling exceptions, maintaining a competitive edge, and operating effectively.

When I was at IBM Research I knew of a research Fellow who would monitor the keystrokes made by his administrative assistants while they were typing, and would spend quite a bit of time with them working on utilizing the fewest keystrokes possible. This is intolerance of waste taken to the extreme, where a few burning trees are saved while sacrificing the forest. Although it is quite easy to disdain waste made at this low level, this is not the kind of intolerance we are discussing here.

While the canonical definition of waste is anything that does not add value, what exactly constitutes waste for the Exception-Tolerant Organization? As previous posts have mentioned, exceptions happen to organizations, to customers, and to people every single day. Those organizations that cannot handle the exceptions and improve from them will fight a constant perception of lower value in the eyes of their customers.

For those who understand the benefits of mapping a value stream, waste would be anything that slows the velocity of movement through the value stream. The inability to swiftly handle exceptions can bring this velocity down to near zero. Being Exception-Tolerant can keep the value stream moving at the pace of innovation.

Waste for the ETO, then, is the set of roadblocks to handling exceptions: inflexibility combined with the stifling of innovation.

Inflexibility is tough to exhibit while being Exception-Tolerant, but is even tougher to recognize in an Exception-Intolerant organization. After all, there are certainly industries and products where strict and rigid standards must be constantly and consistently adhered to (watch-making, food preparation, chip manufacturing, aerospace), but these should not be mistaken for inflexibility. Chip manufacturers, despite adhering to the use of 30+ year old computer code, have been able to find avenues of flexibility in many ways over the past few years - from lower-voltage conduits to hyper-threading to multi-core pipelines.

Stifling of innovation is perhaps the touchiest subject when it comes to organizations who may be looking to become Exception-Tolerant. The urge to suppress innovation is strong in risk-averse environments, and is often exercised in such forms as:

- the boss consistently saying "No"
- a departmental control group demanding that a potential innovative division look and operate the exact same way as the existing divisions
- new ideas whose implementations are unduly laden with process, direct-to-archive documentation, or extreme executive input

The stifling of innovation is often a by-product of culture, and the effects may not be felt in world markets until long after the key culture-makers have moved on. Strong cultures that stifle innovation are not often discovered by the outside until the effect in the marketplace becomes obvious. ETOs, even ones with strict standards, have a tough time saying "No". Inflexible organizations find it all too easy to say "No" and close doors on new ideas and customer needs by stifling innovation.

But with all the good work that we as people do, and all the work that we endeavor to do, how do we let ourselves get into a mode of inflexibility and resistance to innovation? The hard truth here is that we most often become inflexible and stifle innovation when we put self-interests above working together as a team. There is good reason that this struggle is often labeled as the war within.

But the other side of this hard truth is that ETOs are given credit by their customers for handling the exceptions and increasing the overall value proposition. And when credit is given in this fashion, ETOs give their people, who have worked together as a team, every opportunity to take the credit and reap the rewards - thus making team achievement in the best self-interest.

Friday, December 19, 2008

Valuating Technology System Delivery

A well-worn but still prevailing theory of system delivery states that a system can be delivered according to three main factors:

- the system is of high quality and functionality (a "good" system)
- the system runs speedily under all conditions and usage loads (a "fast" system)
- the system can be built and delivered inexpensively (a "cheap" system)

The conventional wisdom is that a system can be delivered possessing AT MOST two of the three above factors. In other words:

- a good and fast system cannot be delivered cheaply
- a good and cheap system will not run speedily
- a fast and cheap system will not be a good system

You may or may not hold this view yourself, but many still do. Of course, it is quite possible (and rather beneficial) to have all three factors represented positively in a system delivery. But you need to look at the entire lifecycle of a system to understand this, not just the planning and construction stages.

Many under-weigh, or ignore completely, the usage and maintenance portions of a system's lifecycle when evaluating TCO and ROI. If the system's features, functionality, and construction are underrepresented before the system is delivered, higher costs will be made apparent later in the lifecycle:

- the system's users will invent workarounds to their normal business processing, just to accommodate the system's lack of performance or mis-matched view of the world. This usually leads to weaker controls and gaps in business processing, which leads to anywhere from higher operating costs to lost revenue.

- the system performs poorly or contains only a few compelling features, but just barely enough to be useful. While there could be no justifiable loss in cost, the cost to keep-alive the system plus the cost to improve and re-architect the system could easily be more than double the initial funds promised. Building a more desirable and well-performing system might have cost more up front, but still could have been less than the funds needed later for repairs.

- the system performs so poorly, or contains such a lack of compelling features, that it will not be used at all. This represents a loss equal to the cost of construction of the system, plus the difference in current operating costs that may have been saved, plus the cost of lost opportunity to improve and be competitive.

So how do you deliver a system that is good, fast, and cheap? Stating the system's costs and its benefits over the ENTIRE system lifecycle, both with equal precision, will help you arrive at acceptable definitions for "good", "fast", and "cheap" for your organization. But without the equality in precision, someone is bound to be disappointed upon delivery of the system.

So which lifecycle phases and cost factors are you accounting for in your system delivery analysis? Do you feel that you have all the bases covered? And how is the precision measured?

Act Like An Owner

Does your organization ask you to act, and think, like an owner? Does your organization ask this of you when you come aboard?

What would an Owner do in your organization? Would an Owner carry a vision? Innovate? Seek improvement and new opportunities?

Would an Owner cut corners? Cut costs? Maintain the status quo?

Would only an Owner understand the key drivers to the business? Or could anyone else in the organization achieve such an understanding?

Does your organization educate and allow you to understand the key drivers to its business?

Does your organization allow you to cut corners and costs? To maintain the status quo?

Does your organization allow you to carry a vision? Innovate? Seek improvement and new opportunities?

Your organization may ask you to act like an owner - but does it allow it?

The Exception-Tolerant Organization - Part 4

As explained in my introductory post on the Exception-Tolerant Organization, ETOs practice daily Risk Management. Another concept that sounds quite simple, but organizations find it difficult to sustain due to these key factors:

- lack of inertia and momentum
- lack of systematized and/or automated assistance
- key-man risk and blame

Lack of inertia and momentum comes from not instilling review activities into the organization's daily operations. Just asking these fundamental questions on a daily basis goes a long way towards an effective review process:

- What did I accomplish yesterday?
- What am I working on today?
- What obstacles or problems am I facing?
- How confident am I that the current goals will be accomplished on-time?

For many organizations that rely on reviewing large amounts of performance or other time-sensitive data, having a systematized and automated collation and presentation of this data is key to establishing a daily momentum. The systems and processes supporting them should have the appropriate fail-safes and redundancies to handle exceptional processing cases so that timely delivery (and thus Risk Management momentum) can be maintained.

Automated assistance is also key for the human side as well. In reviewing an organization's performance or exception conditions, leaning on automation can clear our heads of the mundane and manual steps, and allows us to focus on the exception conditions. Plus, what Director of Risk Management wants to arrive at work at 7 am just to push a button, or wait for reports 1 through 10 to finish running and printing before tackling the issues of the day?

Being a Director of Risk Management is a tough proposition to satisfy for an organization. Not only must you have volumes of data and information at your disposal, you must have the prescience to understand what is going to happen seconds from now in both your own building and halfway around the world. A Risk Management Director could be hailed as the heroic steward of the ship that avoids the icebergs of a competitive landscape, or could easily be the goat when the ship is steered right when perhaps it should have turned left.

For an organization, it is easy to both funnel singular responsibility and cast blame on one person when things blow up. But why do this, when after the blame subsides, the organization is still in an unfavorable situation when a risk materializes? Risk Management is one of those cross-cutting activities that the entire organization can practice effectively on a daily basis via continuous review and improvement. Allow your Risk Management Directors to get the support of the entire organization, and the Director will support the viability and competitive health of the organization in return.

Saturday, November 15, 2008

The Exception-Tolerant Organization - Part 3

As explained in my introductory post on the Exception-Tolerant Organization, ETOs emphasize their strengths and compensate for weaknesses by working together and communicating openly. If the concept sounds simple, that's because it is simple. But with all of the advanced ways, means, and tools we have at our disposal to collaborate and communicate openly, we still find in today's world places and situations where these fundamentals are simply not done.

Working together involves several cultural practices within an organization:

- Total team involvement from the start
- Proactively seeking the improvement of the organization
- Staying open to new ideas and possibilities
- Shared goals without harmful personal agendas

The first point is easier done than said, but the other three involve a resistance factor to change within our organizations. ETOs have the cultural grounding to foster and promote all four points above. As discussed in Part 2, ETOs systematically implement the communication pathways and processes to support these points.

But as is often the case in life, the greater challenge in overcoming obstacles can come from within ourselves, especially on the last point above. We don't always realize that we are more in competition with ourselves than we are at odds with others, and many times we are better served improving ourselves and overcoming our own shortcomings. Instead, we often put up walls where requirements and miracle solutions are volleyed back and forth, neither ever really satisfying the concerns or goals of the parties on both sides of the wall. But when we take down the walls and mprove ourselves, our best foot is then truly placed forward for the benefit of our organizations. Our strengths become the organization's strongest capabilities to be deployed across the greatest spectrum of benefit. And our weaknesses are compensated for by our continuous self-improvement, by identifying risks early, and especially by communicating openly.

Communicating openly, at its most basic level, involves face-to-face discussion, debate, and sometimes conflict - something that people can tend to avoid, in both their personal and professional lives. But for open communication to be effective, an environment needs to be created by the organization where intense debate and disagreement are tolerated, and conflicts can be satisfying to resolve. The organization should make it clear that it is okay to disagree and debate issues, but the result of each debate should still include a clear decision to move forward along a certain path of action.

As an exercise, have everyone in your organization begin to answer these questions out loud daily. Ideally, everyone related to a project or a core business of your organization should be in the same room (or on conference if necessary):

- What did I accomplish yesterday?
- What am I working on today?
- What obstacles or problems am I facing?
- How confident am I that the current goals will be accomplished on-time?

For the last question, use some kind of rating scale. Example: have each person rate their confidence from 1 to 5, with 5 being the most confident. Any answers below a 4 should be a cause for concern and those concerns should be addressed openly. The Agile practice called Scrum advocates the daily use of these questions.

Communicating openly also involves the appropriate use of communication tools. Communications involving urgency or time sensitivity should use direct methods: direct-line phone, internet, and video calls; direct text messages/pages; and of course face-to-face conversation. The use of email and instant messaging, while becoming increasingly mobile and location-independent, should not be relied on for urgent time-sensitive communication. These communication methods often have either multiple inboxes or streams/threads of communication occurring simultaneously, have lengthy queues of messages attached to them, or depend on having your communication device successfully "subscribe" to that message stream. The number of inboxes/threads and the size of the inbox queues are things that you cannot guarantee to be small enough so that your urgent message is received timely. Eliminate this frustration up front by identifying early a reachable line of communication to use for urgent matters.

Some very simple and fantastic exercises in working together and communicating openly (many taking 10 minutes or less to complete with a noisy room full of people) can be found here. While some of these are focused on Agile principles and practices, many of these deal with the core issues of communication and collaboration, and may help to expand your thinking. My thinking was certainly expanded after participating in some of the exercises. My thanks to Michael De La Maza for bringing these to my attention.

Monday, November 10, 2008

Challenging Our Assumptions

At a recent party I sat with a technology team whose members were dejected from having their latest project terminated by their company. This team spoke fondly about their months of development, their early adoption of cloud computing methods and technologies, and even their conferences and consultations with members of NASA regarding the use of advanced technologies. They shook their heads and couldn't understand why the system didn't succeed.

This team had hailed the project as the next great data transmission hub for the team's company. They specified communication protocols and APIs, worked with the application developers to foster understanding, and even reviewed a few sample data streams to understand the problem space. The project seemed to have everything going for it. But during construction of this hub, the team made a fateful design assumption regarding the separation of data streams within communication channels.

This assumption was that data streams could be separated by a data pattern that would "never be seen" in the actual live data itself. The team was warned by the application developers and certain managers that the possibility of this data pattern showing up in live data was in fact likely. Still, the team moved forward with completion of the system, performed some successful testing with a thin time-slice of data transmissions over a period of several days, and then went live with the system.

Within three hours of the system going live, this data pattern delimiter showed up in several places in the live data. As a result, the system was delivering incomplete data and crossing data transmission streams. Processes that were dependent on this new data transmission hub had to be reverted back to their old transmission methods. Beyond changing the data pattern delimiter to something else that might "never be seen" in the actual live data, the team did not have an alternative solution to this problem. And thus a short time later, the project was terminated.

Unfortunately for this team, a fundamental design assumption was made early on that contradicted directly with the problem space that their solution would address. The lack of challenge, investigation, and thorough testing of this assumption sealed the project's fate. While we are often proud of our assumptions, we are made better when our assumptions are challenged and tested at the earliest possible opportunity.

Thursday, November 6, 2008

Project and Technology Development Methodologies

Rather than be the trillionth blogger writing on the subject of Agile, refactoring, Test-Driven-Development, domain-driven-design, and other practices and paradigms, I'll lay out some pragmatic principles to keep in mind when executing projects, no matter which development practices you adopt. In the rush to absorb the hottest trends, we should keep in mind the tried-and-true principles that still work:

Start with iterative and incremental. Project phases are well and good, but in order to complete a project successfully, the path to completion needs to be traversed. And yes, you must have an idea of your destination. But will you wait until all requirements are specified to the last detail before moving on to any architecture or design concerns? Will you have a design project phase that takes into account the current state but fails to connect with the changes in the world three months later when the design phase is completed?

If you break down your project phases into iterations, with a guaranteed feedback session at the end of each iteration, you will be much more likely to keep up with changing requirements and conditions. You will also be able to identify failing efforts sooner, before they cause damage later.

Start as soon as you can. One of the common reasons for project slippage is that some critical component or project dependency was not available earlier in the project cycle. Was it because the need for this component was not forseen? No, it was just not available at the appropriate time. One way to mitigate this is to start working on these components after just enough baseline requirements have been gathered. You may spend a little more up front to support this initial development, but it will cost you far more in the long run should you end up in an ill-timed slippage scenario. The ways it may cost you range from longer development cycles to mis-timed market entry to project personnel turnover. These are big-ticket costs.

Military leaders are trained to make decisions with 40% to 70% of the information necessary, and regularly provide feedback as more information is available. You can do the same on projects and be very effective.

Do your Risk Management. As stated so eloquently in the fantastic treatise on software project management Waltzing With Bears, "Risk Management is Project Management for adults." Listing and valuing project risks is something that can be done to 100% completion up front. But the great (and sometimes painfully realized) thing about Risk Management is that 100% completion is often not enough. Here is a great place to be iterative and incremental in providing feedback loops on the state of risks, mitigating circumstances, and changing requirements. If you build the iterations and feedback loops into your risk management on a project, the rest of your project practice will need to follow suit just to keep up.

You'll also be more inclined to tackle project deliverables that resolve the greatest risks first. The sooner these risks are resolved, the more accurately you can publish a date range for delivery.

Test early and continuously. Your project exists in some form from start to finish. It may start on paper, and live in hardware and software at the finish. But in whatever form it exists, it can be tested. As new requirements are gathered and formulated, test cases can be drafted and pitted against the project's assumptions. As new hardware is installed, its images and monitoring agents can be configured and proven. As new code is written, test cases can be written before or alongside the code.

Don't wait for a project to be 50% or 75% completed to start drafting up test cases and setting up a test facility. Automate much of your tests if you can, so that they can be run at least once a day. Automation is again a case of a little more effort up front, saving much more cost and headache later.