Showing posts with label Software Process. Show all posts
Showing posts with label Software Process. Show all posts

Wednesday, July 30, 2014

The Data Driven Quality Mindset

"Success is not delivering a feature; success is learning how to solve the customer's problem." - Mark Cook, VP of Products at Kodak

I've talked recently about the 4th wave of testing called Data Driven Quality (DDQ). I also elucidated what I believe are the technical prerequisites to achieving DDQ. Getting a fast delivery/rollback system and a telemetry system is not sufficient to achieve the data driven lifestyle. It requires a fundamentally different way of thinking. This is what I call the Data Driven Quality Mindset.

Data driven quality turns on its head much of the value system which is effective in the previous waves of software quality. The data driven quality mindset is about matching form to function. It requires the acceptance of a different risk curve. It requires a new set of metrics. It is about listening, not asserting. Data driven quality is based on embracing failure instead of fearing it. And finally, it is about impact, not shipping.

Quality is the matching of form to function. It is about jobs to be done and the suitability of an object to accomplish those jobs. Traditional testing operates from a view that quality is equivalent to correctness. Verifying correctness is a huge job. It is a combinatorial explosion of potential test cases, all of which must be run to be sure of quality. Data driven quality throws out this notion. It says that correctness is not an aspect of quality. The only thing that matters is whether the software accomplishes the task at hand in an efficient manner. This reduces the test matrix considerably. Instead of testing each possible path through the software, it becomes necessary to test only those paths a user will take. Data tells us which paths these are. The test matrix then drops from something like O(2n) to closer to O(m) where n is the number of branches in the code and m is the number of action sequences a user will take. Data driven testers must give up the futile task of comprehensive testing in favor of focusing on the golden paths a user will take through the software. If a tree falls in the forest and no one is there to hear it, does it make a noise? Does it matter? Likewise with a bug down a path no user will follow.

Success in a data driven quality world demands a different risk curve than the old world. Big up front testing assumes that the cost to fix an issue rises exponentially the further along the process we get. Everyone has seen a chart like the following:

clip_image001

In the world of boxed software, this is true. Most decisions are made early in the process. Changing these decisions late is expensive. Because testing is cumulative and exhaustive, a bug fix late requires re-running a lot of tests which is also expensive. Fixing an issue after release is even more expensive. The massive regression suites have to be run and even then there is little self hosting so the risks are magnified.

Data driven quality changes the dynamics and thus changes the cost curve. This in turn changes the amount of risk appropriate to take at any given time. When a late fix is very expensive, it is imperative to find the issues early, but finding issues early is expensive. When making a fix is quick and cheap, the value in finding a fix early is not high. It is better to lazy-eval the issues. Wait until they become manifested in the real world before a fix is made. In this way, many latent issues will never need to be fixed. The cost of finding issues late may be lower because broad user testing is much cheaper than paid test engineers. It is also more comprehensive and representative of the real world.

Traditional testers refuse to ship anything without exhaustive testing up front. It is the only way to be reasonable sure the product will not have expensive issues later. Data driven quality encourages shipping with minimum viable quality and then fixing issues as they arise. This means foregoing most of the up front testing. It means giving up the security blanket of a comprehensive test pass.

Big up front testing is metrics-driven. It just uses different metrics than data driven quality. The metrics for success in traditional testing are things like pass rates, bug counts, and code coverage. None of these are important in data driven quality world. Pass rates do not indicate quality. This is potentially a whole post by itself, but for now it suffices to say that pass rates are arbitrary. Not all test cases are of equal importance. Additionally, test cases can be factored at many levels. A large number of failing unimportant cases can cause a pass rate to drop precipitously without lowering product quality. Likewise, a large number of passing unimportant cases can overwhelm a single failing important one.

Perhaps bug counts are a better metric. In fact, they are, but they are not sufficiently better. If quality if the fit of form and function, bugs that do not indicate this fit obscure the view of true quality. Latent issues can come to dominate the counts and render invisible those bugs that truly indicate user happiness. Every failing test case may cause a bug to be filed, whether it is an important indicator of the user experience or not. These in turn take up large amounts of investigation and triage time, not to mention time to fix them. In the end, fixing latent issues does not appreciably improve the experience of the end user. It is merely an onanistic exercise.

Code coverage, likewise, says little about code quality. The testing process in Windows Vista stressed high code coverage and yet the quality experienced by users suffered greatly. Code coverage can be useful to find areas that have not been probed, but coverage of an area says nothing about the quality of the code or the experience. Rather than code coverage, user path coverage is a better metric. What are the paths a user will take through the software? Do they work appropriately?

Metrics in data driven quality must reflect what users do with the software and how well they are able to accomplish those tasks. They can be as simple as a few key performance indicators (KPIs). A search engine might measure only repeat use. A storefront might measure only sales numbers. They could be finer grained. What percentage of users are using this feature? Are they getting to the end? If so, how quickly are they doing so? How many resources (memory, cpu, battery, etc.) are they using in doing so? These kind of metrics can be optimized for. Improving them appreciably improves the experience of the user and thus their engagement with the software.

There is a term called HiPPO (highest paid person's opinion) that describes how decisions are too often made on software projects. Someone asserts that users want to have a particular feature. Someone else may disagree. Assertions are bandied about. In the end the tie is usually broken by the highest ranking person present. This applies to bug fixes as well as features. Test finds a bug and argues that it should be fixed. Dev may disagree. Assertions are exchanged. Whether the bug is ultimately fixed or not comes down to the opinion of the relevant manager. Very rarely is the correctness of the decision ever verified. Decisions are made by gut, not data.

In data driven quality, quality decisions must be made with data. Opinions and assertions do not matter. If an issue is in doubt, run an experiment. If adding a feature or fixing a bug improves the KPI, it should be accepted. If it does not, it should be rejected. If the data is not available, sufficient instrumentation should be added and an experiment designed to tease out the data. If the KPIs are correct, there can be no arguing with the results. It is no longer about the HiPPO. Even managers must concede to data.

It is important to note that the data is often counter-intuitive. Many times things that would seem obvious turn out not to work and things that seem irrelevant are important. Always run experiments and always listen to them.

Data driven quality requires taking risks. I covered this in my post on Try.Fail.Learn.Improve. Data driven quality is about being agile. About responding to events as they happen. In theory, reality and theory are the same. In reality, they are different. Because of this, it is important to take an empiricist view. Try things. See what works. Follow the bread crumbs wherever they lead. Data driven quality provides tools for experimentation. Use them. Embrace them.

Management must support this effort. If people are punished for failure, they will become risk averse. If they are risk averse, they will not try new things. Without trying new things, progress will grind to a halt. Embrace failure. Managers should encourage their teams to fail fast and fail early. This means supporting those who fail and rewarding attempts, not success.

Finally, data driven quality requires a change in the very nature of what is rewarded. Traditional software processes reward shipping. This is bad. Shipping something users do not want is of no value. In fact, it is arguably of negative value because it complicates the user experience and it adds to the maintenance burden of the software. Instead of rewarding shipping, managers in a data driven quality model must reward impact. Reward the team (not individuals) for improving the KPIs and other metrics. These are, after all, what people use the software for and thus what the company is paid for.

Team is the important denominator here. Individuals will be taking risks which may or may not pay off. One individual may not be able to conduct sufficient experiments to stumble across success. A team should be able to. Rewards at the individual level will distort behavior and reward luck more than proper behavior.

The data driven quality culture is radically different from the big up front testing culture. As Clayton Christensen points out in his books, the values of the organization can impede adoption of a new system. It is important to explicitly adopt not just new processes, but new values. Changing values is never a fast process. The transition may take a while. Don't give up. Instead, learn from failure and improve.

Wednesday, July 9, 2014

Try.Fail.Learn.Improve

Try.Fail.Learn.Improve. That has been the signature on my e-mail for the past few months. It is intended to be both enlightening and provocative. It emphasizes that we won't get things right the first time. That it is okay to fail as long as we don't fail repeatedly in the same way. Try.Fail.Learn.Improve is a process that needs to be constantly repeated. It is a way of life.

When I first used this phrase, someone responded that it was too strongly worded. Perhaps I should say "Try, Learn, Succeed" instead. But that doesn't convey the true value of the phrase. I specifically chose the word Fail because I wanted to emphasize that we would get things wrong. I avoided the word succeed because I wanted to convey that the process would be a long one.

Try. The essence of getting anything done is to start. In the world of software and especially systems software, we are always doing something unknown. We are not building the nth bridge or even the nth website. As such, the answers are not known up front. How could they be? Thus we can't say, "Do." That implies a known course of action. Try is more accurate. Make a hypothesis about what might work and try it out. Run the experiment.

Fail. Most of the time--not just sometimes--what is tried will fail. It is important to be able to recognize when we fail. Trying something that cannot fail is also doing something from which we cannot learn. Only with the possibility of failure can learning be had. Failure should be expected. "Embrace Failure" is advice I gave early into my new role. Much traditional software has viewed success as the only metric. The downside is that failure was punished. When something is punished, it will diminish. People will shy away from it. Punishing failure will disincentivize people from taking risk. The lack of risk means a lack of failure and a lack of learning. Given that we don't know the correct path to take, this lack of learning ensures a lack of success.

Learn. Einstein is said to have defined insanity as doing the same thing over and over again and expecting a different result. It is possible to fail and not learn. This is usually the result of blame. In a failure-averse culture, admitting you were wrong has severe repercussions. Failure then is not admitted. It is not embraced. Instead, it is blamed on something external. "We would have succeeded if only the marketing team had done their job. The essence of learning is understanding failure. Why did things go differently than predicted? In the scientific method, this helps to set up the next experiment.

Improve. Once we have failed and learned something about why we failed, it is time to try again. Device the next experiment. If what we tried did not work, what else might work? Revise the hypothesis and begin the cycle again. Try the next thing. If at first you don't succeed--and you won't--try, try again.

Succeed. Eventually, if we improve in this iterative fashion, we will succeed. This will take a while. Do not expect it to happen after the first, second, or even third turn of the crank. Success often comes in stages. It is not all or nothing.

If success is such an elusive item and will take many cycles to achieve, is it possible to get to success faster? Yes. Designing better experiments can help. If we understand well enough to do them. Blind experimentation will take too long. But many times we don’t know enough to design great experiments. What then? This is the most likely situation. In it, the best solution is to learn to run more experiments. Reduce the time each turn of the crank takes. Tighten the loop. Design smaller experiments that can be implemented quickly, run quickly, and understood quickly.

One last word on success. It is impossible to succeed if you don't know what success looks like. Be sure to understand where you are trying to go. If you don't, how can you know if an experiment got you closer? You can't learn if you can't fail and you can't fail if you don't know what you are measuring for. As Yogi Berra said, "If you don't know where you are going, you'll end up somewhere else."

So get out there and fail. It's the only way to learn.

Sunday, February 7, 2010

Plan Intentionally

I previously wrote about being intentional, but focused mostly on intentionality in execution.  Being intentional is also important in planning.  When planning a new product or the implementation of a feature, it is important to explicitly consider all aspects.  It can be a temptation to move past the hard problems too quickly.  “We’ll get back to that later.”  Doing this can be disastrous.  It is important to note that the decision on how to solve a hard problem *will* be made.  It can be made up front in a thoughtful way, or it can be made on the fly at a later date, but it will be made.  Putting off the hard decisions merely makes it more likely that they will be made on the fly.

This advice may seem obvious, but people don’t always follow it.  For critical-path decisions, people know better than to leave them for later.  It is the difficult peripheral issues that are sometimes left undecided.  The boundaries between two modules might be such a place.  The decisions on how to implement the module will certainly be decided, but might the way it is extended by or interacts with others be left for later? 

I recall working on a feature for WindowsMe (this feature never shipped) which allowed video playback to be conditioned on some criteria.  Perhaps the parental levels were too high or the rights were not present to play a piece of video.  In that case, this feature would recognize this fact and stop playback.  The developer responsible for this feature had created a complex infrastructure with plug-in providers and a great signaling mechanism.  I was brought in to test the project late and took over from someone else.  After looking at the documentation and playing around with it for a bit, I thought I understood what was going on, but there was one part that confused me.  I went to the developer and asked him what happened when he sent the message that playback should stop.  Who listened to this event?  His response was that he didn’t know.  This hadn’t been specified.  Stop the presses!  Here we had a conditional playback system that was just shouting at the wind, hoping someone would hear it and do the right thing.  The developer of this system as well as the developer of the playback pipeline had both written fine pieces of software, but one detail hadn’t been intentionally planned and the end to end feature would not work.  Admittedly this is an extreme case, but similar things happen on a less obvious scale if people are not careful to plan intentionally.

It isn’t always that the decisions are considered unimportant.  It is that time runs out.  When a hard decision is bypassed because it takes too long to decide, the chances of circling back to it before time pressure says coding has to start is high.  It is better to take the time to make the hard decision up front and leave the easier decisions to be made on the fly.  Acknowledge that there may not be enough time to do all the planning desired.  There never is.  Then decide on the most critical items first.  Be intentional about what is and is not going to get planned.  This way the unplanned items end up being those that are most conducive to being planned on the fly.

In short: Tackle the hard issues early.  If you wait, they will be decided in a default/easy way.  If the issue was hard to decide up front, it will never be decided well in the midst of coding.

Wednesday, May 27, 2009

Five Books To Read If You Want My Job

This came out of a conversation I had today with a few other test leads.  the question was, “What are the top 5 books you should read if you want my job?”  My job in this case being that of a test development lead.  At Microsoft that means I lead a team (or teams) of people whose job it is to write software which automatically tests the product. 

  • Behind Closed Doors by Johanna Rothman – One of the best books on practical management that I’ve run across.  1:1’s, managing by walking around, etc.
  • The Practice of Programming by Kernighan and Pike– Similar to Code Complete but a lot more succinct.  How to be a good developer.  Even if you don’t develop, you have to help your team do so.
  • Design Patterns by Gamma et al – Understand how to construct well factored software.
  • How to Break Software by James Whittaker – The best practical guide to software testing.  No egg headed notions here.  Only ideas that work.  I’ve heard that How We Test Software at Microsoft is a good alternative but I haven’t read it yet.
  • Smart, and Gets Things Done by Joel Spolsky – How great developers think and how to recruit them.  Get and retain a great team.

 

This is not an exhaustive list.  There is a lot more to learn than what is represented in these books, but these will touch on the essentials.  If you have additional suggestions, please leave them in the comments.

Sunday, December 21, 2008

Why You Get Nothing Done When You Have So Much Free Time

Interesting musings on a subject I can attest to be true.  Why is it we get so much done when we're on a tight schedule but then fail to get anything done when we have a long vacation?  The same applies to work too.  Give someone a long time to get a project done and it will still come in late.  Give them a short time and it will be done earlier.  The author attributes this to Innumeracy.  That is, the inability of humans to understand large numbers.  When we see a huge amount of time, we don't understand the true magnitude or true limit of that size.  It is this same effect that often casues lottery winners to go broke.  They have what they think is a huge amount of money and don't understand when the finiteness catches up to them.  Likewise when a person has a seemingly lot of time available, they fill it with too much stuff and end up getting nothing (important) done.  The author has a suggestion for solving this.  He suggests applying the 80-20 rule.  Target the top 20% of the tasks you have to do.  The other 80% only if they can fill in the gaps.  Basically this boils down to prioritizing.  When we have a seemingly infinite amount of time in front of us, prioritizing doesn't feel necessary and so low-priority tasks dominate.  Perhaps this is why Agile works so much better than Waterfall.  By imposing numerous (arbitrary) deadlines mid-project, it forces prioritization. 

Friday, August 15, 2008

Code Review Options

There are many ways to conduct a code review.  Here are a few ways I've seen it done and the advantages and disadvantages of each.

Over-the-Shoulder Reviews

Walk over to someone's office or invite them to yours and walk them through the code.  This has the advantage that the author can talk the reviewer through the code and answer questions.  It has the disadvantage that it pressures the reviewer to hurry.  He won't want to admit he doesn't understand and will sign off on something without really validating it.  It also doesn't work for teams that are not co-located.

E-mail Based Reviews

This model requires some method of packaging the changes.  This is then e-mailed to another person for review.  The reviewer than then open the code in WinDiff or a similar tool and review just the deltas.  This tends to be efficient and allows the reviewer ample time to read and understand the code.  The downside is that the reviewer can take fail to prioritize the review correctly and it can sit around for a long time.

Centralized Reviews

Code reviews are sent (via an e-mail based system usually) to a centralized team of trained code reviewers.  These are usually the senior coders.  It has the advantage of helping to maintain high quality and uniformity in the reviews.  In a team trying to change coding styles or increase the quality of their code, this can be  a great way to do so.  The downside is that this team becomes a bottleneck.  If only a small portion of the team must review all the code, they can fall behind and reviews can back up.  Being part of the central review team can also be a drain on the most efficient coders.  Instead of writing code, they spend a large portion of their time reviewing code.

Decentralized Reviews

In this format, everyone in the team participates.  Any person can review any other person's code.  This form scales most easily with team size but can require some tools if it is to work on a large team.  Otherwise it becomes too hard to tell which people are participating and which are not.  The advantage of this form is that everyone is exposed to both sides.  This allows junior members to learn by reading code written by more experienced members.  The disadvantage is that the quality of the reviews is uneven.  Different members will have different abilities.  Over time this should even out as the team grows together.

Group Reviews

So far I've been speaking of reviews which involve only two people.  The reviewer and the reviewee.  There is another model in which a group gathers together in a room and the reviewer leads them through the code.  This model puts a lot of eyeballs on the code and is likely to find the issues better than any individual reviewer, but it also costs a lot of eyeballs.  This model uses the most manpower and is the least efficient reviewing model.  It suffers the same problems as the over-the-shoulder reviews except exacerbated.  The likelihood of people not being able to follow is high.  Additionally, the review will likely go slowly as people bring up questions about various points.  Finally, it may bog down into argument between the reviewers.  This model can be improved if participants review the code ahead of time and come prepared to only talk about their issues.  If they reviewers aren't reading the code in the meeting, it can be efficient.  Sometimes this is called a code inspection.

What We Use

There is no single right way to conduct reviews.  I won't pretend to proscribe a single manner of conducting them that will fit all organizations.  I will, however, tell you what I've found to work.  My team uses a distributed e-mail model.  We have a tool which will take the diff of the latest code in the repository and the code on your box and package it up.  The reviewee then sends mail to a mailing list asking for someone to review his or her code.  A reviewer volunteers and responds letting everyone know he is taking the review.  The reviewer then reads the code, sends his comments to the reviewee.  The reviewee then makes any required changes and eventually checks in the code.

Wednesday, August 13, 2008

Code Review Rights and Responsibilities

Code reviews are an important part of any project's health.  They are able to find and correct code defects before making it into the product and thus spare everyone the pain of having to find them and take them out.  They are also relatively cheap.  Assuming your team wants to implement code reviews, it is important to lay out the expectations clearly.  Code reviews can be contentious without some guidelines.  What follows are the rights and responsibilities as I see them in the code review process.

Participants:

Reviewer - Person (or persons) reviewing the code.  It is their job to read the code in question and provide commentary on it.

Reviewee - Person whose code is being reviewed.  They are responsible for responding to the review and making necessary corrections.

Must Fix Issues:

A reviewer may comment on many aspects of the code under review.  Some of the comments will require the reviewee to change his code.  Others will be merely recommendations.  Must fix issues are:

  • Bugs - The code doesn't do what it was intended to do.  It will crash, leak memory, act erroneously, etc.
  • Potential maintenance issues - The code is not broken, but is written in such a way that it will be hard to maintain.  Examples might be magic numbers, poorly named variables, lack of indirection, lack of appropriate comments, etc.
  • Coding standard violations - If the group has a coding standard, it must be followed.  Deviations from it must be fixed when pointed out.

Recommendations:

Other items are merely recommendations.  The reviewer can comment on them, but the comments are only advisory.  The reviewee is under no obligation to fix them.  These include:

  • Architectural recommendations - The reviewer thinks there is a better way to accomplish the goal.  Seriously consider changing these, but the reviewee can say no if he disagrees.
  • Style issues - The reviewer wouldn't have done it that way.  Fascinating.

Code Ownership:

In my teams there is no ownership of code.  Some people touch certain pieces of code most often and may even have written the code initially.  That doesn't give them special rights to the code.  The person changing the code now does not need to get the initial writer's permission to make a change.  He would be a fool not to consult the author or current maintainer because they will have insights that will help make the fix easier/better, but he is under no obligation to act upon their advice. 

Tuesday, July 15, 2008

10 Pitfalls of Using Scrum in Games Development

Interesting article about using scrum to manage game development.  Many of the pitfalls are true beyond games development.  The article is well balanced and has advice for how to overcome the pitfalls.  I don't agree with all of the advice, but it is thought provoking.  For example, the article makes a good point that daily standup meetings can be disruptive to the thought process.  It therefore recommends using an electronic means of tracking people that can be filled in at leisure.  I think it too quickly dismisses the collaborative effect of a standup meeting and overplays the disruptive nature.  Sure, it's a disruption, but so is lunch.  Schedule the two together.  :)  I've found that for many projects a daily meeting is unnecessary and instead meet only 2 or 3 times per week. Less disruption, same benefits.


 

Tuesday, May 20, 2008

Get Rid Of Your Security Blankets

A while ago I took  a class on Scrum and Agile Project Management.  During the discussion on Scrum, it became apparent to me that there are several unchallenged assumptions in many peoples' minds that make accepting Scrum difficult.  People assume that Scrum/Agile takes away something they have, but in reality they don't have it.  People assume they have the assurance of a fixed schedule and proper documentation.  In reality, they have neither.  They are like security blankets.  The thought of them makes people feel safe, but the reality isn't that they help.

The Agile Manifesto prefers working software over comprehensive documentation.  That doesn't mean no documentation, but it does mean limiting the output of documentation.  Some projects create reams of paper before writing any code.  Their managers gain comfort from the idea that everything is planned well in advance.  Unfortunately, that comfort is usually founded on hope, not reality.  Projects that do a lot of documenting up front run into one of two problems.  Either the project is in a straight-jacket and can't react when the plans are wrong or it does react and the documentation gets out of date.  I have seen many a project with a lot of documentation that is completely useless because it hasn't been updated.  The trick to getting good documentation is to use it.  As with many things in life, if it isn't being measured, it won't be accurate.  Moving to an agile project model may reduce the overall amount of documentation, but it shouldn't reduce the amount of useful documentation.

But how will the team know what to do?  Won't there be miscommunication without documentation?  No.  Not if things are done right.  If "customers" are close enough to give feedback on iterations and teams work to agreed-upon interfaces and integrate often, these problems can be dealt with. 

The Manifesto also calls for responding to change over following a plan.  Agile projects work on iterations.  At the end of each iteration, the project will be in a working state and closer to the final goal.  The difficulty many managers have is that Agile projects won't commit to getting all N features done by M date.  With a more traditional waterfall model, marketing knows when the product will ship and what features will be present.  It makes life a lot easier.  Except that it doesn't.  The dates are not realistic.  With an Agile project, the team is just admitting that they don't know how long everything will take.  This means everyone can react to the reality of the project rather than making plans based on unrealistic expectations that will be shattered later.  The expected ship date is just a security blanket.  Agile makes it clear you are giving this up (or giving up control of what features will be ready), but doesn't actually make the project any less predictable.  It's merely exposing the unpredictability that is innate in software development.

Many of the objections to an agile or lean software project are based on perceived loss of control.  The trouble is, that control is not real.  Losing something you don't really have isn't actually a loss. 

Tuesday, January 1, 2008

Two Software Development Worlds

I was recently listening to an interview with Joel Spolsky.  The main subject is interviewing and hiring, but in the course of the interview Joel touches on an interesting point.  He says that there are two major types of software:  Shrinkwrap and Custom (listen around the 40 minute mark).  These have very different success metrics and thus should be created in different ways.


Custom


This might also be called internal or in-house software.  It is software that is written with one customer in mind and will only be run on one system.  This sort of software makes up, Joel claims, 70% of all software being written today.  Think about the intranets, inventory management software, etc. that IT groups everywhere are creating.


Joel makes the point that because custom software is only ever going to be used for one purpose and in a restricted environment, there is a steep falloff in return on investment.  After the software "works", there is very little advantage to making it better.  It is at this point that development should stop.


Another important point is that there really isn't any competition in internal software.  Companies won't generally fund two groups to write the same software.  This has implications for the definition of success.  For internal software, there will be a list of requirements.  If those requirements are met, the project is a success.  Being a little faster, slightly easier to use, or having one more feature doesn't make a project any more successful.


Shrinkwrap


Shrinkwrap software is software that is created for sale to others "in the wild."  This is software that is written for a general user category.  It might be shrink-wrapped software like Windows or Office or it might be web-based like Salesforce.com.  The important point is that it will be used in many environments by many different people.


Unlike custom software, the return on investment for shrinkwrap software is greater.  When the program is functional, it has only met the minimum bar for entry.  It must be much better (more robust, more features, more usable) before it can successfully compete.  Each feature, bug fix, etc. helps to increase market share.


These differences have implications for the type of person required, the sort of teams, and perhaps even the development methodologies that can be employed.  I'll revisit some of these implications in a future post.


Joel talks about some of these implications here.

Monday, November 12, 2007

Always Question the Process

Let me recount a story from the television show Babylon 5.  In one episode there is the description of guard posted in the middle of an empty courtyard.  There is nothing there to protect.  When one of the characters, Londo, questions why, he finds that no one, not even the emperor, knows why.  After doing some research, Londo discovers that 200 years before, the emperor's daughter came by the spot at the end of winter.  The first flower of the spring was poking up through the snow.  Not wanting anyone to step on the flower, she posted a guard there.  She then forgot about the flower, the guard, and never countermanded her order.  Now, 200 years later, there was still a guard posted but with nothing to protect.  There had been nothing to protect for 200 years.


This demonstrates the unfortunate power of process.  It often takes on a life of its own.  Those creating the complex system of rules expect it to be followed.  Once written down though, people stop thinking about why it was done.  Instead, they only expect it to be carried out.  This often leads to situations where work is being done for the sake of process instead of the outcome.  It is from this situation that bureaucracy gets its sullied reputation (well, that and the seeming ineptitude of many bureaucrats).  Process can easily become inflexible.  This is especially true in the technology industry where process is embedded in the code of intranet sites and InfoPath forms.


I encourage you to constantly revisit your process.  Question it.  Why do you do things the way you do?  Is there still a reason for each step?  If you don't know, jettison that step.  Simplify.  You should have just enough process to get the job done, but no more.  Once again, I'll re-iterate.  You hire smart people.  You pay them to think.  Let them.


This isn't to say that all process is bad.  Having common ways of accomplishing common tasks is efficient.  If a process truly makes things more efficient, it should be kept.  If not, it should be killed.  What was at one time efficient probably isn't any more.  Be vigilant.

Friday, November 9, 2007

Keep Process Simple

Year ago one of our Software Test Engineers was tasked with documenting our smoke* process.  It should have been something simple like:



  1. Developer packages binaries for testing
  2. Developer places smoke request on web page
  3. Tester signs up for smoke on web page
  4. Tester runs appropriate tests
  5. Tester signs off on fix
  6. Developer checks in

Instead it turned into a ten page document.  Needless to say, I took one glance at the document and dismissed it as worthless.  As far as I know, no one ever followed the process as it was described in that document.  We all had a simple process like the six steps I laid out above and we continued to follow it.


When tasked with creating a process for a given task, the tendency is to make the process complex.  It's not always a conscious effort, but when you start taking into account every contingency, the flow chart gets big, the document gets long, and the process becomes complex.  To make matters worse, when a problem happens that the process didn't prevent, another new layer of process is added on top.  Without vigilance, all process becomes large and unwieldy.


The problem with a large, complex process is that it quickly becomes too big to keep in one's mind.  When a process is hard to remember, it isn't followed.  If it takes a flow chart to describe an everyday task, it is probably too big.  It is far better to have a simple, but imperfect process to one that is complete.  The simple process will be followed out faithfully.  The complete one will at best be simplified by those carrying it out.  Worse, it may be fully ignored.  It is deceptive to think that a complex process will avoid trouble.  More often than not, it will only provide the illusion of serenity.


I have discovered that process is necessary, but it must be simple.  It is best to keep it small enough that it is easily remembered.  The best way to do this is to document what is done, not what you want to be done.  People have usually worked out a good solution.  Document that.  Don't try to cover all the contingencies.  It will only hide the real work flow.  Documented process should exist to bring new people up to speed quickly.  It should be there to unify two disparate systems.  It should not be there to solve problems.  You hire smart people, let them do that.  Expect them to think.


My final advice on this topic is do not task people with creating process.  It is tempting to say to someone, "Go create a process for creating test collateral."  Tasking someone with creating process is a surefire way to choke the system on too much process.  The results won't be pretty.  Nor will they be followed. 


 


*Smoke testing is a process involving testing fixes before committing them to the mainline build.  Its purpose is to catch bugs before they affect most users.

Thursday, November 8, 2007

The Need for a Real Build Process

Jeff Atwood at Coding Horror has a good post about how "F5 is not a build process."  In it, he explains how you need a real centralized build process.  F5 (the "build and debug" shortcut key in Visual Studio) on a developer's machine is not a built process.  At Microsoft, we have a practice of regular, daily builds.  We use a centralized content management system which everyone checks their code into.  At least daily, a "build machine" syncs to the latest source code and builds it all.  There are three main advantages of this system. 

First, it makes sure that the code is buildable on a daily basis.  If someone checks in code which causes an error during the build process, it shows up quickly.  We call this "breaking the build" and it's not something you want to be caught doing.  If you break the wrong build, you can get a call early in the morning asking you to come in immediately and fix it.

Second, it ensures that there is always a build ready for testing.  This has the added benefit of providing one central spot for everyone to install from.  If the build process is just on individual developers' machines, it is not uncommon for different people to be testing a build from radically different sources and thus have conflicting views of the product.  If you find a bug and someone says "Oh, it's fixed on my machine with these private bits" that is a sign of trouble unless they fixed it in the last 24 hours.

Finally, it ensures there is a well-understood manner of building the product.  If the builds are not centralized, there is no documented way of building the product.  Being able to build then becomes a matter of having the right tribal knowledge.  Build this project, then copy these files here, then build this directory, then this one.  Having a centralized and well-understood build process is a sign of a mature project.

If you are looking to improve your build process, there are plenty of tools out there to help.  The oldest and probably most-used tool is make.  It's also probably the hardest to use.  It has a lot of power, but it pretty quirky.  I've heard good things about Ant but I haven't used it.  It seems to be taking the place of make in a lot of projects.  The latest Visual Studio has a new build tool called MSBuild.  Again, I've heard good things but I haven't used it.

Monday, October 22, 2007

Helping Groups Succeed

or What to do when you aren't in control but neither is the leader.

A while back I wrote about providing clarity as a leader.  As part of that essay I mentioned some techniques for keeping groups on track.  Those are well and good if you are the leader, but what if you aren't?  What if the leader of your group didn't read my post and is making a mess of things.  It's common for someone to be in a position of leadership but not be leading.  This usually results in meetings that are contentious, long, and don't bear fruit.  If they do produce anything, it comes at a tremendous price.  What should you do if you are caught in such a meeting?  Below are some techniques that will help.

First, it is important to get a good feel for what is causing the failure.  If the cause of failure can be understood, then the solution can be derived from there.  Many meetings fail because there is no shared vision.  There are two very important items that must be shared by all participants for a meeting to be successful.  First, there must be a shared vision of what the outcome should be.  What is the specific decision the group is trying to make?  What is the goal of the design?  How detailed does the design need to be?  Second, there must be a shared vision of the rules.  It must be understood how you are going to make the decision.  Without this shared vision, there will be a lot of time spent talking past each other, driving toward different agendas, following rabbit trails, etc.

Given that situation, what are the things a person can do to help the meeting succeed?  There are two primary tactics that I've seen work.  First, help form consensus.  Second, help bring things back on track.

Forming consensus involves several actions.  It means listening carefully to what is being said.  If two or more people are coming from the same or similar places, point this out and try to get other members of the group to agree.  A poor leader will let these similar voices get lost in the noise.  Stepping back from one's own agenda to try to point out and support the development of consensus is important.  It can help to move the meeting forward.  If there is no consensus forming, try stepping back from the immediate decision.  Is there a more fundamental decision that, when decided, could help constrain the current decision to more tractable territory?  If so, lead the group to that other decision.

Once consensus is formed, it is common for non-germane conversations to take place.   An interesting, but not relevant topic may be discussed.  Someone may bring up a new point on a decision already made.  If these are the case, it is important that someone bring the group back on track.  Point out that the conversation is straying and then bring up a point that is on topic.  Do so in a friendly manner.  You don't want to be seen as bossy. 

There is one other tactic which can work but doesn't always.  That is, grab the power.  Whoever is controlling the pen or the keyboard has a lot of power.  Offer to take the notes or write on the white board.  This gives you the opportunity to have influence on what gets written.  If you see consensus starting to form, just write it down.  If there is something tangential, don't.

Monday, August 20, 2007

Tacit Approval Often Isn’t

Most of us have found ourselves in situations where we need someone’s approval to get something done, but we can’t seem to get them to respond.  It would be okay if they said no.  It would be better if they said yes.  We just need an answer yet we can’t get one.  One tactic is to just go ahead and do what you wanted.  This has the tendency to come back and bite us if things go wrong though.  A slightly better tactic is to send mail that says something like this:



On the issue at hand, I recommend taking the following actions.  If I don’t hear from you by such and such a date, I’ll move ahead with my recommendations.


This has the benefit of a paper trail. When things go wrongly, you’ll be able to point to this mail and say, “You had a chance to voice your opinion and didn’t.”  I’ll call this strategy getting tacit approval.  The approval is implied.


For some types of decisions, this is enough.  I’ve seen it employed well in situations where one party is being obstructionist via delay or where someone has authority but doesn’t really have a stake in the outcome.  It can work well when trying to get architectural approval for your design.  In the situations where tacit approval works well, you are always the active party and you merely need permission to move forward.


There is another situation where this is often employed and almost always to the tune of failure.  These are situations where you are not the active party.  Instead, you need someone else to do the work.  You’ll define what it is, but you are reliant upon their active participation for things to get done.  A good example would be if you are a release manager and need people to do certain work before the product releases.  You may send out mail explaining what is required and asking for comments.  If, however, you hear nothing back, you didn’t just receive approval.  This is true even if you say “If you disagree, you need to object by this date.” 


The problem is often that people are just too busy.  Too much e-mail is sent.  Too many requirements are put forward by too many disparate groups.  If you don’t hear anything back, it more likely means the message wasn’t received than that it was tacitly approved.  Assuming that silence means approval sets you up for failure.  I’ve seen this happen.  One group I worked with sent out instructions for how to interact with them.  If we didn’t like this, we had to disagree by a certain date.  They did this at a time when everyone was busy doing something else. Then, months later when it came time to finally pay attention to their part of the product, everyone came back with complaints.  They just thought we did.

The solution is to seek expressed approval instead.  Sometimes this can be hard.  The first requests for assent fall on seemingly deaf ears.  If you want to make sure your decision sticks, you need to persist.  If someone has not expressly stated that they agree with your proposal, the chances that they will take actions to being it to fruition are between zero and none.  It is worth the time and effort required to get explicit buy-in when you require the active participation of another party.

Monday, August 13, 2007

Scrum Meetings for Test

A year and a half ago I talked about how I was running scrum meetings with my team.  Since then, we've refined the process but have consistently held scrums on a regular basis.  Note that I'm not running a full Scrum system with sprints and product backlogs and such but rather just adopting the scrum meetings from that system.  Currently we have a team of 8 test developers.  We meet twice a week for 1/2 hour.  The format is simple. We go around the room and each person answers three questions:



  1. What did I do since last time?
  2. What will I be doing next?
  3. Is there anything blocking my progress?

Doing this helps me keep the pulse of the team and--more importantly--helps the team keep its own pulse.  It also encourages the team to act as a team.  It is easy in software to put yourself into a silo.  You have a task and you disappear into an office for a few weeks to get it done.  You might talk to your manager about it, but you don't talk to people outside your dependency list.  The disadvantage of this approach is that you don't get help from others.  In a scrum meeting, everyone learns what everyone else is doing.  If someone has experience in something someone else is struggling with, they offer their assistance.  In this way, the team starts supporting itself and the overall output increases.


Along the way, we did things wrong.  We learned from our mistakes.  Here is at least some of that knowledge:



  • Scrum is disruptive.  Programming is a matter of building up a mental map of the problem and then writing down the solution.  Once someone has this map built up they can work efficiently.  Having to change to another function is akin to swapping out the pages of the map.  Trying to start back up again requires paging everything back in which is slow.  Unfortunately, the human backing store isn't always stable and some paged out data gets lost.

  • Don't run scrum too often.  During a time where you are burning down bugs, meeting daily can be useful.  During the rest of the time, meeting daily is too often.  There isn't enough new to report and, worse, it tends to become disruptive.  Perhaps someone who has done daily scrum during the development phase can explain how this is avoided.

  • Scrum can't be seen as judgmental.  I found that without some calibration, team members felt that they were being judged by their progress.  If they didn't have something new to report, they felt it would be held against them.  Because of this, they didn't want to show up.  The solution was making it very clear that scrum meetings were all about status.  Being open is much more important than the level of productivity any individual was able to demonstrate.  The purpose isn't to take notes for the next review.  Being explicit about this helped.

  • Don't get bogged down in details.  The natural tendency of engineers when faced with a problem is to solve it.  This is good, but scrum isn't the place for solving problems.  It is the place for surfacing them.  Solutions should be derived outside the meeting.  Keep the meetings to their scheduled time limits.  Don't allow discussions to get into too many details.  Instead, take a note and have a followup discussion later.

Thursday, July 19, 2007

Hofstadter's Law

Good advice for all project managers.  Hofstadter's Law:

It always takes longer than you expect, even when you take into account Hofstadter's Law.

Wednesday, July 4, 2007

The Three Stonecutters

Lots of interesting quotes in Dreaming in Code.  This one is the story of three stonecutters.  Each is asked what he is doing.  The first answers that he is, "making a living wage."  The second says, "I am doing the best job of cutting stones in the entire country."  The third, "I am building a cathedral."  Each of the three represent employees you are likely to run across in your days as a manager.


The first represents the employee that's merely putting in his time.  He's the person who works 9-5.  He'll do what is asked, but when the job calls for that extra effort, he'll probably stop short.  There is a place for these employees in an organization.  As long as you can lay out achievable goals and set their direction, they'll serve you well.  Don't give them the highly critical piece though.  They might not come through if it takes a lot of extra effort or extra thought to deliver.


The second often represents a problem.  It's good that they care about the quality of their work, but they have their perspective wrong.  As a manager, you have a particular task at hand.  That task requires certain elements to be created.  Each of those elements needs to be of at least a particular quality.  However, being highly above that quality doesn't really help.  Take the example of the person building a cathedral.  The foundation stones need to be cut and they need to be straight.  However, being perfectly smooth is probably not required.  If it takes twice as long to make a perfect stone as it does an acceptable stone, that extra time and effort is wasted.  Likewise, I've had programmers deliver way more than is required.  They're often quite proud of it.  I once had someone deliver a test months overdue.  He was excited because it delivered all of these extra options beyond what I had asked.  It had been a lot of hard work getting it to all work right.  Unfortunately, I didn't have a need for most of that extra functionality.  It was wasted.  Watch for these types.  They don't have the right priorities.  Programming is always done to accomplish a task.  If programming is seen as the task, the project will suffer. 


At least some of the problems Chandler faced come from this issue.  The repository, the widget model, the UI all became issues where the perfection of the parts was seen as more important than the overall product.  So much time was spent getting the small things "right" that the large things were ignored.  What makes or breaks a piece of software is often the integration of the parts.  This second stonecutter worries about his stone, not the stones around him.


The third stonecutter is the one you want to maximize on your team.  This is the person who sees the big picture.  He not only knows what has been asked of him but also why.  Because he understands why, he is able to make intelligent decisions.  Knowing that your stone is going into a cathedral means you know when you need to cut the stones to one level of precision over another.  This is the sort of employee you just point in a direction and then stand out of the way.  They'll knock down whatever walls get in their way.  You won't even have to ask.

Sunday, July 1, 2007

Avoid 3-Card Combinations

I used to play collectible card games.  I attended Whitman College during Richard Garfield's tenure there as a math professor so I got into Magic: The Gathering near its inception.  For those of you who don't play these, the basic system goes something like this.  You buy packs of cards--not unlike baseball cards--which you trade and assemble into decks.  During gameplay, you draw cards from your deck and place them into play.  The rules for playing them differ from game to game but they almost always involve text on the cards that explains their behavior.  The abilities granted by each card vary greatly and they are most powerful in the way they interact with other cards. 

Sometimes you can find combinations which are nearly unbeatable.  If the game is well designed, such killer combinations usually require get 3 (or more) of the exact right cards.  Most of the time you can only have a limited number of any one card in a deck.  Economics often precludes it even if the rules don't.  Thus the chance that you will have the 3 cards you need in your hand (or in play) at the same time is very low.  When we were playing, we had a saying that went something like "don't build your deck around 3-card combinations."  While these killer combos were game-winners if they came out, they were so rare that you usually lost.  A much better strategy was to build a deck around simpler concepts that required only two cards at a time to pull off.

That's nice, but what does this have to do with computers and programming?  It is a good analogy for the way many people manage projects.  If you consider each dependency as a card you need to draw and shipping as the killer combination, it becomes obvious that the winning strategy is not to take on too many dependencies.

If you assume that the new framework or programming language will be mature enough and that your maintenance work won't take very long and that the team you're relying on will deliver their part on time, you've just built your deck around a 3-card combination.  Sure, it will be a spectacular product and take the market by storm when you ship it.  If you ship it.

The world of software development is full of math people.  You'd think we would pay more attention to the probabilities of success.  Somehow we think we are in a reality distortion bubble though.  Math doesn't apply to us.  Sure, it will be hard, but we're better than average.  Things will fall into play.  We'll live happily ever after.  Except, we don't, it does, we aren't, they won't, and we won't.

The moral of this tale:  Try to accomplish a little less and build on last year's framework.  Your work-life balance will thank you.  So will your marketing people.

Wednesday, June 27, 2007

Trade Accuracy for Understanding

I found myself giving this advice to two people today.  It came in the context off preparing a presentation for upper management.  The desire was to communicate an understanding of what (and why) we are creating a piece of technology.  The difficulty was in trying to convey the information without overwhelming the audience.  This can be tricky.  Engineers are especially bad at this.  Why?

Engineers know a lot about what they are working on.  It's fair to say that they know more about what they work on than almost anyone else does.  Engineers are also taught to be accurate.  Being inaccurate gets you in trouble when you are designing something.  You can't be "close enough" in the weight-bearing characteristics of a bridge.  You can't be "close enough" to the specification when writing a class driver.  You have to be accurate.

Managers of engineers, on the other hand, know less.  It's not that we're dumber, it is that we are more diversified.  Consider the mind to be able to hold a finite capacity of knowledge.  It can either be filled with a lot of knowledge on a few topics or a little knowledge on a lot of topics.  Engineers are the former, managers the latter.

Therein lies the rub.  Engineers need to convey some subset of their knowledge to management.  However, management does not have the same understanding.  Management doesn't need it and probably can't afford it. 

When you ask an engineer to summarize, he or she will try to be very succinct but not lose any information.  This actually amplifies the problem.  Now you have the same amount of technical detail but with less explanation.  That's not a solution.

Instead, the solution must be lossy.  You have to throw away information in order to convey the main point.  Sometimes when you throw away that information, things get a little distorted.  That is to say, inaccurate.  This grates on most engineers.  However, it is exactly what is needed.

Let's use an example from another realm of life.  There is a story that George Washington cut down a cherry tree when he was young.  He was so honest that he went and told his father what he had done.  Is the story accurate?  We don't know.  Probably not fully.  That's okay though, we're trying to teach a moral about honesty and to describe the character of the United States' first president.  Those goals are both accomplished.  Trying to explain how it probably wasn't a cherry tree or it wasn't his fathers or how he took a month to tell or even how in other dealings in life he wasn't always honest may be accurate, but they distort the true picture.

Now let's look at the real world case.  We're trying to convey why we should test audio systems for their output level.  The engineer wants to say something like, "Full-Scale Output Level (or just Output Level for short) on a PC is the amplitude of the analog signal that comes out of the jack/speakers when a digital full-scale waveform is applied to the codec."  There's a great detailed description here.  Instead, to describe this quickly one might just say, "Output level is volume and if it isn't high enough, the volume won't be high enough."  That's not fully accurate, but it conveys the necessary information better.  If someone is really interested in the subject, there is plenty of time to go into detail. 

The important thing is to convey a kernel of truth in a way someone can easily latch onto.  Conveying something wholly false is bad, but so is conveying the truth in so much detail that it can't be grasped.  Often times it is necessary to trade accuracy for understanding.