Wednesday, June 4, 2014

A Brief History of Test

In the exploration of quality, it is important to understand where software testing came from, where it is today, and where it is heading. We can then compare this trajectory to the goal of ensuring quality and see whether a correction is necessary or if we're going the right direction.


I have been involved in software testing for the past 16 and one half years. To give you some perspective, when I started at Microsoft, we were just shipping Windows 98. This gives me a lot of perspective on the history of software testing. What I give below is my take that that history. Others may have experienced it differently.


There are three major waves of software testing and we're beginning to approach a 4th. The first wave was manual testing. The second wave was automated testing. The third wave was that of tooling. It is important to note that each wave does not fully supplant the previous wave. There is still a need for manual testing even in the tooling phase. There is a need for automated testing even in the coming 4th wave.


The first wave was manual testing. Sometimes this is also called exploratory testing. It is often carried out by people carrying the Quality Assurance (QA) or Software Test Engineer (STE) title. Manual testing is just what it sounds like. It is people sitting in front of a keyboard or mouse and using the product. In its best form it is freeform and exploratory in nature: a tester trying to understand the user and carrying out operations trying to break the software. This is where the lore of testing comes from. The uber-tester who can find the bug no one else can imagine. In its worst form, this is the rote repetition of the same steps, the same levels, over and over again. This is also the source of legends, but not good ones. At its best, this form of testing is highly connected to quality. It is all about the user and his (or her) experience with the product.


Manual testing can produce great user experiences. As I understand it from friends who have gone there, this is the primary method of testing at Apple. The problem is, manual testing doesn't scale. It can also be mind-numbing. In the era of continuous integration and daily builds, the same tests have to be carried out each day. It becomes less about exploring and more about repeating. Manual testing is great for finding the bugs initially, but it is a terribly inefficient regression testing model.


It gets even worse when it comes time for software maintenance. At Microsoft, we support our software for a long time. Sometimes a really long time. Windows XP shipped in 2001 and is just now becoming unsupported. Consider for a moment how many testers it would take to test XP. I'll just make up a number and say it was 500. It wasn't, but that's good enough for a thought exercise. Every time you release a fix for XP, you need 500 people to run through all the tests to make sure nothing was broken. But the 500 original testers are probably working on Vista so you need an additional 500 people for sustained engineering. Add Windows 7, Windows 8, and Windows 8.1 and you now need 2,500 people testing the OS. Most are running the same regression tests every time which is not exciting so you end up losing all your good people. It just doesn't work.


This leads us to the second wave. The first wave involved hundreds of people pressing the same buttons every day. It turns out that computers are really good at repetitive tasks. They don’t get bored and quit. They don't get distracted and miss things. Thus was born test automation. Test automation in its most basic form is writing programs to do all of the things that manual testers do. They can even do things testers really can't. Manual testing is great for a user interface. It's hard to manually test an SDK. It turns out that it is easy to write software that can exercise APIs. This magic elixir allowed teams to cover much more of the product every day. We fully drank the Kool-Aid.


We set off to automate everything. There was nothing automation couldn't do and so all STEs were let go. Everyone become a Software Design Engineer in Test (SDET). This is a developer who, rather than writing the operating system, writes tests for the automation. Some of this work is mundane: calling the Foo API with a 1, a 25, and a MAX_INT. Other parts can be quite challenging. Consider how you would test the audio playback APIs. It is not enough to merely call the APIs and look at the return code. How do you know the right sound was played, at the right volume, and without crackling? Hint: it's time to break out the FFT.


Not everything is kittens and roses in the world of automation. Machines are great at doing what they are told to do. They don't take breaks or demand higher pay. However, they do only what they are told to do. They will only report bugs they are told to look for. One of my favorite bugs to talk about involved media player. In one internal build, every time you clicked next track on a CD (remember those things?) the volume would jump to maximum. While a test could be concocted to look for this, it never would be. Test automation happily reported a pass because indeed the next track started playing. It turns out that once you have run a test application for the first time, it is done finding new bugs. It can find regressions, but it can't notice the bug it missed yesterday.


The points toward the second problem with test automation. It requires very complete specifications. Because the tests can't find any bugs they weren't programmed to find, they need to be programmed to find everything. This requires a lot more up-front planning so the tests can cover the full gamut of the system under test. This heavy reliance upon specifications begins to distance testing from the needs of the user and thus we move away from testing quality and toward testing adherence to a spec.


The third problem with test automation is that it can generate too much data. Machines are happy to churn out results every day and every build. I knew teams whose tests would generate millions of results for each build. Staying on top of this becomes a full time job for many people. Is this failure a bug in the test? Is it a bug in the product? Was it an environmental issue (network down, bad installation, server unavailable)?


One other problem that can happen is that the work can grow faster than test developers can keep up. It is easy for a developer to write a little code which creates a massive amount of new surface area. Consider the humble decorator pattern. If I have 4 UI objects in the system, the tester needs to write 4 sets of tests. Now if the developer creates a decorator which can apply to each of the objects, he only has to write one unit of code to make this work. This is the advantage of the pattern. However, the tester has to write 4 sets of tests. The test surface is growing geometrically compared to the code dev is writing. This is unsustainable for very long.


This brings us to the third wave of testing. This wave involves writing software that writes tests. I call this the tooling phase. Rather than directly writing a test case, it is possible to write a tool that, given some kind of specification, can emit the relevant test cases automatically.  Model Based Testing is one form of this tooling. The advantage of this sort of tooling is that it can adapt to changes. Dev added one decorator to the system? Add one new definition in your model and tests just happen.


There are some downsides to the tooling approach. In fact, there are enough downsides that I've never experienced a team that adopted it for all or even most of their testing. They probably exist, but they aren't common. At most, this tooling approach was used to supplement other testing. The first downside is the oracle problem. It is easy enough to create a model of the system under test and generate hundreds or thousands of test cases. It is another thing entirely to understand which of these test cases pass and which fail. There are some problem domains where this is a tractable problem. Each combination or end state has an easily discernible outcome. In others, it can be exceptionally difficult without re-creating all of the logic of the system under test. The second is that the failures can be very hard to reason about in terms of the user. When the Bar API gets this and that parameter while in this state, it produces this erroneous result. Okay. But when would that ever happen in the real world?


Tooling approaches can solve the static nature of testing mentioned above. Because it is mathematically impractical to do a complete search of the state space of any non-trivial application or API, we are always limited to a subset of all possible states for testing. In traditional automation, this subset it fixed. In the tooling approach, the subset can be modified each time with random seeds, longer exploration times, or varying weights. This means each run can expose new bugs. This can be used to good effect. Given some metadata about an API and rules on how to call it, a tool can be created to automatically explore the API surface. We did this to good effect in Windows 8 when testing the Windows Runtime surface.


Sometimes it can have unexpected and even comical outcomes. I recall a story told to me my a friend. He wrote a tool to explore the .Net APIs and left it to run overnight. The next morning he came in to reams of paper on his desk. It turns out that his tool had discovered the print APIs and managed to drain every sheet of paper from every printer in the building. At Microsoft every print job has a cover sheet with the alias of the person doing the printing so his complicity was readily apparent. Some kind soul had gathered all of his print jobs and placed them in his office.


The tooling approach to testing exacerbates two of the problems of automation. It creates even more test results which then have to be understood by a human. It also moves the testing even further away from our definition of quality. Where is the fitness for a function taken into account in the tooling approach?


There is a problem developing in the trajectory of testing. We, as a discipline, have moved steadily further from the premise of quality. We'll examine this in more detail in the next post and start considering a solution in the one after that.

Thursday, May 29, 2014

What is Quality?




Most of my career so far has focused on software testing in one form or another.  What is testing if not verifying the quality of the object under test?  But what does the word quality really mean?  It is hard to define quality, but I will argue that a good operating definition is fitness for a function.  In the world of software then, the question test should be answering is whether the software is fit for the function at hand.


The book Zen and the Art of Motorcycle Maintenance tackles the question of what quality is head on.  Unfortunately, it doesn't give a clear answer.  The basic conclusion seems to be that quality is out there and it drives our behavior.  It's a little like Plato's theory of Forms.  This is interesting philosophically, but not practically.  There are some parts which are more pragmatic.  One passage sticks out to me.  As might be suspected from the title, there is some discussion of motorcycle maintenance in the book (but not much).  At one point the Narrator character is on a trip with his friend John Sutherland.  The Narrator has an older bike while John has a fancy new one.  The Narrator understands his bike.  John chooses not to learn about his and needs a mechanic to do anything to it.  Quality then is that relationship between the operator and the bike.  The more they understand it and can fix it, the higher quality the relationship.  In other words, the more the person can get out of it without needing to go to someone else, the higher the quality.


Christopher Alexander wrote about architecture, yet he is quite famous in the world of software design.  His books talk about patterns in buildings and spaces and how to apply them to get specific outcomes.  The "Gang of Four" translated his ideas from physical space to the virtual world in their book, Design Patterns.  Alexander is interesting not just in his discussion of patterns, but also of quality.  What makes a design pattern good, in his mind, is its fitness for a purpose.  He says, "The form is the part of the world over which we have control, and which we decide to shape while leaving the rest of the world as it is. The context is that part of the world which puts demands on this form; anything in the world that makes demands of the form is context. Fitness is the relation of mutual acceptability between these two." (Notes on the Synthesis of Form)


Both authors are making an argument that quality then is not something that can be determined in a vacuum.  One cannot merely look at a device or a piece of software and make an assessment of quality.  In the motorcycle case, the fancy bike would probably look to be of higher quality.  It was the relationship with the owner that made the chopper of higher quality.  With software, it is similar.  One must understand the users and the use model before a determination of quality can be made.


Let's look at a few examples.  Think of the iPhone.  It has a simple interface.  While it has gained complexity over time (it's nearly on version 8!), it is still quite limited compared to a traditional computer.  It has constrained input options, preferring only touch.  The buttons are big and the screens not dense.  Because of this, the apps tend to be simple and single-purpose.  There is no multitasking to be found.  It didn't even have cut and paste when it appeared on the scene.  Yet the iPhone is considered to be high quality.  Its audience doesn't expect to be doing word processing on it.  They want to check e-mail, "do" Facebook, and play games.  It is exquisitely suited for this purpose.


At the other extreme, consider a workstation running Autocad.  Autocad has thousands of functions, many windows open at once, and requires extreme amounts of processing power and memory.  It's user interface is quite cluttered compared to that of most iPhone apps and it is not easily approachable by mere mortals.  Yet it too is considered high quality.  Its users expect power over everything else.  They need the ability to render in 3D and model physics.  It serves this purpose better than anything else in the market.  The simplicity and prettiness of the iPhone interface limits utility and is unwanted in this domain.  The domain of CAD is one of capability over beauty and efficiency over discoverability.


Too often in the world of software we ignore this synthesis of user and device.  Instead we focus on correctness.  The quality of the software is judged based on how correctly it implement a spec.  This is an easier definition to interpret.  It is more precise.  There is a right and a wrong answer.  Either the software matches the spec or it does not.  With a fitness definition, things are more murky.  Who is the arbiter of good fit?  How bad does it need to be before it is a bug?  It can be alluring to follow the sirens of precision, but that comes at a cost.  I will talk about that cost next time.


 




Tuesday, April 10, 2012

How to answer a programming interview question

I spent a few hours on Friday doing mock interviews for CS undergrads.  The idea was to help them experience the interview process without the pressure of having a job on the line.  The session was interactive with lots of stopping for advice.  I found myself explaining the following to most of the candidates.  I hope you find some value in it.

  1. Explain the algorithm that you will use to solve the problem.  Use visuals where you can.  This will accomplish two things.  First, it will make sure you understand what you are going to do before you start coding.  This in turn will reduce the number of corrections and backtracks you have to make while coding.  Second, it will give the interviewer a chance to help you correct course early.  This means getting onto the right track before you spend a lot of time coding a solution to the wrong problem.
  2. Write the code.  Use a real programming language.  Tell the interviewer what language you are using and then write in it.  Don’t sweat the syntax.  Focus on the flow of the code.  Using a real language is important because it is too easy to gloss over the important things with pseudo-code.  Try to write the code in one pass.  You already have the algorithm figured out so you shouldn’t need to backtrack and change your code.
  3. Walk through the code with real data.  Before you declare yourself “done”, take the time to walk through the code.  Take a real example and explain how each line of code operates on it.  This will a) give you a chance to find bugs and b)demonstrate to the interviewer that the code works.  Pay careful attention here.  I’ve seen a number of candidates with big bugs in their code explain how they want the code to work instead of how it really works.  If you don’t pay attention, you won’t find the bugs.
  4. If you are interviewing for a test developer position, test your code.  You’ll be asked to do so anyway, you might as well get credit for taking the initiative.  Even if you aren’t interviewing as an SDET, it is a good idea to show that you are thinking through the failure cases.

Monday, April 9, 2012

Jack Tramiel, founder of Commodore, dies at age 83

Jack Tramiel was the founder of Commodore International which produced the Commodore 64 and the Amiga computers.  It was also the company that made the once ubiquitous 6502 processor which powered the Apple // and the Commodore 64.  The Commodore 64 was the best selling computer of all time and the Amiga (which came after his tenure, but from his company) was almost a decade ahead of its time.  I learned to program with Basic on the Commodore 128 and spent a lot of my formative years using the Amiga. 


Jack Tramiel was a ruthless business person who drove a very hard bargain.  His relentless push for prices drove the first major round of the home computer revolution.  He died on Monday at age 83.  He had a great legacy and will be missed.  A great book on Commodore and Jack Tramiel's role in it is Commodore: A Company on the Edge.

Tuesday, March 20, 2012

Behind the Scenes of Windows 8

Larry Osterman returns with another installment of his Behind the Scenes... series.  This time with Windows 8.  Larry is a developer for the team I work on.  If you haven't caught it yet, take the time to read how the team developing Windows Runtime experienced this release.

Tuesday, March 13, 2012

Successfully Interviewing for a Developer Job

Having recently completed another round of campus campus interviews, several things stand out as advice that could be useful to those of you aspiring to get jobs in the software industry as a Developer or Test Developer.

  • Always describe what you are doing and thinking – If you are given a programming question and you interact only with the paper or the white board for the next 10 minutes, you did poorly even if you got the answer right.  You failed if you didn’t get the answer right.  The interviewer wants to know not only that you can get the answer but also how you think.  If you have the right thought processes, you might pass the interview even if you don’t get the answer.  Often times people get extra credit for explaining what their options are and why they are picking one or another.
  • When asked about a project or job, describe what *you* did – Explaining that the team wrote a location-aware notification system for Android phones doesn’t win you any points if you only did the graphics.  If you wrote the notification database and not the location APIs, talk about the database.  The interviewer will likely probe the details of the project.  Emphasizing the exciting parts you didn’t have a part in only leads to your having so say “I don’t know.  I didn’t work on that.”
  • Be able to speak to the details of every project on your resume – If you put it there, it is to demonstrate your knowledge of an area.  If you don’t actually have that knowledge any more, it doesn’t help you.  In fact, it makes you look less knowledgeable.  You don’t have to know every detail of the puzzle game you wrote in that AI class 3 years ago, but you should know enough to talk about the algorithm you used.  Spend the time reviewing and perhaps even practicing talking about each one before showing up.
  • If you don’t know, don’t bluff – It is often okay to say “I don’t know…” as long you can also say “…and this is how I would find out.”  It is certainly better than bluffing.  The person interviewing you has probably done a lot of interviews and a lot of programming.  They will likely know that you don’t really know.  Getting caught bluffing almost certainly kills your chances of being hired.
  • Think through your code before you start writing – Having to backtrack a lot and change your code is not a good thing.  It is better than being wrong, but it is less good than the next person who will answer the same question without having to backtrack.  Having to add a lot of extra variables or loops or having to change your algorithm substantially demonstrates that you didn’t really understand the problem to start with.  It also makes your code really hard to read because unlike in a text editor, everything doesn’t shift down when you insert a line.  Take the time to think through your algorithm before you start writing.  Even better, explain it to the interviewer.  That way they can correct you if you are on the wrong track.
  • Time matters – Sometimes it is not enough to get the solution correct.  Taking a long time to get there, especially if you have to backtrack, shows you do not have sufficient mastery.  Some problems are easy and the interviewer expects you to get them right the first time and without a lot of delay.  Finding a value in a binary search tree or reversing a string fall into this category.
  • Be friendly – Believe it or not, your interviewer is human.  He or she will be more inclined to give you the benefit of the doubt if they like you.  Being friendly, making eye contact, and being upbeat all help with this.  If you answer questions with short sentences that do not further the conversation, never make eye contact, and have a sullen attitude, you will fail a close interview.
  • Ask clarifying questions – Don’t just jump in and start programming.  Think about what you are being asked to do.  Many times the question will be intentionally vague.  Ask questions about the constraints, the expected use, the interface, etc.  If the question was intentionally vague, not asking questions will be a negative.  Asking intelligent questions shows that you understand the question enough to notice the edges.  That will earn you points.

Have other advice?  Please leave it in the comments.

    Friday, March 9, 2012

    If you ever wondered why Vim uses hjkl for arrow keys...

    I tend to use Vim as my editor of choice.  Even when using Visual Studio, I do so with the ViEmu plugin.  I have always wondered why the directional keys were hjkl instead of jkl;.  The latter are the home keys for the right hand.  The former are not and overload the right index finger.  Thanks to Hacker News, I now know the answer.  The terminal which Bill Joy wrote the original Vi on was an ADM-3A and that had arrows drawn on the hjkl keys (see the link for a picture).