Thursday, May 15, 2008

Design Principle: Don't Repeat Yourself

There's a design principle I neglected to mention in my initial list but which certainly merits attention.  That principle is this:  whenever possible, don't repeat yourself (DRY).  Put another way, do things one time, in one place rather than having the same or similar code scattered throughout your code base.

There are two main reasons to follow the DRY principle.  The first is that it makes change easier.  The second is that it helps substantially when it comes time for maintenance.

I was once told by an instructor in a design patterns class that the Y2K problem wasn't caused by using 2 digits to represent 4.  Instead, it was caused by doing so all over the place.  While the reality of the situation is a little more complicated than that, the principle is true.  If the date handling code had been all in one place, a single change would have fixed the whole codebase rather than having to pull tons of Cobol coders out of retirement to fix all the business applications.

When I first started working on DVD, I inherited several applications that had been written to test display cards.  They had started life as one large application which switched behavior based on a command-line switch to test various aspects of DirectDraw that DVD decoders relied upon.  Some enterprising young coder had decided we should have separate executables for each of these tests so he made 6 copies of the code and then modified the command-line parser to hard-code the test to be run.  The difficulty here is that we now had 6 copies of the code.  Every time we found a bug in one application, we would have to go make the same fix in the other 5.  It wasn't uncommon for a bug to be fixed in only 4 places.

Shalloway's Law:  “When N things need to change and N>1, Shalloway will find at most N-1 of these things.”

This principle applies to everything you write, not just to copying entire applications.  When you find yourself writing the same code (or substantially the same code) in 2 or more places, it is time to refactor.  Extract a method and put the duplicated code in that method.  When the code is used by more than one application, extract the code into a function call that you put into a shared library.  This way, whenever you want to change something, a change in one place will enhance all callers.  Also, when something is broken, the fix will automatically affect all callers.

Note that this is a principle and not a law.  There are times when substantially similar code is just different enough that it needs to be duplicated.  Consider the alternatives first though.  Can templates solve the problem?  Would a Template Method work?  If the answer is no to everything, then duplicate the code.  You'll pay the price in increased maintenance, but at least you'll be aware of what you are getting yourself into.  It might not be a bad idea to put a comment in the code to let future maintainers know that there's similar code elsewhere that they should fix.

Wednesday, May 7, 2008

Keep Your Eyes on the True Goal

We recently went through mid-year reviews and I found myself repeating similar advice several times.  Here it is:  Be very clear about what your goal is and continuously reassess whether you are on track to get there.

To make this more concrete, consider a plane trip from Seattle, WA to to Austin, TX.  If you look at a map, you can see that the heading to Austin is something like 135 degrees from Seattle and the flying time might be 7 hours.  Does the pilot take off, set the compass to 135 degrees and set an alarm for 7 hours later?  No.  If he does, he could be in Little Rock by the time he wakes up.  Why is that?  Wind will push the plane around.  If the pilot doesn't correct for the wind, he won't end up where he expected.  Instead, the pilot constantly monitors his location and recalculates the necessary heading and speed to get to Austin on time.

There are two tendencies I've noticed among developers which can be rectified by following the example of the pilot.  The first is the tendency to fail to monitor progress.  The second is to focus on the trip, not the destination.

There are three aspects to the tasks most developers are asked to accomplish.  They are called upon to implement some functionality by a specified time and in a quality fashion.  When a developer decides up front what work needs to be done and then plows through it without reassessing along the way, he runs a high risk of failing to deliver.  There are too many unknowns (and unkunks) for a plan created at the onset to be successful.  Without reassessing on a regular basis, the developer won't recognize that he is behind and thus will not be able to correct.  Like the pilot who lands in Little Rock instead of Austin, the developer who does not evaluate progress along the way will get to the end of the project and realize he still has a month of work to do.  This is one of the secrets of Scrum.  It forces reassessment on regular intervals.  Another technique I have found to work is to break up your tasks over the project (or milestone) into granular work items.  Each item should be a few days at most.  Then, at least weekly, mark off progress against them.  Keep track not of how much time you've spent so far but rather how much time is left and compare that to the amount of work left.  When work remaining exceeds time left, it is time for a course correction.  Because quality cannot be sacrificed, either move out the delivery date (if possible) or cut features.

A different problem is when a developer mistakes the assigned tasks for the goals.  An example comes from the test development side of the world.  Let's say the goal is to test a new video manipulation API.  This Core Video or DirectX Video Acceleration or something.  To test whether this works correctly, it is determined that it makes sense to write a simulator.  The test developer starts writing this simulator.  Along the way, he confuses his real goal (test the API) with his task (write the simulator).  If, at the end of the milestone he has a great simulator but hasn't had time to actually use it to test the API, has he succeeded?  In his mind, he has.  Unfortunately, he is like the pilot who confused "Fly at 135 degrees for 7 hours" as his goal instead of "Land in Austin."  What is the solution?  Always keep your eyes on the real goal.  If the simulator is taking too long to complete, consider alternatives.  Is there a way to cut a feature and still test most of the API?  Would it be better to cut the losses and approach from another direction? 

Constant course correction can only work if the pilot knows his true goals.  Having the wrong goal in mind or not recognizing when he has flown off course is disastrous for a pilot.  Likewise, not noticing you are behind or that you aren't going to accomplish the real goals of your project is disastrous for a developer.  Work with your manager (or program manager) to understand what the reason for your project is and always keep that in mind.  This way, you'll be able to correct course before it becomes a problem.  Don't mistake your tasks for your goal.

 

Note:  I'm not a pilot.  Some of what I say about planes is inevitably incorrect.

Monday, April 28, 2008

A Microsoft-Yahoo Takeover Primer

Marc Andreessen has a great blog post today laying out the possibilities in the Microsoft-Yahoo talks.  Unlike most posts on the subject, this one isn't trying to guess what might happen.  Instead, it lays out the options and the forces affecting those options.  What is a proxy battle?  How would it take place?  Who are the investors we're talking about?  What is a tender offer?  How is it affected by a poison pill?  If you are following the subject, check out his post.  It's a good primer for the rest of the pontificating on the subject.

Prefer Composition Over Inheritance

It's probably about time to bring my "Design Principles To Live By" series to a close.  This is the last scheduled topic although I have one or two more I may post.

Let's begin with some definitions:

Composition - Functionality of an object is made up of an aggregate of different classes.  In practice, this means holding a pointer to another class to which work is deferred.

Inheritance - Functionality of an object is made up of it's own functionality plus functionality from its parent classes.

For most non-trivial problems, there will be similar code needed by multiple classes.  It is not a wise idea to put the same code in more than one place (a topic for another day).  There are two strategies in object-oriented programming which attempt to solve the problem of duplicate code.  The one most popular in the early days was inheritance.  Shared functionality was implemented in a base class which allowed each child class to inherit that functionality.  A child would just not implement foo() and the parent would do the work.  This works, but it is not very flexible.

Suppose that the shared functionality is some kind of encryption algorithm.  Each child class will only inherit from one base class.  What if there is a for different encryption algorithms?  It would be possible to have multiple base classes, say AESEncryptionBase and DESEncryptionBase, but this necessitates multiple copies of the child classes--one for each base class.  With more than 2 base classes, this become untenable.  It also becomes very difficult to change out the encryption routine at runtime.  Doing so means creating a new object and copying the contents of the old object to it.

Another difficulty is the distortion of otherwise clean class hierarchies.  Each child should have an "is-a" relationship with its parent.  Is a music file and AESEncryptionBase?  No.  Here is a particularly telling examples from Smalltalk.  In Squeak (the dominant open-source Smalltalk implementation), Semaphore inherits from LinkedList.  Is Semaphore a linked list?  No.  A linked list is used in the implementation, but a sempahore is not a specialization of linked lists.

A better approach is to contain the new functionality via composition.  A class should contain instances of objects it needs to utilize functionality from.  In the music file case, it would have a pointer to an EncryptionImpl class which might be AES, DES, or ROT13.  The class hierarchy will stay smaller and the music file implementation does not even need to be aware of which encryption method it is using.  In the Semaphore case, Semaphore would contain a LinkedList object which it would use to do the work.  Clients of Semaphore would not be expecting LinkedList functionality.  Extraneous methods would not need to be disabled.  Composition would also allow for more flexibility later.  If an implementation based on a heap or a prioritized queue were found to be advantageous, they could be without clients of Semaphore knowing. 

Think twice before inheriting functionality.  There are times when it is a good idea such as when there is a logical default behavior and only some child classes need to over-ride it, but if the intent is to utilize the functionality rather than expose it to child class callers, composition is almost always the right decision.

Friday, April 25, 2008

A History of Filesystems

Ars Technica has a very interesting article about the history of filesystems.  They cover all the major systems including FAT (MS-DOS), HFS (Mac), NTFS (NT), Ext2/3 (Linux), and many others like the Amiga.  They also cover upcoming systems like ZFS.  If you have interest in the systems space, check it out.

Monday, April 21, 2008

Know That Which You Test

Someone recently related to me his experience using the new Microsoft Robotics Studio.  He loaded it up and proceeded through one of the tutorials.  To make sure he understood, he typed everything in instead of cutting and pasting the sample code.  After doing so, he compiled and ran the results.  It worked!  It did exactly what it was supposed to.  The only problem--he didn't understand anything he had typed.  He went through the process of typing in the lines of code, but didn't understand what they really meant.  Sometimes testers do the same thing.  It is easy to "test" something without actually understanding it.  Doing so is dangerous.  It lulls us into a false sense of security.  We think we've done a good job testing the product when in reality we've only scratched the surface.

Being a good tester requires understanding not just the language we're writing the tests in, but also what is going on under the covers.  Black-box testing can be useful, but without a sense of what is happening inside, testing can only be very naive.  Without breaking the surface, it is nearly impossible to understand what the equivalency classes are.  It is hard to find the corner cases or the places where errors are most likely to happen.  It's also very easy to miss a critical path because it wasn't apparent from the API.

There are three practices which help to remedy this.  First, program in the same language as whatever is being tested.  A person writing tests written in C# against a COM interface will have a hard time beginning to understand the infrastructure beneath the interface.  It can also be difficult to understand the frailties of a language different than the one being coded in.  Each language has different weaknesses.  Thinking about the weaknesses of C++ will blind a person to the weaknesses of Perl.  Second, use code coverage data to help guide testing.  Examining code coverage reports can help uncover places that have been missed.  If possible, measure coverage against each test case.  Validate that each new case adds to the coverage.  If it doesn't, the case is probably covering the same equivalency class as another test.  Third, and perhaps most importantly, become familiar with the code being testing.  Read the code.  Read the specs.  Talk to the developers. 

Friday, April 11, 2008

Slow blogging season

I apologize for the very light blogging of late.  I've been busy working on the project for my latest class at the University of Illinois.  CS classes really take a lot of time at the end of the semester.  At the beginning you just have reading, homework, and lectures.  At the end they pile a project on top of that.  Depending on the class, that can mean a lot of work.  This isn't the worst, but I'm adding code to a large codebase which means a lot of time spent understanding it and a little time coding.  It's a lot simpler to write a project from scratch than to add functionality to something large.  As I'm taking a 500-level OS class, we're modifying an OS (Windows CE 6.0) and thus the code base is pretty big. 

It's coming together and I hope to get back to blogging more soon.  I just took a Microsoft class for Senior SDETs and have a lot of interesting ideas to blog about...