Wednesday, May 28, 2014

Code coverage -- meaningless number or the Holy Grail?

Code coverage is a term which is often brought up when discussing TDD. Code coverage is the percentage of the lines of code that are executed by unit tests.



Many teams have a specific goal for the code coverage. For example, a typical goal is to have at least 80% code coverage. Others claim that code coverage is a meaningless number because you can have 100% code coverage with nonsense tests that have no value. So which coverage should one aim for?

Coverage is not a measure of quality

You can have 100% test coverage, and yet the tests may be worthless. A single test which simply executes every line in the production code without asserting on anything will yield a 100% coverage, but it won't detect any defects. Clearly, coverage can't be used to measure test quality. So what is it good for?

Coverage is an alarm bell

Although a 100% coverage does not provide much information, the opposite case does. If you have a test coverage of 20%, it means that almost none of the production code is tested. Sound the alarm and do something!

Hence, code coverage should be used as an indication that something is wrong and not as a measure of test quality. 

So which coverage should we aim for?

Personally, I rarely measure code coverage. If you do proper TDD (write test first and never touch production code before you have a failing test), your coverage will naturally be close to 100% because you already test for all the desired behavior. Some test runners decorate the source code as you type to indicate whether it's covered by tests, so untested code will be screaming at you.

If a team wants to run periodical code coverage analysis in order to enable the lack-of-tests alarm bell, I would aim for close to 100% on testable code and 0% on untestable code. Testable and untestable code should ideally be separated into different modules of the application, but that's another story which deserves a separate blog post at some point in the future...

Monday, May 19, 2014

Test-driven development of plug-ins

How can we do test-driven development in a plug-in environment where our production code lives inside of an application? This may be the case if we are developing plug-ins for applications like Microsoft Office, Adobe Photoshop or Schlumberger Petrel. There are two challenges with this:
  1. NUnit (or any other test runner of choice) may not be able to start the host application in order to execute the plug-ins
  2. Starting the host application may take a while. If it takes 30 seconds to start the application, it's hard to get into the efficient TDD cycle that I describe in the blog post "An efficient TDD workflow".

Abstraction and isolation

Abstraction, inversion of control and isolation are common strategies when we develop code which is dependent on the environment. The idea is to make abstractions to the environment and thereby omitting it when executing the tests. It's not always possible, though. Sometimes our plug-in interacts heavily with and is dependent upon the behavior of the environment.

So what do we do then? If we can't isolate them, join them!

Unit test runner as a plug-in

Both challenges above can be solved by creating your own test runner as a plug-in inside of the host application! Instead of letting NUnit start the host application (slow and perhaps not even possible), the host application is running the tests.

So instead of doing a slow TDD cycle like this:


We want to move the host application startup out of the cycle like this:


So how can we do this? Create a test runner inside the host application... as a plug-in!

Let's have a look on one specific case: a test runner as a plug-in in Petrel. Petrel faces challenge #2 mentioned above. It can run in a "unit testing mode" where NUnit tests can start Petrel and run the Petrel-dependent production code, but startup takes 30-60 seconds.

If you are using NUnit, this is quite simple. NUnit provides multiple layers of test runners, depending on how much you want to customize its behavior. A very rudimentary implementation which allows the user to select a test assembly, list tests and execute them could look like this:

public partial class TestRunnerControl : UserControl
  {
    private readonly Dictionary<string, List<string>> _assembliesAndTests = new Dictionary<string, List<string>>();

    private string _pluginFolder;

    public TestRunnerControl()
    {
      InitializeComponent();

      this.runButton.Image = PlayerImages.PlayForward;

      FindAndListTestAssemblies();
    }

    private void FindAndListTestAssemblies()
    {
      var pluginPath = Assembly.GetExecutingAssembly().Location;
      _pluginFolder = Path.GetDirectoryName(pluginPath);

      AppDomain tempDomain = AppDomain.CreateDomain("tmpDomain", null, new AppDomainSetup { ApplicationBase = _pluginFolder });
      tempDomain.DoCallBack(LoadAssemblies);
      AppDomain.Unload(tempDomain);

      foreach (var testDll in _assembliesAndTests.Keys)
      {
        this.testAssemblyComboBox.Items.Add(testDll);
      }
    }

    private void LoadAssemblies()
    {
      foreach (
        var dllPath
          in Directory.GetFiles(_pluginFolder)
            .Where(f => f.EndsWith(".dll", true, CultureInfo.InvariantCulture) && f.Contains("PetrelTest")))
      {
        try
        {
          Assembly assembly = Assembly.LoadFrom(dllPath);

          var dllFilename = Path.GetFileName(dllPath);

          try
          {
            var typesInAssembly = assembly.GetTypes();

            foreach (var type in typesInAssembly)
            {
              var attributes = type.GetCustomAttributes(true);

              if (attributes.Any(a => a is TestFixtureAttribute))
              {
                if (!_assembliesAndTests.ContainsKey(dllFilename))
                {
                  _assembliesAndTests[dllFilename] = new List<string>();
                }

                _assembliesAndTests[dllFilename].Add(type.FullName);
              }
            }

            Ms.MessageLog("*** Found types in " + assembly.FullName);
          }
          catch (Exception e)
          {
            Ms.MessageLog("--- Could not find types in " + assembly.FullName);
          }
        }
        catch (Exception e)
        {
          Ms.MessageLog("--Could not load  " + dllPath);
        }
      }
    }  

    private void TestAssemblySelected(object sender, EventArgs e)
    {
      this.testClassComboBox.Items.Clear();

      var testAssembly = this.testAssemblyComboBox.SelectedItem as string;
      if (!string.IsNullOrEmpty(testAssembly))
      {
        foreach (var testClass in _assembliesAndTests[testAssembly])
        {
          this.testClassComboBox.Items.Add(testClass);
        }
      }
    }

    private void runButton_Click(object sender, EventArgs e)
    {     
      if (!CoreExtensions.Host.Initialized)
      {
        CoreExtensions.Host.InitializeService();
      }

      var results = RunTests();

      ReportResults(results);
    }

    private void ReportResults(TestResult results)
    {
      var resultsToBeUnrolled = new List<TestResult>();
      resultsToBeUnrolled.Add(results);

      var resultList = new List<TestResult>();
      while (resultsToBeUnrolled.Any())
      {
        var unrollableResult = resultsToBeUnrolled.First();
        resultsToBeUnrolled.Remove(unrollableResult);

        if (unrollableResult.Results == null)
        {
          resultList.Add(unrollableResult);
        }
        else
        {
          foreach (TestResult childResult in unrollableResult.Results)
          {
            resultsToBeUnrolled.Add(childResult);
          }
        }
      }

      int successCount = resultList.Count(r => r.IsSuccess);
      int failureCount = resultList.Count(r => r.IsFailure);
      int errorCount = resultList.Count(r => r.IsError);

      string successString = string.Format("{0} tests passed. ", successCount);
      string failureString = string.Format("{0} Tests failed. ", failureCount);
      string errorString = string.Format("{0} Tests had error(s). ", errorCount);

      string summary = successString + failureString + errorString;

      this.resultSummaryTextBox.Text = summary;

      this.resultSummaryTextBox.Select(0, summary.Length);
      this.resultSummaryTextBox.SelectionColor = Color.FromArgb(80, 80, 80);

      if (successCount > 0)
      {
        this.resultSummaryTextBox.Select(0, successString.Length);
        this.resultSummaryTextBox.SelectionColor = Color.DarkGreen;
      }

      if (failureCount > 0)
      {
        this.resultSummaryTextBox.Select(successString.Length, failureString.Length);
        this.resultSummaryTextBox.SelectionColor = Color.Red;
      }

      if (errorCount > 0)
      {
        this.resultSummaryTextBox.Select(successString.Length + failureString.Length, errorString.Length);
        this.resultSummaryTextBox.SelectionColor = Color.Red;
      }

      this.resultSummaryTextBox.Select(0, summary.Length);
      this.resultSummaryTextBox.SelectionAlignment = HorizontalAlignment.Center;

      this.resultSummaryTextBox.Select(0, 0);

      // Set grid results
      this.resultsGridView.Rows.Clear();

      int firstErrorIdx = -1;

      foreach (var result in resultList)
      {
        var testName = result.Name;
        var image = result.IsSuccess ? GeneralActionImages.Ok : StatusImages.Error;

        int idx = this.resultsGridView.Rows.Add(testName, image);

        if (firstErrorIdx == -1 && !result.IsSuccess)
        {
          firstErrorIdx = idx;
        }
      }

      if (firstErrorIdx != -1)
      {
        this.resultsGridView.FirstDisplayedScrollingRowIndex = firstErrorIdx;
      }
    }

    private TestResult RunTests()
    {
      var testAssembly = this.testAssemblyComboBox.SelectedItem as string;
      var testClass = this.testClassComboBox.SelectedItem as string;

      TestPackage testPackage = new TestPackage(Path.Combine(_pluginFolder, testAssembly));
      TestExecutionContext.CurrentContext.TestPackage = testPackage;

      TestSuiteBuilder builder = new TestSuiteBuilder();
      TestSuite suite = builder.Build(testPackage);

      var testFixtures = FindTestFixtures(suite);
      var desiredTest = testFixtures.First(f => f.TestName.FullName == testClass);
      var testFilter = new NameFilter(desiredTest.TestName);
      TestResult result = suite.Run(new NullListener(), testFilter);

      return result;
    }

    private IEnumerable<TestFixture> FindTestFixtures(Test test)
    {
      var testFixtures = new List<TestFixture>();

      foreach (Test child in test.Tests)
      {
        if (child is TestFixture)
        {
          testFixtures.Add(child as TestFixture);
        }
        else
        {
          testFixtures.AddRange(FindTestFixtures(child));
        }
      }

      return testFixtures;
    }
  }

Note that this code sample is not complete, but it shows how to find and run the tests. Use the test runner control in a plug-in inside the host application, and it will look like this:


Whenever a test is failing, Visual Studio will break at the offending NUnit Assert statement:


So how does this allow for an efficient code-test cycle? By using Visual Studio's Edit & Continue! Whenever you want to edit the production code or the test code, press pause in Visual Studio, edit as needed, and run the tests again. Hence, you can write code and (re)run tests at will without restarting the host application.

It's not as efficient as writing proper unit tests that execute in milliseconds, but it's far better than waiting for the host application on every test execution.

Wednesday, May 14, 2014

An efficient TDD workflow

How do you do your everyday TDD work? There are many ways to skin a cat, but I like to do it in a cyclic fashion:


I start by writing a very simple failing test. I then implement enough of the production code to just make this test pass. I then add more details to the test which make the test fail again. I implement more of the production code so that it passes again. Rinse and repeat until your tests are covering all the use cases.

Let's say that I want to write a Matrix class with an inversion method. I'd start with a super simple test case where I assert that the inverse of the identity matrix I is an identity matrix. I'd then add more and more cases:

  1. Create test which asserts that the inverse of the unity matrix I is a unity matrix. Production code returns an empty matrix, so it will fail
  2. Fix the production code so that it always returns I. The test will pass.
  3. Add new test case which asserts that the inverse of a different 2x2 matrix is correct. The test will fail
  4. Fix the production code so that it calculates the inverse of any 2x2 matrix. The test will pass
  5. Add new test case with a 3x3 matrix. The test will fail
  6. Fix the production code so that it calculates the inverse of any size matrix. The test will pass.
  7. Add new test case which asserts that calculating the inverse of a singular matrix throws an ArithmetricException. The test will fail because it tries to divide by zero
  8. Fix the production code so that it handles singular matrices correctly. The test will pass.
...and the show goes on until I am satisfied. Note that I'm not writing the entire test and then the entire production code method. Both the test and the production code evolve in parallel -- but I always have a failing test before doing anything with the production code.

Furthermore, I like to have the production code and the test code side-by-side like this so that I don't need to switch back and forth between the files:


If you are using a good test runner like NCrunch, you don't even need to run the tests manually. They are run automatically as you type, and you will have instant feedback when the test fails or passes!


Thursday, April 24, 2014

Sounds great, but what is the cost?

My latest posts have been very developer-centric. This post should address a broader audience, including those with suits and ties. Yes, managers and CEOs, I am talking to you! ;-)

Test-driven development is not only about writing code, however. Software development is about delivering value to a customer and earning money. Profit is income minus costs. One of the most common arguments against TDD is "it takes too long time" or "it is too expensive". So what is the cost of doing test-driven development?

It's not easy to put a price tag on test-driven development. We can get an idea about the alternative (not doing TDD) by looking on the Total Cost to Market, however. This is the total cost of a change in an application from idea to customer deployment. The total cost includes

  • Developing the feature
  • Testing
  • Bug fixing
  • Cost of risk that there are defects
  • Customer deployment
The cost incurred by actually typing the code is just a fraction of the Total Cost to Market.

Total cost to market

How much does it cost to develop a given feature?


In the beginning of a development phase, this is a fairly simple calculation: it's the number of hours spent on coding multiplied by the developers' hourly rate.

As time progresses and the product grows, this price increases because we add the risk of introducing bugs that are not discovered. A bug which is introduced late in the development cycle is more costly because there is an added risk that this bug (or a new bug that is introduced by the bug fix) is never discovered.

It becomes even worse after feature freeze when beta testing has started. If a bug is discovered after 90% of the testing has finished, how do we know that the bug fix did not introduce a new bug affecting the previously tested features? Do we test everything over again? Now that's an additional cost that contributes to the exponential cost graph!

And now the nightmare begins: what if the bug is discovered by the customer instead of our testers? That adds the additional cost of lost customer satisfaction and re-deployment. We can't easily put a number on this which can be put into a budget, but we certainly want to avoid that!

Fail early

The moral is; we want to fail early. We want to discover bugs as early as possible in the development cycle. The developer should get an alarm bell as soon as he creates the defect and not when a beta tester (or even worse, the customer) discovers it.

Bugs will occur, so we want to find and fix them while the cost is low. In other words, we want the bug to contribute as little as possible to the Total Cost to Market.

Good unit tests is a good first line of defense against bugs. Surely it's not a guarantee that you will deliver bug-free software, but it's a very efficient tool to deliver software with fewer bugs. Little research is done to actually quantify the economical benefit of TDD, but this study done on a few projects in IBM and Microsoft indicates that TDD roughly gave 25% increased developer time and reduced the number of bugs to a third: http://research.microsoft.com/en-us/groups/ese/fp17288-bhat.pdf

Conclusion

Did I answer my own question about the cost of doing TDD? No, I can't provide you with the number of dollars to put into your budget spreadsheet. However, it's pretty obvious that we can reduce the Total Cost to Market by discovering and fixing bugs as early as possible. That is a task which is far too important to rely only on manual testing.

Monday, March 24, 2014

NUnit TestCase and TestCase source

How do you test different input values for the method under test? Let's say that you want to test a class that converts sentences into single camel-case words. As an example, "what does the fox say" should be converted to "WhatDoesTheFoxSay".

There are many scenarios that need to be tested. There can be one or more words in the sentence. We should also test for an empty string. This can lead to many almost identical tests:


[Test]
public void MakeCamelCase_CalledWithEmptyString_ReturnsEmptyString()
{
  // Arrange
  var input = string.empty;

  // Act
  var result = StringTools.MakeCamelCase(input);

  // Assert
  Assert.AreEqual(string.empty, result);
}

[Test]
public void MakeCamelCase_CalledWithOneWord_ReturnsThatWord()
{
  // Arrange
  var input = "ab";

  // Act
  var result = StringTools.MakeCamelCase(input);

  // Assert
  Assert.AreEqual("Ab", result);
}

[Test]
public void MakeCamelCase_CalledWithTwoWords_ReturnsCamelCase()
{
  // Arrange
  var input = "ab cd";

  // Act
  var result = StringTools.MakeCamelCase(input);

  // Assert
  Assert.AreEqual("AbCd", result);
}


...and the list goes on. So how can we avoid repeating those almost identical tests?

TestCase

In NUnit, the test attribute TestCase comes to the rescue! Simply use one single test and provide multiple test cases as inputs:


[TestCase(string.empty, string.empty)]
[TestCase("ab", "Ab")]
[TestCase("ab cd", "AbCd")]
[TestCase("ab cd ef", "AbCdEf")]
public void MakeCamelCase_CalledWithString_ReturnsCamelCase(string input, string expectedResult)
{
  // Act
  var result = StringTools.MakeCamelCase(input)

  // Assert
  Assert.AreEqual(expectedResult, result);
}


Now these three tests plus an additional test case with three words are collapsed to one single easy-to-read test.

Note that the [TestCase] attribute can take any number of parameters, including input and expected output. The test itself takes those parameters as input.

With NUnit 2.5 or newer, the input and result parameters can be made a bit more readable by using the named attribute parameter Result:


[TestCase(string.empty, Result = string.empty)]
[TestCase("ab", Result = "Ab")]
[TestCase("ab cd", Result = "AbCd")]
[TestCase("ab cd ef", Result = "AbCdEf")]
public string MakeCamelCase_CalledWithString_ReturnsCamelCase(string input)
{
  // Act
  var result = StringTools.MakeCamelCase(input)

  return result;
}

TestCaseSource

All this is great, but [TestCase] has a limitation: it can only take constant expressions as parameters. If we were to test a mathematical algorithm on an input class like a Vector, we couldn't have used [TestCase]. There is another option, though: the TestCaseSource attribute:


[Test, TestCaseSource("VectorDotProductCases")]
public void DotProduct_CalledWithAnotherVector_ReturnsDotProduct(Vector lhs, Vector rhs, double expectedResult)
{
  // Act
  var dotProduct = lhs.DotProduct(rhs);

  // Assert
  Assert.AreEqual(expectedResult, dotProduct);
}

private static readonly object[] VectorDotProductCases =
{
  new object[] { new Vector(1,2,3), new Vector(0,0,0), 0 },
  new object[] { new Vector(1,2,3), new Vector(4,5,6), 32 },  
};

This is slightly less readable than using TestCase, but still it's better than replicating the tests for each set of input data.


Friday, February 28, 2014

Two readability tips - naming and organizing tests

I have mentioned earlier that readability is one of the pillars of good unit tests. Making readable unit tests is not trivial. It takes practice, but there are some good practices that can be applied to get a good start.

As an example, let's consider a test which verifies that a vector dot product is calculated correctly.

Naming

How do you name your unit tests? You should be able to get an idea of what a test verifies by just reading the name. Test names like this tell nothing about the tests:

[Test]
public void VectorDotProductTest1()
{
   ...
}

[Test]
public void VectorDotProductTest2()
{
   ...
}

This is more descriptive, but reading the names is still a bit akward:

[Test]
public void TestThatTheVectorDotProductIsZeroWhenOneOfTheVectorsIsZero()
{
   ...
}

[Test]
public void TestThatTheVectorDotProductIsDoneCorrectlyWhenVectorsAreNonZero()
{
   ...
}

Having a standard naming pattern like this makes it easier:

ItemUnderTest_Scenario_ExpectedBehaviour

As you can see, the name is divided into three parts, divided by underscores. The three parts are
  • Item under test: The item, usually a property, method or constructor, which is being tested
  • Scenario: The scenario, e.g. input data, parameters or other prerequisites
  • Expected behaviour: The expected result or outcome of the test

The previous examples would then be:

[Test]
public void DotProduct_OneVectorIsZero_ReturnsZero()
{
   ...
}

[Test]
public void DotProduct_VectorsAreNonZero_ReturnsCorrectProduct()
{
   ...
}



As you can see, dividing the names into sections following a defined pattern makes the names much easier to read.

Organizing the tests

How do you organize your tests? It should be perfectly clear to the user which part of the test is doing setup and preparation (arrangement), which part does the action which is being tested (act) and which part is doing the assertion (assert).

A common and recommended way to achieve this is to follow the "Arrange,Act,Assert" (AAA) pattern. The test is strictly divided into Arrange, Act and Assert sections like this:

[Test]
public void DotProduct_VectorsAreNonZero_ReturnsCorrectProduct()
{
  // Arrange
  var lhsVector = new Vector(1, 2, 3);
  var rhsVector = new Vector(4, 5, 6);

  // Act
  double result = lhsVector.DotProduct(rhsVector);

  // Assert
  Assert.AreEqual(26, result);
}


In this test, the AAA pattern makes it perfectly clear and obvious to the reader which part is doing what. Moreover, this pattern makes it easier to maintain the test. It's much harder to maintain a test where the asserts are spread all over the place.

If you find that it's tempting to add asserts in between the Arrange and Act sections, it's a sign that it's worth considering refactoring the test and/or the production code... or just take deep breath and have a coffee.


Monday, February 24, 2014

"We are too busy"

This seems to be the most common argument against test-driven development... ;-)