---
title: "When Organizing Big Data Content, It’s Okay To Be Messy"
date: 2013-08-07
description: "$( document ).ready(function() { // Handler for .ready() called. $("
canonical_url: https://idratherbewriting.com/2013/08/07/big-data-when-its-okay-to-be-messy/
---
# When Organizing Big Data Content, It’s Okay To Be Messy
Lowdermilk acknowledges that the idea has been challenged by some, including [Jared Spool](http://dl.acm.org/citation.cfm?id=634236) and others (see [When and Why 5 Test Users Isn't Enough](http://nri.netraker.com/nrinfo/research/FiveUsers.pdf)). And to be fair, Nielsen's theory is more nuanced. Nielsen says if you have 15 research participants, it's best to divide up the participants into 3 smaller studies of 5 each and then iterate your design with each small group, thereby maximizing the value of your research.

But the exact sample size isn't my point here. I assume the UX or HCI proponents latched onto the 5-users-only theory as a means of simplifying user research. Rather than asking developers to orchestrate a massive user research study involving hundreds of users and then quantifying scores of data points, and then tabulating the responses in some meaningful and intelligent way to arrive at a conclusion, you just have to gather 5 users.

My point, rather, is that people use samples rather than wholes because the full data is too massive, unwieldy, and difficult to gather and manage. However, with the trend toward big data, the convention of working with samples rather than the full spectrum of data may be a tradition of the past we discard.

## Enter Big Data

I've recently been listening to [Big Data: A Revolution That Will Transform How We Live, Work, and Think](https://www.amazon.com/books/dp/0544002695), which is quite an interesting book. The authors, Viktor Mayer-Schonberger and Kenneth Cukier, explain some of the difficulties of collecting information from whole groups. For example, they relate the challenges in census taking and in collecting information about flu epidemics.

Apparently it took more than a decade to collect information for the 1880 census. By that time the data was gathered, the results were already out of date. Similarly, during the H1N1 outbreak, the Center for Disease Control tried to take and process incoming reports to track the spread of the outbreak. Reports were slow. Instead, Google found that it could predict the path of the outbreak in near real-time by correlating millions of search queries about H1N1 and location (see [Google Flu Trends](http://www.google.org/flutrends/us/#US)).

Mayer-Schonberger and Cukier assert that we aren't restricted to using samples (and hoping they're representative) in order to make analyze information. With millions of people liking posts on Facebook, millions more uploading and tagging videos in Youtube, millions posting micro tweets about what they think or what's happening, and the incredible processing power of computers, it's possible to grab all this data and analyze it for insights.

You don't have to capture a 5% sample and hope that it's representative enough to predict the whole. Mayer-Schonberger and Cukier don't make any mention of Nielsen's 5-users-only theory about UX research, but I'm guessing that scenarios that depend on massive extrapolation from an extremely small sample will become obsolete in favor of big data crunching scenarios.

How can you tell if your prototype works well or not? Just as big data might listen for engine hum vibrations and other noise patterns to predict end-of-life factors for engines, web researchers might capture raw keystrokes and other usability patterns from eye-tracking movements to mouse clicks and navigation patterns (most of which are captured in the browser) and then correlate this massive amount of information to evaluate different prototypes.

In other words, the 102,000 eye twitches, 55,000 keystrokes, 3,500 mouse clicks, 27,000 page loads, 3500 browser crashes, 3,000 page freezes, 45,000 seconds on each pages, and 133,00 different paths through site collected from 100 users might tell you something different from a couple of users saying, "Gee, I'm not sure if I like this screen or not" and "I like the color of this button."

In fact, you could run usability studies post-release as well, capturing a ton of information each time people use a product. Researchers can conduct sentiment analysis with the millions of social media posts and other feedback to determine overall trends about what "sucks" and what "rocks."

## Finding Insights in Big Data

What can you do with all of this data you collect? Discover unanticipated insights, apparently.

According to Mayer-Schonberger and Cukier, once you start analyzing big data, you start seeing surprises that you don't expect. For example, [Oren Etzioni](http://homes.cs.washington.edu/~etzioni/bio.html), a big data researcher, discovered that airline prices don't always increase as the airfare departure data gets closer (see [Farecast --> Bing/Travel](https://www.bing.com/travel/)). There are some interesting patterns that tend to arise when you look at large amounts of information.

Some researchers number crunched big data from sumo matches and discovered irregularities that clued people to match cheating. Others analyzed credit card transactions to identify irregularities in usage that tipped off officials to fraud (in New Jersey).

Mayer-Schonberger and Cukier assert that rather than looking for causal patterns that help you predict results based on underlying factors, big data moves us more toward probability. We can say that given the occurrence of X, it's likely that Y will also be present -- regardless of what's causing X or Y.

One of the most intriguing application is with DNA sequencing. Given X pattern, it's likely that Y disease results, and so on. Never mind what causes X or Y -- what matters is that X leads to a likelihood of Y when you analyze millions of data points.

## Messy Is Okay Because It's More Accurate

How does big data apply to the tech comm profession? Mayer-Schonberger and Cukier explain that neatly classifying all the data into precise buckets is not really possible with big data. Instead, big data uses a more messy and imprecise tagging methodology (for example, tagging photos on Yahoo's Flickr site).

Why tagging? The authors explain that in the days before big data, sampling was the only means of analyzing the information. And with the small sample, we could be careful to exclude outliers, to assess and classify all the sample data in very clean, neat ways. You could put all the books in a library into various card catalogs sorted by title, author, or subject.