I welcome followers of my old blog -- politome.com -- to this new blog: TheDataCop.blogspot.com. Wiith the demise of intrade.com, a real-money prediction market based in Ireland but serving primarily American investors, I have decided to close down my politome.com blog which had been dedicated to the analysis of prediction markets. With this blog I will still bring attention to prediction markets (such as www.ipredict.co.nz) when events warrant, but my new focus will be on my ongoing research interests (e.g., econometrics, public mood/opinion in the Middle East/Asia, Bayesian statistics, and forecasting) and share with you any insights this research might generate, especially with regards to statistical methods and model building. And I will also occasionally react to analytics I see being presented in other contexts.
Which brings us to today's topic...BIG DATA. Or, as I like to call it, MESSY DATA. While my experience with BIG DATA is largely limited to analyzing Twitter-based data and other social media data sources, I have developed some opinions on the topic, particularly regarding its misapplication when trying to answer many important business problems.
My first observation is that BIG DATA is hardly new. To my ears, it does sound a bit more like a buzz word (think: Lean Six Sigma or Total Quality Management) than some meaningful advancement in analytic methods. I contrast BIG DATA to something like, Bayesian statistics, which is not a buzz word but in fact a significant methodological advancement. So, as college campuses scramble to add BIG DATA courses and degrees to their course catalogues, I would remind them that analyzing BIG DATA in academic settings has been going for decades. I knew colleagues analyzing millions of records (e.g., U.S Census data) way back in the 1980s. They didn't call it BIG DATA. They called it, "I'm analyzing U.S. Census data." And they faced many of the problems that challenge analysts today analyzing large, messy data sources. For example, computers sometimes have a hard time calculating variance (and other statistical measures) in the presence of millions of data cases. In traditional statistics, calculating variance is simple: Summing over all cases the difference between a data value and the mean and dividing that sum by the total number of cases. Not a big deal when you have a sample size of 1,000. Remember that many of the classical statistics that we use today were developed using relatively small sample sizes (e.g., Fisher's Iris Dataset had 150 cases). However, it is a very big deal for computers to calculate variance (and other statistical measures) when your analytic base size is over a million cases. Hence, BIG DATA has given birth to platforms like Hadoop (among many other software solutions) to handle this problem. Just take the problem, slice it up into much smaller clusters, do the calculations, and then put it all back together into one single analysis. That is the essence of BIG DATA analysis. An over-simplification, perhaps, but a good summary I do believe.
Second, I find the field of BIG DATA (and its cousin: data mining) to be dominated by computer scientists and other IT-trained folks. Their talents are obviously important when you want to warehouse, manipulate and query billions of data records on the fly without crashing your network. But this emphasis on the technology behind BIG DATA instead of the intrinsic meaning behind the data has led to a lot companies shifting millions of dollars away from strategic research to BIG DATA analytics that often ignore good theoretical practice (particularly the primacy of psychological factors in understanding human behavior) and focus instead on "kitchen-sink" modeling methods that utilize plentiful and easily available data rather than relevant data. I believe social science has never been more important than in the meaningful application of BIG DATA to solving real-life business problems,
Third, and related to the previous point, is my observation that data that businesses tend to collect are not explicitly designed to answer specific research questions. Instead, they are more often defined by the requirements and conveniences of data architects and database administrators who rarely have direct training or experience in social scientific inquiry. What can happen in this environment is that the variables analysts need to answer critical business questions are not directly measured forcing researchers instead to utilize "proxy variables" that may or may not correlate strongly with what the researchers really want (or, better yet, need) in their statistical analyses. That is a serious flaw in BIG DATA
Second, I find the field of BIG DATA (and its cousin: data mining) to be dominated by computer scientists and other IT-trained folks. Their talents are obviously important when you want to warehouse, manipulate and query billions of data records on the fly without crashing your network. But this emphasis on the technology behind BIG DATA instead of the intrinsic meaning behind the data has led to a lot companies shifting millions of dollars away from strategic research to BIG DATA analytics that often ignore good theoretical practice (particularly the primacy of psychological factors in understanding human behavior) and focus instead on "kitchen-sink" modeling methods that utilize plentiful and easily available data rather than relevant data. I believe social science has never been more important than in the meaningful application of BIG DATA to solving real-life business problems,
Third, and related to the previous point, is my observation that data that businesses tend to collect are not explicitly designed to answer specific research questions. Instead, they are more often defined by the requirements and conveniences of data architects and database administrators who rarely have direct training or experience in social scientific inquiry. What can happen in this environment is that the variables analysts need to answer critical business questions are not directly measured forcing researchers instead to utilize "proxy variables" that may or may not correlate strongly with what the researchers really want (or, better yet, need) in their statistical analyses. That is a serious flaw in BIG DATA
Finally, BIG DATA is inherently tactical data. Which is to say, it rarely provides they type of strategic information that is most critical to long-term business decisions. Today, U.S. businesses disproportionately shifting analytic resources to BIG DATA are more vulnerable than ever to strategic surprise. Yes, every business should care about its tactical issues -- What was the impact of my company's latest email marketing campaign? Are my customers using the website in an efficient and effective manner? Many consumer-oriented websites are a mess and make it difficult for users to find the information they need or want.. The value of BIG DATA to understand how to identify these problems and improve this mess is unquestioned (http://www.sciencedirect.com/science/article/pii/S0378720605000169). BIG DATA can tell you a lot about your inventory and supply chain processes. BIG DATA can help you anticipate resource requirements thereby minimzing costly process breakdowns. Nonetheless, this is tactical information. Companies using BIG DATA analytics are collecting BIG DATA on themselves and rarely, if ever, on their competitors. Tactical information is information located at the point of direct interaction between the consumer and the company. It is important information but it is not strategic information. Strategic information occurs at a higher level of analysis. From the perspective of the company, it is the contextual environment vis-a-vis their competitors that defines the strategic space. From the consumer perspective, the strategic space is defined by the psychological needs and aspirations of the consumer. BIG DATA is not equipped to understand this level of analysis. BIG DATA focuses on behavioral and demographic information.
On a lighter focus, I found this graphic built from an analysis of Twitter traffic during the 2014 World Cup.
What this analysis tell us is that when Germany scored a goal against Brazil in the World Cup semifinal match, Twitter traffic exploded. OK. I confess. The above graphical depiction of Twitter traffic states the obvious. And that is the problem I've seen with most BIG DATA analyses I've seen. BIG DATA generally peddles in the obvious. It takes a lot of information, summarizes it in attractive graphical forms, but in the end provides little valuable insight. BIG DATA is the "Three's Company" of the information age. It's fun to look at but it doesn't add up to much.
Frankly, I found this Twitter analysis much more interesting (even if it is equally as obvious as the previous graphic):
Yes, when anything involving Germany happens on the world stage, people invariably march out references to Nazis. Thank you Twitter for making that clear.
On my next blog post, I will address a topic much more significant in my view: Bayesian statistical methods. I will address it in the context of interest in public opinion in Afghanistan. I will state upfront: I am believer in the ways of Bayesian. But I am new to its implementation and, frankly, not completely coherent in its theoretical underpinnings.
But, hopefully, you will gain something from my own struggles with the Bayesian appoach to statistics.
No comments:
Post a Comment