Machines vs. Models, Noise vs. Signal

An excerpt from Nassim Taleb’s forthcoming book, Antifragile, was posted yesterday on the Farnam Street blog. In “Noise and Signal,” Taleb says that “In business and economic decision-making, data causes severe side effects —data is now plentiful thanks to connectivity; and the share of spuriousness in the data increases as one gets more immersed into it. A not well discussed property of data: it is toxic in large quantities—even in moderate quantities…. the best way… to mitigate interventionism is to ration the supply of information, as naturalistically as possible. This is hard to accept in the age of the internet. It has been very hard for me to explain that the more data you get, the less you know what’s going on, and the more iatrogenics you will cause.”   Continue reading

Posted in Artificial Intelligence, Big Data, Machine Learning | Leave a comment

Facebook’s IPO and the Laws of Big Data

Without using any predictive analytics tools, I confidently predict that Facebook’s IPO will give rise to more vocal demands for people to “get a cut” of its—and other social media companies’—profits. People deserve, so the argument goes, a share of any profits derived from mining the social data pool which they have so willingly helped create. Occupy Facebook, anyone?

But before you set up a tent in Menlo Park, consider this proposition: The value of personal data is zero. Personal data is not worth much if it’s kept personal and a sample of one is good for answering a very limited set of questions. Personal data gains value when it is shared, when it is combined with and compared to other data.  Continue reading

Posted in Big Data, Data Science | Leave a comment

The Reality of Big Data: Findings from Recent Surveys

Big data tools and technologies emerged first from the companies the Web gave birth to–Google, Facebook, Yahoo, and Amazon. No wonder that the term has become associated primarily with the ability to process and analyze large sets of unstructured, web-generated data, for consumer- and market-related activities such as targeted advertising or improving customer loyalty.   Continue reading

Posted in Misc | Leave a comment

Big Data Will Make IT the New Intel Inside

Tim O’Reilly famously declared in 2005: “Data is the next Intel Inside.” It well may be that big data—the organizational skill of using data as the key driver of performance—will make the IT function the new Intel Inside, the most strategic component of any organization.

Based on the success of Google, Amazon, and eBay at the time, O’Reilly correctly asserted that database management was a core competency of Web 2.0 companies and that “control over the database has led to market control and outsized financial returns.” The upcoming—and outsized—Facebook IPO is a testament to O’Reilly’s foresight, saying in 2005 that “data is the Intel Inside of [Web 2.0] applications, a sole source component in systems whose software infrastructure is largely open source or otherwise commodified.”  Continue reading

Posted in Big Data, Data Science | Leave a comment

Upcoming Big Data and Data Science Events

From Data to Knowledge

May 7-11, University of California, Berkeley

Continue reading
Posted in Misc | Leave a comment

A Very Short History of Data Science

data-science-jobs
Source: http://compsocsci.blogspot.com/

I’m in the process of researching the origin and evolution of data science as a discipline and a profession. Here are the milestones that I have picked up so far, tracking the evolution of the term “data science,” attempts to define it, and some related developments.  I would greatly appreciate any pointers to additional key milestones (events, publications, etc.).

[An updated version of this timeline is at Forbes.com]

1974 Peter Naur publishes Concise Survey of Computer Methods in Sweden and the United States. The book is a survey of contemporary data processing methods that are used in a wide range of applications. It is organized around the concept of data as defined in the IFIP Guide to Concepts and Terms in Data Processing, which defines data as “a representation of facts or ideas in a formalized manner capable of being communicated or manipulated by some process.“ The Preface to the book tells the reader that a course plan was presented at the IFIP Congress in 1968, titled “Datalogy, the science of data and of data processes and its place in education,“ and that in the text of the book, ”the term ‘data science’ has been used freely.” Naur offers the following definition of data science: “The science of dealing with data, once they have been established, while the relation of the data to what they represent is delegated to other fields and sciences.”

Continue reading
Posted in Big Data, Data Science | Leave a comment

New Research Reports on Big Data

Two new research reports on big data flash out its early impact on enterprise IT. Continue reading

Posted in Big Data, Data Science | Leave a comment

Domain Expertise vs. Machine Learning: The Debate Continues

“The startup’s three co-founders have backgrounds in engineering and data science, but not weather, and there are no meteorological models involved. By keeping weather predictions within a two-hour window, they believe statistics are sufficient.”–Mashable in “Can Statistics Predict Weather Without Meteorologists? This App Thinks So” on Ourcast, a  new app that uses real-time radar data and crowdsourcing to predict how weather at a given location will change within the next two hours.  Continue reading

Posted in Data Science, Machine Learning, Statistics | Leave a comment

Data Science is so 1996!

 

Source: A History of the International Federation of Classifi cation Societies

Data Science is so 1996!
Posted in Data Science | Leave a comment

Graduate Programs in Big Data Analytics/Data Science

Updated list here

Bentley University

M.S. in Marketing Analytics

DePaul University

M.S. in Predictive Analytics

Continue reading
Posted in Big Data, Data Science | Leave a comment