Big Data Observations: The Science of Asking Questions

“I am a firm believer that without speculation there is no good and original observation”—Charles Darwin

“It is the theory that determines what we can observe”—Albert Einstein

“I suspect, however, like as it is happening in many academic fields, the NSA is sorely tempted by all the data at its fingertips and is adjusting its methods to the data rather than to its research questions. That’s called looking for your keys under the light”—Zeynep Tufekci

“Large open-access data sets offer unprecedented opportunities for scientific discovery—the current global collapse of bee and frog populations are classic examples. However, we must resist the temptation to do science backwards by posing questions after, rather than before, data analysis. A scant understanding of the context in which data sets were collected can lead to poorly framed questions and results, and to conclusions that are plain wrong. Scientists intending to make use of large composite data sets need to work closely with those responsible for gathering the data. Standard scientific principles and practice then demand that they first frame the important questions, then design and execute the data analyses needed to answer them”—David B. Lindenmayer and Gene E. Likens

“The wonderful thing about being a data scientist is that I get all of the credibility of genuine science, with none of the irritating peer review or reproducibility worries… I thought I was publishing an entertaining view of some data I’d extracted, but it was treated like a scientific study… I’ve enjoyed publishing a lot of data-driven stories since then, but I’ve never ceased to be disturbed at how the inclusion of numbers and the mention of large data sets numbs criticism”—Pete Warden

Posted in Big Data, Data Science | Leave a comment

Visualising the Road to Becoming a Data Scientist

Source: Swami Chandrasekaran

Posted in Data Science | Leave a comment

The Digital Marketing Landscape: 2 Views

Gartner Digital Marketing Transit Map

Source: Gartner

marketing_technology_landscape_2012

Source: chiefmartec.com

Posted in Misc | Leave a comment

A Practical Introduction to Data Science Skills (Video)

Google’s Michael Manoochehri at DataEDGE 2013 presenting an introduction to  data analysis and suggestions for how to become a data scientist (his notes for the presentation are here).
[youtube=http://www.youtube.com/watch?v=rpwZ_i-9U0o&w=560&h=315]

Posted in Big Data, Data Science | Leave a comment

Big Data Quotes: Einstein, Come Back When You’ve Got Data

“Big data is what happened when the cost of storing information became less than the cost of making the decision to throw it away”—George Dyson (quoted by Tim O’Reilly)

“If the engineers have their way, every idea, memory, and feeling—the recorded consciousness of a single lifetime—will be stored in the cloud… ‘Information overload’ once referred to the difficulty of absorbing intelligently the data produced by others. Now we face the peril of choking on our own…By remembering everything, we may become haunted by our pasts and immobilized by digital distractions—or we may gain new powers to prevent the bad and promote the good”—G. Pascal Zachary

“[I]n a world where massive datasets can be analysed to identify patterns not easily identified using simpler analogue methods, what happens to genius of the Einstein variety?

Genius is about big ideas, not big data. Analysing the attributes and characteristics of anything is guaranteed to find some patterns. It is inherently a theoretical exercise, one that requires minimal thought once you’ve figured out what you want to measure. If you’re not sure, just measure everything you can get your hands on. Since the number of observations — the size of the sample — is by definition huge, the laws of statistics kick in quickly to ensure that significant relationships will be identified. And who could argue with the data?

Unfortunately, analysing data to identify patterns requires you to have the data. That means that big data is, by necessity, backward-looking; you can only analyze what has happened in the past, not what you can imagine happening in the future. In fact, there is no room for imagination, for serendipitous connections to be made, for learning new things that go beyond the data. Big data gives you the answer to whatever problem you might have (as long as you can collect enough relevant information to plug into your handy supercomputer). In that world, there is nothing to learn; the right answer is given…

What if Albert Einstein lived today and not 100 years ago? What would big data say about the general theory of relativity, about quantum theory? There was no empirical support for his ideas at the time — that’s why we call them breakthroughs.

Today, Einstein might be looked at as a curiosity, an ‘interesting’ man whose ideas were so out of the mainstream that a blogger would barely pay attention. Come back when you’ve got some data to support your point”—Sidney Finkelstein

Posted in Big Data | Leave a comment

Big Data Quotes: Disruptive Innovation?

“By definition, big data cannot yield complicated descriptions of causality. Especially in healthcare. Almost all of our diseases occur in the intersections of systems in the body. For example, there is a drug that is marketed by Elan BioNeurology called TYSABRI. It was developed for MS [multiple sclerosis]. It turns out that of the people who have MS a proportion respond magnificently to TYSABRI. And others don’t. So what do you conclude from this? Is it just a mediocre drug? No. It is that there is one disease but it manifests itself in different ways. How does big data figure out what is the core of what is going on?”–Clayton Christensen

Continue reading

Posted in Big Data | Leave a comment

Big Data, Small World: Kirk Borne at TEDxGeorgeMasonU (Video)

Kirk Borne is Professor of Astrophysics and Computational Science in the George Mason University School of Physics, Astronomy, and Computational Sciences (SPACS). Turns out he is the father of the term “unknown unknowns” – things we do not know we don’t know – popularized by former secretary of defense Donald Rumsfeld and later by Avinash Kaushik as “the unique space in which big data analysts should actually play.”

 

Posted in Big Data | Leave a comment

On Data Janitors, Engineers, and Statistics

Big Data Borat tweeted recently that “Data Science is 99% preparation, 1% misinterpretation.” Commenting on the 99% part, Cloudera’s Josh Wills says: “I’m a data janitor. That’s the sexiest job of the 21st century. It’s very flattering, but it’s also a little baffling.” Kaggle, the data-science-as-sport startup, takes care of the “1% misinterpretation” part by providing a matchmaking service between the sexiest of the sexy data janitors and the organizations requiring their hard-to-find skills. It charges $300 per hour for the service, of which $200 go to the data janitor (at least in the case of Shashi Godbole, quoted in the Technology Review article). Kaggle justifies its mark-up by delivering “the best 0.5% of the 95,988 data scientists who compete in data mining competitions,” the top of its data science table league, the ranking of data scientists based on their performance in Kaggle’s competitions, presumably representing  sound interpretation and top-notch productivity.

Kaggle’s co-founder Anthony Goldbloom tells The Atlantic’s Thomas Goetz that the ranking also represents a solution to a “market failure” in assessing the skills and relevant experience of the new breed of data scientists: “Kaggle represents a new sort of labor market, one where skills have been bifurcated from credentials.” Others see this as the creation of a new, $300 per hour, guild. In “Data Scientists Don’t Scale,” ZDnet’s Andrew Brust says that “’Data scientist’ is a title designed to be exclusive, standoffish and protective of a lucrative guild… The solution… isn’t legions of new data scientists. Instead, we need self-service tools that empower smart and tenacious business people to perform Big Data analysis themselves.”

Continue reading

Posted in Data Science, Statistics | Leave a comment

Big Data Friday

“I can’t express how infuriated I am that my credit history, phone activity, and online browsing habits are being systematically collected and archived without my knowledge by undisclosed organizations that aren’t trying to sell me products”–Area Man

“PRISM 1.0 was a little glitchy, and now that we’ve smoothed out the bugs, well, your privacy, especially inside your own home, will be a thing of the past. The technology is so good that it will basically be as if a member of the NSA is standing right behind you at all times”– NSA director General Keith B. Alexander announcing PRISM 2.0

“NSA email me with job offer. Offer say ‘To accept, nod once. To decline, nod twice’”–BigDataBorat

Posted in Misc | Leave a comment

Big Data in Context (Infographic)

HUMANIZING BIG DATA
Posted in Big Data, Infographics | Leave a comment