Big Data: A Revolution that Will Transform How We Live, Work, and Think

Viktor Mayer-Schönberger and Kenneth Cukier, authors of the just-published Big Data: A Revolution that Will Transform How We Live, Work, and Think,  reacted sharply when I asked them if they are cheerleaders for big data, as one reviewer implied. ”We are messengers of big data, not its evangelists,” said Cukier. Added Mayer-Schönberger: “The reviewer did not read the book.”

I did. Big Data is an excellent introduction for general audiences to what has become a topic of conversation everywhere, faster than any other technology-driven buzzword in recent memory. To those who may react to “big data” as today’s incarnation of “big brother,” Mayer-Schönberger and Cukier offer a comprehensive and highly readable overview of the benefits and risks associated with big data, which they define as “the ability of society to harness information in novel ways to produce useful insights or goods and services of significant value.”    Continue reading

Posted in Big Data | Leave a comment

Big Data Analytics and Data Science at Netflix (Video)

Chris Pouliot, the Director of Analytics and Algorithms at Netflix: “…my team does not only personalizations for movies, but we also deal with content demand prediction. Helping our buyer down in Beverly Hills figure out how much do we pay for a piece of content. The personalization recommendations for helping users find good movies and TV shows. Marketing analytics, how do we optimize our marketing spin. Streaming platform, how do we optimize the user experience once I press play. There’s a wide range of data, so theres a lot of diversity. We have a lot of scale, a lot of challenging problems. The question then is, how do we attract great data scientists that can just see this as a playground, a sandbox of really exciting things. Challenging problems, challenging data, great tools, and then just the ability to have fun and create great products.”
[youtube http://www.youtube.com/watch?v=pJd3PKm9XUk]

Posted in Big Data, Data Science | Leave a comment

LinkedIn’s Daniel Tunkelang on How to Interview a Data Scientist

Tunkelang: The O’Reilly Strata Conference brings together an incredible community of people working on big data. This year, I decided to do something different for my presentation. Rather than talk about science or technology, I addressed the practical problem of interviewing the candidates to build teams of data scientists.

[slideshare id=16798687&w=427&h=356&sc=no]

Posted in Big Data, Data Science | Leave a comment

Cool Data Scientists on Campus

Geek Chic

Hal Varian:  “Data availability is going to continue to grow. To make that data useful is a challenge. It’s generally going to require human beings to do it.”

Source: Carl Bialik, “Data Crunchers Now the Cool Kids on Campus,” The Wall Street Journal, March 1, 2013

See my list of graduate programs in data science and big data analytics

Posted in Big Data, Data Science, Statistics | Leave a comment

Vincent Granville’s 66 job interview questions for data scientists

 

  1. What is the biggest data set that you processed, and how did you process it, what were the results?
  2. Tell me two success stories about your analytic or computer science projects? How was lift (or success) measured?
  3. What is: lift, KPI, robustness, model fitting, design of experiments, 80/20 rule?
  4. What is: collaborative filtering, n-grams, map reduce, cosine distance?
  5. How to optimize a web crawler to run much faster, extract better information, and better summarize data to produce cleaner databases?
  6. How would you come up with a solution to identify plagiarism?
  7. How to detect individual paid accounts shared by multiple users?
  8. Should click data be handled in real time? Why? In which contexts?
  9. What is better: good data or good models? And how do you define “good”? Is there a universal good model? Are there any models that are definitely not so good?
  10. What is probabilistic merging (AKA fuzzy merging)? Is it easier to handle with SQL or other languages? Which languages would you choose for semi-structured text data reconciliation?

To see the other 56 questions assessing “the technical horizontal knowledge of a senior candidate at a high level” go here 

Posted in Data Science | Leave a comment

The Big Data Explosion (Infographic)

Lotsa data in this Infographic about data growth

Continue reading

Posted in Big Data, Infographics | Leave a comment

Data Science at Netflix with Elastic MapReduce

[youtube http://www.youtube.com/watch?v=oGcZ7WVx6EI]

Posted in Data Science | Leave a comment

DJ Patil at LeWeb, December 2012

[youtube http://www.youtube.com/watch?v=J_CYKk8q1Ao]

Summary of the presentation by Ben Rooney here

Update: Ben Rooney interviews DJ Patil

[youtube http://www.youtube.com/watch?v=0LtzMhr0ZCM]

Posted in Data Science | Leave a comment

Past Courses in Big Data Analytics and Data Science: Content Online

Past Courses

in Big Data Analytics and Data Science

Content Online

Analyzing Big Data with Twitter (UC Berkeley, School of Information) (Fall 2012)

Introduction to Data Science (Columbia University, Statistics Department) (Fall 2012

Introduction to  Data Science (UC Berkeley, Computer Science) (Spring 2011)

Posted in Big Data, Data Science | Leave a comment

Big Data Quotes of the Week: December 1, 2012

“Let us cultivate the mathematical sciences with ardor, without wanting to extend them beyond their domain; and let us not imagine that one can attack history with formulas, nor give sanction to morality through theories of algebra or the integral calculus”–Augustin-Louis Cauchy, 1821, quoted by Matthew Jones, Columbia University

“…the common language of business is not going to be Chinese or Spanish. It’s going to be math”–Michael Rhodin, IBM

“The future is going to be owned by people who are comfortable in the quant world but have deep business knowledge”–Christine Poon, Max M. Fisher College of Business, Ohio State

“[One false promise that some proponents of Big Data hold out is that somehow vast oceans of digital data can be sifted for nuggets of pure enterprise gold.] It is not going to happen magically. The software only finds correlations, not causations. In order to find causal relationships you have to do work. If you take any sufficiently large data sets, you are going to find correlations. You need a human in the loop to work out which are important”–Stephen Sorkin, Splunk

Continue reading
Posted in Big Data | Leave a comment