Big Data Quotes

“Big data is like teenage sex: everyone talks about it, nobody really knows how to do it, everyone thinks everyone else is doing it, so everyone claims they are doing it…”—Dan Ariely

“I’m a data janitor. That’s the sexiest job of the 21st century. It’s very flattering, but it’s also a little baffling”–Josh Wills, a senior director of data science at Cloudera

“Given enough data, everything is statistically significant”–Douglas Merrill

Posted in Big Data | Leave a comment

The Real World of Big Data (Infographic)

Click image to see a larger version

The Real World of Big Data via Wikibon Infographics

Posted in Big Data, Infographics | Leave a comment

3 Big Data Milestones

If you were asked to name the top three events in the history of the IT industry, which ones would you choose? Here’s my list:

June 30, 1945: John Von Neumann published the First Draft of a Report on the EDVAC, the first documented discussion of the stored program concept and the blueprint for computer architecture to this day.

May 22, 1973: Bob Metcalfe “banged out the memo inventing Ethernet” at Xerox Palo Alto Research Center (PARC).

March 1989: Tim Berners-Lee circulated “Information management: A proposal” at CERN in which he outlined a global hypertext system.

[Note: if round numbers are your passion, you may opt—without changing the substance of this condensed history—for the ENIAC proposal of April 1943, Ethernet in 1973, and CERN making the World Wide Web available to the world free of charge in April 1993, so that 2013 marks the 70th, 40th, and 20th anniversaries of these events.]

Why bother at all to look back? And why did I select these as the top three milestones in the evolution of information technology?

Most observers of the IT industry prefer and are expected to talk about what’s coming, not what’s happened. But to make educated guesses about the future of the IT industry, it helps to understand its past. Here I depart from most commentators who, if they talk at all about the industry’s past, divide it into hardware-defined “eras,” usually labeled “mainframes,” “PCs,” “Internet,” and “Post-PC.”

Another way of looking at the evolution of IT is to focus on the specific contributions of technological inventions and advances to the industry’s key growth driver: digitization and the resulting growth in the amount of digital data created, shared, and consumed. Each of these three events represents a leap forward, a quantitative and qualitative change in the growth trajectory of what we now call big data.

The industry was born with the first giant calculators digitally processing and manipulating numbers and then expanded to digitize other, mostly transaction-oriented activities, such as airline reservations.  But until the 1980s, all computer-related activities revolved around interactions between a person and a computer. That did not change when the first PCs arrived on the scene.

The PC was simply a mainframe on your desk. Of course it unleashed a wonderful stream of personal productivity applications that in turn contributed greatly to the growth of enterprise data and the start of digitizing leisure-related, home-based activities. But I would argue that the major quantitative and qualitative leap occurred only when work PCs were connected to each other via Local Area Networks (LANs)—where Ethernet became the standard—and then long-distance via Wide Area Networks (WANs). With the PC, you could digitally create the memo you previously typed on a typewriter, but to distribute it, you still had to print it and make paper copies. Computer networks (and their “killer app,” email) made the entire process digital, ensuring the proliferation of the message, drastically increasing the amount of data created, stored, moved, and consumed.

Connecting people in a vast and distributed network of computers not only increased the amount of data generated but also led to numerous new ways of getting value out of it, unleashing many new enterprise applications and a new passion for “data mining.” This in turn changed the nature of competition and gave rise to new “horizontal” players, focused on one IT component as opposed to the vertically integrated, “end-to-end solution” business model that has dominated the industry until then. Intel in semiconductors, Microsoft in operating systems, Oracle in databases, Cisco in networking, Dell in PCs (or rather, build-to-order PCs), and EMC in storage have made the 1990s the decade in which “best-of-breed” was what many IT buyers believed in, assembling their IT infrastructures from components sold by focused, specialized IT vendors.

The next phase in the evolution of the industry, the next quantitative and qualitative leap in the amount of data generated, came with the invention of the World Wide Web (commonly mislabeled as “the Internet”). It led to the proliferation of new applications which were no longer limited to enterprise-related activities but digitized almost any activity in our lives. Most important, it provided us with tools that greatly facilitated the creation and sharing of information by anyone with access to the Internet (the open and almost free wide area network only few people cared or knew about before the invention of the World Wide Web). The work memo I typed on a typewriter which became a digital document sent across the enterprise and beyond now became my life journal which I could discuss with others, including people on the other side of the globe I have never met.  While computer networks took IT from the accounting department to all corners of the enterprise, the World Wide Web took IT to all corners of the globe, connecting millions of people. Interactive conversations and sharing of information among these millions replaced and augmented broadcasting and drastically increased (again) the amount of data created, stored, moved, and consumed. And just as in the previous phase, a bunch of new players emerged, all of them born on the Web, all of them regarding “IT” not as specific function responsible for running the infrastructure but as the essence of their business, data and its analysis becoming their competitive edge.

We are probably going to see soon—and maybe already are experiencing—a new phase in the evolution of IT and a new quantitative and qualitative leap in the growth of data. The cloud—a new way to deliver IT, big data—a new attitude towards data and its potential value, and The Internet of Things (including wearable computers such as Google Glass)—connecting billions of monitoring and measurement devices quantifying everything—combine to sketch for us the future of IT.

[Originally published on Forbes.com]

Posted in Big Data | Leave a comment

Big Data Friday: Borasky’s Law

  • Murphy’s Law: Anything that can go wrong, will go wrong.
  • O’Toole’s Corollary: Murphy was an optimist.
  • Sturgeon’s Law: 95 percent of everything is crap.
  • Mencken’s Law: Nobody ever went broke underestimating the intelligence of the American public.

Borasky’s Law: Sturgeon and Mencken were optimists, too.

Source: What Hath Von Neumann Wrought?

Posted in Misc | Leave a comment

The Big Data Landscape Revisited

Bruce Reading, CEO of VoltDB, has an interesting and original take on the big data landscape.

Last year, Dave Feinleib published the Big Data Landscape, “to organize this rapidly growing technology sector.” One prominent data scientist told me “it’s just a bunch of logos on a slide,” but it has become a popular reference point for categorizing the different players in this bustling market. Sqrrl, a big data start-up, published recently its own version of Feinleib’s chart, its “take on the big data ecosystem.” Sqrrl’s eleven big data “buckets” are somewhat different from Feinleib’s, demonstrating a lack of agreement, understandable at this stage, on what exactly are the different segments of the big data market and what to call them. Furthermore, Sqrrl positions itself “at the intersection of four of these boxes” which raises questions about the accuracy of its positioning  of other big data companies inside just one or two boxes.

Another interesting recent attempt to make sense of the big data landscape comes from The 451’s Matt Aslett in the form of a “Database Landscape Map.” Taking its inspiration from the map of the London Underground and a content technology map from the Real Story Group, it charts the links between an ever-expanding database market and the data storing/organizing/mining technologies and tools (Hadoop, NoSQL, NewSQL…) that now form the core of the big data market.

Which brings me to Bruce Reading, VoltDB, and their take on the big data landscape. “It’s a very noisy market,” Bruce told a packed room at a recent VoltDB event. “It’s like shopping in a mall at Christmas time when there’s a lot of noise and a lot of information about a lot of technologies. We are trying to work with the marketplace to understand what you are trying to accomplish. Instead of using market maps based on technologies, we are looking at use cases.”

“Use case” is technology-speak for the list of requirements for achieving a specific goal, requirements that are embodied in the software that allows the user to achieve that goal. In other words, specialized software focused on addressing some unique need. VoltDB is focused on time (or data velocity) and believes, to quote Bruce, that “the whole world is trying to get as close to real-time as possible because that’s where the greatest value is of a single point of data.” Or, in the words of VoltDB’s website, companies are “devising new ways to identify and act on fast-moving, valuable data,” and VoltDB helps them “narrow the ‘ingestion-to-decision’ gap from minutes, or even hours, to milliseconds.” Which is why they see the “Data Value Chain” like this:Big-Data-Landscape

And describe the “Database Universe” like this:

This is the first attempt I’ve seen to map big data technologies based on what these technologies are trying to achieve and the type of data involved–is it unique (an individual item) or is it a part of a collection of data?–along three dimensions: Time, the value of the data, and application complexity.

The insight behind these charts is that the value of an individual piece of data goes down with time and the value of a collection of data goes up with time. Maybe this should be called “Stonebraker Law.” Mike Stonebraker is the database legend (forty years and counting) behind VoltDB and other big data startups. You can watch him, Bruce, and John Piekos, VoltDB’s  VP of Engineering, here.

[Originally published on Forbes.com]

Posted in Big Data | Leave a comment

The Big Data Debate: Correlation vs. Causation

In the first quarter of 2013, the stock of big data has experienced sudden declines followed by sporadic bouts of enthusiasm. The volatility—a new big data “V”—continues this month and Ted Cuzzillo summed up the recent negative sentiment in “Big data, big hype, big danger” on SmartDataCollective:

“A remarkable thing happened in Big Data last week. One of Big Data’s best friends poked fun at one of its cornerstones: the Three V’s. The well-networked and alert observer Shawn Rogers, vice president of research at Enterprise Management Associates, tweeted his eight V’s: ‘…Vast, Volumes of Vigorously, Verified, Vexingly Variable Verbose yet Valuable Visualized high Velocity Data.’ He was quick to explain to me that this is no comment on Gartner analyst Doug Laney’s three-V definition. Shawn’s just tired of people getting stuck on V’s.”

Indeed, all the people who “got stuck” on Laney’s “definition,” conveniently forgot that he first used the “three-Vs” to describe data management challenges in 2001. Yes, 2001. If big data is a “revolution,” how come its widely-used “definition” is based on a dozen year-old analyst note?

Continue reading

Posted in Big Data, Data Science | Tagged | Leave a comment

Scenarios for the Future of the IT Industry

In November 1998, I sent to my then-colleagues at EMC an email with the subject line “The Demise of Dell.” I wrote:

“My fail-proof crystal ball just talked to me again: By the end of 2000, Dell’s market cap (today at $80B) will be cut in half.

Dell’s only strength, as we all know, is in low-cost distribution. Distribution (of everything) is going to undergo a radical change in the near future because of the Internet.  There will be new players in the PC market that will figure out how to sell PCs over the Internet at half the cost of Dell’s distribution infrastructure. On top of that, the corporate PC market will grind to a halt and we may even see a slight drop in PC revenues in the year 2000. On the consumer side, appliances is where the action will be—led by new players. “

After I sent my email, Dell’s stock went on to almost double to a peak of just over $56 in March 2000. It closed yesterday at $14.09, about half of where it was in late 1998.

Continue reading

Posted in Predictions | Leave a comment

Who’s Big in Big Data? (Infographic)

Who's Big in Big Data? (Infographic)

Source: Datasift

Posted in Big Data, Infographics | Leave a comment

Big Data: A Revolution that Will Transform How We Live, Work, and Think

Viktor Mayer-Schönberger and Kenneth Cukier, authors of the just-published Big Data: A Revolution that Will Transform How We Live, Work, and Think,  reacted sharply when I asked them if they are cheerleaders for big data, as one reviewer implied. ”We are messengers of big data, not its evangelists,” said Cukier. Added Mayer-Schönberger: “The reviewer did not read the book.”

I did. Big Data is an excellent introduction for general audiences to what has become a topic of conversation everywhere, faster than any other technology-driven buzzword in recent memory. To those who may react to “big data” as today’s incarnation of “big brother,” Mayer-Schönberger and Cukier offer a comprehensive and highly readable overview of the benefits and risks associated with big data, which they define as “the ability of society to harness information in novel ways to produce useful insights or goods and services of significant value.”    Continue reading

Posted in Big Data | Leave a comment

Big Data Analytics and Data Science at Netflix (Video)

Chris Pouliot, the Director of Analytics and Algorithms at Netflix: “…my team does not only personalizations for movies, but we also deal with content demand prediction. Helping our buyer down in Beverly Hills figure out how much do we pay for a piece of content. The personalization recommendations for helping users find good movies and TV shows. Marketing analytics, how do we optimize our marketing spin. Streaming platform, how do we optimize the user experience once I press play. There’s a wide range of data, so theres a lot of diversity. We have a lot of scale, a lot of challenging problems. The question then is, how do we attract great data scientists that can just see this as a playground, a sandbox of really exciting things. Challenging problems, challenging data, great tools, and then just the ability to have fun and create great products.”
[youtube http://www.youtube.com/watch?v=pJd3PKm9XUk]

Posted in Big Data, Data Science | Leave a comment