What’s the Big Data? 12 Definitions

Last week I got an email from UC Berkeley’s Master of Information and Data Science program, asking me to respond to a survey of data science thought leaders, asking the question “What is big data”? I was especially delighted to be regarded as a “thought leader” by Berkeley’s School of Information, whose previous dean, Hal Varian (now chief economist at Google, answered my challenge fourteen years ago and produced the first study to estimate the amount of new information created in the world annually, a study I consider to be a major milestone in the evolution of our understanding of big data.

The Berkeley researchers estimated that the world had produced about 1.5 billion gigabytes of information in 1999 and in a 2003 replication of the study found out that amount to have doubled in 3 years. Data was already getting bigger and bigger and around that time, in 2001, industry analyst Doug Laney described the “3Vs”—volume, variety, and velocity—as the key “data management challenges” for enterprises, the same “3Vs” that have been used in the last four years by just about anyone attempting to define or describe big data.

The first documented use of the term “big data” appeared in a 1997 paper by scientists at NASA, describing the problem they had with visualization (i.e. computer graphics) which “provides an interesting challenge for computer systems: data sets are generally quite large, taxing the capacities of main memory, local disk, and even remote disk. We call this the problem of big data. When data sets do not fit in main memory (in core), or when they do not fit even on local disk, the most common solution is to acquire more resources.”

In 2008, a number of prominent American computer scientists popularized the term, predicting that “big-data computing” will “transform the activities of companies, scientific researchers, medical practitioners, and our nation’s defense and intelligence operations.” The term “big-data computing,” however, is never defined in the paper.

The traditional database of authoritative definitions is, of course, the Oxford English Dictionary (OED). Here’s how the OED defines big data: (definition #1) “data of a very large size, typically to the extent that its manipulation and management present significant logistical challenges.”

But this is 2014 and maybe the first place to look for definitions should be Wikipedia. Indeed, it looks like the OED followed its lead. Wikipedia defines big data (and it did it before the OED) as (#2) “an all-encompassing term for any collection of data sets so large and complex that it becomes difficult to process using on-hand data management tools or traditional data processing applications.”

While a variation of this definition is what is used by most commentators on big data, its similarity to the 1997 definition by the NASA researchers reveals its weakness. “Large” and “traditional” are relative and ambiguous (and potentially self-serving for IT vendors selling either “more resources” of the “traditional” variety or new, non-“traditional” technologies).

The widely-quoted 2011 big data study by McKinsey highlighted that definitional challenge. Defining big data as (#3) “datasets whose size is beyond the ability of typical database software tools to capture, store, manage, and analyze,” the McKinsey researchers acknowledged that “this definition is intentionally subjective and incorporates a moving definition of how big a dataset needs to be in order to be considered big data.” As a result, all the quantitative insights of the study, including the updating of the UC Berkeley numbers by estimating how much new data is stored by enterprises and consumers annually, relate to digital data, rather than just big data, e.g., no attempt was made to estimate how much of the data (or “datasets”) enterprises store is big data.

Another prominent source on big data is Viktor Mayer-Schönberger and Kenneth Cukier’s book on the subject. Noting that “there is no rigorous definition of big data,” they offer one that points to what can be done with the data and why its size matters:

(#4) “The ability of society to harness information in novel ways to produce useful insights or goods and services of significant value” and “…things one can do at a large scale that cannot be done at a smaller one, to extract new insights or create new forms of value.”

In Big Data@Work, Tom Davenport concludes that because of “the problems with the definition” of big data, “I (and other experts I have consulted) predict a relatively short life span for this unfortunate term.” Still, Davenport offers this definition:

(#5) “The broad range of new and massive data types that have appeared over the last decade or so.”

Let me offer a few other possible definitions:

(#6) The new tools helping us find relevant data and analyze its implications.

(#7) The convergence of enterprise and consumer IT.

(#8) The shift (for enterprises) from processing internal data to mining external data.

(#9) The shift (for individuals) from consuming data to creating data.

(#10) The merger of Madame Olympe Maxime and Lieutenant Commander Data.

#(11) The belief that the more data you have the more insights and answers will rise automatically from the pool of ones and zeros.

#(12) A new attitude by businesses, non-profits, government agencies, and individuals that combining data from multiple sources could lead to better decisions.

I like the last two. #11 is a warning against blindly collecting more data for the sake of collecting more data (see NSA). #12 is an acknowledgment that storing data in “data silos” has been the key obstacle to getting the data to work for us, to improve our work and lives. It’s all about attitude, not technologies or quantities.

What’s your definition of big data?

See here for the compilation of Big data definitions from 40+ thought leaders.

[Originally published on Forbes.com]

Posted in Big Data | Leave a comment

The nature of data (Infographic)

The nature of data (Infographic)
Posted in Big Data, Infographics | Leave a comment

The Internet of Things: Why now and how big?

The Internet of Things: Why now and how big?

Now that it has been established that the Internet of Things is the most hyped “emerging technology” today, and that the term—and the associated technologies—is far from being new, the only question to be answered is Why the sudden surge in interest in 2014?

That’s the question I put to a number of tech luminaries earlier this year. Bob Metcalfe, inventor of the Ethernet and now Professor of Innovation at University of Texas at Austin, is familiar with the sudden prominence of technologies, coming after lengthy incubation periods. Metcalfe points to scribbles like me as the main culprit: “It’s a media phenomenon. Technologies and standards and products and markets emerge slowly, but then suddenly, chaotically, the media latches on and BOOM!—It’s the year of IoT.” Hal Varian, Chief Economist at Google, believes Moore’s Law has something to do with the newfound interest in the IoT: “The price of sensors, processors, and networking has come way down.  Since WiFi is now widely deployed, it is relatively easy to add new networked devices to the home and office.”

Janus Bryzek, known as “the father of sensors” (and a VP at Fairchild Semiconductor), thinks there are multiple factors “accelerating the surge” in interest. First, there is the new version of the Internet Protocol, IPv6, “enabling almost unlimited number of devices connected to networks.” Another factor is that four major network providers—Cisco, IBM, GE and Amazon—have decided “to support IoT with network modification, adding Fog layer and planning to add Swarm layer, facilitating dramatic simplification and cost reduction for network connectivity.” Last but not least, Bryzek mentions new forecasts regarding the IoT opportunity, with GE estimating that the “Industrial Internet” has the potential to add $10 to $15 trillion (with a “T”) to global GDP over the next 20 years, and Cisco  increasing to $19 trillion its forecast for the economic value created by the “Internet of Everything” in the year 2020.  “This is the largest growth in the history of humans,” says Bryzek.

These mind-blowing estimates from companies developing and selling IoT-related products and services, no doubt have helped fuel the media frenzy. But what do the professional prognosticators say? Gartner estimates that IoT product and service suppliers will generate incremental revenue exceeding $300 billion in 2020. IDC forecasts that the worldwide market for IoT solutions will grow from $1.9 trillion in 2013 to $7.1 trillion in 2020.

Other research firms focus on slices of this potentially trillion-dollar market such as connected cars, smart homes, and wearables. Here’s a roundup of estimates and forecasts for various segments of the IoT market:

ABI Research:  The installed base of active wireless connected devices will exceed 16 billion in 2014, about 20% more than in 2013. The number of devices will more than double from the current level, with 40.9 billion forecasted for 2020. 75% of the growth between today and the end of the decade will come from non-hub devices: sensor nodes and accessories. The chart above is from ABI’s research on smart cars.

Acquity Group (Accenture Interactive): More than two thirds of consumers plan to buy connected technology for their homes by 2019, and nearly half say the same for wearable technology. Smart thermostats are expected to have 43% adoption in the next five years (see chart below).

IoT_Accenture_Adaptation Graph

IHS Automotive: The number of cars connected to the Internet worldwide will grow more than sixfold to 152 million in 2020 from 23 million in 2013.

Navigant Research: The worldwide installed base of smart meters will grow from 313 million in 2013 to nearly 1.1 billion in 2022.

Morgan Stanley: Driverless cars will generate $1.3 trillion in annual savings in the United States, with over $5.6 trillions of savings worldwide.

Machina Research: Consumer Electronics M2M connections will top 7 billion in 2023, generating $700 billion in annual revenue.

On World: By 2020, there will be over 100 million Internet connected wireless light bulbs and lamps worldwide up from 2.4 million in 2013.

Juniper Research: The wearables market will exceed $1.5 billion in 2014, double its value in 2013–

Endeavour Partners: As of September 2013, one in ten U.S. consumers over the age of 18 owns a modern activity tracker. More than half of U.S. consumers who have owned a modern activity tracker no longer use it. A third of U.S. consumers who have owned one stopped using the device within six months of receiving it.

Originally published on Forbes.com

Posted in Misc | Leave a comment

Josh Wills on Machine Learning in a Business Setting

[youtube https://www.youtube.com/watch?v=IgfRdDjLxe0?rel=0]
Academic machine learning is all about optimization. Machine learning in a business setting is all about understanding: “My focus is always on how do I understand what the system is doing, come up with new hypotheses about this very complex system, test them, and then use what I’ve learned from those tests to find new ways to improve the system.”

An overview of Cloudera’s current data science tools, including Oryx and Spark for building and serving machine learning models, Gertrude for multivariate testing, and Impala for ludicrously high-performance SQL queries against HDFS.

Josh Wills is Cloudera’s Senior Director of Data Science

Posted in Data Science, Machine Learning | Leave a comment

What Happens on the Web in 60 Seconds (Infographic)

What Happens on the Web in 60 Seconds (Infographic)

Source: Qmee

Posted in Infographics, Statistics | Leave a comment

What is the Internet of Things? (Infographic)

What is the Internet of Things? (Infographic)
Posted in Infographics, Internet of Things | Leave a comment

The Landscape of the Internet of Things

ENCHANTED OBJECTS

Source: Entrepreneur and Media Lab researcher David Rose talks ‘enchanted objects’

The book on Amazon: Enchanted Objects: Design, Human Desire, and the Internet of Things

Posted in Internet of Things | Leave a comment

A Very Short History Of The Internet Of Things

There have been visions of smart, communicating objects even before the global computer network was launched forty-five years ago. As the Internet has grown to link all signs of intelligence (i.e., software) around the world, a number of other terms associated with the idea and practice of connecting everything to everything have made their appearance, including machine-to-machine (M2M), Radio Frequency Identification (RFID), context-aware computing, wearables, ubiquitous computing, and the Web of Things. Here are a few milestones in the evolution of the mashing of the physical with the digital.

1932                                    Jay B. Nash writes in Spectatoritis: “Within our grasp is the leisure of the Greek citizen, made possible by our mechanical slaves, which far outnumber his twelve to fifteen per free man… As we step into a room, at the touch of a button a dozen light our way. Another slave sits twenty-four hours a day at our thermostat, regulating the heat of our home. Another sits night and day at our automatic refrigerator. They start our car; run our motors; shine our shoes; and cult our hair. They practically eliminate time and space by their very fleetness.”

January 13, 1946              The 2-Way Wrist Radio, worn as a wristwatch by Dick Tracy and members of the police force, makes its first appearance and becomes one of the comic strip’s most recognizable icons.

1949                                    The bar code is conceived when 27 year-old Norman Joseph Woodland draws four lines in the sand on a Miami beach. Woodland, who later became an IBM engineer, received (with Bernard Silver) the first patent for a linear bar code in 1952. More than twenty years later, another IBMer, George Laurer, was one of those primarily responsible for refining the idea for use by supermarkets.

1955                                    Edward O. Thorp conceives of the first wearable computer, a cigarette pack-sized analog device, used for the sole purpose of predicting roulette wheels. Developed further with the help of Claude Shannon, it was tested in Las Vegas in the summer of 1961, but its existence was revealed only in 1966.

October 4, 1960               Morton Heilig receives a patent for the first-ever head-mounted display.

1967                                    Hubert Upton invents an analog wearable computer with eyeglass-mounted display to aid in lip reading.

October 29, 1969             The first message is sent over the ARPANET, the predecessor of the Internet.

January 23, 1973              Mario Cardullo receives the first patent for a passive, read-write RFID tag.

June 26, 1974                    A Universal Product Code (UPC) label is used to ring up purchases at a supermarket for the first time.

1977                                    CC Collins develops an aid to the blind, a five-pound wearable with a head-mounted camera that converted images into a tactile grid on a vest.

Early 1980s                        Members of the Carnegie-Mellon Computer Science department install micro-switches in the Coke vending machine and connect them to the PDP-10 departmental computer so they could see on their computer terminals how many bottles were present in the machine and whether they were cold or not.

1981                                    While still in high school, Steve Mann develops a backpack-mounted “wearable personal computer-imaging system and lighting kit.”

1990                                    Olivetti develops an active badge system, using infrared signals to communicate a person’s location.

September 1991              Xerox PARC’s Mark Weiser publishes “The Computer in the 21st Century” in Scientific American, using the terms “ubiquitous computing” and “embodied virtuality” to describe his vision of how “specialized elements of hardware and software, connected by wires, radio waves and infrared, will be so ubiquitous that no one will notice their presence.”

1993                                    MIT’s Thad Starner starts using a specially-rigged computer and heads-up display as a wearable.

1993                                    Columbia University’s Steven Feiner, Blair MacIntyre, and Dorée Seligmann develop KARMA–Knowledge-based Augmented Reality for Maintenance Assistance. KARMA overlaid wireframe schematics and maintenance instructions on top of whatever was being repaired.

1994                                    Xerox EuroPARC’s Mik Lamming and Mike Flynn demonstrate the Forget-Me-Not, a wearable device that communicates via wireless transmitters and records interactions with people and devices, storing the information in a database.

1994                                    Steve Mann develops a wearable wireless webcam, considered the first example of lifelogging.

September 1994              The term ‘context-aware’ is first used by B.N. Schilit and M.M. Theimer in “Disseminating active map information to mobile hosts,” Network, Vol. 8, Issue 5.

1995                                    Siemens sets up a dedicated department inside its mobile phones business unit to develop and launch a GSM data module called “M1” for machine-to-machine (M2M) industrial applications, enabling machines to communicate over wireless networks. The first M1 module was used for point of sale (POS) terminals, in vehicle telematics, remote monitoring and tracking and tracing applications.

December 1995                MIT’s Nicholas Negroponte and Neil Gershenfeld write in “Wearable Computing” in Wired: “For hardware and software to comfortably follow you around, they must merge into softwear… The difference in time between loony ideas and shipped products is shrinking so fast that it’s now, oh, about a week.”

October 13-14, 1997       Carnegie-Mellon, MIT, and Georgia Tech co-host the first IEEE International Symposium on Wearable Computers, in Cambridge, MA.

1999                                    The Auto-ID (for Automatic Identification) Center is established at MIT. Sanjay Sarma, David Brock and Kevin Ashton turned RFID into a networking technology by linking objects to the Internet through the RFID tag.

1999                                    Neil Gershenfeld writes in When Things Start to Think: “Beyond seeking to make computers ubiquitous, we should try to make them unobtrusive…. For all the coverage of the growth of the Internet and the World Wide Web, a far bigger change is coming as the number of things using the Net dwarf the number of people. The real promise of connecting computers is to free people, by embedding the means to solve problems in the things around us.”

January 1, 2001                David Brock, co-director of MIT’s Auto-ID Center, writes in a white paper titled “The Electronic Product Code (EPC): A Naming Scheme for Physical Objects”: “For over twenty-?ve years, the Universal Product Code (UPC or ‘bar code’) has helped streamline retail checkout and inventory processes… To take advantage of [the Internet’s] infrastructure, we propose a new object identi?cation scheme, the Electronic Product Code (EPC), which uniquely identi?es objects and facilitates tracking throughout the product life cycle.”

March 18, 2002                Chana Schoenberger and Bruce Upbin publish “The Internet of Things” in Forbes. They quote Kevin Ashton of MIT’s Auto-ID Center: “We need an internet for things, a standardized way for computers to understand the real world.”

April 2002                          Jim Waldo writes in “Virtual Organizations, Pervasive Computing, and an Infrastructure for Networking at the Edge,” in the Journal of Information Systems Frontiers: “…the Internet is becoming the communication fabric for devices to talk to services, which in turn talk to other services. Humans are quickly becoming a minority on the Internet, and the majority stakeholders are computational entities that are interacting with other computational entities without human intervention.”

June 2002                          Glover Ferguson, chief scientist for Accenture, writes in “Have Your Objects Call My Objects” in the Harvard Business Review: “It’s no exaggeration to say that a tiny tag may one day transform your own business. And that day may not be very far off.”

January 2003                    Bernard Traversat et al. publish “Project JXTA-C: Enabling a Web of Things” in HICSS ’03 Proceedings of the 36th Annual Hawaii International Conference on System Sciences. They write: “The open-source Project JXTA was initiated a year ago to specify a standard set of protocols for ad hoc, pervasive, peer-to-peer computing as a foundation of the upcoming Web of Things.”

October 2003                    Sean Dodson writes in the Guardian: ”Last month, a controversial network to connect many of the millions of tags that are already in the world (and the billions more on their way) was launched at the McCormick Place conference centre on the banks of Lake Michigan. Roughly 1,000 delegates from across the worlds of retail, technology and academia gathered for the launch of the electronic product code (EPC) network. Their aim was to replace the global barcode with a universal system that can provide a unique number for every object in the world. Some have already started calling this network ‘the internet of things’.”

August 2004                      Science-fiction writer Bruce Sterling introduces the concept of “Spime” at SIGGRAPH, describing it as “a neologism for an imaginary object that is still speculative. A Spime also has a kind of person who makes it and uses it, and that kind of person is somebody called a ‘Wrangler.’ … The most important thing to know about Spimes is that they are precisely located in space and time. They have histories. They are recorded, tracked, inventoried, and always associated with a story…  In the future, an object’s life begins on a graphics screen. It is born digital. Its design specs accompany it throughout its life. It is inseparable from that original digital blueprint, which rules the material world. This object is going to tell you – if you ask – everything that an expert would tell you about it. Because it WANTS you to become an expert.”

September 2004              G. Lawton writes in “Machine-to-machine technology gears up for growth” in Computer: “There are many more machines—defined as things with mechanical, electrical, or electronic properties­—in the world than people. And a growing number of machines are networked… M2M is based on the idea that a machine has more value when it is networked and that the network becomes more valuable as more machines are connected.”

October 2004                    Neil Gershenfeld, Raffi Krikorian and Danny Cohen write in “The Internet of Things” in Scientific American: “Giving everyday objects the ability to connect to a data network would have a range of benefits: making it easier for homeowners to configure their lights and switches, reducing the cost and complexity of building construction, assisting with home health care. Many alternative standards currently compete to do just that—a situation reminiscent of the early days of the Internet, when computers and networks came in multiple incompatible types.”

October 25, 2004             Robert Weisman writes in the Boston Globe: “The ultimate vision, hatched in university laboratories at MIT and Berkeley in the 1990s, is an ‘Internet of things’ linking tens of thousands of sensor mesh networks. They’ll monitor the cargo in shipping containers, the air ducts in hotels, the fish in refrigerated trucks, and the lighting and heating in homes and industrial plants. But the nascent sensor industry faces a number of obstacles, including the need for a networking standard that can encompass its diverse applications, competition from other wireless standards, security jitters over the transmitting of corporate data, and some of the same privacy concerns that have dogged other emerging technologies.”

2005                                    A team of faculty members at the Interaction Design Institute Ivrea (IDII) in Ivrea, Italy, develops Arduino, a cheap and easy-to-use single-board microcontroller, for their students to use in developing interactive projects. Adrian McEwen and Hakim Cassamally in Designing the Internet of Things: “Combined with an extension of the wiring software environment, it made a huge impact on the world of physical computing.”

November 2005               The International Telecommunications Union publishes the 7th in its series of reports on the Internet, titled “The Internet of Things.”

June 22, 2009                    Kevin Ashton writes in “That ‘Internet of Things’ Thing” in RFID Journal: “I could be wrong, but I’m fairly sure the phrase ‘Internet of Things’ started life as the title of a presentation I made at Procter & Gamble (P&G) in 1999. Linking the new idea of RFID in P&G’s supply chain to the then-red-hot topic of the Internet was more than just a good way to get executive attention. It summed up an important insight—one that 10 years later, after the Internet of Things has become the title of everything from an article in Scientific American to the name of a European Union conference, is still often misunderstood.”

Thanks to Sanjay Sarma and Neil Gershenfeld for their comments on a draft of this timeline.

[Originally posted on Forbes.com]

Posted in Internet of Things | Leave a comment

Neil Gershenfeld on Turning Data into Things and Things into Data (Video)

[youtube https://www.youtube.com/watch?v=L0RDrSKenGo]

Neil Gershenfeld, Director of MIT’s Center for Bits and Atoms, at the 2014 Solid Conference: Analog telephone calls degraded with distance; digitizing communications allowed errors to be detected and corrected, leading to the Internet. Analog computations degraded with time; digitizing computing again allowed errors to be detected and corrected, leading to microprocessors and PCs. Manufacturing today remains analog; although the designs are digital, the processes are not. Neil Gershenfeld presents emerging research on digitizing fabrication by coding the construction of functional materials, and explores its implications for programming the physical world.

Gershenfeld wrote in his 1999 book, When Things Start to Think: “Beyond seeking to make computers ubiquitous, we should try to make them unobtrusive…. For all the coverage of the growth of the Internet and the World Wide Web, a far bigger change is coming as the number of things using the Net dwarf the number of people. The real promise of connecting computers is to free people, by embedding the means to solve problems in the things around us.”

Recently, Gershenfeld published (with JP Vasseur) “As Objects Go Online” in Foreign Affairs:

“Although the Internet of Things is now technologically possible, its adoption is limited by a new version of an old conflict. During the 1980s, the Internet competed with a network called BITNET, a centralized system that linked mainframe computers. Buying a mainframe was expensive, and so BITNET’s growth was limited; connecting personal computers to the Internet made more sense. The Internet won out, and by the early 1990s, BITNET had fallen out of use. Today, a similar battle is emerging between the Internet of Things and what could be called the Bitnet of Things. The key distinction is where information resides: in a smart device with its own IP address or in a dumb device wired to a proprietary controller with an Internet connection. Confusingly, the latter setup is itself frequently characterized as part of the Internet of Things. As with the Internet and BITNET, the difference between the two models is far from semantic. Extending IP to the ends of a network enables innovation at its edges; linking devices to the Internet indirectly erects barriers to their use…

The size and speed of the Internet have grown by nine orders of magnitude since the time it was invented. This expansion vastly exceeds what its developers anticipated, but that the Internet could get so far is a testament to their insight and vision. The uses the Internet has been put to that have driven this growth are even more surprising; they were not part of any original plan. But they are the result of an open architecture that left room for the unexpected. Likewise, today’s vision for the Internet of Things is sure to be eclipsed by the reality of how it is actually used. But the history of the Internet provides principles to guide this development in ways that are scalable, robust, secure, and encouraging of innovation.

The Internet’s defining attribute is its interoperability; information can cross geographic and technological boundaries. With the Internet of Things, it can now leap out of the desktop and data center and merge with the rest of the world. As the technology becomes more finely integrated into daily life, it will become, paradoxically, less visible. The future of the Internet is to literally disappear into the woodwork.”

 

Posted in Internet of Things | Leave a comment

Here Comes the Next Bubble: #IoT (Video)

[youtube https://www.youtube.com/watch?v=zG2dvxSKEGU]

Bubblino is a Twitter-monitoring, bubble-blowing Arduino-bot.

He watches twitter for a chosen keyword and every time he finds a new mention then he blows bubbles.

Posted in Internet of Things | Leave a comment