Mobile Drives Big Data: Ericsson Mobility Report

MOBILE SUBSCRIPTIONS OUTLOOK
MOBILE TRAFFIC
MOBILE APPLICATION TRAFFIC OUTLOOK

The complete report is here and here

Posted in Big Data | Leave a comment

Sources and Types of Big Data (Infographic)

INTELLIGENCE BY VARIETY
Posted in Big Data | Leave a comment

Big Data: Who, Why, and How (Infographic)

The who, why and how of BIG DATA

“Early adopters of Big Data analytics have gained a significant lead over the rest of the corporate world. Examining more than 400 large companies, we found that those with the most advanced analytics capabilities are outperforming competitors by wide margins.”

Source: Bain & Company

Posted in Big Data, Statistics | Leave a comment

DataKind’s Jack Porway on Data Science

[youtube=http://www.youtube.com/watch?v=Mm1RplOU0cQ&w=560&h=315]

“If you leave an excited data scientist on his own to solve a problem, he’s going to solve his own problem – which is usually parking his car, or finding a bar to drink at. The trick that we worked on was actually less about data and more about translation, about finding a way for data scientists to speak the language of the people who were trying to solve the big problems… the biggest [challenge] is actually the framing of the problem: really finding the question. As any good data scientist will tell you, it’s not so much about the data, it’s the question you start with”–Jack Porway, DataKind

More here

Posted in Data Science | Leave a comment

What Will Make You a Big Data Leader?

What Will Make You a Big Data Leader?

The IBM Institute for Business Value’s 2013 analytics survey surveyed 900 business and IT executives from 70 countries. “Leaders” (19% of the sample) were respondents self identified as “substantially outperforming their market or industry peers” in a question used by the IBM Institute for Business Value for years across a wide variety of surveys.

The full report is here

Posted in Big Data, Statistics | Leave a comment

The Digitization of IT

In many companies today, the “consumerization of IT” is turning into the “Digitization of IT.” The spreading of consumer technologies and services into the workplace is being expanded into a larger set of IT practices, borrowed from Silicon Valley innovators and adapted to the needs of enterprises in a variety of industries.

The old IT was analog IT: A single-purpose function designed to automate specific business activities, provide support and governance, and “keep the trains running on time.” The new IT is digital: Multi-purpose, extremely flexible, weaved into every aspect of the business, and gushing with unexplored and previously unknown opportunities.

The digitization of IT means that the IT organization is both stable and innovative, fault tolerant and fast learning, reliable and experimental. It solves the paradox of “safe is risky, stable is dangerous.” It promotes a culture of constant change which ensures resilience, and experimentation which safeguards continuity. Yes, you can have the best of both worlds.

Continue reading

Posted in Misc | Leave a comment

SAS CTO on Big Data and Big Compute

“One of my biggest challenges,” Keith Collins told me recently, “is helping SAS understand how to communicate to IT organizations. We present workloads which look odd and different. IT does not know how to have an SLA (Service Level Agreement) around them.  We take all of the compute and I/O capacity that they can give us.”

SAS, the largest independent vendor in the business intelligence market, used to be a prime example of “shadow IT,” the purchasing of information technology tools by business users without the knowledge and approval of the central IT organization. But this is changing in the era of big data. The collection and analysis of data are becoming a very large part of many business activities and the IT organization is asked to provide support, even leadership, in tying together these disparate efforts.

Collins is SVP and CTO at SAS, where he has spent almost 30 years, helping the company grow with the market through a number of phases (and buzzwords)—statistical analysis, decision-support, data mining, knowledge and risk management, business intelligence, and business analytics.  Now SAS is helping its customers, including CIOs and their IT teams, address the challenges of big data. Collins has seen this movie before: “People are all hyped up about Hadoop.  But what is it, really? It is big and wide record sizes, big block sizes, designed specifically for high-volume, sequential processing. Just like a SAS data set in 1968… The only difference between a SAS data set and Hadoop is that now the disks are cheap enough that you can do replication.”  The following is an edited transcript of our conversation.

Gil Press:  Indeed, many people talk about Hadoop as a replacement for tape.

Keith Collins:  We love that people get that as a pattern now, because it really helps them understand SAS.  So it is a really good time for us to have the conversation with IT about it. But they are still struggling.  They see it as “what is my next big data repository?”  They do not see it as “this is my next big way to answer questions.”

Continue reading

Posted in Big Data | Leave a comment

5 Minutes on the Myth of the Data Science Platypus (Video)

[youtube=http://www.youtube.com/watch?v=9f-XXR9j6m8&w=420&h=315]

“Data science is in danger of being a fad. Data scientists need to build a reputation for providing actual value”–Kim Stedman

Posted in Data Science | Leave a comment

The CIO Interview: Annabelle Bexiga, TIAA-CREF

“Innovation is everyone’s job,” Annabelle Bexiga, EVP and CIO at TIAA-CREF told me recently. “The most mundane thing,” says Bexiga, “even stacking servers in the data center, can be innovative if you can think of a different way of doing it.”

Contrary to repeated predictions heralding the end of IT innovation, IT is now synonymous with the ever-changing technological landscape of all aspects of our lives. It is also synonymous, for the most part, with business innovation, as IT transforms all business activities from operations to manufacturing to customer relations.

At TIAA-CREF, the IT organization is innovating in support of the growth and expansion of the business. Founded in 1918 to provide retirement services to university faculty, TIAA-CREF is expanding to provide a wider range of financial services and establish a growing presence in other not-for-profit sectors, including health care, research, cultural organizations, and the public sector.  It is already one of the largest pension funds in the U.S., with $520 billion of assets under management, serving 3.9 million active and retired individuals, in addition to institutional investors, retirement plan sponsors, and financial planners.   Continue reading

Posted in Digitization | Leave a comment

The OED, Big Data, and Crowdsourcing

The term “big data” was included in the most recent quarterly online update of the Oxford English Dictionary (OED). So now we have a most authoritative definition of what recently became big news: “data of a very large size, typically to the extent that its manipulation and management present significant logistical challenges.”

Beyond succinct definitions, the enchanting beauty of the OED, at least for those who love words and their history, lies in the collection of quotations illustrating the forms and uses of each word from the earliest known instance of its occurrence to more recent ones.

As someone who has been somewhat preoccupied with uncovering the historical antecedents for our present day usage of the term big data (see A Very Short History of Big Data), I was delightfully surprised to find out that the OED team has discovered that the earliest use of the term happened in 1980, seventeen years before the publication of the first paper in the ACM digital library to use (and define) “big data.” Sociologist Charles Tilly wrote in a 1980 working paper surveying “The old new social history and the new old social history” that “none of the big questions has actually yielded to the bludgeoning of the big-data people.” While the context is the increasing use of computer technology and statistical methods by historians, it is clear that Tilly used the term not to describe specifically the magnitude of the data but as a flourish of the pen following the words “big questions.” The meaning of the sentence would not change if he used only the word “data.”

While I’m quite sure that Tilly did not have in mind big data as it is defined by the OED itself, the context of his discussion is very relevant to today’s debates regarding big data and data science. In the section of the article from which the “big data” quote is taken, Tilly paraphrases the discussion in a 1979 paper by historian Lawrence Stone of the use of quantitative methods in historical research and attempts to make it a “science.”

Stone’s criticism of “cliometricians,” whose “special field is economic history,” reads like a description of the work of many “quants”—in Wall Street, academia, or government—in the forty-five years since he issued his warning: “[Their] great enterprises are necessarily the result of team-work, rather like building the pyramids: squads of diligent assistants assemble data, encode it, programme it, and pass it through the maw of the computer, all under the autocratic direction of a team-leader. The results cannot be tested by any of the traditional methods since the evidence is buried in private computer-tapes, not exposed in published footnotes. In any case the data are often expressed in so mathematically recondite a form that they are unintelligible to the majority of historical profession. The only reassurance to the bemused laity is that the members of this priestly order disagree fiercely and publicly about the validity of each other’s findings.”

Anticipating today’s doubts about the effectiveness of big data and concerns about the ratio of signal to noise, Stone concludes “in general, the sophistication of the methodology has tended to exceed the reliability of the data, while the usefulness of the results seem—up to a point—to be in inverse correlation to the mathematical complexity of the methodology and the grandiose scale of data-collection.” (For a recent enthusiastic embrace of the application of data science to the humanities and a rebuttal.

As Tilly hinted in the title to his paper, the new on many occasions is a very familiar old. Just scratch the surface and you find that the “revolution”—a word which we now tend to use liberally to describe any technological development—nicely delivers us to some place in the past while providing a soothing sense of moving forward. Indeed, the first sense of the word “revolution” in the OED is “The action or fact, on the part of celestial bodies, of moving around in an orbit or circular course” or simply “The return or recurrence of a point or period of time.”

Another word added to the OED online in the recent update affirms the notion that (almost) everything old is new again. While “crowdsourcing” was coined by Jeff Howe in 2006, this “new” (revolutionary?) practice launched the OED a century and a half ago:

In July 1857 a circular was issued by the ‘Unregistered Words Committee’ of the Philological Society of London, which had set up the Committee a few weeks earlier to organize the collection of material to supplement the best existing dictionaries. This circular, which was reprinted in various journals, asked for volunteers to undertake to read particular books and copy out quotations illustrating ‘unregistered’ words and meanings—items not recorded in other dictionaries—that could be included in the proposed supplement. Several dozen volunteers came forward, and the quotations began to pour in.

The volume of the “unregistered” material was such that in January 1858, The Philological Society decided that “efforts should be directed toward the compilation of a complete dictionary, and one of unprecedented comprehensiveness.” It took a while, but in April 1879, the newly-appointed editor James Murray issued an appeal to the public, asking for volunteers to read specific books in search of quotations to be included in the future dictionary. Within a year there were close to 800 volunteers and over the next three years, 3,500,000 quotation slips were received and processed by the OED team.

Was this the first big-data-crowdsourcing project?

Posted in Big Data, Data Science | Leave a comment