The CIO Interview: Annabelle Bexiga, TIAA-CREF

“Innovation is everyone’s job,” Annabelle Bexiga, EVP and CIO at TIAA-CREF told me recently. “The most mundane thing,” says Bexiga, “even stacking servers in the data center, can be innovative if you can think of a different way of doing it.”

Contrary to repeated predictions heralding the end of IT innovation, IT is now synonymous with the ever-changing technological landscape of all aspects of our lives. It is also synonymous, for the most part, with business innovation, as IT transforms all business activities from operations to manufacturing to customer relations.

At TIAA-CREF, the IT organization is innovating in support of the growth and expansion of the business. Founded in 1918 to provide retirement services to university faculty, TIAA-CREF is expanding to provide a wider range of financial services and establish a growing presence in other not-for-profit sectors, including health care, research, cultural organizations, and the public sector.  It is already one of the largest pension funds in the U.S., with $520 billion of assets under management, serving 3.9 million active and retired individuals, in addition to institutional investors, retirement plan sponsors, and financial planners.   Continue reading

Posted in Digitization | Leave a comment

The OED, Big Data, and Crowdsourcing

The term “big data” was included in the most recent quarterly online update of the Oxford English Dictionary (OED). So now we have a most authoritative definition of what recently became big news: “data of a very large size, typically to the extent that its manipulation and management present significant logistical challenges.”

Beyond succinct definitions, the enchanting beauty of the OED, at least for those who love words and their history, lies in the collection of quotations illustrating the forms and uses of each word from the earliest known instance of its occurrence to more recent ones.

As someone who has been somewhat preoccupied with uncovering the historical antecedents for our present day usage of the term big data (see A Very Short History of Big Data), I was delightfully surprised to find out that the OED team has discovered that the earliest use of the term happened in 1980, seventeen years before the publication of the first paper in the ACM digital library to use (and define) “big data.” Sociologist Charles Tilly wrote in a 1980 working paper surveying “The old new social history and the new old social history” that “none of the big questions has actually yielded to the bludgeoning of the big-data people.” While the context is the increasing use of computer technology and statistical methods by historians, it is clear that Tilly used the term not to describe specifically the magnitude of the data but as a flourish of the pen following the words “big questions.” The meaning of the sentence would not change if he used only the word “data.”

While I’m quite sure that Tilly did not have in mind big data as it is defined by the OED itself, the context of his discussion is very relevant to today’s debates regarding big data and data science. In the section of the article from which the “big data” quote is taken, Tilly paraphrases the discussion in a 1979 paper by historian Lawrence Stone of the use of quantitative methods in historical research and attempts to make it a “science.”

Stone’s criticism of “cliometricians,” whose “special field is economic history,” reads like a description of the work of many “quants”—in Wall Street, academia, or government—in the forty-five years since he issued his warning: “[Their] great enterprises are necessarily the result of team-work, rather like building the pyramids: squads of diligent assistants assemble data, encode it, programme it, and pass it through the maw of the computer, all under the autocratic direction of a team-leader. The results cannot be tested by any of the traditional methods since the evidence is buried in private computer-tapes, not exposed in published footnotes. In any case the data are often expressed in so mathematically recondite a form that they are unintelligible to the majority of historical profession. The only reassurance to the bemused laity is that the members of this priestly order disagree fiercely and publicly about the validity of each other’s findings.”

Anticipating today’s doubts about the effectiveness of big data and concerns about the ratio of signal to noise, Stone concludes “in general, the sophistication of the methodology has tended to exceed the reliability of the data, while the usefulness of the results seem—up to a point—to be in inverse correlation to the mathematical complexity of the methodology and the grandiose scale of data-collection.” (For a recent enthusiastic embrace of the application of data science to the humanities and a rebuttal.

As Tilly hinted in the title to his paper, the new on many occasions is a very familiar old. Just scratch the surface and you find that the “revolution”—a word which we now tend to use liberally to describe any technological development—nicely delivers us to some place in the past while providing a soothing sense of moving forward. Indeed, the first sense of the word “revolution” in the OED is “The action or fact, on the part of celestial bodies, of moving around in an orbit or circular course” or simply “The return or recurrence of a point or period of time.”

Another word added to the OED online in the recent update affirms the notion that (almost) everything old is new again. While “crowdsourcing” was coined by Jeff Howe in 2006, this “new” (revolutionary?) practice launched the OED a century and a half ago:

In July 1857 a circular was issued by the ‘Unregistered Words Committee’ of the Philological Society of London, which had set up the Committee a few weeks earlier to organize the collection of material to supplement the best existing dictionaries. This circular, which was reprinted in various journals, asked for volunteers to undertake to read particular books and copy out quotations illustrating ‘unregistered’ words and meanings—items not recorded in other dictionaries—that could be included in the proposed supplement. Several dozen volunteers came forward, and the quotations began to pour in.

The volume of the “unregistered” material was such that in January 1858, The Philological Society decided that “efforts should be directed toward the compilation of a complete dictionary, and one of unprecedented comprehensiveness.” It took a while, but in April 1879, the newly-appointed editor James Murray issued an appeal to the public, asking for volunteers to read specific books in search of quotations to be included in the future dictionary. Within a year there were close to 800 volunteers and over the next three years, 3,500,000 quotation slips were received and processed by the OED team.

Was this the first big-data-crowdsourcing project?

Posted in Big Data, Data Science | Leave a comment

Big Data Observations: The Science of Asking Questions

“I am a firm believer that without speculation there is no good and original observation”—Charles Darwin

“It is the theory that determines what we can observe”—Albert Einstein

“I suspect, however, like as it is happening in many academic fields, the NSA is sorely tempted by all the data at its fingertips and is adjusting its methods to the data rather than to its research questions. That’s called looking for your keys under the light”—Zeynep Tufekci

“Large open-access data sets offer unprecedented opportunities for scientific discovery—the current global collapse of bee and frog populations are classic examples. However, we must resist the temptation to do science backwards by posing questions after, rather than before, data analysis. A scant understanding of the context in which data sets were collected can lead to poorly framed questions and results, and to conclusions that are plain wrong. Scientists intending to make use of large composite data sets need to work closely with those responsible for gathering the data. Standard scientific principles and practice then demand that they first frame the important questions, then design and execute the data analyses needed to answer them”—David B. Lindenmayer and Gene E. Likens

“The wonderful thing about being a data scientist is that I get all of the credibility of genuine science, with none of the irritating peer review or reproducibility worries… I thought I was publishing an entertaining view of some data I’d extracted, but it was treated like a scientific study… I’ve enjoyed publishing a lot of data-driven stories since then, but I’ve never ceased to be disturbed at how the inclusion of numbers and the mention of large data sets numbs criticism”—Pete Warden

Posted in Big Data, Data Science | Leave a comment

Visualising the Road to Becoming a Data Scientist

Source: Swami Chandrasekaran

Posted in Data Science | Leave a comment

The Digital Marketing Landscape: 2 Views

Gartner Digital Marketing Transit Map

Source: Gartner

marketing_technology_landscape_2012

Source: chiefmartec.com

Posted in Misc | Leave a comment

A Practical Introduction to Data Science Skills (Video)

Google’s Michael Manoochehri at DataEDGE 2013 presenting an introduction to  data analysis and suggestions for how to become a data scientist (his notes for the presentation are here).
[youtube=http://www.youtube.com/watch?v=rpwZ_i-9U0o&w=560&h=315]

Posted in Big Data, Data Science | Leave a comment

Big Data Quotes: Einstein, Come Back When You’ve Got Data

“Big data is what happened when the cost of storing information became less than the cost of making the decision to throw it away”—George Dyson (quoted by Tim O’Reilly)

“If the engineers have their way, every idea, memory, and feeling—the recorded consciousness of a single lifetime—will be stored in the cloud… ‘Information overload’ once referred to the difficulty of absorbing intelligently the data produced by others. Now we face the peril of choking on our own…By remembering everything, we may become haunted by our pasts and immobilized by digital distractions—or we may gain new powers to prevent the bad and promote the good”—G. Pascal Zachary

“[I]n a world where massive datasets can be analysed to identify patterns not easily identified using simpler analogue methods, what happens to genius of the Einstein variety?

Genius is about big ideas, not big data. Analysing the attributes and characteristics of anything is guaranteed to find some patterns. It is inherently a theoretical exercise, one that requires minimal thought once you’ve figured out what you want to measure. If you’re not sure, just measure everything you can get your hands on. Since the number of observations — the size of the sample — is by definition huge, the laws of statistics kick in quickly to ensure that significant relationships will be identified. And who could argue with the data?

Unfortunately, analysing data to identify patterns requires you to have the data. That means that big data is, by necessity, backward-looking; you can only analyze what has happened in the past, not what you can imagine happening in the future. In fact, there is no room for imagination, for serendipitous connections to be made, for learning new things that go beyond the data. Big data gives you the answer to whatever problem you might have (as long as you can collect enough relevant information to plug into your handy supercomputer). In that world, there is nothing to learn; the right answer is given…

What if Albert Einstein lived today and not 100 years ago? What would big data say about the general theory of relativity, about quantum theory? There was no empirical support for his ideas at the time — that’s why we call them breakthroughs.

Today, Einstein might be looked at as a curiosity, an ‘interesting’ man whose ideas were so out of the mainstream that a blogger would barely pay attention. Come back when you’ve got some data to support your point”—Sidney Finkelstein

Posted in Big Data | Leave a comment

Big Data Quotes: Disruptive Innovation?

“By definition, big data cannot yield complicated descriptions of causality. Especially in healthcare. Almost all of our diseases occur in the intersections of systems in the body. For example, there is a drug that is marketed by Elan BioNeurology called TYSABRI. It was developed for MS [multiple sclerosis]. It turns out that of the people who have MS a proportion respond magnificently to TYSABRI. And others don’t. So what do you conclude from this? Is it just a mediocre drug? No. It is that there is one disease but it manifests itself in different ways. How does big data figure out what is the core of what is going on?”–Clayton Christensen

Continue reading

Posted in Big Data | Leave a comment

Big Data, Small World: Kirk Borne at TEDxGeorgeMasonU (Video)

Kirk Borne is Professor of Astrophysics and Computational Science in the George Mason University School of Physics, Astronomy, and Computational Sciences (SPACS). Turns out he is the father of the term “unknown unknowns” – things we do not know we don’t know – popularized by former secretary of defense Donald Rumsfeld and later by Avinash Kaushik as “the unique space in which big data analysts should actually play.”

 

Posted in Big Data | Leave a comment

On Data Janitors, Engineers, and Statistics

Big Data Borat tweeted recently that “Data Science is 99% preparation, 1% misinterpretation.” Commenting on the 99% part, Cloudera’s Josh Wills says: “I’m a data janitor. That’s the sexiest job of the 21st century. It’s very flattering, but it’s also a little baffling.” Kaggle, the data-science-as-sport startup, takes care of the “1% misinterpretation” part by providing a matchmaking service between the sexiest of the sexy data janitors and the organizations requiring their hard-to-find skills. It charges $300 per hour for the service, of which $200 go to the data janitor (at least in the case of Shashi Godbole, quoted in the Technology Review article). Kaggle justifies its mark-up by delivering “the best 0.5% of the 95,988 data scientists who compete in data mining competitions,” the top of its data science table league, the ranking of data scientists based on their performance in Kaggle’s competitions, presumably representing  sound interpretation and top-notch productivity.

Kaggle’s co-founder Anthony Goldbloom tells The Atlantic’s Thomas Goetz that the ranking also represents a solution to a “market failure” in assessing the skills and relevant experience of the new breed of data scientists: “Kaggle represents a new sort of labor market, one where skills have been bifurcated from credentials.” Others see this as the creation of a new, $300 per hour, guild. In “Data Scientists Don’t Scale,” ZDnet’s Andrew Brust says that “’Data scientist’ is a title designed to be exclusive, standoffish and protective of a lucrative guild… The solution… isn’t legions of new data scientists. Instead, we need self-service tools that empower smart and tenacious business people to perform Big Data analysis themselves.”

Continue reading

Posted in Data Science, Statistics | Leave a comment