The Web at 25: The Value of Open

The Internet started as a network for linking research centers. The World Wide Web started as a way to share information among researchers at CERN. Both have expanded to touch today a third of the world’s population because they have been based on open standards.

Creating a closed and proprietary system has been the business model of choice for many great inventors and some of the greatest inventions of the computer age. That’s where we were headed towards in the early 1990s: The establishment of global proprietary networks owned by a few computer and telecommunications companies, whether old (IBM, AT&T) or new (AOL). Tim Berners-Lee’s invention and CERN’s decision to offer it to the world for free in 1993 changed the course of this proprietary march, giving a new—and much expanded—life to the Internet (itself a response to proprietary systems that did not inter-communicate) and establishing a new, open platform, for a seemingly infinite number of applications and services.

As Bob Metcalfe told me in 2009: “Tim Berners-Lee invented the URL, HTTP, and HTML standards… three adequate standards that, when used together, ignited the explosive growth of the Web… What this has demonstrated is the efficacy of the layered architecture of the Internet. The Web demonstrates how powerful that is, both by being layered on top of things that were invented 17 years before, and by giving rise to amazing new functions in the following decades.”

Metcalfe also touched on the power and potential of an open platform: “Tim Berners-Lee tells this joke, which I hasten to retell because it’s so good. He was introduced at a conference as the inventor of the World Wide Web. As often happens when someone is introduced that way, there are at least three people in the audience who want to fight about that, because they invented it or a friend of theirs invented it. Someone said, ‘You didn’t. You can’t have invented it. There’s just not enough time in the day for you to have typed in all that information.’ That poor schlemiel completely missed the point that Tim didn’t create the World Wide Web. He created the mechanism by which many, many people could create the World Wide Web.”

“All that information” was what the Web gave us (and what was also on the mind of one of the Internet’s many parents, J.C.R. Licklider, who envisioned it as a giant library). But this information comes in the form of ones and zeros, it is digital information. In 2007, 94% of storage capacity in the world was digital, a complete reversal from 1986, when 99.2% of all storage capacity was analog. The Web was the glue and the catalyst that would speed up the spread of digitization to all analog devices and channels for the creation, communications, and consumption of information.  It has been breaking down, one by one, proprietary and closed systems with the force of its ones and zeros.

Metcalfe’s comments were first published in ON magazine which I created and published for my employer at the time, EMC Corporation. For a special issue (PDF) commemorating the 20th anniversary of the invention of the Web, we asked some 20 members of the Inforati how the Web has changed their and our lives and what it will look like in the future. Here’s a sample of their answers:

Guy Kawasaki: “With the Web, I’ve become a lot more digital… I have gone from three or four meetings a day to zero meetings per day… Truly the best will be when there is a 3-D hologram of Guy giving a speech. You can pass your hand through him. That’s ultimate.”

Chris Brogan: “We look at the Web as this set of tools that allow people to try any idea without a whole lot of expense… Anyone can start anything with very little money, and then it’s just a meritocracy in terms of winning the attention wars.”

Tim O’Reilly: “This next stage of the Web is being driven by devices other than computers. Our phones have six or seven sensors. The applications that are coming will take data from our devices and the data that is being built up in these big user-contributed databases and mash them together in new kinds of services.”

John Seely Brown: “When I ran Xerox PARC, I had access to one of the world’s best intellectual infrastructures: 250 researchers, probably another 50 craftspeople, and six reference librarians all in the same building. Then one day to go cold turkey—when I did my first retirement—was a complete shock. But with the Web, in a year or two, I had managed to hone a new kind of intellectual infrastructure that in many ways matched what I already had. That’s obviously the power of the Web, the power to connect and interact at a distance.”

Jimmy Wales: “One of the things I would like to see in the future is large-scale, collaborative video projects. Imagine what the expense would be with traditional methods if you wanted to do a documentary film where you go to 90 different countries… with the Web, a large community online could easily make that happen.”

Paul Saffo: “I love that story of when Tim Berners-Lee took his proposal to his boss, who scribbled on it, ‘Sounds exciting, though a little vague.’ But Tim was allowed to do it. I’m alarmed because at this moment in time, I don’t think there are any institutions our there where people are still allowed to think so big.”

Dany Levy (founder of DailyCandy): “With the Web, everything comes so easily. I wonder about the future and the human ability to research and to seek and to find, which is really an important skill. I wonder, will human beings lose their ability to navigate?”

Howard Rheingold: “The Web allows people to do things together that they weren’t allowed to do before. But… I think we are in danger of drowning in a sea of misinformation, disinformation, spam, porn, urban legends, and hoaxes.”

Paul Graham: “[With the Web] you don’t just have to use whatever information is local. You can ship information to anyone anywhere. The key is to have the right filter. This is often what startups make.”

How many startups and grown-up companies today are entirely based on an idea first flashed out in a modest proposal 25 years ago? And there is no end in sight for the expanding membership in this club, now also increasingly including the analogs of the world. All businesses, all governments, all non-profits, all activities are being eaten by ones and zeros. Tim Berners-Lee has unleashed an open, ever-expanding system for the digitization of everything.

We also interviewed Berners-Lee in 2009. He said that the Web has “changed in the last few years faster than it changed before, and it is crazy to for us to imagine this acceleration will suddenly stop.” He pointed out the ongoing tendency to lock what we do with computers in a proprietary jail: “…there are aspects of the online world that are still fairly ‘pre-Web.’ Social networking sites, for example, are still siloed; you can’t share your information from one site with a contact on another site.” But he remained both realistic and optimistic, the hallmarks of an entrepreneur: “The Web, after all, is just a tool…. What you see on it reflects humanity—or at least the 20 percent of humanity that currently has access to the Web… No one owns the World Wide Web, no one has a copyright for it, and no one collects royalties from it. It belongs to humanity, and when it comes to humanity, I’m tremendously optimistic.”

The Pew Research Center is marking the 25th anniversary of the Web in a series of reports. Berners-Lee says in a press release issued today by the World Wide Web Consortium: “I hope this anniversary will spark a global conversation about our need to defend principles that have made the Web successful, and to unlock the Web’s untapped potential. I believe we can build a Web that truly is for everyone: one that is accessible to all, from any device, and one that empowers all of us to achieve our dignity, rights and potential as humans.”

See also Berners-Lee post on Google’s official blog: “…today is a day to celebrate. But it’s also an occasion to think, discuss—and do. Key decisions on the governance and future of the Internet are looming, and it’s vital for all of us to speak up for the web’s future. How can we ensure that the other 60 percent around the world who are not connected get online fast? How can we make sure that the web supports all languages and cultures, not just the dominant ones? How do we build consensus around open standards to link the coming Internet of Things? Will we allow others to package and restrict our online experience, or will we protect the magic of the open web and the power it gives us to say, discover, and create anything? How can we build systems of checks and balances to hold the groups that can spy on the net accountable to the public? These are some of my questions—what are yours?”

Posted in Misc | Leave a comment

The Web at 25: Tim Berners-Lee on the Web of Data

In 2009, on the occasion of the 20th anniversary of the Web, Jason Rubin and I talked to Tim Berners-Lee about his invention and its future, the Semantic Web, which he described as “the Web of data.”

Twenty years on, the World Wide Web has proven itself both ubiquitous and indispensible. Did you anticipate it would reach this status, and in this time frame?

Tim Berners-Lee: I think while it’s very tempting for us to look at the Web and say, “Well, here it is, and this is what it is,” it has, of course, been constantly growing and changing—and it will continue to do so. So to think of this as a static “This is how the Web is” sort of thing is, I think, unwise. In fact, it’s changed in the last few years faster than it changed before, and it’s crazy for us to imagine this acceleration will suddenly stop. So yes, the 20-year point goes by in a flash, but we should realize that, and we are constantly changing it, and it’s very important that we do so.

I believe that 20 years from now, people will look back at where we are today as being a time when the Web of documents was fairly well established, such that if someone wanted to find a document, there’s a pretty good chance it could be found on the Web. The Web of data, though, which we call the Semantic Web, would be seen as just starting to take off. We have the standards but still just a small community of true believers who recognize the value of putting data on the Web for people to share and mash up and use at will. And there are other aspects of the online world that are still fairly “pre-Web.” Social networking sites, for example, are still siloed; you can’t share your information from one site with a contact on another site. Hopefully, in a few years’ time, we’ll see that quite large category of social information truly Web-ized, rather than being held in individual lockdown applications.

You mentioned a “small community” of people who see the value of the Semantic Web. Is that a repeat occurrence of the struggle 20 years ago to get people to understand the scope and potential impact of the World Wide Web?

It’s remarkably similar. It’s very funny. You’d think that once people had seen the effect of Web-izing documents to produce the World Wide Web, doing likewise with their data would seem the next logical step. But for one thing, the Web was a paradigm shift. A paradigm shift is when you don’t have in your vocabulary the concepts and the ideas with which to understand the new world. Today, the idea that a web link could connect to a document that originates anywhere on the planet is completely second nature, but back then it took a very strong imagination for somebody to understand it.

Now, with data, almost all the data you come across is locked in a database. The idea that you could access and combine data anywhere in the world and immediately make it part of your spreadsheet is another paradigm shift. It’s difficult to get people to buy into it. But in the same way as before, those who do get it become tremendously fired up. Once somebody has realized what it would be like to have linked data across the world, then they become very enthusiastic, and so we now have this corps of people in many countries all working together to make it happen.

Do you see the Semantic Web as enabling greater collaboration between and among parties, as opposed to the point-to-point or point-to-many communication that seems more prevalent in the current Web?

The original web browser was a browser editor and it was supposed to be a collaborative tool, but it only ran on the NeXT workstation on which it was developed. However, the idea that the Web should be a collaborative place has always been a very important goal for me. I think harnessing the creative energy of people is really important. When you get people who are trying to solve big problems like cure AIDS, fight cancer, and understand Alzheimer’s disease, there are a huge number of people involved, all of them with half-formed ideas in their minds. How do we get them communicating so that the half of an idea in one person’s head will connect with half of an idea in somebody else’s head, and they’ll come up with the solution?

That’s been a goal for the Web of documents, and it’s certainly a goal for the Web of data, where different pieces of data can be used for all kinds of different things. For example, a genomist may suspect that a particular protein is connected to a certain syndrome in a cell line, search for and find data relating to each area, and then suddenly put together the different strains of data and discover something new. And this is something he can do with the owners of the respective pieces of data, who might never have found each other or known that their data was connected. So the Web of data will absolutely lead to greater collaboration.

Is your vision of the Semantic Web one in which data is freely available, or are there access rights attached to it?

A lot of information is already public, so one of the simple things to do in building the new Web of data is to start with that information. And recently, I’ve been working with both the U.K. government and the U.S. government in trying not only to get more information on the Web, but also to make it linked data. But it’s also very important that systems are aware of the social aspects of data. And it’s not just access control, because an authorized user can still use the right data for the wrong purpose. So we need to focus on what are the purposes for accessing different kinds of data, and for that we’ve been looking at accountable systems.

Accountable systems are aware of the appropriate use of data, and they allow you to make sure that certain kinds of information that you are comfortable sharing with people in a social context, for example, are not able to be accessed and considered by people looking to hire you. For example, I have a GPS trail that I took on vacation. Certainly, I want to give it to my friends and my family, but I don’t necessarily wish to license people I don’t know who are curious about me and my work and let them see where I’ve been. Companies may want to do the same thing. They might say, “We’re going to give you access to certain product information because you’re part of our supply chain and you can use it to fine-tune your manufacturing schedule to meet our demand. However, we do not license you to use it to give to our competition to modify their pricing.”

You need to be able to ask the system to show you just the data that you can use for a given task, because how you wish to use it will be the difference in whether you can use it. So we need systems for recording what the appropriate use of data is, and we need systems for helping people use data in an appropriate way so they can meet an ethical standard.

Ultimately, what is one of the most significant things the Semantic Web will enable?

One thing I think we’ll be able to do is to write intelligent programs that run across the Web of data looking for patterns when something went wrong—like when a company failed, or when a product turned out to be dangerous, or when an ecological catastrophe happened. We can then identify patterns in a broad range of data types that resulted in something serious happening, and that will allow us to identify when these patterns recur, and we’ll be better able to prepare for or prevent the situation.

I think when we have a lot of data available on the Web about the world, including social data, ecological data, meteorological data, and financial data, we’ll be able to make much better models. It’s been quite evident over the last year, for example, that we have a really bad grasp of the financial system. Part of the reason for that might be that we have insufficient data from which to draw conclusions, or that the experts are too selective in which data they use. The more data we have, the more accurate our models will be.

After 20 years, what about the Web—either its current or future capabilities—excites you the most?

One of the things that gets me the most excited are the mash-ups, where there’s one market of people providing data and there’s a second layer of people mashing up the data, picking from a rich variety of data sources to create a useful new application or service. A classic example of a mash-up is when I find a seminar I want to go to, and the web page has information about the sponsor, the presenter, the topic, and the logistics. I have to write all that down on the back of an envelope and then go and put it in my address book; I have to put it in my calendar; I have to enter the address in my GPS—basically, I have to copy this information into every device I use to manage my life, which is inefficient and time-consuming. This is because there is no common format for this data to become integrated into my devices.

Now, the vision of Semantic Web is that the seminar’s web page has information pointed at data about the event. So I just tell my computer I’m going to be attending that seminar and then, automatically, there is a calendar that shows things that I’m attending. And automatically, an address book I define as having in it the people who have given seminars that I’ve attended within the last six months appears, with a link to the presenter’s public profile. And automatically, my PDA starts pointing towards somewhere I need to be at an appropriate time to get me there. All I need to do is say, “I’m going to that seminar,” and then the rest should follow.

The Web is such a mélange of useful, noble content and stuff that runs the gamut from the mundane to the grotesque. Do you think humanity is using this incredible invention of yours appropriately?

Yes. The Web, after all, is just a tool. It’s a powerful one, and it reconfigures what we can do, but it’s just a tool, a piece of white paper, if you will. So what you see on it reflects humanity—or at least the 20 percent of humanity that currently has access to the Web.

As a standards body, the W3C is not interested in policing the Web or in censoring content, nor should we be. No one owns the World Wide Web, no one has a copyright for it, and no one collects royalties from it. It belongs to humanity, and when it comes to humanity, I’m tremendously optimistic. After 20 years, I’m still very excited and extremely hopeful.

[First published in ON magazine]

Posted in Misc | Leave a comment

Jeff Kelly and Dave Vellante from Wikibon on the Big Data Market (Video)

[youtube http://www.youtube.com/watch?v=ef_yQvtC9aA?list=PLenh213llmcYiyiYRzkku1MgwvrL_TIGN]
From Silicon Angle:

Our coverage kicked into high gear after the release of Wikibon’s third annual Big Data Vendor Revenue and Market Forecast, which author Jeff Kelly stopped by to discuss with hosts John Furrier and Dave Vellante.

Companies spent $18.6 billion on analytics in 2013, according to the report, up 58 percent over the previous 12 months and about two and a half times more than in 2011. Kelly estimates that the industry will pass the $28 billion mark by the end of this year and achieve revenues of over $50 billion in 2017, a massive increase he credits to rapidly growing demand for emerging solutions such as NoSQL databases as well as more established technologies that are proving valuable in extracting insights.

Posted in Big Data | Leave a comment

Jake Flomenberg from Accel Partners on the Big Data Market (Video)

[youtube http://www.youtube.com/watch?v=SHOw-2IWHZE]

From VentureBeat:

In a new video, Jake Flomenberg of Accel Partners lays out his view of the big data market and the investing opportunities he’s excited about. He’s talking with another data expert: Stefan Groschupf, the chief executive of well-funded big data startup Datameer.

Flomenberg knows what he’s talking about: He worked on sales, marketing, and product problems at hot data startup (and likely IPO candidate) Cloudera.

He’s one person who works with Accel’s big data fund. He managed to get in on hot data-transformation startup Trifacta, as well as marketing-focused Origami Logic and log-management company Sumo Logic.

 

 

Posted in Big Data | Leave a comment

Doing Data Science at Manheim

As ones and zeros eat the world, data is the new product and data science is the new process of innovation.

The International Institute for Analytics predicts that in 2014 companies in a variety of industries will increasingly use analytics on the data they have accumulated to develop new products and services. NewVantage Partners’ most recent Big Data Survey reports that 68% of executives felt that “new product innovations” was the greatest value to their organization from big data. In releasing the Accenture Technology Vision 2014, Accenture’s CTO Paul Daugherty said that “Digital is rapidly becoming part of the fabric of [large enterprises’] operating DNA and they are poised to become the digital power brokers of tomorrow.”

The best example of this trend I’ve encountered recently came from an industry one does not necessarily associate with data crunching and analysis—the vehicle remarketing industry, better known as used cars auctions. In 2012, Manheim, a subsidiary of Cox Enterprises, handled nearly 8 million used vehicles, facilitating transactions representing more than $50 billion in value.  With annual revenues of more than $2.5 billion, Manheim offers its services in 14 countries, from physical and online auction channels to financing, transportation, and mobile solutions. Manheim’s research and consulting arm, Manheim Consulting, provides market intelligence and publishes the monthly Used Vehicle Value Index and the annual Used Car Market Report (see here for the 2014 version).

Manheim has provided for free this type of analysis, seeing it as part of the value it offers to auto dealers who are members of its network.  But now it has moved into using its deep knowledge of the used car market and its analytics expertise to offer a new, fee-based service.  Shifting the analytics team from supporting the business to generating revenues, “we’ve decided to look at how we can help dealers in managing the risk associated with their inventory,” T. Glenn Bailey told me.

Bailey is Senior Director of Enterprise Product Planning at Manheim, and his responsibilities include market segmentation, forecasting, and optimization.  He and his team started testing last year a new service called DealShield. The idea came from the financial markets, specifically put option contracts. Just like a put option protects the buyer from a decline in the price of a stock below a specified price, so does DealShield offer a guarantee that Manheim will buy a car back from the dealer, within a certain time frame, for what they paid for it plus the fee they paid.  “It is as if they never bought the car,” says Bailey.

Manheim’s market knowledge and analytics skills give it confidence in its estimates of the value of a car and what they would be able to offer for it if it comes back to them. “We see a lot of value in it,” says Bailey, “because one of the things dealers like to have is liquidity. They use wholesale financing to buy used cars and typically repay the loan within seven to fourteen days. The inventory that’s sitting out there is money that is tied up. DealShield allows them to get out of that car and get their money back in a certain period of time.”

To do their analysis, the Manheim team uses tools that have served this purpose for years, demonstrating that for certain types of analysis and data you can do data science without using any of the new big data technologies. The data is collected and stored in an IBM DB2 database and the analysis is done using a variety of SAS analytics tools.  “The need to combine data from different sources is why we moved into a SAS cloud,” says Bailey. “I wanted our analyst team to be focused on the analytics and not worry about the administrative side.”

Speaking of the analyst team, Bailey says that “we are in the same market for analytics and data science talent with everybody.” In the competition for these hard-to-find professionals, Bailey looks for creativity, communications skills and willingness to learn the business. “In my experience,” he says, “it is fairly easy to tell if you have the technical chops.” He spends most of the time when he interviews people trying to determine if they are creative and can come up with new ideas on how to apply analytics tools to the data to find new insights. “Reversing the flow of cause and effect,” Bailey calls it. “Maybe optimization can tell us where to send a vehicle to maximize value.”

In addition to looking for “people that can bring technology to the business,” Bailey also looks for people who are comfortable with “getting with the business itself.” He calls it “putting on the polo shirt,” spending time with the dealers and getting engaged with them to understand their business first-hand.  This practical bent does not stop with the hiring of the right people but continues with establishing the right work environment and a “fail-fast” culture. “In some sense,” says Bailey, ”failure is rewarded because it means you are testing this thing out.” When they developed DealShield, “we had a chart that over a 2-month period showed all the things that failed. If it doesn’t work, kill it.”

In addition to being the first knowledge-based service that is expected to bring in a new revenue stream, DealShield breaks new ground for Manheim because it is the first time the company actually owns cars (when they come back from the dealer), not just acting as a middle-man. That became an opportunity for an analyst on Bailey’s team to hone further her knowledge of the business.  “She is now responsible for selling the cars. She is setting the auction, the floor price, where to run the auction,” says Bailey.

Doing data science means engaging with the business, inventing new data-based products, even becoming an integral part of revenue stream for the business.

[Originally published on Forbes.com]

Posted in Data Science | Leave a comment

Satya Nadella, New Microsoft CEO, on the Digitization of Everything (Video)

[youtube http://www.youtube.com/watch?v=LQ8Hiss2EkE?list=PL6RKqpezpCYp5-XtFje-AeT4XHOs0X4ZK]
Satya Nadella on the Microsoft Blog:

On Tuesday at LeWeb’13 in Paris, I joined Om Malik on stage to talk with thousands of entrepreneurs, startups and large companies about technology and where we’re headed as an industry. The theme of this year’s conference is The Next 10 Years, and we spent our time talking about the new ideas we see from startups around the world and how their work is shaping the future.

Today software intermediates — and digitizes — many of the things we do in business, life and our world. New technologies help businesses engage with customers in more meaningful ways, connect us to our friends and families, and allow us to see, interact with and share our world in ways never before possible. But we’re only at the beginning.

Over the next 10 years we will reach a point where nearly everything will become digitized. This will be made possible by an ever-growing network of connected devices, incredible computing capacity from the cloud, insights from big data and intelligence from machine learning. Developers, with access to these technologies and simplified development frameworks, will create new applications and services that help us transform what we do at work, and life, into digital equivalents.

Historically, there’s been a lot of attention focused on the digitization happening in Silicon Valley, but what’s even more interesting to me is what’s happening around the world. In every corner of the globe, new innovations are bringing this digitization of everything one step closer, and that’s incredibly important as this transformation should be — it must be a global phenomenon for it to reflect the needs of our distributed and diverse world. And below are some great examples of the work that is already underway:

In Israel, AKOL is taking a digital approach to modern food agriculture. Through its new platform, AKOLogic, it will enable local officials to monitor fruit, vegetable, dairy and poultry production in real time. This will increase public safety by allowing for faster identification of the source for spoiled food, significantly increase awareness lead time of potential food supply issues, and allow third-world farmers to cost effectively comply with first-world standards and regulations, thereby helping them gain access to new markets.

In China, Beijing Rendering Company used the cloud to create a new digital world for its action movie, “Young Detective Dee: Rise of the Sea Dragon.” With no real-world outdoor or underwater filming, the company saved an estimated 90 percent in production costs and brought in more than $100 million at the box office.

In Paris, Kompass recently moved its worldwide contact database to the cloud. As a result it was able to launch in 69 countries simultaneously, stay focused on user experience and service development, and keep resources flat. Now it’s gaining new insights and value due to an increased ability to segment, analyze, update, maintain and visualize its databases.

Temenos and 91JinRong are changing the face of finance in Africa and China by moving in-person financial options online to improve service, gain access to more customers and enter new markets. Sparsha Learning Technologies is creating customized interactive online solutions for K–12 and higher education that help educators reduce cost, connect with more students and extend their programs globally. askem is turning real-world Q&As into digital pictures. And startups like SkyGiraffe, qunb, SmartNotify, Lokad and Buddy are building new services to help businesses build enterprise mobile apps, create better data visualizations, connect people with the right message at the right time, improve commerce through big data, and provide easy-to-use back-end services for developers.

Ushered in by innovations from startups and investment from the enterprise in this new era, every business will be a digital business, everyone can be a developer and nearly everything analog will be digitized. We’re investing in this new era through programs like Microsoft BizSpark and Microsoft Ventures, and we’re committed to helping our customers get there with our enterprise cloud products and services. Whether you’re just getting started or further along your journey, Microsoft will be here to help.

Posted in Digitization | Leave a comment

Design Thinking for Dummies (Data Scientists)

[slideshare id=30767715&style=border: 1px solid #CCC; border-width: 1px 1px 0; margin-bottom: 5px; max-width: 100%;&sc=no]

Data scientists often face ambiguous challenges and, as a group, should use and make use of the design process to address these challenges. These slides briefly make the case for using the design process.
Posted in Data Science | Leave a comment

Why Ones and Zeros Are Eating the World

30 years ago today, Steve Jobs unveiled the Macintosh. More accurately, The Great Magician took it out of a bag and let it talk to us. The Macintosh, as I learned from first-hand experience in 1984, was a huge leap forward compared to the PCs of the time. But I couldn’t have written and published the previous words and shared a digitized version of Jobs’ performance so easily, to a potential audience of 2.5 billion people, without two other inventions, the Internet and the Web.

45 years ago this year (October 29, 1969), the first ARPANET (later to be known as the Internet) link was established between UCLA and SRI. 25 years ago this year (March 1989), Tim Berners-Lee circulated a proposal for “Mesh” (later to be known as the World Wide Web) to his management at CERN.

The Internet started as a network for linking research centers. The World Wide Web started as a way to share information among researchers at CERN. Both have expanded to touch today a third of the world’s population because they have been based on open standards. The Macintosh, while a breakthrough in human-computer interaction, was conceived as a closed system and did not break from the path established by its predecessors: It was a desktop/personal mainframe. One ideology was replaced by another, with very little (and very controlled) room for outside innovation. (To paraphrase Search Engine Land’s Danny Sullivan, the big brother minions in Apple’s “1984” Super Bowl ad remind one of the people in Apple stores today).

This is not a criticism of Jobs, nor is it a complete dismissal of closed systems. It may well be that the only way for his (and his team’s) design genius to succeed was by keeping complete ownership of their proprietary innovations. But the truly breakthrough products they gave us—the iPod (and iTunes), and especially the iPhone (and “smartphones”)—were highly dependent on the availability and popularity of an open platform for sharing information, based on the Internet and the Web.

Creating a closed and proprietary system has been the business model of choice for many great inventors and some of the greatest inventions of the computer age. That’s where we were headed towards in the early 1990s: The establishment of global proprietary networks owned by a few computer and telecommunications companies, whether old (IBM, AT&T) or new (AOL). Tim Berners-Lee’s invention and CERN’s decision to offer it to the world for free in 1993 changed the course of this proprietary march, giving a new—and much expanded—life to the Internet (itself a response to proprietary systems that did not inter-communicate) and establishing a new, open platform, for a seemingly infinite number of applications and services.

As Bob Metcalfe told me in 2009: “Tim Berners-Lee invented the URL, HTTP, and HTML standards… three adequate standards that, when used together, ignited the explosive growth of the Web… What this has demonstrated is the efficacy of the layered architecture of the Internet. The Web demonstrates how powerful that is, both by being layered on top of things that were invented 17 years before, and by giving rise to amazing new functions in the following decades.”

Metcalfe also touched on the power and potential of an open platform: “Tim Berners-Lee tells this joke, which I hasten to retell because it’s so good. He was introduced at a conference as the inventor of the World Wide Web. As often happens when someone is introduced that way, there are at least three people in the audience who want to fight about that, because they invented it or a friend of theirs invented it. Someone said, ‘You didn’t. You can’t have invented it. There’s just not enough time in the day for you to have typed in all that information.’ That poor schlemiel completely missed the point that Tim didn’t create the World Wide Web. He created the mechanism by which many, many people could create the World Wide Web.”

“All that information” was what the Web gave us (and what was also on the mind of one of the Internet’s many parents, J.C.R. Licklider, who envisioned it as a giant library). But this information comes in the form of ones and zeros, it is digital information. In 2007, when Jobs introduced the iPhone, 94% of storage capacity in the world was digital, a complete reversal from 1986, when 99.2% of all storage capacity was analog. The Web was the glue and the catalyst that would speed up the spread of digitization to all analog devices and channels for the creation, communications, and consumption of information.  It has been breaking down, one by one, proprietary and closed systems with the force of its ones and zeros.

Metcalfe’s comments were first published in ON magazine which I created and published for my employer at the time, EMC Corporation. For a special issue (PDF) commemorating the 20th anniversary of the invention of the Web, we asked some 20 members of the Inforati how the Web has changed their and our lives and what it will look like in the future. Here’s a sample of their answers:

Guy Kawasaki: “With the Web, I’ve become a lot more digital… I have gone from three or four meetings a day to zero meetings per day… Truly the best will be when there is a 3-D hologram of Guy giving a speech. You can pass your hand through him. That’s ultimate.”

Chris Brogan: “We look at the Web as this set of tools that allow people to try any idea without a whole lot of expense… Anyone can start anything with very little money, and then it’s just a meritocracy in terms of winning the attention wars.”

Tim O’Reilly: “This next stage of the Web is being driven by devices other than computers. Our phones have six or seven sensors. The applications that are coming will take data from our devices and the data that is being built up in these big user-contributed databases and mash them together in new kinds of services.”

John Seely Brown: “When I ran Xerox PARC, I had access to one of the world’s best intellectual infrastructures: 250 researchers, probably another 50 craftspeople, and six reference librarians all in the same building. Then one day to go cold turkey—when I did my first retirement—was a complete shock. But with the Web, in a year or two, I had managed to hone a new kind of intellectual infrastructure that in many ways matched what I already had. That’s obviously the power of the Web, the power to connect and interact at a distance.”

Jimmy Wales: “One of the things I would like to see in the future is large-scale, collaborative video projects. Imagine what the expense would be with traditional methods if you wanted to do a documentary film where you go to 90 different countries… with the Web, a large community online could easily make that happen.”

Paul Saffo: “I love that story of when Tim Berners-Lee took his proposal to his boss, who scribbled on it, ‘Sounds exciting, though a little vague.’ But Tim was allowed to do it. I’m alarmed because at this moment in time, I don’t think there are any institutions our there where people are still allowed to think so big.”

Dany Levy (founder of DailyCandy): “With the Web, everything comes so easily. I wonder about the future and the human ability to research and to seek and to find, which is really an important skill. I wonder, will human beings lose their ability to navigate?”

Howard Rheingold: “The Web allows people to do things together that they weren’t allowed to do before. But… I think we are in danger of drowning in a sea of misinformation, disinformation, spam, porn, urban legends, and hoaxes.”

Paul Graham: “[With the Web] you don’t just have to use whatever information is local. You can ship information to anyone anywhere. The key is to have the right filter. This is often what startups make.”

How many startups have flourished on the basis of the truly great products Apple has brought to the world? And how many startups and grown-up companies today are entirely based on an idea first flashed out in a modest proposal 25 years ago? And there is no end in sight for the expanding membership in the latter camp, now also increasingly including the analogs of the world. All businesses, all governments, all non-profits, all activities are being eaten by ones and zeros. Tim Berners-Lee has unleashed an open, ever-expanding system for the digitization of everything.

We also interviewed Berners-Lee in 2009. He said that the Web has “changed in the last few years faster than it changed before, and it is crazy to for us to imagine this acceleration will suddenly stop.” He pointed out the ongoing tendency to lock what we do with computers in a proprietary jail: “…there are aspects of the online world that are still fairly ‘pre-Web.’ Social networking sites, for example, are still siloed; you can’t share your information from one site with a contact on another site.” But he remained both realistic and optimistic, the hallmarks of an entrepreneur: “The Web, after all, is just a tool…. What you see on it reflects humanity—or at least the 20 percent of humanity that currently has access to the Web… No one owns the World Wide Web, no one has a copyright for it, and no one collects royalties from it. It belongs to humanity, and when it comes to humanity, I’m tremendously optimistic.”

[Originally published on Forbes.com]

Posted in Digitization | Leave a comment

Tim O’Reilly on Open Data

[brightcove vid=3058495035001&exp3=1971702156001&surl=http://c.brightcove.com/services&pubid=1971571337001&pk=AQ~~,AAABywrPJyk~,MP34hwWOTrPs3yLiJKkINM_zsiFWIvnW&lbu=http://www.mckinsey.com/videos/video?vid=3058495035001%26plyrid=2399849255001%26Height=270%26Width=480&w=300&h=225]

Source: McKinsey

“We should define a little bit what we mean by “open,” because there’s open as in it’s open source. Anybody can take it and reuse it in whatever way they want. And I’m not sure that’s always necessary. There’s a pragmatic open and there’s an ideological open. And the pragmatic open is that it’s available. It’s available in a timely way, in a nonpreferential way, so that some people don’t get better access than others.

And if you look at so many of our apps now on the web, because they are ad-supported and free, we get a lot of the benefits of open. When the cost is low enough, it does in fact create many of the same conditions as a commons. That being said, that requires great restraint, as I said earlier, on the part of companies, because it becomes easy for them to say, “Well, actually we just need to take a little bit more of the value for ourselves. And oh, we just need a bit more of that.” And before long, it really isn’t open at all.”

Posted in Misc | Leave a comment

2013 Data Science Salary Survey: Open source tools correlate with higher salary

“In our report, 2013 Data Science Salary Survey, we make our own data-driven contribution to the conversation. We collected a survey from attendees of the Strata Conference in New York and Santa Clara, California, about tool usage and salary…

What did we find?

In a sentence: those who use data tools make more.

More specifically, the tools that correlate with higher salary are scalable and generally open source; they are often script-based or built for machine learning.  Those attendees who tend to use one such tool tend to use others––that is, these tools form a ‘cluster’ in terms of usage among our sample.  Perhaps just as interesting is that some of the traditional, popular tools such as Excel and SAS were not used as widely as R and Python. This might be food for thought for those data analysts who have thus far resisted learning how to code or moving beyond query-based data tools.”

Source: 2013 Data Science Salary Survey 

Posted in Big Data, Data Science | Leave a comment