Google: Machine Learning and Deep Neural Networks Explained (Video)

[youtube https://www.youtube.com/watch?v=bHvf7Tagt18?rel=0]

*Greg and Chris did an AMA on Friday, September 25th to answer people’s deep learning questions. Check out their answers here: https://goo.gl/jpbMy9

*To read more about machine learning, neural nets, and the like – check out the Google Research blog:http://googleresearch.blogspot.com/ and Chris’s blog: http://colah.github.io/

Posted in Artificial Intelligence, Machine Learning | Tagged , | Leave a comment

Evolution of Computer Storage 1956-2015

Evolution-of-a-Terabyte-of-Data-7DayShop-800px

Posted in Misc | Tagged | Leave a comment

68% of Americans have smartphones, 45% have tablet computers, other devices not growing

Pew_Device Ownership

Today, 68% of U.S. adults have a smartphone, up from 35% in 2011, and tablet computer ownership has edged up to 45% among adults, according to newly released survey data from the Pew Research Center. Smartphone ownership is nearing the saturation point with some groups: 86% of those ages 18-29 have a smartphone, as do 83% of those ages 30-49 and 87% of those living in households earning $75,000 and up annually.

Posted in Misc | Tagged , , | Leave a comment

The Economist’s Data Editor on Data Fetishism

Ken Cukier

Ken Cukier

“We fetishize data, we think that data is the answer. It’s far from the truth. In fact, it’s ridiculous, because the data is only a simulacrum of reality in the same way that a map is not a territory. And so while we need to use information and data to make decisions as we need to do, the data is always unfaithful, always unreliable, it always misleads, and you have to torture it until it confesses”–Kenneth Cukier, Data Editor, The Economist

Source: Economist Radio, “Arthur Miller and Modern-Day Witch-Hunts”

Posted in Big Data | Tagged | Leave a comment

FinTech Startups:The Landscape of Blockchain Companies in Financial Services

Blockchain Startups

Source: Startup Management 

HT: Leaders in Pharmaceutical Business Intelligence

The Economist:

Bitcoin fanatics are enthralled by the libertarian ideal of a pure, digital currency beyond the reach of any central bank. The real innovation is not the digital coins themselves, but the trust machine that mints them—and which promises much more besides.

Whatever you think of the cryptocurrency, the “blockchain” is a trust machine that may yet take its place alongside double-entry book-keeping and the limited-liability company as a way of oiling the wheels of commerce.

Posted in Misc | Tagged , , | Leave a comment

Driverless Cars: A Misguided 20th Century Idea

Our Robots

IEEE Spectrum:

A vision of fully autonomous, self-driving cars allowing human owners to nap or read in the car seems to come from the future. But David Mindell, a historian and electrical engineer at MIT, says that the idea of such fully autonomous vehicles roaming the streets represents a more rigid vision left over from the last century. Mindell casts some doubt over the current course along which Google and other huge tech companies are racing to build self-driving cars that don’t require any human supervision.

In his new book, released this month, titled, “Our Robots, Ourselves: Robotics and the Myths of Autonomy” (Viking/Penguin), Mindell envisions a future in which humans are kept in the loop for (mostly) self-driving cars and other robotic technologies, rather than taking them completely out of the equation…

Spectrum: What do you think of the current focus of Google and other tech companies pursuing self-driving cars?

Mindell: Overall, robotics is still focused on full autonomy as the ultimate goal. Researchers should be working on a “perfect five” with trusted, transparent, flexible collaboration between people and autonomous systems. (The “perfect five” refers to the middle of a scale for automation that ranges from very low at level 1, to fully autonomous at 10; the concept is based on the work of Tom Sheridan, professor of mechanical engineering at MIT.)

Such systems should have the ability to turn on autonomy when it can be helpful. Autonomy can reduce human workload and fatigue, but humans should still be present in the system. That’s an empirical argument based on everything we’ve seen in the last 40 years of autonomous systems. People are always thinking that full autonomy is just around the corner. But there are 30 to 40 examples in the book, and in every one, autonomy gets tempered by human judgment and experience.

Spectrum: You’ve said that the best way forward involves a mix of humans, remotely-controlled systems and autonomous robots. Do you think the future you’re hoping for is the one we’re likely to see?

Mindell: I’m hoping the likely future is the one I’m arguing for. There is a quote in the book from the chief of BMW saying “People buy our cars because people like driving them; we’d be crazy to cut them out of the loop.” I think the world is ready for a more nuanced approach to robotics.

Posted in Robotics | Tagged | Leave a comment

Connected Cars: A History of Security Vulnerabilities

Driverless_Cars_Vulnerabilities

Chris Poulin, IBM, on Tech Crunch:

A Short History Of Car Vulnerability Research

In 2010, researchers from the University of Washington and University of California, San Diego published a seminal paper proving that once an attacker has physical access to a vehicle, they can compromise every component, from the entertainment system to the electronic control units (ECUs) that operate the engine, brakes and even the steering wheel in modern cars that self-park and sport lane-departure correction.

This research showed that an attacker could use connection points between vehicle systems as an entry point to inject arbitrary commands on the controller area network (CAN) bus to perform activities such as disabling all the engine’s cylinders, locking up one brake pad and disabling all brakes — even when the car was traveling at 40 miles per hour. The researchers even created a CAN bus analysis and packet injection tool, dubbed CarShark.

But the automakers weren’t phased by the research; their view was that an attacker would have to be jacked into your car in order to execute an attack.

In response, these same researchers undertook another study in 2011 to further prove their point, this time centered on how to remotely gain access to the vehicle. The paper enumerated the attack surfaces, including channels that provide remote access: Bluetooth, in-vehicle Wi-Fi, telematics, remote keyless entry and RFID immobilizers, dedicated short-range communications (DSRC) used to communicate between vehicles and the road infrastructure, global positioning (GPS), satellite radio and even tire pressure monitor sensors.

The researchers took the play from the punt to the end zone by remotely compromising a vehicle, then using the techniques they created in their first paper to gain complete control of the car. They even claimed they could compromise the telematics unit by simply playing an audio file over the mobile carrier’s network.

Using another vector, the researchers wrote a mobile phone Trojan that gave them remote access to a driver’s or passenger’s mobile phone, and when paired with a vehicle’s telematics unit, exploited a vulnerability in the Bluetooth firmware. They effectively used the mobile phone as a springboard to pwn the vehicle.

The researchers also compromised a typical diagnostics computer used by many service shops so that when it was connected to the diagnostics port on a vehicle, the computer would infect the vehicle with malware allowing the attackers to control it. In a zombie apocalypse scenario, the researchers even wrote software that could turn cars into a rolling “bot” army that reports back to a command and control (C&C) channel through which a criminal could issue commands.

It would seem that these researchers had proven conclusively that connected vehicle security required retooling, and that the consequences could have a major impact on customer confidence and safety. However, without details on the specific vehicles involved in the research, nor publicly disclosed proof of concept instructions, the automotive industry made little public noise about the research.

In fairness, the auto industry may have rallied war rooms and devised plans to amp up security in their automotive products; however, the automotive industry is tight-knit and guards new designs and technology closely. Further, modern automobiles are complex marvels of engineering, and the process of retooling the mechanics and software has to be undertaken slowly, carefully and over a period of many years. Bear in mind that from inception, a new automobile typically takes 5-7 years before it hits the mass market.

And yet, to the general public — and especially to researchers — the silence implied apathy on the part of the automakers. Some in the industry may not fully recognize the broader implications of these results. For example, I spoke to the design manager on the topic of the tire pressure monitoring system (TPMS) vulnerability and he responded with: “So what? All you could do is light up an amber LED on the dashboard.”

Which would be true if all TPMS receivers only had a wire loop that went to the LED in question; however, it’s likely that most of the automakers connect the TPMS receiver to other parts of the in-vehicle network, if for no other reason than to send that data as telemetry back to the predictive maintenance analytics running in the cloud. But let’s not get hung up on the TPMS system: The vehicle threat surface is as broad as the African savanna is to a big game poacher.

Enter Charlie Miller and Chris Valasek, whose 2013 Today Show vehicle hack elicited a collective gasp from the public. Automakers pointed out that such a hack would be unfeasible in real life, as the dashboard is dismantled and there’s a guy sitting in your back seat with a laptop. As is the way with such stories, other shiny objects and celebrity reality television soon overwrote that chunk of the public’s short-term memory, and drivers slid behind the wheel with nary a thought of cars gone wild.

In 2015, Miller and Valasek were back. The widely publicized video of these researchers remotely hacking into a vehicle on the road and ultimately sending it into a ditch struck a chord with the general public that research to date had yet to reach.

To put this in perspective, Recorded Future, which collects intelligence from more than 600,000 sources, including social media and underground forums, queried their data warehouse for mentions of connected vehicle security. As displayed [above], there was a fair amount of chatter when the CarShark exploit was announced, then it exploded around the two Valasek and Miller exploits. The red “bubbles” show the amount of references by date and the milestones are called out. Additionally, references to announced or publicly speculated future events are plotted at the bottom of the chart.

Posted in Robotics | Tagged , | Leave a comment

The Dell-EMC Merger and the Googlization of IT

Yes, Joe Tucci is a great salesman and Michael Dell is the ultimate entrepreneur, but it is Google that is really behind the $67 billion merger. Tucci: “The waves of change we now see in our industry are unprecedented and, to navigate this change, we must create a new company for a new era.” In other words, we must survive in the digital natives era, ushered in by Google, and magnified by the likes of Amazon and Facebook.

To understand what Tucci calls “the new world order,” let’s take a quick tour of the old one, to better understand how the digital natives forced Dell and EMC into the largest tech acquisition in history. Dell and EMC were the two most successful U.S. stocks in the 1990s, appreciating more than any other stock over that booming decade.  They rode on a new tidal wave of digital data, unleashed by the advent of the PC and the networking of PCs in 1980s.

As a result, between 1990 and 2000, the structure of the IT industry has changed for the first—and so far, the last—time, expanding to include large vendors focused on one layer of the IT stack: Intel in semi-conductors, EMC in storage, Cisco in networking, Microsoft in operating systems, Oracle in databases. IBM—the dominant player in the previous era of vertically-integrated, “one-stop-shopping” IT vendors—saved itself from the fate of DEC, Wang, and Prime (all, like EMC, based in Massachusetts) by focusing on services.

The restructured IT industry, and specifically, the focused, “best-in-class” vendors, answered a pressing business need. Digitization and the rapid growth of data unleashed new business requirements and opportunities that called for new ways to sell and buy IT.

There were new business needs for storing much larger volumes of data, mining the data for new market insights, and providing better service to customers by making increasingly “mission-critical” computer systems available 24/7. New IT buyers, such as executives in leading-edge IT departments, business executives impatient with their IT departments, or IT  executives that were asked to take over the out-of-control IT systems acquired by the business units, eschewed the vertically-integrated IT vendors in favor of the new focused competitors, embracing enthusiastically the new “mix and match” IT mentality.

The 2000s were a decade of more-of-the-same with the industry and IT buyers recuperating for a long time from Y2K and the implosion of the dot-com bubble, and going through two recessions. IBM (minus its PC business) and HP (plus Compaq, a successful, focused, PC vendor, like Dell) were the only large “one-stop-shopping” vendors to survive (Sun Microsystems did not). Dell tried, not too successfully, to expand its business beyond PCs to become a one-stop-shopping enterprise IT vendor.

But IT was not the same. Yet another wave of digital data was unleahsed by the advent of the World Wide Web (a.k.a. “the Internet). Unlike the previous wave, this one gave rise to “digital natives,” a new breed of companies with new business models based on Web domination (i.e., mastering online advertising) and data mining (i.e., indexing, recommendations, linking, etc.).  It also gave rise to a new breed of IT buyers.

In the early 200os, Google’s business presented unprecedented IT requirements for performance, availability and scalability (IT jargon for “we have lots of data to store, process, and shuttle around”). They could buy computer storage, servers and networks from existing IT vendors but the cost was prohibitive. More important, Google’s engineers, as someone who was there at the time told me, always thought they could do a better job than anyone else. So they went ahead and built their own IT infrastructure, stringing together “commodity” (off-the shelf) hardware components, and developing innovative software to manage it.

In a recently published paper, Google’s engineers described their approach to “overcoming the cost, operational complexity, and limited scale endemic to datacenter networks a decade ago.” This was the latest in a long string of influential papers that Google has published (starting, I think, in 2006), sharing with the world its experience and expertise in building an IT infrastructure for the 21st century. Moreover, it also released some of the code it has developed as open software, available for free for anyone dealing with similar IT requirements.

Other digital natives were the first to benefit from Google’s academic-like “publish or perish” mentality. They developed Google’s ideas further or came up with their own solutions, taking a page from Google’s business model—it’s a business where IT matters a lot, IT is a core competency. A prominent example is Hadoop, originally developed at Google as a solution to a storage bottleneck standing in the way of analyzing or manipulating large amounts of data, developed further by Yahoo engineers and released by them as open source software, eventually to become a foundational technology for big data analytics.

Facebook, absorbing some top Google engineering talent, went on further to invent an IT infrastructure handling not only petabytes of data every day but also providing an online service to more than 1 billion people worldwide. And it went further than Google in influencing how IT is done everywhere, by establishing the Open Compute Project, with companies such as Goldman Sachs, Bank of America, and Fidelity as members.

Amazon not only built an IT infrastructure for the 21st century, but went even further than Google and Facebook by making it available to the world for a fee, establishing the concept of IT-on-demand or cloud computing on a solid footing. In the process, it has convinced many digital natives, such as Netflix, to run their entire demanding IT infrastructure on Amazon Web Services.  Now, Amazon is ready to take over the enterprise IT market, making clear at AWS:reinvent 2015 that it is going after the legacy IT business.

This is the supply side of the equation that forced Dell and EMC into this merger. But the demand side is no less important. Just like in the early 1990s, when cheaper hardware and software allowed business executives to do their own computing, by-passing the central IT department, we see today the rise of business executives building their fame and fortunes by buying computer services directly from cloud computing providers.

But the Googlization or Amazonization of IT is not limited to business executives.  It is impossible to overstate the impact Google and other digital natives had on IT executives. The new breed of IT executives is ready to “mix and match,” to buy “best-of-breed,” to experiment with off-the-shelf hardware and open source software.

All of this explains why Dell and EMC are merging but also hints at the enormous challenges they will have in convincing IT buyers to buy into their “back-to-the-future” strategy, that a business model that stopped working in the 1990s is the answer to winning in “a new world order.” All the Google-derived talk about “software-defined-everything” and “converged infrastructure” may not be enough for IT buyers looking to take charge of what is increasingly becoming, if not a core competency, a competitive differentiator and a new source of revenues for many companies. All businesses are now digital businesses and their IT requirements are starting to resemble those Google encountered a decade ago.

IBM, HP, Oracle, and Cisco also need to articulate why “one-stop-shopping” is the way forward for IT buyers. Their task is not made easier by the industry’s influential opinion makers, such as Gartner. In its recent Symposium, Gartner told the more than 8,000 CIOs and senior IT executives in attendance to choose as partners “digital accelerators” such as Amazon and Google, not “digital inhibitors” such as Dell and EMC.

Gartner, however, put VMware, the crown jewel in the EMC “federation,” somewhat ahead of the legacy vendors. Will the company that made cloud computing a reality (there will be no cloud computing without server virtualization) save the biggest technology-industry takeover ever?

Originally published on Forbes.com

Posted in Misc | Tagged , , , , , | Leave a comment

Survey: The Hunt for Unicorn Data Scientists Boosts the Salaries of Predictive Analytics Professionals

Burtch1

Base Salaries for Individual Contributors

Burtch2

Base Salaries for Managers

Unicorn Data Scientists (upgraded from “sexy data scientists”) are hard to find and are paid more than $200,000 per year. A new survey finds that the rising data science tide lifts the compensation of all other data analytics professionals, even if they don’t know how to code.

The Burtch Works Study: Salaries for Predictive Analytics Professionals is based on interviews with 1,757 data analytics professionals conducted over the 12 months ending April 2015 by executive recruiting firm Burtch Works. It is a unique source of information in that it does not rely on self-reporting or data provided by human resources departments. It also provides insights into how the demand for data scientists impact the salaries of other data analytics professionals because it excludes data scientists, covered in a separate Burtch Works study, published earlier this year (I wrote about that study here).

Burtch Works defines predictive analytics professionals as those who can “apply sophisticated quantitative skills to data describing transactions, interactions, or other behaviors of people to derive insights and prescribe actions.” Data scientists are a subset of this group—they have the “computer science skills necessary to acquire and clean or transform unstructured or continuously streaming data, regardless of its format, size, or source.”

The additional computer science skills put data scientists on top in terms of compensation regardless of their levels of experience and managerial responsibilities but predictive analytics professionals are keeping up, seeing their salaries and bonuses rise. For example, the median base salary for the most experienced individual contributors rose from $115,250 last year to $125,000 this year and for managers managing teams of ten or more the median base salary rose from $225,000 to $235,000.

Predictive analytics professionals continue to benefit from the increasing demand and short supply for their quantitative analysis skills. The median base salary of individual contributors varies from $76,000 for those at level 1 (0 to 3 years of experience) to $125,000 for those at level 3 (9+ years of experience). The median bonus received varies from $8,100 to $18,100, depending on job level.

The median base salary of managers varies from $125,500 for those at level 1 (1 to 3 reports) to $235,000 for those at level 3 (10+ reports). The median bonus received by managers varies from $23,000 to $75,000 depending on job level.

More and more people are attracted by the demand for data analytics professionals and the potential to become a unicorn. Data recently released by the National Center for Education Statistics, according to Phys.org, shows bachelor’s degrees in statistics grew 17% from 2013 to 2014. This marks 15 consecutive years the number of undergraduates in statistics has risen, increasing by more than 300% since the 1990s. In addition, from 2000 to 2014, master’s and doctorate degrees in statistics also grew significantly at 260% and 132%, respectively.

“The Bureau of Labor Statistics projects job growth for statisticians will increase 27% between 2012 and 2022, outpacing the projected 11% rate for all other occupations. The number of graduates in statistics each year—approximately 2,000 bachelor’s degrees, 3,000 master’s degrees and 575 doctorate degrees—seems unlikely to match this demand,” says Phys.org.

Originally published on Forbes.com

Posted in Data Science, Predictive analytics | Tagged , , | Leave a comment

How to Become a Unicorn Data Scientist and Make More than $240,000

What makes a good data scientist? And if you are a good data scientist, how much should you expect to get paid?

Owen Zhang, ranked #1 on Kaggle, the online stadium for data science competitions, lists his skills on his Kaggle profile as “excessive effort,” “luck,” and “other people’s code.” An engineer by training, Zhang says in this ODSC interview that data science is finding “practical solutions to not very well-defined problems,” similar to engineering. He believes that good data scientists, “otherwise known as unicorn data scientists,” have three types of expertise. Since data science deals with practical problems, the first one is being familiar with a specific domain and knowing how to solve a problem in that domain. The second is the ability to distinguish signal from noise, or understanding statistics. The third skill is software engineering.

Not having formal education in statistics or software engineering, Zhang explains that he acquired his data science skills by competing in Kaggle and learning from its community. No doubt being very good at learning on your own is a required skill, to say nothing about hanging out with the right people, preferably unicorn data scientists. Galit Shmueli, Professor of Business Analytics at NTHU, told rjmetrics that her one piece of advice for data scientists just getting started is to “attend a conference or two, see what people are working on, what are the challenges, and what’s the atmosphere.”

Recent data shows that unicorn data scientists can make more than $240,000 annually. This according to the 2015 Data Science Salary Survey where O’Reilly Media’s John King and Roger Magoulas report the results of a survey of 600 “data practitioners” (reflecting the recency of the term, only one-quarter of the respondents have job titles that explicitly identify them as “data scientists”).

The median annual base salary of the survey sample is $91,000, and among U.S. respondents is $104,000, similar to last year’s results. 23% said that it would be “very easy” for them to find another position.

Keep in mind that “23% of the sample hold a doctorate degree,” and additional 44% hold a master’s. The word “sample” here means, as it does in almost all other surveys today, “the people that wanted to answer our survey.” But unlike other survey report authors, King and Magoulas make sure to issue this warning: “We should be careful when making conclusions about survey data from a self-selecting sample—it is a major assumption to claim it is an unbiased representation of all data scientists and engineers… the O’Reilly audience tends to use more newer, open source tools, and underrepresents non-tech industries such as insurance and energy.”

Still, we can learn quite a lot about the background and skills required for admission into this well-paid group of data masters. Two-thirds of respondents had academic backgrounds in computer science, mathematics, statistics, or physics.

Beyond the initial training, it is important to keep abreast of the ever-changing landscape of data science tools: “It seems likely that in the long run knowing the highest paying tools will increase your chances of joining the ranks of the highest paid,” say King and Magoulas. And the most recent additions to the data science tool pantheon provide the greatest boost to salaries: “…learning Spark could apparently have more of an impact on salary than getting a PhD. Scala is another bonus: those who use both are expected to earn over $15,000 more than an otherwise equivalent data professional.”

The bad news is that the more time spent in meetings (even for non-managers), the more money a data scientist makes. Another widely discussed unpleasant part of the job—data cleaning—is the #2 task on which data scientists spend the most time, with 39% of survey participants spending at least one hour per day on this task. The good news is that exploratory data analysis is what occupies them most, with 46% spending one to three hours per day on this task and 12% spending four hours or more.

More data on the skills employed by practicing data scientists comes from an AnalyticsWeek survey of 410 data professionals. In Optimizing Your Data Science Team, Bob E. Hayes reports that respondents were asked to indicate their level of proficiency for 25 different skills.” Solving problems with data,” says Hayes, “requires expertise across different skill areas: 1) Business, 2) Technology, 3) Programming, 4) Math & Modeling and 5) Statistics. Proficiency in each skill area is related to job role.”

All of these skills may not present themselves in a single data scientist but it’s possible to assemble all of them by putting together a top-notch data science team. In “Tips for building a data science capability” from consulting firm Booz Allen Hamilton, we learn that “rather than illuminate a single data science rock star, it is important to highlight a diversity of talent at all levels to help others self-identify with the capability. It is also a more realistic version of the truth. Very rarely will you find ‘magical unicorns’ that embody the full breadth of math and computer science skills along with the requisite domain knowledge. More often, you will build diverse teams that when combined provide you with the ‘triple-threat’ (computer science, math/statistics, and domain expertise) model needed for the toughest data science problems.”

The concept of a data science team, combining various skills and educational backgrounds, is high on the agenda of the 175-year-old American Statistical Association (ASA) which is probably looking in dismay at the oodles of funds going to establishing new data science programs and research centers at American universities, to say nothing about the salaries of data scientists as opposed to the salaries of statisticians.

The ASA issued a “policy statement” on October 1, reminding the world that statistics is one of the three disciplines “foundational to data science” (the other two being database management and distributed and parallel systems, providing a “computational infrastructure”). The statement concludes with “The next generation [of statisticians] must include more researchers with skills that cross the traditional boundaries of statistics, databases and distributed systems; there will be an ever-increasing demand for such ‘multi-lingual’ experts.”

In other words, if you aspire to a $200,000+ salary, better call yourself a data scientist and start coding.

Posted in Data Science | Tagged , , | Leave a comment