{"id":1868,"date":"2015-01-06T14:58:10","date_gmt":"2015-01-06T20:58:10","guid":{"rendered":"http:\/\/redmonk.com\/dberkholz\/?p=1868"},"modified":"2015-01-07T09:07:39","modified_gmt":"2015-01-07T15:07:39","slug":"time-for-sysadmins-to-learn-data-science","status":"publish","type":"post","link":"https:\/\/redmonk.com\/dberkholz\/2015\/01\/06\/time-for-sysadmins-to-learn-data-science\/","title":{"rendered":"Time for sysadmins to learn data science"},"content":{"rendered":"<p>At PuppetConf 2012, I had an epiphany when watching a talk by Google&#8217;s Jamie Wilkinson where he was live-hacking monitoring data in R. I can&#8217;t recommend his talk highly enough \u2014 as an analytics guy, this blew my mind:<\/p>\n<p><iframe loading=\"lazy\" src=\"\/\/www.youtube.com\/embed\/eq4CnIzw-pE\" width=\"560\" height=\"315\" frameborder=\"0\" allowfullscreen=\"allowfullscreen\"><\/iframe><\/p>\n<p>Since then, one thing has become clear to me:\u00a0<strong>As we scale applications and start thinking of servers as <a href=\"http:\/\/www.slideshare.net\/randybias\/pets-vs-cattle-the-elastic-cloud-story\">cattle rather than pets<\/a>, coping with the vast amounts of data they generate will require increasingly advanced approaches.\u00a0<\/strong>That means over time, monitoring will require the integration of statistics and machine learning in a way that&#8217;s incredibly rare today, on both the tools and people\u00a0sides of the equation.<\/p>\n<p>It&#8217;s\u00a0clear that the analysis paralysis induced by the wall of dashboards doesn&#8217;t work. We&#8217;ve moved to an approach defined largely by alerting on-demand with tools like Nagios, Sensu, and PagerDuty. Most of the data is never viewed unless there&#8217;s a problem, in which case you investigate much more deeply than you ever see in any overview or dashboard.<\/p>\n<p>However, <strong>most alerting remains broken.<\/strong> It&#8217;s based on dumb thresholds rather than anything even the slightest bit smarter. You&#8217;re lucky if you can get something as advanced as alerting based on percentiles, let alone standard deviations or their robust alternatives (black magic!). With log analysis, it&#8217;s considered great if you can even manage basic pattern-matching to group together repetitive entries. Granted, this is a big step forward from manual analysis, but we&#8217;re still a long way from the moon.<\/p>\n<p>This needs to change. As scale and complexity increase with companies moving to the cloud, to microservice architectures, and to transient containers, <strong>monitoring needs to\u00a0go back to school for its Ph.D. to cope with this new generation of IT.<\/strong><\/p>\n<p>Exceptions are few and far between, often as add-ons that many users haven&#8217;t realized exist \u2014 for example Prelert (first for Splunk, now available as a standalone API engine too), or Bischeck for Nagios. Etsy open-sourced\u00a0the <a href=\"https:\/\/codeascraft.com\/2013\/06\/11\/introducing-kale\/\">Kale<\/a> stack, which does some of this, but it wasn&#8217;t widely adopted. More recently Numenta announced Grok, its own foray into anomaly detection, which looks\u00a0quite impressive. And today, Twitter <a href=\"https:\/\/blog.twitter.com\/2015\/introducing-practical-and-robust-anomaly-detection-in-a-time-series\">announced<\/a> another R-based tool in its anomaly-detection suite. Many of you may be surprised to hear that, completely on the other end of the tech spectrum, IBM&#8217;s monitoring tools can do some of this too.<\/p>\n<p>On the system-state\u00a0side, we&#8217;re seeing more entrants helping deal\u00a0with related problems like configuration drift including Metafor, ScriptRock, and Opsmatic. They take a variety of approaches at present. But it&#8217;s clear that in the long term, a great deal of intelligence will be required behind the scenes because it&#8217;s incredibly difficult to effectively visualize web-scale systems.<\/p>\n<p>The tooling of the future applies techniques like adaptive thresholds that vary by day, time, and more; predictive analytics; and anomaly detection to do things like:<\/p>\n<ul>\n<li>Avoid false-positive alerts that wake you up at 3am for no reason;<\/li>\n<li>Prevent eye strain from staring at hundreds of graphs looking for a blip;<\/li>\n<li>Pinpoint problems before they would hit a static threshold, like an instance\u00a0gradually running out of RAM; and<\/li>\n<li>Group together alerts from a variety of applications and systems into a single logical error.<\/li>\n<\/ul>\n<p>DevOps or not, I&#8217;m running into more people and bleeding-edge vendors who are bringing a &#8220;data science&#8221; approach to IT. This is epitomized by attendees to Jason Dixon&#8217;s <a href=\"http:\/\/monitorama.com\/\">Monitorama<\/a> conference. Before long, it will be unavoidable in modern infrastructure.<\/p>\n<p>Want to get started? You could do a lot worse than Coursera&#8217;s <a href=\"https:\/\/www.coursera.org\/specialization\/jhudatascience\/1\">data-science specialization<\/a>.<\/p>\n<p><span style=\"color: #999999;\"><em><strong>Disclosure<\/strong>: Prelert, Splunk, IBM, and ScriptRock are clients. Puppet Labs has been. Etsy, Metafor, Nagios Inc, Numenta, Opsmatic, Twitter, and PagerDuty are not.<\/em><\/span><\/p>\n<div class=\"acc_license\"><a href=\"http:\/\/creativecommons.org\/licenses\/by-sa\/3.0\/\"><img decoding=\"async\" src=\"http:\/\/i.creativecommons.org\/l\/by-sa\/3.0\/88x31.png\" alt=\"by-sa\" \/><\/a><\/div><!--<rdf:RDF xmlns=\"http:\/\/creativecommons.org\/ns#\" xmlns:dc=\"http:\/\/purl.org\/dc\/elements\/1.1\/\" xmlns:rdf=\"http:\/\/www.w3.org\/1999\/02\/22-rdf-syntax-ns#\"><Work rdf:about=\"\"><license rdf:resource=\"http:\/\/creativecommons.org\/licenses\/by-sa\/3.0\/\" \/><\/Work><License rdf:about=\"http:\/\/creativecommons.org\/licenses\/by-sa\/3.0\/\"><requires rdf:resource=\"http:\/\/creativecommons.org\/ns#Attribution\" \/><permits rdf:resource=\"http:\/\/creativecommons.org\/ns#Reproduction\" \/><permits rdf:resource=\"http:\/\/creativecommons.org\/ns#Distribution\" \/><permits rdf:resource=\"http:\/\/creativecommons.org\/ns#DerivativeWorks\" \/><requires rdf:resource=\"http:\/\/creativecommons.org\/ns#ShareAlike\" \/><requires rdf:resource=\"http:\/\/creativecommons.org\/ns#Notice\" \/><\/License><\/rdf:RDF>-->","protected":false},"excerpt":{"rendered":"<p>At PuppetConf 2012, I had an epiphany when watching a talk by Google&#8217;s Jamie Wilkinson where he was live-hacking monitoring data in R. I can&#8217;t recommend his talk highly enough \u2014 as an analytics guy, this blew my mind: Since then, one thing has become clear to me:\u00a0As we scale applications and start thinking of<\/p>\n","protected":false},"author":6,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"spay_email":"","footnotes":"","jetpack_publicize_message":"","jetpack_is_tweetstorm":false},"categories":[6,7,8,42,9],"tags":[],"class_list":["post-1868","post","type-post","status-publish","format-standard","hentry","category-cloud","category-data-science","category-devops","category-docker","category-ibm"],"jetpack_featured_media_url":"","jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p23Tsn-u8","_links":{"self":[{"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/posts\/1868","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/comments?post=1868"}],"version-history":[{"count":0,"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/posts\/1868\/revisions"}],"wp:attachment":[{"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/media?parent=1868"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/categories?post=1868"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/redmonk.com\/dberkholz\/wp-json\/wp\/v2\/tags?post=1868"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}