{"id":5373,"date":"2025-10-27T16:54:48","date_gmt":"2025-10-27T16:54:48","guid":{"rendered":"https:\/\/redmonk.com\/jgovernor\/?p=5373"},"modified":"2025-10-27T16:54:48","modified_gmt":"2025-10-27T16:54:48","slug":"why-log-data-management-is-a-thing","status":"publish","type":"post","link":"https:\/\/redmonk.com\/jgovernor\/why-log-data-management-is-a-thing\/","title":{"rendered":"Why Log Data Management Is a Thing"},"content":{"rendered":"<p><a href=\"http:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-scaled.jpeg\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-5374\" src=\"http:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-1024x696.jpeg\" alt=\"picture of 3 men on a ship with a cord, measuring ship log. Ship log or chip log is a navigation tool to estimate the speed of the ship throwing a wooden chip overboard at the end of a line marked by knots during a predetermined interval of time\" width=\"1024\" height=\"696\" srcset=\"https:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-1024x696.jpeg 1024w, https:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-300x204.jpeg 300w, https:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-768x522.jpeg 768w, https:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-1536x1043.jpeg 1536w, https:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-2048x1391.jpeg 2048w, https:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-480x326.jpeg 480w, https:\/\/redmonk.com\/jgovernor\/files\/2025\/10\/ship-log-stock-923x627.jpeg 923w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400;\">With all the buzz around Observability over the last few years it\u2019s easy to imagine that when it comes to logs, metrics and traces, it\u2019s game over. Just stick all the data you need in a database, or these days a data lakehouse, and start building queries and dashboards. Easy.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Glibness aside, Observability tools vendors generally claim they can manage all your metrics and telemetry data in a single coherent store, with consistent access mechanisms.\u00a0<\/span><span style=\"font-weight: 400;\">But as telemetry data has exploded so have costs. This is especially true when we want to correlate the data in terms of business needs- the problem with things like customer number, user or product ID, or IP address is that they are inherently high cardinality. Columns with many unique values drive up costs of memory and compute, and these costs get passed on to customers.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Where it used to be that folks complained Splunk was expensive, these days we hear the same about Datadog. Datadog, long seen as the darling of the APM space, rather than a \u201clegacy player\u201d is now seen as expensive. In 2025 this is a Datadog weakness &#8211; this issue comes up repeatedly in customer conversations. It\u2019s not that people don\u2019t value Datadog &#8211; they rave about the user experience. But costs are a concern.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">There is a huge opportunity here around cost management, notably in the emerging log data management (LDM) space. Organisations are concerned with costs of storage, and cardinality.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">So what is Log Data Management and why is it useful?\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The bottom line is that log management is indeed a data management problem. Data sources continue to fragment, with every new platform the organisation uses. Modern Observability is not about instrumentation but data, especially in the open standards world of Open Telemetry. But we\u2019re not yet living in a world where every piece of your infrastructure is using OTel. There is plenty of telemetry in different systems that needs to be integrated, collated and transformed before it\u2019s truly useful- similarly to ETL in the data warehousing space.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">So we\u2019re faced with at least two factors that need to be addressed &#8211; cost of storage, and complexity of the data landscape.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">You might justifiably claim if your tool is primarily used for troubleshooting by developers that you don\u2019t actually need to store all the unique events in your log stream, but a lot of organisations with a strong focus on security and compliance, such as those in regulated industries, do indeed want to store all the telemetry.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">With log data management you\u2019re concentrating on integrating data from a range of sources, often leaving it in place, but with pipeline routing and data refinery capabilities to allow you to manage all of your log data cost as cost effectively as possible.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The company that arguably best represents this view of the market right now is Cribl. It doesn\u2019t position itself as a replacement for Observability or even log management vendors, but rather as an adjunct to them.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Cribl Stream, formerly called LogStream, is about sending data to multiple places, rather than collecting it into one. It can enrich or redact data, and is used by customers to reduce volume for existing ingest platforms. Storing every AWS CloudTrail or Windows XML event gets very expensive quickly.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Cribl is taking a similar approach with Cribl Search. Rather than consolidating data in one place before searching, the platform will search at the edge, so that you can search across multiple contexts, with data left in place, using a unified query language. Federated search is <\/span><i><span style=\"font-weight: 400;\">really<\/span><\/i><span style=\"font-weight: 400;\"> hard, but the philosophy of \u201csearch-in-place\u201d is a great play for customers that don\u2019t want to buy another centralisation promise.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Chronosphere moved from using CrowdStrike as an OEM log provider to launching their own log control product in June 2025. Chronosphere\u2019s platform is all about letting users control the volume of data they are storing, but because they don\u2019t price on ingest users can have more visibility into what they need to keep. When Chronopshere launched their log product, CEO Martin Mao told RedMonk:<\/span><\/p>\n<blockquote><p><span style=\"font-weight: 400;\">One of the big gaps is you don&#8217;t know which and which sections of your logs to reduce and how to reduce them. And the way we solve that, and this is one of the reasons why we built our own back end, is we actually have to analyze all of how you use the logs in terms of dashboards, alerts and things like that. Then we turn the analytics into suggestions to feed a telemetry pipeline, and you can reduce your data volume.\u00a0<\/span><\/p><\/blockquote>\n<p><span style=\"font-weight: 400;\">Startup <\/span><a href=\"https:\/\/www.controltheory.com\/\"><span style=\"font-weight: 400;\">Control Theory<\/span><\/a><span style=\"font-weight: 400;\"> was created to give companies more operational and cost control over their logs. Co-founder Bob Quillen <a href=\"https:\/\/redmonk.com\/blog\/2025\/06\/30\/rmc-bob-quillin\/\">argues that<\/a> OTel helped democratize instrumentation and collection of telemetry, but it still led to \u201cfat dumb pipes that dump into a data lake, and then you pay for ingest of that data, indexing it, and retaining it. And we thought, \u2018there\u2019s got to be a better way to do this.\u2019\u201d And so he and his co-founders set out to create a control layer to sit on top of telemetry data to better manage it.<\/span><\/p>\n<p><a href=\"https:\/\/hydrolix.io\/\"><span style=\"font-weight: 400;\">Hydrolix<\/span><\/a><span style=\"font-weight: 400;\">, on the other hand, is tackling the cost problem of logs by delivering an extremely high compression data lake that can stay always hot. Hydrolix\u2019s approach is maybe tangential to the other log data management approaches mentioned here. While other competitors focus on reducing the total volume of logs saved and stored, Hydrolix instead have focused on reigning in costs by building their own proprietary compression methodology. A big reason why they can compress so efficiently is that their solution focuses exclusively on logs, not any other type of telemetry.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Honeycomb, which positions itself as the best solution for querying and analysing high cardinality data at scale, has responded to the need for pipeline-based log data management approaches. It launched the Honeycomb Telemetry Pipeline product. Telemetry Pipeline Manager uses the OpenTelemetry Collector to scrape and collect system logs. The collector supports multiple log formats. Logs can also be refined, the elimination of redundant data. The Refiner also enables the identification of potentially important events &#8211; showing for example, errors or slow requests. The rest of the data is archived, but you can rehydrate full-fidelity logs and traces directly from S3.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Datadog also offers an Observability Pipelines product &#8211; customers save on egress costs by sending only valuable logs to a chosen observability vendor, and then routing other logs to long-term storage such as AWS S3, Azure Blob Storage or Google Cloud. As ever Datadog offers plenty of out of the box functionality, in this case more than 150 predefined parsing rules. to transform logs into structured formats for querying using its Grok parser.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">\u00a0<\/span><span style=\"font-weight: 400;\">Other products to look at include Mezmo (telemetry pipelines and log analysis) &#8211; it actively markets itself round, for example, <\/span><a href=\"https:\/\/www.mezmo.com\/blog\/controlling-datadog-costs-with-telemetry-pipelines\"><span style=\"font-weight: 400;\">reducing datadog spend<\/span><\/a><span style=\"font-weight: 400;\">. Edge Delta also plays in the (telemetry pipeline space)[<\/span><a href=\"https:\/\/edgedelta.com\/comparison\/edge-delta-vs-cribl\/\"><span style=\"font-weight: 400;\">https:\/\/edgedelta.com\/comparison\/edge-delta-vs-cribl\/<\/span><\/a><span style=\"font-weight: 400;\">]<\/span><\/p>\n<p><span style=\"font-weight: 400;\">So telemetry pipelines is definitely part of the solution, but we feel an active and intentional approach to log data management, with a specific focus on, in effect, hierarchical storage, where data be stored in cheaper blog storage, but also quickly rehydrated and made available for querying. The key point here is that log storage is expensive. Every enterprise or SaaS company we talk to feels that pain. Which is where log data management comes in.<\/span><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"font-weight: 400;\">Disclosure: Splunk, Cribl, Chronosphere, Control Theory, Honeycomb, AWS, Microsoft (Azure), and Google Cloud are RedMonk clients.\u00a0<\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>With all the buzz around Observability over the last few years it\u2019s easy to imagine that when it comes to logs, metrics and traces, it\u2019s game over. Just stick all the data you need in a database, or these days a data lakehouse, and start building queries and dashboards. Easy.\u00a0 Glibness aside, Observability tools vendors<\/p>\n","protected":false},"author":5,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"spay_email":"","footnotes":"","jetpack_publicize_message":"","jetpack_is_tweetstorm":false},"categories":[1],"tags":[],"class_list":["post-5373","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"jetpack_featured_media_url":"","jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p9wfjh-1oF","_links":{"self":[{"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/posts\/5373","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/users\/5"}],"replies":[{"embeddable":true,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/comments?post=5373"}],"version-history":[{"count":0,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/posts\/5373\/revisions"}],"wp:attachment":[{"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/media?parent=5373"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/categories?post=5373"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/redmonk.com\/jgovernor\/wp-json\/wp\/v2\/tags?post=5373"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}