Back to Blogs

Why the Log-Cost Playbook Needs to Change

By Rohit Shekhar, Founder, Chiron Published September 25, 2026
Why the Log-Cost Playbook Needs to Change, by Rohit Shekhar

Originally published on LinkedIn.

Reduce indexing, preserve fidelity. Real-time root-cause from streaming telemetry.

The familiar ways to cut log costs depend on knowing what we will need later. AI raises the stakes, and makes a summary computed in flight the lever that helps both the bill and the capability.

An engineering leader can make every familiar move to control log costs and still end up back in the same budget discussion.

A better contract lowers the price. A pipeline sends fewer events. More data goes to cheaper storage. Each decision creates room. Then the system grows, a new team needs different information, or another workflow starts using data that was supposed to sit quietly in an archive. The next saving becomes harder to find.

The pressure is visible in Grafana’s 2026 Observability Survey of more than 1,300 practitioners: cost was the top tool-selection criterion for the third year running, and half of respondents expected spending to rise, mostly from broader adoption rather than vendor price increases.

Every one of those moves rests on an assumption about how volume will grow, which details will matter, or how often the data will be read. The tools keep working while the assumptions lose their fit. I have made most of these moves myself. The problem is not the moves. It is treating another round of them as a strategy, when the better question is how much useful information we can preserve while asking the downstream system to process less. That one question connects the bill, what can be recovered later, and what an AI agent has to work with. The answer depends on how telemetry is represented before the expensive work begins.

The tools keep working while the assumptions underneath them lose their fit. The problem is not the moves. It is treating another round of them as a strategy.

The assumptions hidden inside the savings

A discount is attractive because it changes nothing about how people work. It also changes nothing about the work performed on the data. A different platform may improve efficiency; leaders need to establish which benefit they are buying. Under a linear volume-based charge, halving the unit price is fully offset the year volume doubles.

Under a linear volume-based charge, a 50% discount is fully consumed the year volume doubles. The rate moved. The bill did not.

Figure 1. Price per unit falls from 1.0 times to 0.5 times, volume rises from 1 times to 2 times, and the bill stays at 1.0 times.
Figure 1. A lower rate and a lower bill are different outcomes. Illustrative linear volume-based charge, no tiers or commitments.

Sending less data offers more direct relief. Once the obvious waste is gone, trimming and sampling become judgments about the future. The team decides which details it can afford to lose before it knows every question the business will ask. Pattern-based drop recommendations, Grafana’s Adaptive Logs for one, make that judgment better informed. They still discard, which is the contrast with a consolidation you can reverse. An ordinary event may matter only through its relationship with another. A request ID is worthless on its own and is the only thing that joins a gateway log to a payment retry three services away. Those relationships are invisible to a pipeline that examines each message alone.

A request ID is worthless on its own and is the only thing that joins a gateway log to a payment retry three services away. Those relationships are invisible to a pipeline that examines each message alone.

Shortening retention reduces what is kept. It cannot undo the ingestion and indexing work already performed.

Routing complete originals to cheaper storage preserves the evidence and moves the decision onto a different assumption: that the records will be read infrequently enough, and served efficiently enough, for the saving to hold. Object storage can support efficient queries, and AWS’s Athena guidance shows how much layout and file format matter. A well-designed query system still performs work when used. An archive that becomes part of routine analysis belongs in a different cost calculation from one opened occasionally.

Keeping summaries in the active system helps to the extent those summaries cover the work people need to do. A set of totals can support a dashboard for years, then prove inadequate for a question about one customer, a sequence of events, or a dependency across services.

Rate, bytes, events, retained history, access frequency. Five mechanisms, five assumptions. These can all be sound decisions. The mistake is letting assumptions made for one workload become a permanent cost strategy.

Figure 2. Where each lever acts on the bill: trim, negotiate, retention, summaries, and routing to object storage.
Figure 2. Where each lever acts on the bill. Each one reduces one thing and rests on one assumption.

Each lever acts on one stage of the pipeline and reduces one thing. Retention acts after ingest and indexing have already been paid for. Routing moves originals to a cheaper tier and bets that retrieval stays rare.

AI raises the pressure, and the value of a better representation

AI changes both the activity being observed and the way observations are used.

A single request can involve multiple model calls, tool executions and retries. Recording prompts, responses and tool results adds substantial detail when enabled, as OpenTelemetry’s GenAI guidance describes. The increase depends on what a company builds and records. It is one more reason to question a capacity plan based on today’s workload.

On the other side, AI gives us consumers that keep asking questions. An agent follows a relationship across services, retrieves more evidence and returns to an earlier result. Data that was low priority for a person’s occasional search becomes routine input for an automated workflow. The same Grafana survey found the leading barrier to AI adoption was too much manual input of required context. That is largely a data-platform problem.

AI also creates the opportunity. If the data platform prepares connected information, people and agents reuse that work instead of repeating it. How telemetry is represented decides what they start with, and how much detailed data they need to retrieve.

The missing job: connect related activity before it is indexed

A pipeline that handles messages independently can transform them and choose their destination. Connecting an event to earlier activity in another stream requires shared memory, or state. With state, a processor can maintain an account of an operation across services and over time, and that changes what a compact representation can contain.

Consider an illustrative checkout operation. A basic summary keeps a failure count. A richer view connects the affected operations, retries, relevant dependency conditions and recovery history, with the observations supporting each relationship. That view can drive a dashboard by customer or operation stage and give an AI agent useful context before it touches a raw record. Its relationships guide analysis; they do not, by themselves, prove a cause.

The economic opportunity is to complete more work from that compact view. Repeated activity describing a condition that already exists does not need to grow the representation in proportion to volume. New entities and new history still take space.

One distinction is easy to miss. Compressing a file reduces storage and network. If the data is expanded before indexing, the backend still processes every original record. To avoid that work, the compact representation has to stay useful downstream. Some engines search compressed logs without expanding them; YScope’s CLP is one. They replace the storage and search tier. The argument here is about what the stack you already run has to index.

If the data is expanded before indexing, the backend still processes every original record. To avoid that work, the compact representation has to stay useful downstream.

Compression reduces transfer size, but expanding before indexing restores the original workload. Correlating first lets the backend index a compact record.

Figure 3. Processing each event alone still expands and indexes every record. Correlating first indexes a compact record, at 5 to 10 percent of downstream compute.
Figure 3. Where the work happens. Compression still expands before indexing. Correlating first lets the existing stack index less.

Streaming joins are established technology; Apache Flink supports them, and Cribl documents the memory pressure of aggregation as distinct groups multiply. The pieces exist. The difference is what they are for. A stream processor is a toolkit: the correlation logic, the state management and the output format are yours to build and operate. An observability pipeline is mostly per-event, and its stateful functions produce statistics rather than a record of what happened. What most stacks lack is the job itself: before indexing, connect related activity and prepare information that dashboards, engineers and agents can reuse, in a form the existing stack can index. Every pipeline does part of it, parsing, enriching, routing. Few do the connecting part, because holding state across many streams, late arrivals and long histories costs memory and coordination. The job grows more valuable as consumers multiply, since each one otherwise repeats the correlation. Doing it economically is the engineering challenge. It earns its place only when it removes more downstream work than it adds upstream.

Choose what to preserve: every event, or its operational meaning

The decision begins with what the business needs to preserve.

Where original event detail must remain available, demand a clear recovery guarantee: which inputs are covered, what can be reconstructed, what cannot, and which queries change. The word “lossless” should come with those four answers.

Where selected operational information is enough for routine work, judge the summary by the questions it answers. If complete originals are archived, detailed retrieval remains available. If they are discarded, no summary recreates every detail.

Both choices can be right inside one organization. The useful improvement is more control over what survives and how much work preserving it requires.

This is the thinking behind Chiron. We built a streaming engine around one idea: correlate first, store the answer. Related streams and telemetry types are correlated as the data arrives, no language model reads the incoming events, and the engine runs in front of the existing observability stack rather than replacing it. It gives a leader two answers to the question this article began with, each a different form of preservation. The checkout example maps onto them directly. Today’s baseline is a counter; Chiron adds a pattern with its variants, or a grouped record.

The counter answers one question. Lossless consolidation reconstructs every event with less indexing. A grouped record preserves incident history and context, not original strings.

Figure 4. One checkout failure as a counter, a pattern with variants, or a grouped record.
Figure 4. One checkout failure, three representations. Illustrative checkout operation. Values are examples, not measurements.

Lossless consolidation preserves the information needed to reconstruct every event, and comes with the four answers. Inputs covered: events that match an active pattern policy. What reconstructs: the original event, from the pattern stored once plus the fields that vary, kept per event, with no separate raw archive required. What does not: events outside a pattern policy, which pass through untouched, and anything handled by the summary path below. What changes: the query text, not the answers. Routed through a query proxy or an automated rewrite that expands consolidated records, a query that ran against original lines returns the same results, at the same cardinality. Individual query performance may need tuning. The work removed is ingestion and indexing in the downstream system, the compute that a better contract or shorter retention never touched.

Rich summaries before reduction or routing preserve selected operational meaning. Related multi-line activity is grouped into a single record that keeps the attributes and relationships, not the original strings. The record is computed from the full incoming telemetry; only then does the chosen sampling, filtering or routing policy apply. Queries on the record’s attributes return the same answers they did on the original lines, and the record carries its event count, so counts match. The original strings are not kept, so a search over them is the one thing this path gives up. Dropping still discards. Routing still preserves originals elsewhere. The active system simply holds a far better account of what happened before the reduction.

Neither applies everywhere. Events with no repeating pattern and no stable join key pass through unchanged. That boundary belongs in any evaluation, ours included.

So does the engine’s own cost. It runs at five to ten percent of the compute of the ingestion pipeline it sits in front of, measured on one terabyte per day of logs against that pipeline’s own compute for the same volume, averaged over a day at default policies. Our current release sits at the low end; the figure rises with how much consolidation a team asks for. Savings depend on the workload and on how much of it the policies cover, which is why we do not publish a single percentage.

Give the next cost decision a longer life

Engineering leaders should expect more than another reduction percentage. Ask three things of every option, including this one. What does it actually reduce? What will it cost to keep reducing it as the workload grows? Which capability does it give up along the way?

Evaluate how your telemetry is represented before accepting another trade-off between cost and visibility.

The familiar tools keep their roles. The sharper expectation is this: evaluate how your telemetry is represented before accepting another trade-off between cost and visibility. The next substantial gain comes from understanding the data before narrowing what the organization keeps readily available, and from doing that work once, in flight, rather than repeating it downstream for every question.

A log-cost strategy should become more capable as the business grows. If every new round of savings requires accepting another blind spot, it is time to reconsider the architecture behind the bill.

About Chiron

Chiron is a streaming engine that runs in front of your existing observability stack and connects related telemetry as it arrives, before anything is indexed. The result is less data downstream and a clearer picture of what happened, for the people and AI agents who have to act on it. Works with the tools you already run. Schedule a demo: https://chironvision.ai/contact