Originally published on Medium.
Incident context before the first query
Chiron prepares incident context before the AI asks its first question — connecting the symptoms, their timing and the services involved. The agent starts with that evidence already assembled.
When a customer reports that checkout is broken, an AI agent must find the relevant failures, establish when they began and connect them across services. Gathering and aligning that evidence can require repeated searches through logs, metrics and traces before the agent can explain the cause.
Chiron performs this correlation as telemetry arrives. By the time an investigation starts, the agent can query a compact view of the affected services, their dependencies and how their condition changed. Our ORCA-bench evaluation examines what this prepared context changes in root-cause accuracy, investigation effort and token use.
We changed what the agent reads. Chiron is ChironVision’s in-cluster observability engine. It correlates OTLP metrics, logs and traces as they arrive and keeps state per request hop, service and flow, with history. That state forms a context graph in which degradation propagates from a hop to service health and on to the user flows that depend on it. The agent queries this state through an API. No LLM runs inside Chiron.
Results across the benchmark
ORCA-bench tests incident investigation in a running e-commerce application with injected faults. An agent receives a user-style problem report and must identify all expected root causes, explain their mechanisms and cite supporting evidence. Our evaluation uses 664 public tickets: 526 incidents and 138 quiet-window controls.
On 526 ORCA-bench incident tickets, Chiron with Claude Sonnet 4.6 named every expected root cause in 46.2% of cases. The published Sonnet + Grafana reference reports 30.9% on 884 incidents. Sonnet with Chiron used 154K tokens per incident trial on average; the paper reports 1.43M per incident trial. This is a tuned setup: Chiron’s digest carried a cause catalogue and investigation instructions fitted to this benchmark, which the paper’s Grafana-supported agents did not have, so it is not a zero-shot, like-for-like comparison. The token accounting and evaluation scope are specified below.
Evaluation scope. Our run uses the public ORCA-bench tickets covered by our telemetry extract, a subset of the same benchmark. The published reference covers the paper’s full ticket set. Fitted application knowledge and query limits differ; tokens are compared per incident trial; How we tested gives the counts and settings.
Footprint and cost. The replay processed 12.5 GB of telemetry and retained 27 MB of RCA state. This compares the input tape with the selected state available to the investigator; raw telemetry is not retained in that state. Reported model cost was about $0.26 per Sonnet ticket and $0.16 per Gemini ticket. Prompt caching means token reduction and billing reduction are different quantities.
The five agents in the paper’s main code-and-telemetry evaluation each spend over half a million tokens per incident trial and stay under 31% RCA accuracy. Sonnet reading Chiron’s state reaches 46.2%, at 154K tokens per incident trial. The paper’s Sonnet without source code reaches 21.4%; our investigators retained source access.
The improvement is not uniform across metrics. RCA depth was 42.6%, below the published Sonnet result of 46.7%. Finding the expected causes does not guarantee a complete explanation with all the required evidence; the case studies and judge analysis examine that remaining work.
The paper’s Sonnet accuracy falls from 58.7% on Easy tickets to 8.6% on Hard, whose expected cause count averages 4.4. In our cohort, Sonnet with Chiron stays between 41.4% and 50.0% across the three difficulty groups. Easy remains below the published reference. Excluding imageSlowLoad raises our Easy accuracy from 48.0% to 54.9%, showing how much this uninstrumented fault affects our result; it does not establish the paper’s performance on an equivalent filtered cohort.
The paper reports about 52 telemetry commands per Sonnet ticket; 32.2% return empty results or errors. Its command shares imply roughly 77 commands in total. On Chiron, Sonnet averages 2.3 state reads and 12.5 total commands per incident ticket. Source investigation remains substantial at 8.2 commands, versus about 11.5 derived from the paper. No Chiron state read returned empty in the reported trajectory analysis. These are descriptive workload comparisons: the paper’s telemetry-command mean includes controls, while our trajectory analysis covers incidents.
Sonnet averaged 12.5 commands per incident with Chiron, including 8.2 source-investigation commands. Prepared context supplied the evidence; checking the causes remained part of the work.
Overall hallucination under the benchmark’s definition is 13.1%, compared with the paper’s 14.8%. Here, hallucination means a non-empty report naming none of the expected causes; it does not count every unsupported statement or extra cause. In our run, 61 of 69 such reports occur on imageSlowLoad tickets. The remaining 457 tickets have a rate of 1.8%, a within-run breakdown. On controls, Sonnet stays quiet on 67.4%, versus the paper’s Sonnet at 57.9% and Opus at 84.1%. The digest labels alerts that are no longer active, giving the agent evidence to distinguish historical trouble from a current incident.
How Chiron prepares evidence for investigation
Chiron uses stateful stream processing with keyed state kept next to the partition that updates it. Signals for an entity are routed to its owning partition, where the engine updates the correlated state as they arrive. The agent then reads that prepared state rather than reconstructing the same relationships across stores for each question. Cascades connect hop state to service health and dependent user flows; snapshots preserve the history needed to investigate when a condition changed.
How the state addresses investigation problems
The paper describes several obstacles to incident investigation. The table connects each obstacle to the evidence Chiron prepares and an observation from our evaluation. It explains how the prepared context is used; isolating the contribution of an individual feature would require an ablation.
The agent receives this information through the incident digest described next. Missing instrumentation, extra background causes and evidence outside the RCA state remain limitations, discussed after the investigations.
Inside the incident digest
Every investigation starts with one incident-context call over the ticket’s window. The digest combines two distinct inputs.
Correlated state, from the engine. Degraded hops, services and flows, backed by live numbers; the cascade linking a hop to service health and user flows; log-signature counts with the tick where they start rising; and hop error, latency and traffic changes with their peaks. Items carry the relevant OTel metric, span or parsed log signature.
Application knowledge, fitted to this shop. A cause catalogue maps symptom patterns to fault flags, for example “Cart AddItem/GetCart/EmptyCart errors → cartFailure”, with evidence guidelines such as “name a flag only with a matching symptom”. The paper’s Grafana agents had neither this catalogue nor these instructions.
The catalogue supplies a mapping, but the agent still has to choose which candidates explain the report. The flags evaluated in a window include flags that were not causes; in Case 1, nine flags appear and one is expected. Rising log tallies help identify active changes, flat tallies help reject candidates, and the cascade links a failing flow to the hop beneath it. Source code remains useful for explaining how the fault produces the symptom.
The first digest brings together rising errors, stopped traffic and shared dependencies, giving the agent several relevant paths to investigate.
Four investigations with prepared incident context
The investigations below show how Chiron’s prepared context supports diagnosis: separating relevant failures from background activity, distinguishing payment errors from stopped traffic, and connecting checkout problems to a shared dependency. In all four selected cases, both Sonnet and Gemini identified every expected cause. Their paths reveal which evidence Chiron supplied upfront and which questions still required source inspection or further checks.
The selection contains one Easy, two Medium and one Hard ticket. These successful investigations illustrate how the context was used; aggregate success rates are reported above. The chart preserves the command order for each Chiron case. Its paper row is an aggregate workload reference, not a run on these tickets. The raw-backend paragraphs describe plausible ways to retrieve equivalent evidence, not measured counterfactual runs.
Case 1. Finding the cart fault among concurrent signals
Easy ticket: “Empty Cart waits about 15 seconds, then HTTP 504”. Expected cause: cartFailure.
Outcome and Chiron’s contribution. Both models identified cartFailure. The first digest placed the cart errors alongside other rising signals and their timing, so the investigation could check which failure explained EmptyCart rather than follow the largest log increase.
What Chiron showed. Three log signatures rising in the window: ad manual GC (+97), cart storage errors with log body “connect to redis” (+9), and payment charge errors (+1). Nine flags were evaluated in the window.
Sonnet confirmed the symptom was live with one state read, then read CartService.cs, ValkeyCartStore.cs and Program.cs end to end. It found that EmptyCart routes to a store wired to badhost:1234, named cartFailure, and set the ad and payment signatures aside as concurrent but unrelated. Gemini hypothesised first and needed two grep retries for what one full read gave Sonnet.
Retrieving equivalent evidence from raw backends could involve log searches tallied by service and message, PromQL ranges to check whether alerts remain active, and traces on the cart route. Chiron placed the rising signatures and their timing in the first response.
Both models named the expected cause. GPT-5.4 scored Sonnet 2 of 3; Gemini scored the Gemini report 3 of 3. These depth scores come from different judges and are not a controlled comparison of explanation quality.
Sonnet used 105K tokens and 12 commands; Gemini used 101K and 13. Here, the compact context gave both models a useful starting point, but their total investigation effort was similar. Selecting the signal relevant to EmptyCart still mattered more than choosing the largest log increase.
Case 2. Separating payment errors from disappearing traffic
Medium ticket: “customers are having trouble buying”. Expected causes: cartFailure, paymentFailure (two events) and paymentUnreachable.
Outcome and Chiron’s contribution. Both models identified every expected cause. The digest separated payment errors from disappearing payment traffic and placed both on the timeline, preserving the clues to two distinct payment failure paths.
What Chiron showed. Cart storage errors (+38) and payment charge errors (+27), both starting at 20:40. Payment Charge traffic dropped to zero at the payment and checkout hops from 21:35. The key distinction is two signatures from one service: Charge errors point to paymentFailure, Charge traffic stopping to paymentUnreachable, and the digest shows both side by side with their times.
Sonnet pulled the GetProduct error-rate series, then read the failure code for the candidate faults and wrote a causal chain for each expected event. Gemini searched each candidate flag and pulled the cart and payment log series, then grouped its diagnosis under “multiple failure-injection flags”.
On raw backends, an investigator could follow the checkout path through PlaceOrder errors, Charge errors versus Charge traffic, hung cart calls and logs from three services, then align their onsets. The digest already separated the two payment signatures and placed them in time.
Both models named every expected cause. Under GPT-5.4, Sonnet received 2 of 3 for each paymentFailure event and 1 of 3 for cartFailure and paymentUnreachable. Finding the full cause set therefore did not imply complete mechanism evidence.
Sonnet used 150K tokens and 15 commands; Gemini used 122K and 14. This case shows the specific value of keeping error rate and traffic volume together: “payments are failing” and “payment traffic has stopped” point to different failure paths, even within the same service.
Case 3. Following checkout failures into the catalog
Medium ticket: “shoppers are seeing checkout problems”. Expected causes: adManualGc, paymentFailure (two events) and productCatalogFailure.
Outcome and Chiron’s contribution. Both models identified every expected cause. Chiron exposed the failing catalog dependency alongside the payment symptoms, giving the investigator a specific connection to check before treating payment errors as the whole explanation.
What Chiron showed. Three failing flow SLAs, all traced to one hop: GetProduct as called from checkout, at 6% errors. PlaceOrder failing 25% at the frontend. Payment charge (+19) and ad GC (+5) signatures rising from 20:10; cart, Kafka and ad-request signatures flat.
This ticket is easy to get wrong: payment errors explain “checkout problems” on their own, and the catalog fault belongs only because checkout calls GetProduct for every cart item. Sonnet read the GetProduct failure path, confirmed the hop with a state read, and wrote down why an error that predates the window still belongs in the answer. It set aside cartFailure and kafkaQueueProblems because their tallies were flat. Gemini spent its follow-ups dating an onset the digest had already flagged, and never wrote that checkout calls the catalog.
The digest also supplies telemetry names to cite and the cause catalogue. Its first chain connects three failing user flows through ProductCatalogServiceHealth to GetProduct. The hop series identifies checkout as the caller, a 6% error rate, and a condition already present when the window opens. Together, those observations give the agent a concrete dependency to investigate in code.
An equivalent investigation through raw backends could use route error rates, traces to find the shared catalog call and its caller, and service logs to date rising signatures and reject flat ones. Chiron brought those relationships and changes into the initial context.
Both models named every expected cause. GPT-5.4 gave Sonnet 3 of 3 on the catalog cause; Gemini withheld mechanism credit from the Gemini report. The different judges limit a direct score comparison, but the reports show the substantive difference: Sonnet explained the checkout-to-catalog dependency, and Gemini did not.
Sonnet used 135K tokens and 15 commands; Gemini used 103K and 11. The catalog cause was visible in the first call, alongside the more obvious payment errors. The remaining task was to explain why that dependency belonged in the answer. This is where prepared context and the investigator’s reasoning meet.
Case 4. Identifying several causes in four commands
Hard ticket: “users are reporting site issues”, over a 24-hour scan. Expected causes: adFailure, adManualGc, cartFailure, paymentFailure (two events) and productCatalogFailure.
Outcome and Chiron’s contribution. Both models identified every expected cause across a full-day window. Chiron assembled the concurrent signatures, traffic changes and dependency evidence upfront; Gemini completed its investigation in four commands and Sonnet in fourteen.
What Chiron showed. Five log signatures with separate onsets, from ad GC at 07:35 to cart storage and payment charge at 22:15. Charge traffic stopped from 20:20. GetProduct p99 rose 14× at the frontend, ListProducts 9.6× and recommendations 5.9×. The order-pipeline SLA failed from 04:40.
Gemini named every expected cause in four commands. The digest exposed distinct symptoms for the causes, making a short investigation possible in this case. Sonnet read the mechanisms, wrote a causal chain for each cause, and treated the recommendation slowdown as downstream of the catalog rather than a separate cache fault. It also named two live background faults outside the expected set.
Retrieving equivalent evidence through raw backends could involve bucketed log searches over 24 hours to place the separate onsets, latency and request-rate series, and traces connecting recommendations to the catalog. The initial digest had already assembled that evidence.
Both models named every expected cause. Under GPT-5.4, Sonnet lost mechanism credit on cartFailure; the judge comparison below examines that disagreement.
Sonnet used 121K tokens and 14 commands; Gemini used 53K and four. More investigation produced a more explicit dependency explanation, but did not guarantee a cleaner cause list. Chiron supplied evidence that supported both paths; each model still decided how far to inspect it and what to include.
Both models identified every expected cause in all four selected investigations. In the selected Hard case, Gemini reached that result in four commands.
What the state does not resolve
Missing instrumentation. imageSlowLoad has no diagnostic metric, log or trace clusters in the benchmark: Sonnet is correct on 3 of 69 such tickets, which also account for 61 of its 69 hallucinations. The supplied telemetry lacks the evidence needed for a grounded diagnosis of this fault. Additional instrumentation, such as real-user monitoring or suitable flag-change events, could expose it.
Background faults. Real faults that do not explain the reported user impact, such as kafkaQueueProblems and adHighCpu, can remain visible in state and get named, as in Case 4. Strict accuracy requires all expected causes but permits extras: 144 of the 243 passing reports name at least one cause outside the expected set. Precision matters alongside complete enumeration.
Parsed logs. The RCA state keeps log signatures rather than verbatim lines; the raw lines go to the customer’s own store losslessly, and the RCA API does not serve them today. GPT-5.4 matched a log body in 180 of 2,897 checks. The judge example below shows one disagreement over a parsed signature; it does not establish how much of the aggregate depth gap comes from log representation. Carrying one raw exemplar per signature is the fix in progress.
What these investigations show
Chiron changes when incident correlation happens: the engine connects signals as telemetry arrives, so each investigation begins with a compact view of faults, timing and dependencies. In this evaluation, 12.5 GB of telemetry produced 27 MB of RCA state. Sonnet averaged 12.5 commands per incident and identified every expected cause on 46.2% of evaluated incidents, compared with the published 30.9% reference. It used 154K tokens per incident trial against the paper’s 1.43M.
The investigations show what changes when incident evidence is connected before the first query. Chiron makes timing, dependencies and concurrent failure signals available together, giving the agent a focused starting point for checking causes. The cases show that both models could use this context to identify the expected causes, while still needing to interpret signals, inspect mechanisms and decide which failures explained the reported impact.
For teams building an AI SRE agent, Chiron offers a way to prepare incident evidence as telemetry arrives and expose that context through an API or MCP. The practical evaluation is how this changes your agent’s diagnosis quality, investigation workload and token use on your own telemetry. Chiron runs inside your cluster, and its ingestion pipeline can also reduce the volume forwarded to a logging vendor. To evaluate your agent, contact us.
Technical analysis and evaluation details
The following sections examine investigator behavior, the test setup and judge sensitivity in detail. They preserve the evidence needed to assess the results and distinguish Chiron’s prepared context from the reasoning performed by each model.
Investigator behaviour in detail
Chiron supplies connected evidence, but the investigator still decides which causes explain the incident and how far to verify them. Sonnet and Gemini used the same prepared context differently. Examining those differences helps identify the work that context preparation addresses and the reasoning that remains with the model.
Both investigators received identical state, tools and instructions. Under the same Gemini judge, Sonnet achieved 58.0% RCA accuracy and Gemini 37.1%. Sonnet used more tokens on average (135K versus 77K). The trajectories help explain how the models used the supplied context differently.
Sonnet 4.6 often follows the digest into source code. Its next move is a state check in 50% of investigations or a full-file read in 44%; it reads 3.1 files per ticket. In the case studies it builds separate causal chains and explains why some candidate faults do not belong. Chiron gives it timing and dependency evidence to check against the implementation.
Gemini 3.1 Pro more often starts with a targeted search. Its next move is a grep in 54% of investigations; it reads 1.1 files and runs 5.8 greps per ticket, and 53% of reports immediately follow a grep. Its diagnoses sometimes group several faults under “multiple flags were enabled”. Candidate names and precomputed onsets can support a short route to an answer, as the four-command Hard case demonstrates.
The models are close on single-cause tickets (51.1% versus 47.8%). With two or more distinct expected causes, Sonnet remains at 61–63%, while Gemini ranges from 25.5% to 37.4%. The trajectories suggest that accepting an early explanation and not following additional paths contributes to Gemini’s misses. The grouped scores alone do not establish that behaviour as the sole cause of the gap.
The largest payment gap is paymentUnreachable: Sonnet names it in 91% of the 130 tickets that expect it, versus Gemini in 41%. The cases show why payment charge errors and stopped payment traffic need separate explanations. Naming the error-producing cause does not necessarily explain the traffic-stoppage path. The mention rates identify a gap to examine in individual trajectories; they do not establish a single reason for every miss.
Chiron supplies the same starting evidence to both models, but the value appears in different parts of their work. Sonnet uses the prepared context for source-level causal explanations and more complete enumeration. Gemini can use the candidate names and onsets for a shorter investigation, but the context does not ensure that it follows every relevant branch. There is no raw-store Gemini baseline in this evaluation, so its measured results describe performance with Chiron rather than a quantified uplift over Gemini without it.
How we tested
ORCA-bench is a public benchmark for this problem. It runs the OpenTelemetry Astronomy Shop, a demo e-commerce system of about twenty services, under load while feature flags inject real faults such as payment failures, cart storage outages and catalog errors. Each of its 1,079 tickets (884 incidents and 195 quiet-window controls) is a user-style report at a point in time. The agent must decide whether an incident happened, find every plausible root cause and its onset, and write a report with a Summary, Timeline, 5 Whys and Remediation. An LLM judge scores the report against per-fault rubrics.
In the paper, the agent investigates through Prometheus, Jaeger and OpenSearch behind Grafana, plus the shop’s source code. The best published results are 25.3% on Medium tickets (Sonnet 4.6) and 10.0% on Hard (Opus 4.7). Sonnet 4.6 reaches 30.9% strict accuracy overall and uses 1.43M tokens per trial.
We used the benchmark’s public tickets and telemetry, and kept the paper’s harness wherever we could.
Tickets: 664 of the 755 tasks in the public Harbor release (a subset of the paper’s 1,079), 526 incidents and 138 quiet-window controls. The other 91 fall after our telemetry extract ends, so we removed them by timestamp before any run.
Telemetry: we replayed the benchmark’s OTLP tape (12.5 GB, 2026–04–19 03:40Z to 04–24 19:30Z) through Chiron once. That produced 27 MB of correlated state. The RCA state does not hold the raw telemetry.
Agent loop: the paper’s Harbor harness and Terminus-2 agent, with the same report template. In place of Grafana, the agent had a helper CLI over Chiron’s API (incident-context, state, series), the shop’s source code and its flag config. Nothing after the ticket’s cutoff was visible. Each ticket ran in a new container and session.
Investigators: Claude Sonnet 4.6, the paper’s model, as the primary run, and Gemini 3.1 Pro as a second run. Both got identical state, tools and instructions.
Judges: Sonnet reports were scored with the official verifier prompts using GPT-5.4, the paper’s judge model, and then re-scored with Gemini 3.1 Pro. Gemini investigator reports were scored with Gemini 3.1 Pro. Our recorded GPT-5.4 setting is default effort; the paper specifies high effort. The comparison therefore uses the same judge model, with different recorded settings.
Tokens: input plus output over every model call in a trial. Across all 664 tickets the run means are 135K for Sonnet and 77K for Gemini. The paper’s 1.43M reference is its incident-trial mean, so we compare it with Sonnet’s incident mean: 154K over the 526 incident tickets (61K on the 138 controls).
What was fitted. The context graph, its cause catalogue (which symptom patterns point to which fault) and the agent instructions were tuned on this benchmark’s data before the run, then frozen. The agent also gets at most three follow-up queries after the first digest. The paper’s Grafana agents had neither the catalogue nor these instructions, so read this as a tuned system rather than a zero-shot result.
How the judge changes the reported result
The same Sonnet reports score 46.2% RCA accuracy under GPT-5.4 and 58.0% under Gemini 3.1 Pro. Depth moves from 42.6% to 60.5%. The aggregate hallucination rates are close (13.1% and 13.3%), but that alone does not show agreement on individual reports or claims. The example below shows the judges accepting different evidence for the same diagnosis.
Investigator and judge sensitivity. The three configurations use the same 664-ticket Chiron cohort. Strict accuracy and depth refer to the 526 incidents; quiet refers to the 138 controls.
One report, two verdicts. For cartFailure in Case 4, the rubric’s expected log line reads “Can’t access cart storage… Wasn’t able to connect to redis”, and the mechanism it expects is EmptyCart routing to a store wired to badhost:1234. Sonnet’s report cited Chiron’s parsed signature, log body “connect to redis” from the cart service.
In this report, GPT-5.4 did not accept the cited log body or mechanism, while Gemini did. That is an observed scoring disagreement, not proof that GPT-5.4 always requires a verbatim string. The mechanism verdict also differs, so the score gap cannot be attributed to log wording alone. Chiron’s RCA state provides parsed signatures and correlated state; retrieving raw exemplars could make some evidence easier to verify, but its effect on scores has not been measured here.
The difference also changes direction on controls: Gemini accepts fewer Sonnet reports as correctly quiet (61.6% versus 67.4%). We lead with GPT-5.4 because it is the paper’s judge model, while disclosing the different recorded effort settings. On identical Sonnet reports, changing the judge model moved RCA accuracy by 11.8 percentage points. Evaluations should identify the judge and its configuration alongside the investigator.
Reproduction details
- Context graph and instructions: frozen before the runs and identical for every ticket; no imageSlowLoad hint.
- State: one OTLP replay of the benchmark tape.
- Agent: Terminus-2; anthropic/claude-sonnet-4-6 and gemini/gemini-3.1-pro-preview; temperature 1; reasoning medium; max_tokens 16,384.
- Judges: official check_prediction.py prompts via gpt-5.4 at the recorded default effort setting and gemini-3.1-pro-preview. The paper specifies GPT-5.4 at high effort, temperature 1 and max_tokens 128,000.
- Reference: Gong et al., ORCA-bench: How Ready Are Language Model Agents for Oncall?, arXiv:2607.28545v3, 29 September 2026. Baseline scores: Tables J.1, J.2, J.4b, J.5 and J.6. Incident-token means: Table J.9. Telemetry commands: Figure 7; command shares: Figure J.2. Metric definitions: Appendix H. Judge configuration: Appendix I. https://arxiv.org/html/2607.28545v3