Resources · 113

Collection outages: reconcile logs and missing events

Separate produced, received and usable events to explain an outage without treating silence as absence of incidents.

· 2 min

Server racks in a data centre Illustration · fictional scene

What this guide helps achieve

  • Define an observable reference
  • Exercise a controlled interruption
  • Compare identifiers and state unknowns

Quick check

  • Reference sequence and reconciliation schema.
  • Outage timeline and boundary behaviour.
  • Discrepancy list and explicit coverage state.

Step-by-step method

  1. 01

    Define an observable reference

    Map application, collector and storage with owners. Prepare a finite fictional event sequence with stable identifiers and no secrets or customer data. Retain event and observation times: collection delay does not automatically mean loss.

    Reference sequence and reconciliation schema.

  2. 02

    Exercise a controlled interruption

    In testing, interrupt the destination and restore it. Observe queue filling, retries, rejections and restart. Persistent queues can limit some losses but remain bounded and depend on storage; define full-queue behaviour. Check that collection does not silently block the service.

    Outage timeline and boundary behaviour.

  3. 03

    Compare identifiers and state unknowns

    Reconcile the reference with events actually readable at the same observation cutoff. Count repeats, delays and absences separately; network acknowledgement does not prove final usability. Without a reliable reference, state unknown coverage rather than zero lost events or zero incidents.

    Discrepancy list and explicit coverage state.

Acceptance case to reproduce

Fictional example: these inputs describe no customer or observed result.

View case inputs
{
    "expected_ids": [
        "E1",
        "E2",
        "E3"
    ],
    "received_ids": [
        "E1",
        "E1",
        "E3"
    ],
    "received_rows": 3,
    "unique_received": 2,
    "missing_ids": [
        "E2"
    ],
    "observation_cutoff": "same_for_both_sets"
}

Expected decision

Fictional example: E1, E2 and E3 are expected, but E1, E1 and E3 arrive. Three rows do not establish completeness: E1 repeats and E2 is missing at the observation cutoff.

OpenTelemetry — Collector resiliency

Your acceptance workbook

Record observations against this guide’s criteria. A record is not certification.

The workbook does not save automatically. Export before leaving.

Management indicators

IndicatorWhat it measuresFirst action
Known missing eventsExpected identifiers absent at destinationSame reference and observation cutoff
Received repeatsReceived rows − unique identifiersDoes not measure unknown losses

Common pitfalls

    Frequently asked questions

    Do three received rows prove three events were retained?

    No. Two rows may repeat an identifier and hide an absence. Compare identifiers, the observation window and pipeline stages.

    Official references

    References consulted: . The method and worksheet propose checks to adapt to your context; they do not constitute certification.