OpenTelemetry is a telemetry framework, not a backend
OpenTelemetry provides APIs, SDKs, automatic instrumentation packages, semantic conventions and the OpenTelemetry Collector for producing and moving traces, metrics and logs. It gives teams a common way to describe telemetry and export it to one or more compatible backends. Storage, query experience, retention and alerting still belong to the chosen backend systems.
Use it to reduce one-off instrumentation and to make service identity consistent across languages and environments. Do not roll it out as a blanket library upgrade with no data plan. Define the questions telemetry must answer, the attributes permitted in each signal, the destinations and the people who will respond when export or collection fails.
Choose signals for a specific investigation path
Good telemetry lets an operator move from a symptom to evidence without guessing field names.
Traces
Use spans to show work across service boundaries and timing. Decide sampling and propagation before treating trace coverage as complete; sampled traces are not an event-for-event audit trail.
Metrics
Use metrics for repeated numerical questions such as rate, errors, duration and saturation. Keep labels bounded and avoid copying high-cardinality request identifiers into metric dimensions.
Logs
Use structured logs where event context matters. Decide which trace or service attributes are useful to correlate, and do not put secrets or unnecessary personal data in log attributes.
Resource identity
Set shared service, environment and deployment attributes so a backend can distinguish a production checkout service from its staging or worker counterpart.
Design the Collector pipeline from input to destination
Each receiver, processor, exporter and extension has an operational cost and data effect.
| Stage | Question | Validation |
|---|---|---|
| Receivers | Which protocols and sources may send telemetry? | A sample workload exports only through the approved listener and authentication path. |
| Processors | What batching, filtering, enrichment or sampling is required? | A test signal shows expected resource attributes and excluded fields are absent. |
| Exporters | Which backends receive each signal and with what credentials? | A controlled backend failure produces a visible error and the intended retry or drop behaviour. |
| Extensions and health | How will the Collector expose health and diagnostics safely? | Operators can check health without making admin endpoints public. |
Place Collectors by traffic and failure boundary
A Collector can run close to an application, as an agent on a host, or as a gateway that receives telemetry from many workloads. The right topology depends on network routes, workload density, credentials, buffering needs, transformations and the failure behaviour you can accept. Do not copy a topology diagram without testing what happens when its central component is unavailable.
Document an illustrative path: application SDK sends OTLP to a local or gateway Collector; processors add permitted resource context and apply policy; exporters send signals to the designated backend. Measure queue growth, memory use, export failures and dropped data under representative load. That is more useful than treating the Collector as invisible plumbing.
Govern instrumentation across teams
Adopt shared names for services, deployments and key operations using current OpenTelemetry guidance. Publish a short local convention for attributes that the organisation adds beyond the standard semantic conventions.
Choose trace sampling from the investigation need and cost boundary. Record head or tail sampling decisions, exceptional paths that must be retained, and the effect on trace-linked logs or metrics.
Version instrumentation packages deliberately. Test automatic instrumentation, context propagation and custom spans in a representative service before rolling the change across every runtime.
Filter data before it becomes telemetry
Telemetry can expose request headers, database statements, account identifiers, file names, network addresses and internal topology. Define prohibited fields and review automatic instrumentation defaults. Use Collector processors or application-side controls to remove or transform data according to policy, but verify the result in the destination backend as well.
Protect OTLP endpoints, Collector administration and exporter credentials. Use TLS and authentication appropriate to the deployment, restrict network reachability and keep secrets outside versioned configuration. A telemetry pipeline often spans development, platform and security teams; unclear ownership is a common way for sensitive data to remain after everyone assumed somebody else removed it.
Keep backend responsibilities explicit
OpenTelemetry can export one signal to more than one destination, but duplication should have a purpose. Write which backend owns traces, metrics and logs, their retention, who can query them and how dashboards or alerts will use them. An exporter configured successfully does not guarantee schema compatibility, useful queries or correct access control.
Test correlation in the place users will investigate: find a service symptom, filter by service and environment, inspect a trace, then locate the related log or metric if that connection is required. If the hand-off fails, fix service naming, propagation or backend mapping before instrumenting more services.
Roll out in service-sized increments
Start with a high-value service whose owners will use the output. Inventory its language, runtime, outbound dependencies, existing metrics and sensitive-data risks. Add basic resource attributes, a small set of traces or metrics, and an approved Collector route. Validate with a known request path and an expected failure condition before adding custom detail.
Expand with a shared onboarding checklist rather than copying bespoke configuration. Each new service should name its owner, environment attributes, signal destinations, sampling choice, data review and rollback path. Keep a record of services that are only partially instrumented so an incident responder does not mistake absent telemetry for healthy behaviour.
Operate the pipeline as a product
Collectors and instrumentation need owners, dashboards and change control of their own.
Collector owner
- Monitors receiver traffic, queues, memory, export errors and dropped telemetry.
- Tests upgrades and configuration changes with representative signals.
- Maintains the receiver and exporter dependency map.
Service owner
- Owns service names, custom attributes and instrumentation correctness.
- Reviews sampling and data fields when the application changes.
- Uses incidents to identify missing signals or misleading context.
Data owner
- Sets sensitive-data rules and backend access boundaries.
- Reviews Collector filtering and destination retention.
- Maintains response steps for exposed telemetry or credentials.
Change record
- Track Collector versions, receivers, processors and exporters.
- Test dropped-signal and back-pressure behaviour before rollout.
- Keep service owners informed when naming or sampling changes.
OpenTelemetry questions before broad adoption
Does OpenTelemetry store telemetry?
No. It provides a standard way to generate, process and export telemetry. Choose compatible backends for storage, query, retention and alerting, then test the complete path.
Where should a Collector run?
Place it according to workload, network, credentials, transformation and failure boundaries. An agent may fit local collection; a gateway may fit shared processing. Validate the selected path under loss or pressure.
Can we capture all trace data?
Possibly for a bounded case, but volume, cost and privacy may make that inappropriate. Set sampling deliberately and communicate its effect on investigations.
What is the first security check?
Inspect actual exported fields from a representative workload. Confirm that secrets, unnecessary personal data and prohibited request content are absent before adding more instrumentation.

