OpenTelemetry Observability Tutorial
Learn OpenTelemetry with the Grafana LGTM stack by instrumenting the same e-commerce demo across Spring Boot, Quarkus, and Python.
What this tutorial is
If you have run a distributed system, you know the failure mode: a request crosses half a dozen services, and when it is slow or quietly wrong, the logs you have are too sparse to explain it, too noisy to search, or costly enough that you hesitate to add more. The thing that would actually help โ the request's full path, in order, across every service, with the time spent in each hop and the cost of serializing data between them โ is the thing you almost never have.
This tutorial builds it, three times over. The same small e-commerce domain is instrumented with OpenTelemetry on Spring Boot, Quarkus, and Python, chapter by chapter, until a single request is a trace you can read end to end โ with metrics that show it in aggregate and logs that correlate back to the same trace, all in a Grafana stack you run yourself. Every chapter shows the same concept across all three stacks side by side, so you can see what each runtime's instrumentation actually buys you, and what it costs. It stays honest about the trade-offs throughout: what auto-instrumentation gives you for free, where you pay for it in overhead or noise, and when one hand-placed span is worth more than a dozen automatic ones.
Foundations
What observability is for, how OpenTelemetry is put together, the Grafana LGTM stack, and getting the tooling in place before any code is instrumented.
The demo application
A small e-commerce domain, the shared infrastructure it runs on, and a bare service skeleton in Spring Boot, Quarkus, and Python before any signal is added.
The signals
Traces, metrics, logs, baggage, context propagation across Kafka, profiling, and the semantic conventions that keep them all readable โ one signal per chapter, three stacks each.
Making signals useful
Correlating traces, metrics, and logs; choosing auto, manual, or hybrid instrumentation; building dashboards; and putting the Collector and sampling in the path.
Production concerns
Cardinality and cost, the service graph and SLOs, a production readiness checklist, and where to go from here.
Appendix
Deploying the stack and the three-language services to OpenShift with CodeReady Containers (CRC).