Tracing AEM Microservices with OpenTelemetry

Distributed systems fail in ways that monolithic stacks never did. A single request might hop from a dispatcher, through a Sling servlet, into an OSGi service, then out to a search index or a personalisation engine, and back again before a single pixel renders. When that pixel arrives three seconds late to a Telstra customer on their phone in Brisbane, nobody thanks you for the architecture diagram. They want the page to load before the kettle finishes.

OpenTelemetry has become the default toolkit for making sense of those journeys. It is a vendor-neutral specification for collecting traces, metrics, and logs, with SDKs that speak to a long list of backends. For teams running AEM as a fleet of microservices rather than a single author-publish pair, it is the cleanest path to answers you can actually act on.

This piece walks through how to wire OpenTelemetry into AEM, what to expect when the spans start flowing, and where the Australian AEM community is heading with the technology. It assumes you are comfortable with Java, Sling, and the usual CI plumbing, and that you have at least one microservice whose latency you would like to understand.

What OpenTelemetry Actually Is

OpenTelemetry, often shortened to OTel, is the merged product of the old OpenTracing and OpenCensus projects. It defines a standard for emitting telemetry data, plus reference SDKs in most major languages. The Java SDK is what matters here, because AEM is fundamentally a Java stack built on OSGi and Sling.

Three signals matter most for AEM work. Traces follow a single request across service boundaries, metrics give you aggregated counts and histograms, and logs remain the verbose, contextual breadcrumb trail. OpenTelemetry's strength is that one consistent set of APIs produces all three, and you can swap the backend without rewriting instrumented code.

The project is governed by the Cloud Native Computing Foundation, which gives it a longevity story that proprietary agents rarely match. That matters when you are pitching observability investment to a Commonwealth Bank architect who has watched three monitoring vendors get acquired in five years.

Why AEM Microservices Need Tracing

AEM started life as a coupled monolith: one JVM held the author, the publish, the workflow engine, and the bundle cache. Modern deployments split these into separate processes. You might run the dispatcher as a sidecar, the personalisation engine as its own pod, and a custom search service in another container. Each split buys you scalability and team autonomy, and costs you the ability to grep a single log file when something goes wrong.

Without distributed tracing, a slow response to a Sydney-based editor looks like an AEM problem. With tracing, you discover the search service is waiting on a Solr commit that is queued behind a noisy neighbour in the ap-southeast-2 region. That kind of insight does not come from CPU graphs or APM summaries, and it is exactly the kind of thing that costs an editorial team their arvo.

The community around AEM has been quietly preparing for this shift. The Adobe partner ecosystem, including firms such as ICF Olson's team, has been pushing reference architectures that treat observability as a first-class concern rather than a late-night afterthought.

Wiring OpenTelemetry into an AEM Service

The Java SDK ships with auto-instrumentation for Servlet, JAX-RS, JDBC, Kafka, and HTTP clients. For an OSGi-based AEM service, the Servlet agent is the biggest win, because every request enters through Sling's main servlet. Drop the agent into the JVM flags, set an OTLP endpoint, and you will see spans before your next deploy finishes. Manual instrumentation is where the value compounds, and you should plan to add it from day one rather than retrofitting later.

Wrap your business-critical code in spans with semantic attributes such as aem.page.path, aem.template, and aem.workflow.step. These become the labels you filter by in the backend. Treat span names like you would treat log lines: human-readable, scoped to the operation, stable across releases. Renaming a span between versions is a fast way to break every dashboard your team has built.

Event-driven patterns are common in modern AEM deployments. JCR observers, custom event handlers, and Kafka-based workflows all benefit from proper context propagation across thread and message boundaries. The CIRCUIT session on event-driven architecture with AEM covers the patterns in detail, and is worth revisiting if you are designing around topics and asynchronous pipelines.

Once the agent is emitting data, the next question is where the spans go. A comparison of the backends most Australian teams evaluate:

Capability OTel Collector + Jaeger OTel Collector + Tempo Dynatrace New Relic
Self-host effort Moderate Low None None
Cost at AEM scale Free, you run it Free, you run it Per-host licence Per-GB ingest
AEM-specific UI None, build your own None, build your own Yes Partial
Vendor lock-in Low Low High Medium
Sampling flexibility Full Full Limited Limited

The right backend depends less on the tracing format and more on who has to keep the lights on. A team in Melbourne running AEM for a state government portal will value self-hosting and predictable cost. A team shipping marketing pages for a national retailer will value turnkey dashboards more than absolute cost control.

Rolling Out Without the Drama

Tracing in production is not free. Every span is a serialised object on the wire, and a busy author cluster can generate millions of them per minute. Sampling is the answer. Tail-based sampling, where the collector decides what to keep after seeing the full trace, gives you the best of both worlds: cheap telemetry collection, rich records of the interesting requests. Head-based sampling is simpler but tends to miss the rare, slow traces you most want to see.

Deploy the agent gradually. Start with a single read-heavy publish instance behind a canary. Watch the data flow, then check that span counts roughly match request counts. Once you trust the pipeline, expand to author and the rest of the fleet. The Australian AEM meetup in Sydney has run sessions on this staged rollout pattern, and the consensus is to never skip the canary step.

Paths worth instrumenting first

  • Sling entry points and any custom servlet that fronts them.
  • External HTTP and JDBC calls leaving the JVM.
  • Long-running workflows, especially the ones editors complain about.
  • Personalisation and search calls, because they usually live in a different service.

Gotchas and Production Realities

Context propagation across executors is the classic trap. AEM's Sling scheduler and CQ event framework both hand work to other threads, and the OpenTelemetry context does not travel automatically. You need to wrap the work explicitly, or you will get orphan spans that look like the request vanished into thin air. The fix is small once you understand it, but the symptom is genuinely confusing the first time you see a slow trace that stops mid-page render.

Span cardinality bites too. If you slap a unique attribute like a request ID on every span, your backend will groan under the cardinality of indexed tags. Stick to low-cardinality attributes for filtering and put anything unique into events or span links. This is a backend-agnostic rule that will save your team a billing surprise at quarter-end.

Things that catch teams off guard

  • The OSGi classloader isolation means some auto-instrumentation needs explicit bundle wiring.
  • Headless dispatchers often bypass the main servlet, so the Servlet agent catches nothing.
  • Sampling decisions must be deterministic per trace, not per span, or you get confusing partial records.

Instrumenting AEM well is a multi-quarter project rather than a Friday afternoon job. The good news is that the spans keep paying back the work you put in, especially when something in the ap-southeast-2 region misbehaves during a long weekend and the on-call engineer is trying to enjoy the footy.

Ready to put a tracer on your AEM fleet? Start with a single service, send its spans to a free Jaeger all-in-one, and see what your requests actually look like. When you are ready to swap backends or compare notes with other Australian AEM engineers, swing past the CIRCUIT sessions, check out a community hack day for the social side of distributed systems work, and bring your trickiest traces along to share.