Using Sling Health Checks to monitor AEM system readiness

Adobe Experience Manager powers mission-critical websites for banks, insurers, retailers, and government agencies across Sydney, Melbourne, and Brisbane. When an AEM author or publish instance drifts towards an unhealthy state, the impact is immediate — broken pages, failed replication, sluggish content authoring. A robust readiness monitoring layer catches these signals early, and Apache Sling Health Checks give AEM engineers a standardised way to expose the live state of every component.

This article walks through how to wire Sling Health Checks into an AEM deployment, configure the right tags for readiness and liveness probes, build targeted probes for Australian enterprise scenarios, and pipe the results into dashboards and alerting tools. Along the route, you will also see how CIRCUIT sessions explored adjacent topics such as integrations with NoSQL stores and content fragment strategy.

How Sling Health Checks fit into AEM

Apache Sling, the underlying REST framework for AEM, ships with a Health Checks module that runs checks on demand and reports the combined state of an instance. Each check is an OSGi service implementing the HealthCheck interface and carries tags that describe what kind of signal it produces. Tags such as critical, ready, live, or external let monitoring agents filter the checks they care about — the foundation of the readiness probe pattern.

When a check executes, it returns a Result object containing a status code (OK, WARN, CRITICAL), a message, and arbitrary data. The combined outcome is exposed through two endpoints: an HTML page that is helpful in the Felix Console during development, and a JSON document that scripts and external probes can consume. Australian teams running multi-region setups between Sydney and Singapore often route critical results straight into an internal status page.

The Health Checks execution model is asynchronous and cached for a short window, so it can be polled frequently without putting pressure on the JVM. A typical publish farm running on AWS Sydney or Azure Australia Central can have dozens of checks per instance — disk space, bundle state, replication queues, search indexes, workflow queues — without measurable overhead.

Configuring tags for readiness and liveness

The most common starting point is to differentiate between checks that confirm an instance is alive and checks that confirm it is ready to serve traffic. The live tag is intended for liveness — should this process be restarted? The ready tag answers a different question — is this instance currently able to serve requests? Combining both tags with critical gives an operator full visibility in a single dashboard.

A sensible configuration for a publish tier in financial services looks like this. Replication agent connectivity and Oak index health are tagged ready and critical, because without them the instance cannot serve accurate results. Bundle state checks are tagged live because a bundle stuck in Installed state usually means a restart will fix things. Disk space and JCR query performance get a ready tag without critical, so they appear in the dashboard but do not page anyone after hours.

To tag an existing check, drop a node under /apps/config/.../org.apache.sling.hc.core.impl.executor.HealthCheckExecutorImpl with the tag mappings, or use the OSGi configuration factory. Engineers in Brisbane and Perth consultancies prefer the factory approach because it lives in source control. Teams in regulated industries often add a custom tag like compliance for checks auditors need to see in quarterly reviews.

Tags worth setting on day one

  • critical — surfaces in dashboards and pages on-call
  • ready — feeds Kubernetes readiness probes and load balancer health checks
  • live — drives liveness probes that restart unhealthy processes
  • external — isolates checks for third-party dependencies from internal health

Exposing results for external monitoring

Once the checks are in place, the JSON endpoint at /system/health returns the full report. The response includes a top-level status and an array of results, each carrying its own status, message, and tag list. A simple curl from any CI runner or Kubernetes pod is enough to validate the instance during a deployment.

Kubernetes users frequently wire this endpoint into a readiness probe. The HTTP status code returns 200 for OK, 200 for WARN (configurable), and 503 for CRITICAL. A typical readinessProbe stanza points at the path with a fifteen-second interval and a ten-second timeout. Australian banks running AEM on AKS in Sydney regions have adopted this pattern, and it has noticeably reduced late-night pages caused by stale pods.

For more sophisticated monitoring, the same JSON feed feeds Datadog, New Relic, or Prometheus exporters. The system/health?tags=ready filter restricts the output to readiness-relevant checks, which is what most teams want to graph. Detailed bundle-level checks still appear in the full feed and are useful for the support rota in Adelaide or Hobart when a query starts misbehaving.

Building custom probes for AEM workflows

Out of the box, Sling offers checks for JVM memory, disk, system properties, and bundle state. Real-world AEM teams need more. A common custom check verifies the replication queue on a publish instance — if pending items exceed a threshold, the check flips to WARN or CRITICAL depending on severity. Another is a workflow queue check, because a stuck backlog on an author instance is a productivity killer for content editors.

Custom checks extend AbstractHealthCheck and return a result from the execute method. The annotation @HealthCheckService registers the service, and tags are applied with @HealthCheckTag. A pattern that works well in Australian retail is a check that pings the downstream search service — if Elasticsearch is unreachable, the publish instance is alive but not ready, so the check is tagged ready rather than live. There is a great walkthrough of pairing AEM with Elasticsearch in the AEM and Elasticsearch for Advanced Search Capabilities session recording.

Another frequent request from analytics teams is a check that validates the analytics cloud configuration endpoint. If the connection breaks, content personalisation silently degrades. The same pattern works for any external integration — MongoDB for unstructured data stores, as discussed in the AEM and MongoDB for Unstructured Data Storage recording, or third-party CDNs.

Common failure scenarios in Australian deployments

Operating AEM at the bottom of the world introduces a few quirks. The first is network latency to Adobe licensing endpoints — the round-trip from Sydney to Adobe's US data centres can stretch past 400 milliseconds during peak US business hours, which causes licensing checks to flap. Engineers at Big Four consultancies in Melbourne have learned to extend the timeout on licensing health checks or to mark them warn-only.

The second is disk pressure from segment store growth. Author instances that drive large content campaigns accumulate TarMK segments quickly, and disk full scenarios are the most common cause of author unavailability. A disk-space check with a low threshold tagged critical prevents the dreaded No space left on device error that brings the entire team in on a Sunday morning.

The third is replication backlog during peak publishing. Australian retailers pushing Black Friday or Click Frenzy campaigns can overwhelm replication queues if the publish tier is under-provisioned. A replication queue check with a sensible threshold gives the operations team a warning an hour before things collapse. Teams that manage DAM-heavy workloads also build custom checks around asset processing queues, because a backlog there blocks business users faster than anything else. Content-heavy Australian sites — from media publishers to festival portals like the Maritime Metal Fest — rely on AEM's authoring layer to deliver time-sensitive announcements, and a single stalled workflow can derail a campaign.

Readiness checks Australian teams ship first

  • Replication queue depth and connectivity on every publish instance
  • TarMK disk space on author instances, thresholded well before the OS warning
  • Workflow queue backlog on author instances, split by model
  • Downstream Elasticsearch connectivity for search-driven sites

Integrating health checks with CI/CD and alerting

Health checks earn their keep when they become part of the deployment pipeline. A smoke-test step that hits /system/health?tags=ready and fails the build if status is not OK catches misconfigured instances before they reach production. The same check runs post-deploy and gives the release manager confidence to flip traffic.

Alerting ties the same endpoint to PagerDuty, Opsgenie, or Slack. The Australian pattern is to wire Datadog monitors to PagerDuty and PagerDuty to a Slack channel — a quiet afternoon handover between Sydney and Melbourne on-call rotations. The integration with development tools is just as relevant. Engineers reviewing the AEM Content Fragments vs Experience Fragments talk will recognise the same instrumentation pattern applied to content authoring flows.

For broader context, the CIRCUIT 2016 programme included sessions on microservices and architecture that explored these patterns in depth. Recordings remain available on the conference archive, and the CSharpCon material helps .NET-heavy shops that are integrating AEM into a polyglot stack. A line of sight from the health endpoint to the business dashboard is what turns monitoring into an operational advantage — fair dinkum, that is the goal.

Watch the CIRCUIT 2016 architecture sessions for walkthroughs of readiness probes, custom checks, and alerting integrations, and start building your AEM monitoring strategy with the patterns covered here. Subscribe for the next instalment on advanced probe design, and share questions from your own production environment at the next developer meet-up.