Keeping AEM Dispatcher caches trustworthy in production
AEM Dispatcher sits between Adobe Experience Manager and visitors, handling request filtering, caching, invalidation and delivery of published content. When it is healthy, users receive fast responses and the publish tier carries a manageable load. When it is misconfigured, an apparently available site can serve stale pages, miss cache opportunities or expose an origin that should remain private.
A useful monitoring approach must check two things together: the state of the cache and the freshness of the content behind it. These AEM Dispatcher health checks should reflect how an Australian site operates across time zones, networks and traffic patterns, rather than relying on a single server ping. The objective is practical evidence that visitors are receiving the right version of each page.
What a meaningful health check should prove
A process check confirms that Apache HTTP Server or the relevant web server is running, the Dispatcher module has loaded and configuration files are valid. Those checks are necessary, but they do not prove that a request is reaching the correct publish instance or that the response is cacheable. A web server can return HTTP 200 while an incorrect virtual host, broken rewrite rule or denied cache directory quietly damages the customer experience.
A stronger probe requests a controlled, published page and records the status code, response time, cache headers, content marker and origin behaviour. The first request may populate the cache; a second request should show the expected cache hit signal, such as an appropriate Age value or Dispatcher-related header configuration. The exact header differs by implementation, so monitoring should focus on an agreed observable rather than assuming every AEM deployment exposes the same wording.
The probe should also test a page that must never be cached, such as a personalised account route, alongside a public page designed for caching. This verifies that filters, permissions and cache rules are working in both directions. A healthy Dispatcher protects private paths while allowing high-volume public content to be served efficiently.
Inspect cache behaviour at the edge
Cache monitoring begins with the filesystem. Dispatcher stores cached files beneath its configured cache root, and a monitoring agent can check disk usage, inode consumption, directory ownership and write permissions. Sudden growth may indicate ineffective invalidation, query-string variation or a crawler generating large numbers of unique URLs. A full volume can cause cache writes to fail even while existing files continue to be served.
Statfiles and invalidation rules deserve specific attention. When an author activates a new article, the flush request should reach the Dispatcher, remove or mark the relevant cached representation and allow the next visitor to receive the current version. A check can publish a test asset, record its version identifier, activate a change and confirm that the edge response changes within the agreed service level.
Cache hit ratio is useful as a trend rather than a pass-or-fail target. A retail site in Sydney may see a different pattern from a professional services site in Adelaide, while campaign landing pages can have naturally low reuse. Pair hit ratio with origin request volume, latency, response size and error rate. A falling hit ratio with rising publish traffic is more meaningful than a low percentage viewed in isolation.
Confirm content freshness and replication
Cache status is only half the picture. The published repository, replication queues and Dispatcher flush agents must work together. A page can be present in the cache and still be wrong because a publish activation failed, a content fragment was not replicated or a flush agent could not connect. Health checks should therefore compare a known content revision at author, publish and edge layers.
For sites that obtain data outside AEM, monitor the integration separately from the page cache. External data may arrive through an API, a scheduled import or a persistence layer. The CIRCUIT discussion of AEM and MongoDB is a useful reminder that replication and external storage introduce their own operational states. A page can be freshly cached while displaying old product, availability or profile data from a failed dependency.
Use synthetic content markers rather than inspecting arbitrary production copy. A small, controlled page can contain a revision number and timestamp that are safe to expose to monitoring. The check should verify that the expected version reaches the edge, then report separately on authoring delay, replication delay, invalidation delay and delivery delay. This breakdown makes incidents faster to diagnose.
Design monitoring for Australian operations
Australian deployments often serve users across Sydney, Melbourne, Brisbane, Perth and regional areas with different network conditions. Run probes from more than one location, ideally through the same CDN or edge paths used by customers. A page that is fast from a Sydney monitoring node may have a different result for Perth visitors if routing, cache locality or an origin connection is poorly configured.
Schedule probes with Australian time zones in mind. AEST and AEDT changes can affect dashboards, alert windows and overnight maintenance jobs, while public holidays and major retail events can alter normal traffic patterns. Before a Boxing Day promotion, an NRL or AFL campaign, or a government service deadline, establish a baseline for cache hit rate and origin load so an alert reflects a real deviation rather than an expected surge.
Data residency and privacy also matter in the local market. Check that diagnostic payloads contain no customer identifiers and that logs sent to an observability platform follow the organisation’s retention and access policies. A public content probe can be modelled on any approved external reference, such as this child development resource, but production checks should use controlled test content and should never submit personal information to an unrelated site.
For a public-facing service, alert on user impact rather than every internal fluctuation. Useful thresholds include repeated five-hundred responses, an unexpected cache miss rate, stale content beyond its service objective, a flush queue that continues to grow, and a measurable increase in origin latency. Send high-priority alerts to the team responsible for AEM, the web tier and the network path, with enough context to identify the failing layer.
Troubleshoot without damaging the cache
When a check fails, begin by classifying the response. A five-hundred error from publish, a four-hundred response created by a Dispatcher filter, a cache permission error and a stale successful response require different actions. Capture the URL pattern, request method, host, response headers, publish instance, Dispatcher log entry and cache file timestamp. This evidence is more valuable than repeatedly clearing the entire cache.
Review recent configuration changes before forcing a purge. Dispatcher filters, allowed client headers, cache rules, rewrite maps and flush-agent settings can all change behaviour. A broad cache deletion may hide the original symptom and create a surge of requests to AEM. Prefer targeted invalidation, configuration validation and a controlled warm-up using approved public paths.
Health checks should be versioned with the deployment and tested in a staging environment. Include cache-control behaviour, selectors, extensions, URL parameters, compressed responses and authentication boundaries in the test set. If the organisation uses AEM as a Cloud Service, align checks with the platform’s supported deployment and observability model rather than copying assumptions from an older on-premises installation.
Teams can use technical conference material to broaden their operational view; the CIRCUIT conference archive includes AEM-focused sessions covering architecture, integrations and related engineering practices. The best monitoring design still comes from translating those ideas into explicit service objectives: how fresh content must be, how quickly invalidation must work and what visitors must experience during a dependency failure.
Define a small set of synthetic pages, instrument both cache and content signals, and run the checks from the locations that matter to your Australian audience. Review the results after each release, campaign and infrastructure change, then make AEM Dispatcher health checks part of the normal production runbook rather than an emergency-only task.