AEM and Hazelcast for resilient publish-cluster data

Adobe Experience Manager publish tiers are often designed as horizontally scaled, largely stateless services. That model simplifies failover and capacity planning, but applications still need fast access to shared data: session state, rate limits, feature flags, temporary workflow results, personalization inputs, and integration responses. A local in-process cache can serve one publish node well while producing inconsistent behavior across the cluster.

Hazelcast provides an in-memory data grid that allows AEM publish instances to share distributed maps, sets, queues, and other data structures. Used carefully, it can reduce repeated calls to external systems and coordinate short-lived state across nodes. It does not replace AEM’s repository, dispatcher, or persistence strategy. Instead, it adds a low-latency coordination layer around them.

The strongest design begins with a clear distinction between durable content and operational data. Content belongs in AEM and should flow through authoring and publication processes. Volatile application state may belong in Hazelcast, provided its lifetime, ownership, consistency requirements, and failure behavior are explicitly defined.

Why a publish cluster needs shared memory

A request can reach any publish instance through a load balancer or dispatcher. If a servlet stores a value in a local Java map, the next request may land on another node and miss that value. The result can be repeated authentication calls, inconsistent counters, uneven throttling, or a user experience that changes after every request.

Hazelcast creates a logical data space accessible by each participating JVM. Distributed maps can hold serialized objects, while replicated or near-cache configurations can improve read performance for carefully selected values. Expiration policies prevent temporary records from growing without control, and entry processors can perform certain updates close to the data.

Shared memory is valuable only when the application can tolerate its operational characteristics. Network partitions, member restarts, serialization errors, and eviction can affect availability. A cache entry must therefore be treated as disposable unless the system has a deliberate backup or persistence design.

A practical AEM cluster architecture

A typical arrangement places multiple AEM publish nodes behind a dispatcher and load balancer. Each node runs the AEM application and a Hazelcast client or member, depending on the chosen topology and operational model. The Hazelcast cluster should use explicit discovery, predictable network rules, and settings that distinguish production members from development instances.

Running a full Hazelcast member inside every AEM JVM can reduce network hops, but it also couples application and data-grid resource consumption. A separate Hazelcast cluster offers clearer scaling and isolation, though it introduces additional deployment and monitoring responsibilities. Hazelcast clients can be a useful boundary when the data grid is managed independently from AEM.

The integration layer should be packaged as an OSGi service rather than scattered across servlets and models. That service can manage map access, serialization, time-to-live policies, error handling, and lifecycle events. Consumers then request a domain-level operation such as “get product availability” instead of depending directly on Hazelcast APIs throughout the codebase.

Choosing data ownership and consistency

The central design question is whether Hazelcast owns a value or merely accelerates access to another system. A cache of an external API response can be rebuilt after a restart. A distributed lock, idempotency key, or request token may need stronger coordination. A shopping basket or business transaction usually requires a durable system of record rather than an in-memory grid alone.

Consistency should be matched to the business operation. Eventual consistency may be appropriate for recommendations, catalog enrichment, or analytics hints. Atomic operations are more important for counters, quotas, and deduplication. If two nodes update the same record, use data structures and operations that make the update indivisible instead of reading, modifying, and writing in separate steps.

Avoid placing large AEM repository objects, resource resolvers, sessions, or non-serializable framework objects in a distributed map. Store compact data transfer objects with explicit versioning. Include tenant, locale, and schema information in keys where necessary, and define invalidation behavior when content is activated or an upstream record changes.

Requirement Suitable approach Main caution
Short-lived API response Distributed map with TTL Stale data and cache stampedes
Cross-node request deduplication Atomic key or lock Lock expiry and failure recovery
Shared rate limiting Atomic counter with expiration Clock and burst behavior
Durable business record External database or service Hazelcast should not be the only copy
Local hot reads Near cache or local cache Invalidation and memory pressure
Publish-triggered invalidation Event-driven removal Missed events need reconciliation

Integrating with AEM services

An OSGi component can obtain a Hazelcast instance through a controlled service reference and expose methods for get, put, remove, and atomic update operations. Configuration should be supplied through OSGi factory or metatype definitions, allowing operators to change cluster names, connection settings, TTL values, and feature switches without rebuilding application code.

Servlets, Sling Models, schedulers, and event handlers should call the integration service through narrow interfaces. A scheduler might prewarm product data, while an activation listener removes entries associated with changed content. The listener should avoid assuming that every invalidation succeeds; retries, dead-letter handling, or periodic reconciliation may be required for important caches.

Failure handling needs to preserve AEM’s primary request path. If Hazelcast is unavailable, the application should have a defined fallback: query the source system, serve a safe default, or return a controlled error. Timeouts should be short enough to protect publish throughput, and circuit breakers can prevent a failing grid from consuming all request threads.

For developers reviewing related AEM architecture and deployment practices, the Docker development guidance offers useful context for creating repeatable local environments. A containerized setup can include multiple publish instances and a Hazelcast service, making membership, failover, and cache behavior easier to test before deployment.

Testing behavior across publish nodes

A single-node test proves very little about a distributed cache. Local testing should include at least two publish processes, a dispatcher or load-balancing layer, and a Hazelcast environment that resembles the intended discovery mechanism. Requests should be deliberately routed across nodes to verify that shared values are visible where expected.

Test node departure during reads and writes, delayed network responses, rolling restarts, duplicate activation events, and expired entries. Check that serialization remains compatible when application versions overlap during a deployment. If rolling releases are part of operations, both versions must understand the data formats present in the grid.

Performance testing should measure more than average latency. Track cache hit rate, miss amplification, p95 and p99 response times, serialization cost, heap usage, network traffic, and the time required to recover after a member leaves. A fast cache that causes frequent garbage collection or blocks request threads is not an improvement.

Past conference material can help teams compare implementation ideas with broader AEM practices; the conference session recordings include technical discussions from the event’s AEM-focused program. Use those examples as architectural input, then validate every integration against the specific AEM and Hazelcast versions in use.

Operating the grid in production

Hazelcast health must be visible alongside AEM health. Monitor cluster membership, partition migration, owned and backup entry counts, heap usage, near-cache statistics, operation latency, rejected tasks, and client connection state. Alerts should distinguish a temporary cache miss from a partition loss that threatens application behavior.

Capacity planning should account for both primary and backup data. Estimate the number of entries, serialized size, expiration rate, peak update volume, and the memory required during rebalancing. Keep headroom for deployments and traffic spikes. Eviction should be a deliberate policy based on data value, not a substitute for an undefined retention model.

Security deserves equal attention. Restrict cluster communication to private networks, authenticate clients and members, protect credentials, and apply encryption where required. Never place personal data, tokens, or secrets in a grid without a documented retention and protection strategy. Logs should identify keys or correlation IDs without exposing sensitive values.

AEM upgrades and Hazelcast upgrades should be tested independently and together. Review compatibility matrices, client protocols, serialization behavior, and OSGi package imports. A rolling restart can expose hidden assumptions about membership, cache ownership, or schema compatibility, so the deployment procedure should include a recovery plan.

Implementation priorities for a dependable rollout

Start with one bounded use case whose correctness can be measured, such as caching a read-heavy integration response or coordinating an idempotency key. Avoid migrating every local cache at once. Establish baseline latency, origin-system load, and error rates before introducing the grid.

Use these implementation priorities:

  • Define whether each value is a cache, coordination record, or business record before choosing a data structure.
  • Set explicit TTL, maximum size, serialization, invalidation, and fallback policies for every map.
  • Keep Hazelcast access behind OSGi services with typed interfaces and centralized timeout handling.
  • Test node loss, network interruption, rolling deployment, stale data, and incompatible serialized values.
  • Monitor hit rate, latency, memory, membership, and fallback volume as production service indicators.

A disciplined rollout makes an in-memory data grid a useful extension to an AEM publish tier rather than an invisible source of state. Document the ownership model, rehearse failure scenarios, and confirm that the application remains safe when the grid is empty or unavailable. Then deploy the smallest valuable capability first, measure it across real traffic patterns, and expand only when the operational evidence supports it.