Custom AEM Replication Agents for Publish Farm Synchronization

Adobe Experience Manager replication is often presented as a simple author-to-publish workflow: an editor activates content, and a replication agent delivers it. That model works for a small environment, but a production publish farm introduces questions about ordering, retries, cache invalidation, topology changes, and partial delivery.

A custom replication design can address those concerns without turning every publication event into a bespoke application. The strongest implementations preserve AEM’s familiar authoring workflow while extending transport, filtering, observability, or coordination where the standard agents no longer fit.

These concerns align closely with the engineering subjects covered by the CIRCUIT conference, including AEM architecture, integrations, Java development, microservices, and systems engineering. Understanding the platform’s replication model is the foundation for making a custom solution reliable.

Why Standard Replication Agents Need Extension

AEM’s standard replication agent can serialize repository content, send it through an HTTP-based transport, and process common activation, deactivation, and deletion commands. Its configuration covers essential settings such as the target endpoint, authentication, trigger behavior, queue processing, and retry intervals. For a straightforward author-publish connection, this is usually enough.

A publish farm changes the problem. Several publish instances may sit behind a load balancer, while a dispatcher or content delivery network retains cached responses. A single activation must therefore reach every required target or be routed through a mechanism that guarantees equivalent state. If one instance is unavailable, the system must distinguish a temporary transport failure from a permanent content or permission error.

Custom behavior is useful when delivery requires tenant-aware routing, signed requests, an internal message broker, content transformation, or a coordinated cache flush. It can also provide richer audit records than the default queue exposes. The goal should be a narrow extension around a stable contract, rather than a replacement for AEM’s activation lifecycle.

Designing The Agent Boundary

The first design decision is where customization belongs. A custom transport handler is appropriate when the content package and replication command are valid but the delivery mechanism differs from the default HTTP transport. A custom replication process may be preferable when a workflow needs to invoke a separate service, publish metadata, or apply controlled filtering before transport.

The extension should remain an OSGi service with explicit configuration and clear dependency boundaries. Keep repository reads, serialization, network delivery, and acknowledgement handling separate where possible. This makes unit testing easier and prevents a slow remote system from consuming unbounded AEM worker resources.

A useful agent configuration includes a stable identifier, target type, endpoint or broker destination, credentials managed outside source code, connection and read timeouts, retry limits, and queue throttling. It should also declare whether delivery is synchronous or asynchronous. Avoid placing business rules inside a servlet or workflow step when the behavior is fundamentally part of replication; doing so can create competing queues and confusing activation results.

Coordinating A Publish Farm

There are two common delivery models. In direct fan-out, the author environment maintains a replication agent for each publish node. This provides direct visibility into every queue, but the number of agents grows with the farm and operational management becomes heavier. In a brokered model, AEM sends one event to a durable intermediary, and a delivery service distributes it to publish instances.

Direct fan-out is easier to understand in a smaller deployment. Brokered distribution can improve resilience and support rolling infrastructure changes, but it introduces another system that must preserve ordering, deduplicate events, and expose delivery status. Neither pattern is automatically consistent; each needs an explicit definition of what “published” means.

Many installations do not require every event to arrive in strict global order. They do require ordering for the same resource, such as a page activation followed by an image update or deletion. Use a resource path, content identifier, or aggregate key as the partition key. This allows unrelated content to move in parallel while preserving causal order for dependent updates.

Comparing Delivery Patterns

The right pattern depends on farm size, failure tolerance, and operational maturity. A decision should account for whether each publish node must be independently verified, whether the organization already operates a message platform, and how quickly stale content must disappear.

Delivery pattern Strengths Main risks Suitable use
Direct agent per publish node Clear target-level status and simple topology Configuration grows with the farm; partial failure needs careful handling Small or medium farms with stable nodes
Shared load-balanced endpoint Fewer agent configurations and easy node replacement Acknowledgement may confirm only one node; hidden skew is possible Farms where a downstream synchronizer verifies fan-out
Durable message broker Strong buffering, scalable consumers, and decoupled deployment Additional operations, deduplication, and ordering responsibilities Large environments with frequent infrastructure changes
Periodic repository comparison Detects drift and repairs missed updates Slower correction and higher repository-read cost Audit, recovery, and disaster-repair processes
Hybrid replication plus reconciliation Fast normal delivery with a safety net More moving parts and monitoring requirements Business-critical sites with strict freshness targets

A reconciliation process should not be treated as a substitute for delivery acknowledgements. It is a repair mechanism that compares expected state with observed state, identifies missing or stale paths, and safely replays eligible operations. Designing this from the beginning is less expensive than reconstructing it after a production incident.

Making Delivery Idempotent And Observable

Retries are unavoidable. A network timeout can occur after the target has stored a package but before the source receives the response. If the retry creates a duplicate side effect, a temporary fault becomes a data integrity problem. Each operation should therefore carry an event identifier, content version, or equivalent idempotency key that the receiving side can recognize.

The receiver should validate the operation, record its processing result, and return an unambiguous response. “Accepted for processing” is different from “applied successfully,” so those states should not be collapsed in logs or dashboards. If a broker is used, the consumer should acknowledge a message only after the target action and durable result recording have completed.

Monitoring should expose queue depth, oldest queued item, retry count, delivery latency, failure category, and target-level freshness. Correlation IDs should follow an activation from authoring through serialization, transport, publish processing, and cache invalidation. Structured logs make it possible to investigate a single page without searching through unrelated Java stack traces.

Testing Failure And Recovery Paths

A replication agent is incomplete until its failure behavior has been tested. Simulate a stopped publish instance, expired credentials, connection resets, slow responses, malformed payloads, and a full target disk. Verify that transient errors retry with bounded backoff, while permanent errors move to a visible dead-letter or blocked state rather than cycling indefinitely.

Test ordering with related page, asset, and deletion events. Test duplicate delivery by replaying the same event identifier. Test a farm node that returns success while its local repository update is incomplete, and confirm that downstream verification can detect the discrepancy. Performance tests should include activation bursts, large assets, and a growing queue, not just isolated requests.

A useful recovery exercise starts with one target falling behind and ends with a controlled resynchronization. The procedure should specify how to pause or drain a queue, identify the missing range, replay operations, clear affected dispatcher entries, and confirm parity. Documenting this path also helps support engineers respond consistently during releases and infrastructure maintenance.

Operational Practices For Sustainable Synchronization

Custom replication is a platform capability, so it needs ownership beyond the initial Java implementation. Keep configuration versioned, separate secrets from code, and record topology changes as deployment events. Review permissions for service users carefully: the agent should have only the repository and transport access required for its role.

Use the event agenda as a useful reference for the wider AEM engineering context, especially when planning discussions around architecture, integrations, and operations. Replication decisions affect authoring, deployment, caching, and monitoring, so they should be reviewed across those teams rather than treated as an isolated AEM configuration task.

Practical operating recommendations include:

  • Define whether success means queue acceptance, target application, or farm-wide convergence.
  • Use idempotency keys and resource-based ordering for every custom delivery path.
  • Monitor each publish target separately, even when traffic passes through a shared endpoint.
  • Add reconciliation and replay procedures before enabling production activation.
  • Load-test queue behavior under bursts, target outages, and large binary transfers.

A well-designed custom agent makes the standard AEM publishing experience more dependable without hiding the complexity of a distributed system. Teams can study related conference recordings and event resources through the mobile app, then apply those architectural lessons to their own author-publish topology.

Treat replication as an observable, recoverable delivery pipeline rather than a single HTTP request. Define the contract, test its failure modes, and document the recovery path before the publish farm becomes business-critical. With those practices in place, custom synchronization can provide predictable freshness, controlled scaling, and a safer foundation for future AEM integrations.