AEM and Kubernetes: Scaling Author and Publish Instances
Adobe Experience Manager (AEM) combines content authoring, asset management, personalisation and publishing in a platform that can support large digital estates. Kubernetes adds a flexible control layer for running those workloads, but placing AEM containers on a cluster is not the same as making the platform cloud-native.
The most effective design separates author and publish responsibilities. Authors need durable storage, predictable application behaviour and carefully managed deployments, while publish tiers need fast horizontal scaling, resilient caching and secure delivery. Treating both environments identically usually creates unnecessary complexity.
For architects and Java engineers assessing this approach, the technical sessions collected on the CIRCUIT conference site provide useful context around AEM integrations, architecture and development practices. The same principles remain relevant when planning an Australian deployment across Sydney, Melbourne, Brisbane or a wider regional audience.
Separate Authoring From Content Delivery
An author instance is a stateful system. It stores pages, metadata, workflows, users, permissions and repository indexes, so scaling it requires more than adding replicas behind a load balancer. AEM authoring typically works best as a carefully sized primary service with persistent repository storage, controlled access and a tested standby or recovery path.
Kubernetes can manage the author environment through StatefulSets, persistent volume claims, secrets and scheduled jobs. This provides repeatable infrastructure, but the repository still needs storage with suitable latency, backup support and failure semantics. A storage class designed for generic application files may perform poorly when repository writes, Lucene indexes and asset operations compete for the same resources.
Publish instances have a different profile. They serve activated content to web visitors and can often scale horizontally because the content is replicated from author and cached by a Dispatcher or content delivery network. A deployment can therefore run several publish pods across availability zones, with readiness probes ensuring that traffic reaches only instances that have completed startup and repository checks.
Australian organisations should also account for geography. A retailer serving customers in Perth, Adelaide and the east coast may need regional caching and carefully selected cloud zones, while a government or healthcare service may require Australian data residency. Kubernetes placement, backup regions and CDN configuration should reflect those obligations rather than being selected solely for lower compute pricing.
Build AEM Containers For Predictable Releases
A reliable AEM container image should be immutable. Application code, OSGi bundles, configuration and required packages should be assembled through a repeatable build pipeline, scanned for vulnerabilities and promoted through environments without manual changes inside running pods. This makes it easier to identify whether a release, configuration update or infrastructure change caused a problem.
Author and publish images may share a base, but they should not be treated as interchangeable. Authoring requires author-specific run modes, development tooling controls and workflow configuration. Publish images should expose only the services required for delivery and should be combined with a hardened Dispatcher layer. Separate images or clearly separated deployment profiles reduce the risk of accidentally exposing author capabilities publicly.
AEM startup can be slow, especially when indexes, bundles or repository migrations are involved. Kubernetes health checks must distinguish between a process that is alive and an instance that is ready for traffic. A liveness probe that restarts a slow-starting pod can create a loop during a busy deployment. Startup probes, generous initial delays and application-level readiness checks are safer choices.
Release automation should include content compatibility checks. Code can be deployed before content activation, but a new component may depend on policies, templates or content structures that do not yet exist. Blue-green or canary strategies are valuable for publish tiers, while author releases generally require a controlled maintenance window and explicit validation of workflows, integrations and permissions.
Scale Publish Capacity Around Real Demand
Horizontal Pod Autoscaling can increase publish capacity when CPU or memory crosses a threshold, but infrastructure metrics alone do not describe the visitor experience. AEM publish pods may be constrained by request latency, garbage collection, repository access, connection pools or Dispatcher cache misses. Useful scaling signals include response time, active requests, error rates and cache-hit ratios.
The cache layer should absorb the majority of repeat traffic. Dispatcher rules need to reflect the site’s invalidation model, query-string behaviour, authentication requirements and personalisation strategy. If every activation clears broad sections of the cache, adding more publish replicas may simply move the bottleneck to origin servers. Targeted invalidation and sensible cache headers usually deliver a larger improvement.
Traffic patterns differ across the Australian market. A national retailer may see sharp increases during Boxing Day promotions, while an education provider may experience predictable peaks around enrolment periods. Media sites can face sudden demand after a major event in Sydney or Melbourne. Load tests should reproduce these patterns, including cache cold starts, asset delivery and activation bursts.
The Kubernetes cluster itself needs room for failure. Use pod anti-affinity or topology spread constraints so all publish replicas do not land on one node. Set resource requests honestly, reserve capacity for rolling updates and use a PodDisruptionBudget to avoid draining too many instances during maintenance. Autoscaling should have sensible upper bounds, because uncontrolled replica growth can overload the repository, network or external services.
Protect Integrations, Storage And Observability
AEM rarely operates alone. Identity providers, product information systems, search platforms, marketing tools, payment services and analytics pipelines may all be called during authoring or delivery. Kubernetes Secrets, external secret managers and network policies can protect credentials, but they do not remove the need for connection timeouts, retry limits and circuit breakers in the application.
Persistent storage requires a recovery design that is tested under pressure. Backups should cover repository data, indexes where appropriate, configuration, Dispatcher rules and deployment manifests. Teams should define recovery point and recovery time objectives, then verify that a restored author environment can publish usable content. A backup that exists but has never been restored is an assumption, not a recovery strategy.
Observability should combine Kubernetes telemetry with AEM-specific signals. Collect container resource usage, pod restarts and node health alongside request latency, replication queues, workflow backlogs, bundle states, repository activity and Dispatcher cache performance. Correlating these signals helps distinguish a failing pod from a slow downstream API or a content activation problem.
Teams attending technical events can also learn from varied engineering perspectives by reviewing the CIRCUIT speakers, particularly when comparing Java, architecture and systems-engineering approaches. In an Australian operation, dashboards should include local business time zones and planned events such as public holidays, end-of-financial-year campaigns and scheduled maintenance windows.
Make Operations Safe For Teams
Kubernetes introduces an operational model that rewards standardisation. Define namespaces, resource quotas, network policies, deployment conventions and ownership boundaries before production. Platform engineers can provide reusable templates, while AEM specialists remain responsible for repository behaviour, application configuration and content delivery rules.
Access to author should normally be private, using corporate identity, VPN or a zero-trust access layer. Publish endpoints need stronger edge protection through a web application firewall, rate limiting, TLS management and DDoS controls. Administrative consoles, debug endpoints and package managers should never be exposed as part of the public delivery path.
Deployment runbooks should describe activation verification, rollback, cache invalidation and integration checks. A rollback of application code may not reverse a repository migration or content change, so each release needs a clear compatibility boundary. Disaster recovery exercises should include loss of a node, loss of a zone, corrupted content and an unavailable external identity provider.
The archived session videos offer a useful way to revisit practical discussions about AEM development and architecture. For current projects, those lessons should be combined with supported product versions, cloud-provider guidance and the organisation’s own service-level requirements.
Signals Worth Tracking
Operational dashboards become more useful when they focus on decisions rather than collecting every available metric.
Publish indicators
- P95 and P99 response latency by URL type
- Dispatcher cache-hit and origin-request rates
- Replication queue depth and activation delay
- Pod restarts, memory pressure and error responses
Author indicators
- Workflow backlog and failed jobs
- Repository response time and storage latency
- Indexing status and long-running queries
- Login failures, author availability and backup status
A scaling policy should be tested with realistic traffic before it is trusted. Measure how quickly new pods become ready, whether caches warm successfully and whether external services can tolerate the increased request rate. Capacity planning should include a failure margin for Australian peak periods, planned campaigns and an entire unavailable node or zone.
The strongest implementation keeps authoring deliberate and durable while making publishing elastic and disposable. With immutable images, appropriate persistence, layered caching, secure integrations and meaningful observability, Kubernetes can provide a dependable operating foundation for AEM rather than another source of hidden failure modes.
Use these principles to assess the current topology, map author and publish dependencies, and test scaling behaviour before the next major campaign. A measured architecture review can turn container adoption into a safer, faster publishing platform for customers across Australia.