AEM Oak Persistence: TarMK or DocumentMK?
Choosing between AEM Oak TarMK and DocumentMK affects performance, availability, scaling, operations, and the shape of an Adobe Experience Manager deployment. The decision is less about selecting a universally superior repository and more about matching persistence to authoring patterns, publishing volume, infrastructure skills, and recovery expectations.
For Australian teams, the context can be especially practical. A platform serving editors in Sydney, Melbourne, Brisbane, and Perth may need predictable latency across regions, clear data-residency decisions, and an operating model that works with local cloud capacity and support coverage. Understanding Oak’s storage options helps architects avoid an expensive platform choice made on assumptions rather than workload evidence.
What Oak persistence controls
Apache Jackrabbit Oak is the repository layer beneath AEM. It stores content, configuration, permissions, workflows, tags, assets, and other repository-managed data. TarMK and DocumentMK are different storage approaches within Oak, so the choice influences how AEM nodes share state and how the repository grows.
TarMK stores repository content in segment tar files on local disk. It is generally the standard choice for a single AEM instance or a farm where each publish instance has its own repository. DocumentMK stores data as documents in a MongoDB or relational database backend, allowing several Oak instances to participate in a clustered deployment.
This distinction should be separated from the dispatcher, CDN, binary data store, and application server. A fast CDN cannot compensate for a repository that is poorly sized, and a database-backed repository does not automatically make every AEM topology highly available.
Why TarMK remains the default
TarMK has a comparatively simple architecture. The repository runs against local storage, which avoids network round trips for most content operations and keeps the dependency list relatively small. With suitable SSDs, memory, and repository maintenance, it can deliver strong performance for authoring and publishing workloads.
The model is particularly comfortable for AEM author environments, development systems, and publish farms. A common pattern is one TarMK author instance with a standby or backup strategy, plus multiple independent TarMK publish instances behind a dispatcher. The publish nodes do not need to share a live repository because content activation distributes changes to each one.
Operational simplicity is a major advantage. Teams can focus on repository sizing, compaction, backups, and monitoring instead of running a database cluster. For a mid-sized Australian organisation with a small platform team, that reduced moving-parts count can matter more than theoretical horizontal scale.
TarMK is not a reason to ignore resilience. Local storage failure, corruption, an unsuccessful compaction, or an unavailable author instance can still affect delivery. Cold standby, reliable backups, tested restores, and documented recovery time objectives are essential parts of a TarMK design.
Where DocumentMK earns its place
DocumentMK is designed for clustered Oak deployments. Several AEM instances can use a common document store, with MongoDB or a supported relational database providing the shared persistence layer. This can support authoring high availability, where another node can continue serving editors if one node fails.
That capability comes at a cost. Network latency between AEM nodes and the database becomes significant, and database transactions, indexes, connection pools, storage performance, and cluster health all become part of the AEM operating picture. A DocumentMK deployment must be engineered as a complete system rather than treated as “TarMK with MongoDB added”.
DocumentMK is most compelling when the business needs active-active authoring, a larger concurrent editor population, or a topology that cannot rely on a single primary author. It may also suit organisations with mature database operations, established MongoDB expertise, and strong automation around failover and observability.
The database should be located close to the AEM nodes. Placing an author node in Sydney and its persistence database in another distant region may introduce latency that overwhelms the advantages of clustering. Cross-region designs require careful testing, supported configurations, and a clear view of consistency and disaster recovery.
Comparing the two persistence models
The right choice depends on the failure model and workload as much as on content volume. A large media library does not automatically require DocumentMK, because binary assets can use a separate data store and publishing can be scaled independently. Conversely, a modest repository may justify DocumentMK if authoring availability is a business-critical requirement.
The following comparison provides a starting point, not a substitute for load testing and an assessment against the specific AEM release and Adobe support matrix.
| Consideration | TarMK | DocumentMK |
|---|---|---|
| Primary storage | Local segment tar files | Shared document store |
| Typical topology | Single author or independent publish farm | Clustered author or shared Oak instances |
| Main strength | Low complexity and strong local performance | Higher authoring availability and horizontal clustering |
| Infrastructure burden | Storage, backup, compaction, monitoring | Database cluster, network, indexes, backups, monitoring |
| Sensitivity | Local disk and instance capacity | Network latency and database health |
| Common fit | Most standard author and publish deployments | Large or highly available clustered deployments |
| Scaling approach | Scale up, add publish nodes, tune topology | Scale cluster within supported architecture |
| Key risk | Single-author interruption without standby | Distributed-system complexity and database contention |
Match persistence to AEM topology
Start with the author tier. If authors work through a single primary instance and an outage can be managed through standby recovery, TarMK is often an efficient fit. If the author service must remain available during node failure, DocumentMK may be appropriate, provided the cluster is genuinely supported and properly operated.
The publish tier follows a different pattern. AEM publish nodes commonly remain independent, with dispatcher and CDN layers distributing requests. This means adding publish capacity does not necessarily require a shared DocumentMK repository. Teams sometimes overengineer the entire platform because they need scale at the edge, when the actual solution is more publish instances, better caching, and a resilient activation process.
Consider authoring behaviour rather than page count alone. Frequent asset ingestion, large package installations, extensive workflows, content fragment operations, and simultaneous editorial activity can create repository and CPU pressure. Measure writes, reads, session counts, workflow queues, replication rates, and activation delays before selecting a persistence model.
A site serving customers around Australia may need a different arrangement from an internal intranet. For example, Melbourne-based authors and a Perth operations team may experience acceptable authoring latency within one well-designed region, while customer traffic is handled globally by the CDN. Keeping those concerns separate usually produces a cleaner architecture.
Account for Australian operating realities
Australian organisations often weigh data sovereignty and regulated workloads alongside technical performance. A team may prefer an Australian cloud region for customer, health, financial, or government-related content, while still using a managed service or support team located overseas. The repository decision should document where content, backups, logs, and database replicas reside.
Network distance also matters in a geographically wide country. Sydney and Melbourne are well connected, but a design that relies on constant repository communication should still be tested from the actual office, VPN, and cloud locations. A “quick look” from a local laptop does not represent production behaviour for editors on a constrained corporate link in Adelaide or Darwin.
Operational staffing is another factor. A DocumentMK platform may be entirely reasonable for a large enterprise with database administrators, 24-hour monitoring, and tested incident procedures. A smaller team may sensibly favour TarMK and invest the saved effort in backups, automation, dispatcher configuration, and recovery drills. In Australian workplace language, the simplest option that meets the SLA is often the better fit, rather than the flashiest architecture.
Teams attending technical events such as CIRCUIT can use the CIRCUIT registration details to find sessions and recordings that add implementation context. Discussions with Java developers, AEM architects, and systems engineers are particularly useful when a design needs to balance product guidance with local hosting and support constraints.
Plan capacity, backup, and recovery
TarMK capacity planning should cover repository growth, segment store performance, compaction overhead, temporary disk space, and backup windows. SSD performance and available memory can have a noticeable effect on repository operations. A backup is valuable only when restore procedures are tested against a realistic outage and the recovered instance passes repository and application checks.
DocumentMK capacity planning includes the Oak nodes and the database layer. Teams need to monitor database latency, disk utilisation, replication or replica-set health, connection pools, query patterns, checkpoints, and cluster membership. Database backups must be coordinated with AEM recovery requirements, including binaries and externalised data where applicable.
Recovery objectives should drive the design. A business asking for a five-minute recovery point and near-continuous authoring availability has a different requirement from one that can restore an author instance overnight. Write down the acceptable data loss, recovery time, failover process, and ownership of each action before choosing a repository.
Also review the AEM version, cumulative fixes, storage configuration, and Adobe documentation. Supported combinations change, and a design that worked for an older release may not be suitable for a newer deployment or managed AEM offering.
Validate the choice before migration
Build a representative test rather than comparing empty repositories. Include page creation, asset uploads, package installation, workflow execution, queries, permissions, replication, indexing, and concurrent editor activity. For DocumentMK, introduce realistic database latency and failover conditions; for TarMK, test compaction, disk pressure, restore time, and standby promotion.
Measure user-facing outcomes as well as infrastructure metrics. Editors care about page activation and asset processing, while customers care about response time and cache effectiveness. Record repository response times, CPU, memory, disk I/O, database latency, error rates, queue depth, and recovery duration across normal and peak periods.
Migration planning should include repository compatibility, indexes, binaries, custom code, workflows, integrations, and rollback. Moving from TarMK to DocumentMK is not simply a storage switch, and moving in the other direction can be equally disruptive. Use a rehearsal environment and define a cutover window that respects Australian business hours and support coverage across time zones.
A sound decision is usually evidence-led: establish the topology, document the SLA, test realistic load, and choose the model your team can operate consistently. TarMK is often the practical baseline; DocumentMK is a deliberate choice for clustering and availability needs that justify its added complexity.
Review the event FAQ for practical details about CIRCUIT resources, then use the conference recordings and architecture discussions to deepen your AEM planning. Bring the comparison into your next design review, validate it against your repository workload, and make persistence a measured platform decision rather than an assumption.