Building flexible AEM search facets with Solr and Elasticsearch

Faceted search helps visitors move from a broad content set to a useful answer. Instead of asking users to refine a query repeatedly, an AEM search interface can expose categories, tags, dates, locations, content types, product attributes, or other indexed properties as selectable filters. The result is a faster path through large repositories and a clearer picture of how content is organized.

AEM provides strong foundations for search through repository indexing, QueryBuilder, Oak indexes, Sling Models, and component-based rendering. Yet custom facets often require more control than a standard full-text query can provide. Teams may need external aggregation engines, custom analyzers, multilingual fields, nested attributes, or search behavior that spans AEM and other systems.

Solr and Elasticsearch can both support this architecture, but the search engine should not be selected in isolation. The important decisions involve the index contract, synchronization strategy, query translation, facet presentation, failure handling, and operational ownership. The engineering discussions featured through the CIRCUIT speaker archive reflect the kind of cross-functional thinking required for this work: search sits between application code, infrastructure, content modeling, and user experience.

Why custom facets matter in AEM

A facet is an aggregated view of a field across the current result set. For example, a search for “camera” might return counts for brand, mount type, price range, availability, and publication status. Selecting a value changes the query, while the remaining counts update to reflect the narrower context.

AEM’s built-in repository search can handle many straightforward use cases, especially when properties are indexed consistently and the query requirements are modest. Custom implementations become valuable when the website needs high-volume filtering, complex aggregations, relevance tuning, or a single search experience across AEM content and external catalog data.

The quality of a facet depends on the data behind it. A field must have a predictable type, stable naming, suitable analyzers, and a clear rule for missing or multi-valued values. A display label such as “Mobile Phones” may require a separate identifier, translation key, or taxonomy path rather than relying on raw repository text.

Shape the index before writing queries

The most reliable implementations begin with an index contract. This document defines which AEM properties become searchable fields, which become facet fields, how values are normalized, and how each source maps to a canonical content identifier. It should also describe permissions, publication state, locale, timestamps, and URL metadata.

Facet fields usually need exact-value treatment. A title may use tokenization and stemming for full-text relevance, while a brand or content type should remain intact so “Sony Mobile” is not split into unrelated terms. Dates should be indexed as dates, numbers as numeric values, and hierarchical categories should be represented in a way that supports the intended drill-down behavior.

AEM content is often authored in a flexible structure, so indexing code should account for missing properties, inherited values, repeated fields, and inconsistent legacy content. A normalization layer can convert repository values into a stable search document. That layer is also an appropriate place to add calculated fields such as audience, region, content maturity, or a flattened taxonomy path.

Compare Solr and Elasticsearch for facet workloads

Both engines can return search hits and aggregations, but their APIs and operational models differ. Solr is built around Apache Lucene and commonly uses explicit schemas, managed fields, Solr cores or collections, and a strong tradition of configuration-driven search. Elasticsearch exposes a JSON-oriented document and aggregation model, with mappings and index templates that fit naturally into application-managed workflows.

The best choice depends on existing expertise, deployment standards, query complexity, and the desired pace of schema evolution. Solr can be attractive when a team values explicit field governance and established search administration. Elasticsearch can be convenient when developers want aggregations, document-oriented APIs, and an ecosystem already used for logs, analytics, or distributed data services.

Concern Solr Elasticsearch
Field governance Explicit schema and managed fields Mappings and index templates
Facet and aggregation model Faceting and JSON Facet API patterns Aggregations in the search request
AEM integration style Connector, custom OSGi service, or indexing pipeline Connector, custom OSGi service, or event-driven pipeline
Schema changes Often deliberate and administrator-led Flexible, but mappings still require discipline
Operational focus Collections, shards, replicas, cores Indices, shards, replicas, cluster health
Strong fit Controlled enterprise search platforms Flexible document and aggregation workloads

Neither engine removes the need for careful AEM integration. Sending repository content to an external index introduces consistency, security, and deployment concerns. A successful proof of concept should therefore test indexing throughput, facet accuracy, incremental updates, deletes, permissions, and recovery rather than measuring only response time.

Connect AEM to the search engine

A common architecture separates four responsibilities: content extraction, index writing, query execution, and result presentation. An OSGi service can transform AEM pages or assets into search documents, while a scheduled job, replication event, workflow step, or message queue triggers updates. The query service translates application filters into Solr or Elasticsearch syntax and returns a domain-level result object to the component.

The query layer should avoid exposing engine-specific request structures to HTL templates or front-end code. Instead, define an application model containing hits, facet groups, selected filters, counts, pagination data, and sort options. This makes the component easier to test and leaves room to change the backend without rewriting the presentation layer.

Incremental indexing is usually preferable to rebuilding everything after every content change. However, a full rebuild remains important for schema changes, corrupted indexes, migration projects, and changes to calculated fields. The process should support aliases or versioned indexes so a new index can be populated and validated before traffic switches to it.

A search result must also respect publication and access rules. Filtering unpublished content at query time is essential for public delivery, while authoring environments may need a different endpoint or permission-aware behavior. If users can search protected content, document-level security must be designed explicitly rather than assumed from AEM repository permissions.

Design facet behavior for real users

A technically correct aggregation can still create a poor experience. Facets should be ordered by business relevance, alphabetically, or by count according to the use case. Counts should communicate the effect of a selection, and selected values should remain visible even when their current count is small. A clear reset action prevents users from becoming trapped in a narrow result set.

High-cardinality fields require special care. Author names, SKUs, timestamps, and arbitrary tags can produce unwieldy facet lists and expensive aggregations. These fields may need search-as-you-type behavior, a limited result window, range buckets, or a different interaction pattern. Numeric and date facets are generally more useful as meaningful ranges than as thousands of individual values.

The query should preserve the distinction between filters that narrow independently and filters that represent alternatives. Selecting several brands typically means “brand A or brand B,” while selecting a brand and a price range means “brand A or B, within this range.” The backend request must express those boolean rules consistently, including when filters are removed or combined with free-text terms.

Performance testing should include realistic combinations rather than a single popular query. Measure cold and warm searches, aggregation-heavy requests, deep pagination, large result sets, and concurrent traffic. Cache stable taxonomy data where appropriate, but avoid caching responses that can expose stale publication status or incorrect permission results.

Operate and monitor the indexing pipeline

Search reliability depends on more than engine availability. AEM authors may publish content successfully while an indexing consumer is delayed, a mapping rejects a document, or a delete event is lost. Track indexing lag, queue depth, failed documents, retry counts, document totals, and the time between publication and search visibility.

Operational visibility should cover both AEM and the search cluster. JMX-based health checks can expose service state, job metrics, thread pools, and queue behavior; the practical value of this approach is illustrated by the guidance on AEM health checks. Alerts should distinguish a temporary latency spike from a sustained synchronization failure.

Use correlation identifiers across authoring events, indexing requests, and search requests. When a page is missing from results, support engineers should be able to determine whether the content was unpublished, rejected during transformation, blocked by permissions, delayed in a queue, or absent from the target index.

Reindexing should be treated as a controlled operational process. Define who can start it, how progress is measured, how the active index is protected, and how the team verifies counts and sample documents before switching traffic. This prevents a routine schema update from becoming an unexpected outage.

Practical implementation recommendations

A durable AEM facet solution benefits from a small set of explicit engineering rules:

  • Keep a versioned index contract for fields, types, analyzers, taxonomy values, and permissions.
  • Separate AEM content transformation from engine-specific request and response handling.
  • Use stable content identifiers and process creates, updates, moves, and deletes as distinct events.
  • Test facet counts against curated content fixtures, including missing, repeated, localized, and deprecated values.
  • Monitor indexing lag and failed documents alongside query latency and search error rates.

Start with one representative search journey rather than indexing every property in the repository. Select a content type with meaningful filters, define the expected user behavior, and measure the complete path from publication to visible facet count. Once that workflow is reliable, expand the index contract based on evidence instead of speculation.

Move from prototype to production search

Building custom search facets with Solr or Elasticsearch is an architecture exercise as much as a query-writing task. The strongest implementations align AEM content modeling, external indexing, aggregation semantics, front-end interaction, and operational monitoring from the beginning. They also make the boundary between repository search and external search explicit, so future teams can maintain the system without reverse-engineering hidden assumptions.

Document the field mappings, event flow, security model, reindex procedure, and failure alerts alongside the code. Then validate the experience with authors, developers, and infrastructure owners before launch. Use the CIRCUIT resources as a technical reference point while turning the design into a tested AEM search capability that can grow with the content platform.