Adobe AEM

AEM Search Optimization & Custom Oak Indexes Complete Guide

28 min read

The definitive, deep-dive guide to AEM search optimization, custom Oak Lucene and property indexes, fixing traversal warnings, reading Explain Query output, and deploying indexes to AEM as a Cloud Service.

AEMOakSearchLucenePerformanceCloud ServiceReference
AEM Search Optimization & Custom Oak Indexes Complete Guide

It was 2:00 PM on a Tuesday, black Friday was exactly three weeks away, and our AEM author instance had completely locked up. The JVM was pinned at 100% CPU, garbage collection was thrashing, and authors were seeing terrifying 504 Gateway Timeout screens. After a frantic thread dump analysis, we found the culprit: a seemingly innocent component that queried for the latest news articles to display in a sidebar. The query had no path restriction, no node type specified, and was trying to order by a custom date property that wasn't indexed. The Oak engine had quietly fallen back to traversing the entire repository. The repository had just crossed 1.5 million nodes, and the query was pulling the whole hierarchy into memory. In production enterprise environments, mastering Oak indexes is not optional—it is the literal line between a sub-100ms response time and a catastrophic outage.

Most teams get this wrong because they treat search in Adobe Experience Manager (AEM) as magic. They write a QueryBuilder statement, it returns results locally on a small dataset, and they push it to production. But understanding what happens between a QueryBuilder map and the actual disk read is the difference between a junior developer and a Staff-level AEM architect.

In this exhaustive guide, we are going to tear down the AEM search infrastructure and rebuild it. We cover exactly how AEM search works under the hood, traversing the path from query execution to Index evaluation. We will dissect the out-of-the-box indexes (like damAssetLucene and cqPageLucene) and determine exactly when to extend them versus when to build from scratch. We will build custom Lucene and Property indexes node-by-node under /oak:index, looking at real JCR XML definitions for different use cases. We will dive deep into index rules, path exclusions, full-text aggregation, and cost overrides. Finally, we will cover the critical differences in deploying custom indexes to AEM as a Cloud Service versus AEM 6.5, how the asynchronous indexing lane actually works, what happens when it falls behind, how to banish the dreaded TraversalWarning with real error messages and fixes, and out-of-band re-indexing strategies.

Before diving into indexes, make sure you understand how queries are constructed and how the underlying repository works. I highly recommend reviewing the Query Builder API Complete Reference and the JCR & Oak Repository Complete Guide. If you are troubleshooting an active issue, keep the AEM Performance Troubleshooting Complete Guide open in another tab. For deployment nuances on the latest architecture, see the AEM Cloud Service Complete Guide, and for a broader backend perspective, reference the AEM Backend Development Complete Guide.

How AEM search works under the hood

When you execute a query in AEM—whether via the QueryBuilder API, the JCR QueryManager, or a raw SQL2 statement—the query does not immediately start scanning the repository. The underlying Apache Jackrabbit Oak repository engine routes the request through a sophisticated Query Engine.

Every query you write using QueryBuilder is ultimately compiled down into JCR-SQL2 or XPath (though XPath is technically legacy, Oak still supports it heavily internally and it often yields better performance).

Once the query is parsed, Oak enters the Query Evaluation Phase.

Oak maintains a list of all index definitions located at /oak:index. When a query arrives, Oak does not just pick the first index it finds. It asks every single index in the repository: "How much would it cost you to execute this query?"

This is called Cost Estimation.

Each index looks at the query and calculates a "Cost" based on:

  1. Does the index cover the requested node type?
  2. Does the index cover the requested properties?
  3. Does the index cover the requested path restriction?
  4. How many nodes does the index currently contain (size)?

The index that returns the lowest cost wins and is selected to execute the query.

If no index returns a reasonable cost, Oak falls back to the absolute worst-case scenario: Traversal. Traversal means Oak will start at the root path of the query (or / if none is provided) and literally walk down the tree, checking every single node one by one to see if it matches. If your repository has 500,000 nodes, traversal will kill your instance.

The Oak Query Engine execution flow

Let's visualize this flow to understand exactly where performance bottlenecks happen.

+---------------------+
|   Query Execution   |  (QueryBuilder, JCR-SQL2, XPath)
+----------+----------+
           |
           v
+---------------------+
|    Query Parser     |  (Validates syntax, extracts constraints)
+----------+----------+
           |
           v
+---------------------+
|   Cost Evaluator    |  (Asks all /oak:index definitions for a cost)
+----------+----------+
           |
           +-----------------------------+
           |                             |
    [Index Found]                 [No Index Found]
           |                             |
           v                             v
+---------------------+       +---------------------+
|  Index Execution    |       |     Traversal       |
| (Fast, Sub-100ms)   |       | (Slow, OOM Errors)  |
+---------------------+       +---------------------+
           |
           v
+---------------------+
|   ACL Evaluation    |  (Checks permissions post-query)
+---------------------+
           |
           v
+---------------------+
|   Result Returned   |
+---------------------+

Notice the ACL Evaluation step at the bottom. Oak indexes do not inherently know about AEM permissions (Access Control Lists). The index might find 1,000 matching nodes, but before returning them to the user, Oak must check if the requesting session actually has jcr:read access to those nodes. This means a query that returns huge results but where the user has limited access can still be slow because of post-query ACL filtering. This is a critical edge case we will discuss later.

Mastering the Explain Query tool

Before you write a single line of XML for a custom index, you must know how to read the Explain Query tool. Without this, you are flying blind.

In AEM, navigate to Tools > Operations > Diagnosis > Query Performance (or /libs/granite/operations/content/diagnosis/tool.html/granite_queryperformance).

Here, you can paste an XPath or SQL2 query and click "Explain". AEM will not actually execute the full query to return data; instead, it outputs the execution plan.

A healthy execution plan looks like this:

[cq:Page] as [a] /* lucene:cqPageLucene(/oak:index/cqPageLucene) +jcr:content/cq:template:[my-template] */

This tells you:

  1. Oak identified the node type as cq:Page.
  2. Oak selected the cqPageLucene index.
  3. Oak is using the property jcr:content/cq:template to filter the results.

A dangerous execution plan looks like this:

[nt:unstructured] as [a] /* traverse "/content/my-site//*" where ([a].[myCustomProperty] = 'value') */

The word traverse means no index picked up your query. Oak is walking the tree. If you see this in production, it is an immediate priority-one issue.

When you see a query failing to use an index, your first step is not to immediately build a new one. Your first step is to figure out why the existing indexes ignored your query. Usually, it's because you are missing a node type (querying nt:base instead of cq:Page) or missing a path restriction.

OOTB indexes: damAssetLucene and cqPageLucene

Adobe ships AEM with several pre-configured Lucene indexes out of the box (OOTB). The two most important ones that handle 90% of your day-to-day authoring and publishing queries are cqPageLucene and damAssetLucene.

cqPageLucene

Located at /oak:index/cqPageLucene, this index handles searches across the site hierarchy. It is explicitly bound to the cq:Page node type. If you look at its index definition, you will see that it inherently indexes properties like jcr:title, jcr:description, and cq:tags. It also has aggregation rules configured so that full-text searches against a page actually search the text inside the page's components.

damAssetLucene

Located at /oak:index/damAssetLucene, this index handles everything in the Digital Asset Manager (DAM). It is bound to dam:Asset. It extensively indexes metadata extracted during asset ingestion (like EXIF data for images or text extracted from PDFs).

ntBaseLucene

Located at /oak:index/ntBaseLucene, this is a massive, highly restricted index. It attempts to index all nodes (nt:base) but only specific properties. AEM uses this internally for references and structural searches. Never attempt to customize or rely on ntBaseLucene for your custom application queries.

The golden rule: Extend vs Create New

A very common dilemma AEM architects face is this: I added a custom property myStatus to my Pages. Should I extend cqPageLucene to index it, or create a brand new custom index?

The rule of thumb:

  1. Extend OOTB indexes ONLY when you need your custom property to participate in AEM's OOTB Authoring searches (e.g., you want authors to be able to search for myStatus in the Omnisearch bar in the AEM UI).
  2. Create New custom indexes for your application-specific queries (e.g., a query executed by your custom OSGi service to render a news feed on the frontend).

Most teams fail because they aggressively modify cqPageLucene and damAssetLucene to support frontend component queries. This bloats the OOTB indexes, making AEM's internal operations sluggish, and risks breaking the Authoring UI during AEM upgrades.

If you must extend an OOTB index in AEM 6.5, you copy the node from /libs/settings/oak:index to /apps/settings/oak:index (or directly modify /oak:index depending on your version), add your properties, and set reindex=true. In AEM as a Cloud Service, extending OOTB indexes requires creating a strict overlay with a -custom-1 suffix, which we will cover in the Cloud Service section.

Generally, for application data, always build a custom, narrowly-scoped index.

Oak provides several types of indexes. Understanding which one to use is critical for performance and disk space management.

1. The Property Index

A property index (type="property") is synchronous and extremely fast. It does not use the Lucene library. It simply creates a B-tree in the repository mapping a specific property value to a node path.

When to use a Property Index:

  • The query is an exact match (=).
  • The query involves only one or two properties.
  • You need the data immediately after it is written (synchronous).
  • The dataset is relatively small and bounded.

When NOT to use a Property Index:

  • You need full-text search (contains()).
  • You need complex sorting (orderby).
  • You are indexing millions of highly varied nodes.

Because property indexes are synchronous, every time a node is saved with that property, the thread saving the node must wait for the index to update. If you create a property index on jcr:lastModified (a property that changes constantly on almost every node), you will crash the repository with write-contention lockups.

2. The Lucene Index

A Lucene index (type="lucene") is asynchronous. It uses the Apache Lucene library. When a node is saved, Oak puts the indexing task into an asynchronous lane (a background thread). The node is saved instantly, and milliseconds later, the indexer picks it up and updates the Lucene files on disk.

When to use a Lucene Index:

  • You need full-text search capabilities.
  • You need complex range queries (>, <).
  • You need robust sorting capabilities.
  • You are indexing multiple properties across large datasets.

The vast majority of your custom indexes will be Lucene indexes. Because they are asynchronous, they do not block authoring write operations, making them safe for heavily modified properties.

3. Elastic Search (AEM as a Cloud Service Only)

In AEM as a Cloud Service, Adobe introduced Elastic Search indexes (type="elasticsearch"). This moves the index completely off the AEM JVM and into a remote Elastic Search cluster managed by Adobe.

Currently, Adobe relies on Elastic Search primarily for the Authoring UI (Omnisearch), while custom application indexes remain Lucene-based. You will generally still write Lucene definitions for your custom code in AEM CS, but the underlying engine may route them differently. Always stick to the standard Oak Lucene definitions unless explicitly building against the advanced Elastic Search APIs.

The anatomy of a custom Lucene index definition

Let's look at a production-grade custom Lucene index definition. Imagine we have a custom node type (or a standard nt:unstructured node) under /var/commerce/products that holds product data. We query it heavily on the frontend.

Here is the precise XML structure you need. We'll break down every property below.

<?xml version="1.0" encoding="UTF-8"?>
<jcr:root xmlns:jcr="http://www.jcp.org/jcr/1.0" xmlns:nt="http://www.jcp.org/jcr/nt/1.0" xmlns:oak="http://jackrabbit.apache.org/oak/ns/1.0"
    jcr:primaryType="oak:QueryIndexDefinition"
    compatVersion="{Long}2"
    type="lucene"
    async="async"
    evaluatePathRestrictions="{Boolean}true"
    includedPaths="[/var/commerce/products]"
    queryPaths="[/var/commerce/products]">
    <indexRules jcr:primaryType="nt:unstructured">
        <nt:unstructured jcr:primaryType="nt:unstructured">
            <properties jcr:primaryType="nt:unstructured">
                <sku
                    jcr:primaryType="nt:unstructured"
                    name="sku"
                    propertyIndex="{Boolean}true"
                    notNullCheckEnabled="{Boolean}true"/>
                <category
                    jcr:primaryType="nt:unstructured"
                    name="category"
                    propertyIndex="{Boolean}true"/>
                <price
                    jcr:primaryType="nt:unstructured"
                    name="price"
                    ordered="{Boolean}true"
                    type="Double"/>
                <description
                    jcr:primaryType="nt:unstructured"
                    name="description"
                    analyzed="{Boolean}true"
                    useInSuggest="{Boolean}true"/>
            </properties>
        </nt:unstructured>
    </indexRules>
</jcr:root>

Breaking down the Root Definition

  • compatVersion="{Long}2": Always set this to 2. Version 1 is legacy.
  • type="lucene": Specifies the engine.
  • async="async": Critical. This puts the indexing on the background thread. If you forget this on a Lucene index, AEM will try to build Lucene synchronously and likely freeze during large write operations.
  • evaluatePathRestrictions="{Boolean}true": This tells Oak that this index cares about paths. If your query says path=/var/commerce/products, this flag ensures Oak uses the path to calculate cost.
  • includedPaths and queryPaths: These restrict the index to only build for nodes under /var/commerce/products, and only respond to queries targeting that path. This prevents index bloat.

Breaking down the Index Rules

Under <indexRules>, you define exactly which node types this index applies to. In this case, we are indexing nt:unstructured. If you are querying cq:Page, your index rule must be named cq:Page.

Inside <properties>, you map the specific JCR properties you want indexed.

Let's look at the property-level flags:

Property FlagTypeDescription
nameStringThe relative path to the property from the node type. E.g., for a cq:Page, the title is at jcr:content/jcr:title.
propertyIndexBooleanSet to true to enable exact matching (=) for this property. If false, you cannot query WHERE sku = '123'.
notNullCheckEnabledBooleanAllows you to query for the existence of a property (e.g., WHERE sku IS NOT NULL). Without this, null checks fail to use the index.
orderedBooleanSet to true if your query uses ORDER BY on this property. Without it, sorting happens in memory and is incredibly slow.
typeStringExplicitly define the property type (e.g., Double, Date). Crucial for range queries (price > 10.0) to work mathematically instead of alphabetically.
analyzedBooleanSet to true for full-text search. This breaks the text into tokens so you can use CONTAINS(description, 'apple').
useInSuggestBooleanUsed for search suggestions/autocomplete features.

Cost Overrides: Forcing Oak's Hand

Sometimes, Oak stubbornly refuses to use your perfectly crafted custom index. You will see in the Explain Query tool that it picked cqPageLucene instead. This happens because Oak estimates the cost of cqPageLucene to be lower than your custom index, often due to node count estimations.

You can force Oak's hand by applying a cost multiplier.

At the root of your index definition, you can add costPerExecution="{Double}0.1" or costPerEntry="{Double}0.1". This artificially lowers the calculated cost during the evaluation phase, heavily biasing Oak toward selecting your custom index. Use this sparingly—if you have to force Oak, it might mean your query lacks a specific path or node type restriction that would naturally isolate your index.

Aggregation and Full-Text Search Configuration

When you run a full-text search on a cq:Page (e.g., fulltext=marketing), you don't just want to search the jcr:content node itself. You want to search the text within all the components on that page (text components, titles, accordions, custom components).

This requires Aggregation. You must tell Lucene to pull text from descendant nodes into the parent node's index entry.

Under your index rule, you add an <aggregates> node. Here is how you aggregate component text up to the page level:

<aggregates jcr:primaryType="nt:unstructured">
    <cq:Page jcr:primaryType="nt:unstructured">
        <include0
            jcr:primaryType="nt:unstructured"
            path="jcr:content"/>
        <include1
            jcr:primaryType="nt:unstructured"
            path="jcr:content/*"/>
        <include2
            jcr:primaryType="nt:unstructured"
            path="jcr:content/*/*"/>
        <include3
            jcr:primaryType="nt:unstructured"
            path="jcr:content/*/*/*"/>
    </cq:Page>
</aggregates>

This configuration aggregates the text of all child components up to three levels deep under jcr:content.

The Gotcha: Be extremely careful here. Over-aggregation leads to massive index bloat on disk. If you set aggregation to */*/*/*/*, and you have deeply nested layout containers, you are duplicating massive amounts of string data into the Lucene index. Keep it as shallow as your component architecture allows.

The TraversalWarning and How to Fix It

If you look in your AEM error.log and see this:

2026-10-01 14:02:11.455 *WARN* [oak-lucene-0] org.apache.jackrabbit.oak.spi.query.Cursors$TraversalCursor Traversed 100000 nodes with filter Filter(query=select [jcr:path], [jcr:score], * from [nt:base] as a where [myStatus] = 'active' and isdescendantnode(a, '/content/my-site')) called by org.apache.jackrabbit.oak.query.QueryImpl.getRow...

Your query is fundamentally broken. By default, Oak sets a hard limit of 100,000 nodes for traversal. If a query hits this limit, it throws an exception and fails to return results.

How to fix it:

  1. Get the exact query from the log snippet.
  2. Run it in the Explain Query tool.
  3. Identify why no index is picking it up:
    • Is it missing a node type constraint? (Always include type=cq:Page or type=dam:Asset instead of nt:base or nt:unstructured if possible).
    • Is it querying a property that isn't indexed?
    • Is it lacking a path restriction? (e.g., searching / instead of /content/my-site).
  4. Modify the Java/HTL code executing the query to add those restrictions.
  5. If the query is as optimized as possible, create or modify a custom Lucene index to explicitly cover the combination of properties and paths.

Real-world scenario: The ACL Traversal Trap

Here is a vicious edge case. You write a query, check the Explain Query tool, and it says it uses your index. Great! But in production, you STILL get a TraversalWarning.

Why? Because of ACLs (Access Control Lists).

Remember the Oak execution flow diagram? ACL evaluation happens after the index returns results. Suppose your query is SELECT * FROM [cq:Page] WHERE [jcr:content/hideInNav] = true. The index finds 200,000 matching pages. But the user executing the query only has read access to 5 of them.

Oak pulls the 200,000 nodes from the index, and then starts iterating through them to check permissions. As it checks permissions, it increments its traversal counter. It hits 100,000 permission checks, throws the TraversalWarning, and dies.

The Fix: Never write queries that rely on ACLs to filter the result set down to a manageable size. Your index properties and path restrictions MUST reduce the result set to a small number before ACL evaluation happens.

Anatomy of a Custom Property Index

We talked about Property Indexes earlier. They are synchronous and fast. Here is what a simple property index looks like in XML. We use this when doing backend OSGi job processing that queries nodes immediately after they are created by another system.

<?xml version="1.0" encoding="UTF-8"?>
<jcr:root xmlns:jcr="http://www.jcp.org/jcr/1.0" xmlns:nt="http://www.jcp.org/jcr/nt/1.0" xmlns:oak="http://jackrabbit.apache.org/oak/ns/1.0"
    jcr:primaryType="oak:QueryIndexDefinition"
    propertyNames="[myCustomStatus]"
    type="property"
    declaringNodeTypes="[nt:unstructured]"
    includedPaths="[/var/my-app/data]"/>

Notice the absence of async="async". This makes it synchronous. It maps the myCustomStatus property for nt:unstructured nodes under /var/my-app/data. Very simple, very fast, but very limited.

The Async Indexing Lane: When Indexes Fall Behind

We established that Lucene indexes are asynchronous. AEM maintains a background thread (the "async indexer") that constantly looks for repository mutations and updates the Lucene files.

Under normal load, the delay between saving a node and the index updating is milliseconds. However, during massive bulk updates—like a catalog import or a rollout of a massive Blueprint—the async indexing lane can fall behind.

When this happens, Authors might complain: "I just updated the page title, but when I search for it, the old title still shows up!"

You can monitor this via the JMX Console.

  1. Go to /system/console/jmx.
  2. Search for IndexStats (specifically AsyncIndexerService).
  3. Look for the Failing flag and the LastIndexedTime.

If the LastIndexedTime is hours in the past, your async lane is blocked. This usually happens because someone triggered a massive synchronous re-index, or a corrupted Lucene file is throwing errors and causing the lane to retry infinitely.

If the async lane is completely locked, you often have to restart the instance or manually intervene using the oak-run tool to bypass the corruption.

Deploying Custom Indexes on AEM as a Cloud Service

Deploying indexes is radically different between AEM 6.5 (On-Premise/AMS) and AEM as a Cloud Service (AEM CS).

In AEM 6.5, you could deploy an index via a standard content package, set reindex="{Boolean}true", and watch the server CPU spike to 100% while it rebuilt on the live instance. Do not do this in Cloud Service.

AEM as a Cloud Service uses a CI/CD Blue/Green deployment model. The repository is split into a mutable (e.g., /content) and immutable (e.g., /apps, /oak:index) space.

When you deploy a custom index to AEM CS via Cloud Manager:

  1. The index definition MUST be in your ui.apps package (because /oak:index is now immutable).
  2. The index definition name MUST be suffixed with -custom-N where N is an integer (e.g., mySiteIndex-custom-1).
  3. Cloud Manager spins up the "Green" (new) environment.
  4. Before switching traffic to the new environment, Cloud Manager extracts the index definition and runs an out-of-band oak-run process to build the index on the Green environment.
  5. Once the index is 100% built and validated, traffic routes to the Green environment.

This ensures zero downtime and zero performance impact on the live (Blue) environment.

The Update Gotcha: If you need to add a new property to your index three months later, you cannot just update the XML and deploy. Because the index is immutable, Cloud Manager will ignore the change or fail the build. You MUST increment the suffix to -custom-2 (e.g., mySiteIndex-custom-2). Cloud Manager sees the new node, builds it out-of-band, and when traffic switches, it uses the new version. The old -custom-1 version is eventually garbage collected.

Extending OOTB Indexes in Cloud Service

If you need to add a property to cqPageLucene in AEM CS, you create a node in your ui.apps package at /oak:index/cqPageLucene-custom-1.

Inside this XML, you do not need to copy the entire OOTB definition! You use the merges feature. You define only the properties you want to add, and Cloud Manager will merge your -custom-1 definition with the OOTB definition during the build process.

Re-indexing Strategies

What happens if an index gets corrupted, or you absolutely must force a rebuild?

In AEM 6.5 On-Premise

Never set reindex=true on a large Lucene index during business hours. It blocks the async indexing lane, meaning NO other indexes will update while the large one rebuilds.

Instead, use the oak-run jar to reindex out-of-band:

  1. Stop AEM.
  2. Run java -jar oak-run.jar index ... pointing to your NodeStore (TarMK or MongoMK).
  3. This builds the Lucene index directly against the disk, utilizing all CPU cores without JVM overhead.
  4. Restart AEM.

Alternatively, for smaller indexes, use the asynchronous reindexing MBean via the JMX console to rebuild the index on a background thread without locking up the repository.

In AEM as a Cloud Service

You don't manage reindexing. Cloud Manager handles it during the Blue/Green deployment phase. If an index goes corrupt (extremely rare in CS due to the architecture), Adobe Support must intervene. If you simply want to force a clean rebuild of your own index, you just increment your index -custom-N version in your Git repository and run a deployment pipeline.

Performance Monitoring Tools

Keep an eye on these specific tools to proactively manage your search infrastructure:

  1. Slow Query Log: AEM automatically logs queries that take longer than a configurable threshold to logs/error.log (or a dedicated queries.log). Watch this like a hawk. Any query taking longer than 100ms needs to be investigated.
  2. JMX MBeans: Navigate to /system/console/jmx and search for "Lucene Index Statistics". This shows you the size of the index on disk, the number of documents, and how far behind the async indexer is.
  3. Index Consistency Check: Available in the Operations Dashboard, this verifies that your index actually matches the node state in the repository.
  4. Developer Console (Cloud Service): In AEM CS, use the Developer Console to execute Explain Query against the stage or production tiers without needing local access to the environments.

Advanced Troubleshooting Guide

When search fails in production, the symptoms can range from empty results to a completely locked JVM. Use this troubleshooting table to quickly diagnose the issue based on the exact log output.

Symptom / Log OutputRoot CauseSolution
Cursors$TraversalCursor Traversed 100000 nodesQuery lacks sufficient restrictions and Oak fell back to traversal, hitting the hard limit.Add a specific path restriction, define a type (e.g., cq:Page), and ensure the properties queried are indexed. Verify with Explain Query.
Explain Query shows cost: 1.0E8Oak evaluated the index but deemed it too expensive, usually because it estimates it would have to scan too many nodes.Add costPerExecution="0.1" to your custom index definition, or narrow the includedPaths to make the index more specific.
Search returns old data / new pages don't appearThe asynchronous indexing lane is blocked or falling behind.Check JMX AsyncIndexerService. If blocked, investigate thread dumps for stuck indexing jobs, or check for massive bulk content updates happening concurrently.
java.lang.OutOfMemoryError: Java heap space during QueryBuilder executionThe query used p.limit=-1 (unlimited results) on a massive dataset, pulling all nodes into memory at once.Never use -1 in production. Always paginate (p.limit=50, p.offset=0) and use p.guessTotal=true for fast total counts.
AEM Cloud Service pipeline fails at the "Build Images" step with an Index ErrorThe index definition in ui.apps is malformed, or you tried to update an existing index without incrementing the -custom-N suffix.Increment the suffix to -custom-2, ensure compatVersion="2" is set, and validate the XML structure locally before pushing.
Query works for admin but fails for normal usersACL Traversal limit hit. The index returns nodes, but Oak throws a traversal error while filtering out nodes the user cannot read.Redesign the query to rely on explicit property filters rather than relying on AEM permissions to filter the result set.
Full-text search contains(*, 'apple') returns nothingThe analyzed="true" flag is missing on the properties, or aggregation is not configured.Add analyzed="{Boolean}true" to the target properties, and if searching across a page, configure the <aggregates> node correctly.

Cheat Sheet & Best Practices

If you want your AEM platform to scale effortlessly, engrave these rules into your team's development standards.

The Do's

  • DO always specify a node type: Never write a query without a type parameter (e.g., type=cq:Page). Querying nt:base forces Oak to evaluate every single node in the repository.
  • DO always specify a path restriction: A query with path=/content/my-site/us/en is exponentially faster than one searching from /.
  • DO paginate your results: Always set p.limit. If you need a total count for UI pagination, use p.guessTotal=true which tells Oak to return an approximation if the result set is massive, saving precious CPU cycles.
  • DO append -custom-1 to all index names: Start this habit even in AEM 6.5, so your codebase is ready for AEM as a Cloud Service migration.
  • DO consolidate indexes: It is computationally cheaper for Oak to maintain one slightly larger Lucene index that covers 5 properties than 5 separate single-property Lucene indexes.
  • DO use includedPaths: If your index is only for commerce product data, set includedPaths="[/var/commerce]". Do not let it index /content or /apps.
  • DO test queries with different user sessions: Always test queries logged in as an author, and logged out as an anonymous user, to catch ACL Traversal traps before they hit production.

The Don'ts

  • DON'T modify out-of-the-box indexes lightly: (e.g., cqPageLucene, damAssetLucene). Always prefer copying them, renaming them, and customizing the copy for application data. Overriding OOTB indexes aggressively is the number one cause of upgrade failures.
  • DON'T index nt:base globally: Creating a global index on nt:base for a common property will bloat your Lucene files to terabytes, crash your disk, and bring down the indexing thread.
  • DON'T use contains(*, 'term'): This is incredibly expensive. Always target specific properties like contains(jcr:content/jcr:title, 'term').
  • DON'T set reindex=true on a live production server: Especially for large indexes. Use oak-run for out-of-band indexing.
  • DON'T rely on indexing for real-time transactional data: AEM is a content management system, not a real-time database. If you are updating a property 50 times a second and trying to query it, you are using the wrong architecture. Offload transactional data to an external database or Elastic Search cluster.

QueryBuilder vs JCR-SQL2: A Performance Perspective

While this guide focuses heavily on indexes, it is impossible to talk about AEM search optimization without discussing the abstraction layers that sit on top of the Oak Query Engine. The most common debate among AEM developers is whether to use the AEM QueryBuilder API or write raw JCR-SQL2 statements.

The Abstraction Tax of QueryBuilder

The QueryBuilder API is inherently an abstraction layer. When you create a map of predicates (path=/content, type=cq:Page, property=jcr:title), the QueryBuilder engine must parse that map, resolve the predicates through OSGi services, and dynamically construct a raw XPath statement. This XPath statement is then handed to the JCR QueryManager, which passes it to the Oak Query Engine.

This abstraction process takes time. In high-throughput, latency-sensitive applications (like an API endpoint serving a headless React application), the milliseconds spent inside the QueryBuilder compilation phase can add up.

If you profile a complex QueryBuilder execution in a tool like JProfiler, you will often see that constructing the query takes almost as long as executing it.

When to use JCR-SQL2

If you know the exact structure of your query, and it does not need to be highly dynamic, writing a raw JCR-SQL2 statement and executing it directly via the QueryManager is significantly faster.

// Fast: Direct JCR-SQL2 execution
String statement = "SELECT * FROM [cq:Page] AS s WHERE ISDESCENDANTNODE(s, '/content/my-site') AND s.[jcr:content/hideInNav] = true";
Query query = queryManager.createQuery(statement, Query.JCR_SQL2);
QueryResult result = query.execute();

By doing this, you completely bypass the QueryBuilder predicate evaluation phase. You hand the query directly to Oak.

When to use QueryBuilder

You should use QueryBuilder when:

  1. The query is highly dynamic (e.g., built from user input in a search form with variable filters).
  2. You need complex facet extraction (QueryBuilder's faceting API is excellent and very difficult to replicate in raw SQL2).
  3. You are building OSGi services that need to share query logic via custom Predicate Evaluators.

The Golden Rule: The engine executing the query (Oak) is the same regardless of whether you use QueryBuilder or JCR-SQL2. A poorly indexed SQL2 query will crash AEM just as fast as a poorly indexed QueryBuilder query. Optimize the index first, then optimize the execution layer.

Utilizing the oak-run Tool for Out-of-Band Indexing

We mentioned oak-run several times as the correct way to handle massive re-indexing tasks, especially in AEM 6.5. Let's look at exactly how to use this tool, as it is a critical skill for any Staff-level AEM engineer.

The oak-run tool is a runnable JAR file provided by Apache Jackrabbit. It allows you to interact directly with the underlying NodeStore (the actual files on disk or the Mongo database) completely bypassing the AEM OSGi framework, the JVM memory limits of the AEM instance, and the HTTP stack.

Why use oak-run?

If you trigger a large re-index via the AEM Web UI (by setting reindex=true), AEM attempts to rebuild the index while simultaneously serving web traffic, running background workflows, and handling replication. This causes massive garbage collection spikes. Worse, if the instance is restarted during the rebuild, the index might become corrupted.

oak-run avoids all of this by running as an isolated JVM process.

How to execute an out-of-band index rebuild

First, you must shut down the AEM instance. You cannot have two JVMs (AEM and oak-run) attempting to lock the TarMK files simultaneously.

Next, download the oak-run JAR that perfectly matches the version of Oak running inside your AEM instance (check the /system/console/bundles UI for the exact version).

Run the following command:

java -Xmx8g -jar oak-run-1.22.4.jar index --reindex --index-paths=/oak:index/cqPageLucene --read-write /path/to/crx-quickstart/repository/segmentstore

Breaking down the command:

  • -Xmx8g: Give the process plenty of memory.
  • index: The sub-command telling oak-run we are performing index operations.
  • --reindex: Forces a full rebuild.
  • --index-paths: Specifies the exact index to rebuild.
  • --read-write: Required to actually save the new Lucene files back into the NodeStore.
  • /path/to/.../segmentstore: The absolute path to your TarMK files.

This process will rip through the repository at maximum disk I/O speed. An index that might take 4 hours to rebuild inside AEM can often be rebuilt in 20 minutes using oak-run.

Conclusion

By deeply understanding the Oak Query Engine and mastering custom index definitions, you eliminate the single most common cause of performance degradation in enterprise AEM deployments. Indexes are not an afterthought—they are a foundational piece of your content architecture. When you design a new feature that requires querying data, the index definition must be designed alongside the component and the OSGi service.

Write tight, explicitly restricted queries, index them properly, respect the asynchronous indexing lane, and your AEM platform will scale to millions of nodes without breaking a sweat.


For more deep dives into AEM backend architecture, check out the AEM Backend Development Complete Guide and the OSGi in AEM Complete Guide. To understand how this fits into the broader modern AEM stack, read the AEM Cloud Service Complete Guide.

Share this article

Discussion

By commenting you agree to the Privacy Policy. Guest comments are reviewed before they appear.

Loading discussion…

Subscribe to the Newsletter

Get the latest articles, tutorials, and tech insights delivered straight to your inbox. No spam, unsubscribe anytime.

Back to Blog