Every GraphQL POST request sent from a client browser directly to an AEM Publish tier is a guaranteed cache miss, a massive CPU hit, and a profound scaling liability. Treating AEM's headless API like a raw database endpoint by firing dynamic queries at it bypasses the CDN, ignores the Dispatcher, and forces the JVM to parse and execute complex ASTs (Abstract Syntax Trees) on the fly for every single page load. Under production traffic, this architectural pattern will invariably cause your Publish instances to spike in CPU usage and ultimately crash. Persisted Queries change the game entirely. By shifting the paradigm from dynamic POSTs to pre-registered, cacheable GET requests, you transform a fragile, heavy backend operation into a highly scalable edge-delivered asset that can serve millions of requests without ever touching the AEM server.
In this comprehensive, deep-dive guide, we will completely map out the AEM Headless GraphQL ecosystem from the ground up. You will learn the exact mechanics of how AEM auto-generates schemas from Content Fragment Models, how to structure incredibly complex queries for maximum efficiency, and the critical implementation details of Persisted Queries. We will explore advanced query syntax including complex filtering, cursor-based pagination, inline references, and variation selection. We will cover Dispatcher configurations, CDN integration, robust error handling, security hardening, and authentication patterns for server-to-server access. Finally, we'll look deeply at the AEM Headless SDK and architect a real-world system using Next.js and Incremental Static Regeneration (ISR).
Before proceeding, ensure you have a firm grasp of AEM's underlying data structures and caching mechanisms by reviewing the Content Fragments & Experience Fragments complete guide and the foundational AEM Architecture guide. For caching specifics, keep the Dispatcher complete guide handy.
How AEM auto-generates GraphQL schemas
Unlike traditional GraphQL middleware servers (like Apollo or Yoga) where backend engineers manually write explicit schema definitions (schema.graphql) and wire up resolvers, AEM operates on a strict auto-generation model. In AEM, your schema is inextricably linked to your Content Fragment Models (CFMs). AEM acts as both the database and the GraphQL engine.
The Trigger and Compilation Process
When you create, modify, or delete a Content Fragment Model in AEM (located under /conf/<tenant>/settings/dam/cfm/models), AEM triggers an internal OSGi event that compiles the model's structural definition into a GraphQL schema. Every field, data type, and reference configured in the model editor maps directly and predictably to a GraphQL scalar, object type, or union.
This compilation happens asynchronously in AEM as a Cloud Service. When a model changes, the schema compilation job runs in the background.
| AEM CFM Data Type | GraphQL Type | Internal JCR Property Type | Description |
|---|---|---|---|
| Single-line text | String | String | Standard text string. |
| Multi-line text | MultiFormatString | String | Returns raw text, plaintext, html, markdown, or structured JSON. |
| Number | Float or Int | Double or Long | Handled natively by GraphQL numeric primitives. |
| Boolean | Boolean | Boolean | True/False checkboxes. |
| Date and Time | Calendar or String | Date | ISO-8601 formatted string or explicit calendar object type. |
| Tags | [String] | String[] | Array of strings representing tag IDs. |
| Fragment Reference | <ModelName>Model | String (path) | A strongly-typed nested reference to another model. |
AEM doesn't just generate basic type representations; it generates an entire Query API surface. For a model named Product, AEM automatically generates two primary entry points:
productByPath: Fetch a single product by its exact JCR path. Returns a single object or null.productList: Fetch a list of products, offering sophisticated filtering, sorting, and pagination arguments. Returns an object containing anitemsarray and pagination metadata.
Configuration Endpoints and JCR Underpinnings
AEM requires you to explicitly configure GraphQL Endpoints. You do not query a single monolith schema spanning the entire repository. AEM supports multiple isolated endpoints to allow strict scoping of content.
The configurations live deep within Oak. You can inspect the configured endpoints via CRXDE at /conf/global/settings/graphql/endpoints (or tenant-specific paths).
/conf
└── /my-tenant
└── /settings
├── /dam/cfm/models/product <-- The CFM definition
├── /dam/cfm/models/category <-- Another CFM definition
└── /graphql/endpoints/default <-- The endpoint config tying them togetherMost production architectures should aggressively avoid the Global Endpoint (/content/cq:graphql/global/endpoint.json). Instead, create a specific configuration tied to your tenant's /conf directory (/content/cq:graphql/my-tenant/endpoint.json). This restricts the GraphQL schema to only the models defined in that specific tenant space. This reduces the schema size (improving compilation time and introspection performance), prevents cross-tenant data leakage in multi-tenant environments, and prevents naming collisions if two tenants have a model named Article.
Production Gotcha: If you add a new field to a Content Fragment Model, existing queries will not break (GraphQL handles missing fields gracefully in queries). However, to actually query the newly added field, you must ensure the endpoint configuration has picked up the schema update. In rare cases on AEM 6.5, if the OSGi event listener drops the compilation trigger, you might need to manually trigger a schema regeneration via JMX.
POST queries vs Persisted Queries: the critical difference
Understanding the difference between dynamic POST queries and Persisted Queries is the most critical concept in AEM headless architecture.
The Danger of Dynamic POST Queries
Standard GraphQL specifications default to using POST requests. A client application (like a React SPA or mobile app) sends an HTTP POST to the endpoint (e.g., /content/cq:graphql/my-tenant/endpoint.json) with a massive JSON payload containing the query string, operation name, and variables.
Why is this an architectural disaster in production?
- Uncacheable at the Edge: CDNs (like Fastly, Akamai, Cloudflare) and the AEM Dispatcher do not cache HTTP
POSTrequests by default. Every POST request bypasses the caching layers completely. - Direct Publish Hits: Because the cache is bypassed, every single user request hits the AEM Publish instance directly. If your site gets 10,000 visitors a minute, your AEM Publish tier receives 10,000 dynamic queries a minute.
- CPU Intensive Parsing: GraphQL is notoriously CPU-heavy. The server must parse the query string into an AST, validate the AST against the schema, execute the resolvers (which means translating the query into Oak JCR queries), and serialize the response. Doing this on the fly for thousands of concurrent requests will result in JVM thread starvation and OutOfMemory errors.
- Security Vulnerabilities: Allowing arbitrary POST queries means malicious actors can intentionally craft deeply nested, expensive queries (Denial of Service via query complexity) to take down your servers.
The Persisted Query Paradigm
The solution is Persisted Queries. A Persisted Query is a GraphQL query that the development team pre-saves on the AEM server ahead of time. Instead of sending the full query string via POST, the client sends a simple HTTP GET request referencing the pre-saved query's name (and passing any variables in the URL).
GET /graphql/execute.json/my-tenant/get-homepage-products;category=electronicsThis changes the entire delivery architecture:
- Cacheable: Because it's a standard HTTP
GETrequest, the CDN and the Dispatcher can cache the JSON response heavily. - Zero Publish Load: Once the first request populates the cache, subsequent requests are served instantly from the CDN edge or Dispatcher. The AEM Publish server does zero work.
- Pre-validated: AEM parses and validates the query when it is saved, not at runtime, saving CPU cycles.
- Secure by Default: Because clients can only execute queries that you have explicitly saved on the server, arbitrary query injection is impossible. The attack surface is eliminated.
Persisted Queries are not an optional optimization in AEM—they are a mandatory requirement for any production deployment.
Managing Persisted Queries via API
Persisted queries are managed via AEM's GraphiQL IDE (built into the AEM UI) or programmatically via a RESTful API. In a modern CI/CD pipeline, persisting queries should be an automated step during deployment, not a manual UI task.
Saving a Query (PUT)
To create or update a persisted query, you issue a PUT request to the persistence endpoint, providing the query payload.
curl -X PUT \
-u admin:admin \
-H "Content-Type: application/json" \
-d '{"query":"query GetProducts($cat: String!) { productList(filter: { category: { _expressions: [{ value: $cat }] } }) { items { title price } } }"}' \
"https://author-p1234-e5678.adobeaemcloud.com/graphql/persist.json/my-tenant/get-products"AEM saves these queries directly into the JCR as nodes under /conf/my-tenant/settings/graphql/persistentQueries/get-products. You must save the queries on Author and then publish them (activate them) to the Publish tier, just like standard content.
Listing Queries (GET)
To verify what queries are deployed to an environment, you can fetch a list:
curl -X GET "https://author-p1234-e5678.adobeaemcloud.com/graphql/persist.json/my-tenant"This returns a JSON object detailing all persisted queries for that configuration, including their exact JCR paths and hashed identifiers (useful for advanced caching strategies).
Advanced Query Syntax and Operations
AEM's GraphQL implementation extends far beyond basic field selection. It includes highly sophisticated filtering, sorting, pagination, and multi-format text retrieval out of the box.
The MultiFormatString and Rich Text
When a Content Fragment Model contains a Multi-line text field (often used for rich text), AEM exposes it as a MultiFormatString. This is critical for headless implementations.
query GetArticle {
articleByPath(_path: "/content/dam/my-tenant/news/release-notes") {
item {
_path
title
body {
plaintext
html
markdown
json
}
}
}
}Never blindly inject the html response into a React component using dangerouslySetInnerHTML. This opens you up to XSS vulnerabilities and prevents you from rendering custom components inside the text. Instead, always request the json format.
The json format returns an Abstract Syntax Tree (AST) representation of the rich text. Modern frontend frameworks can iterate over this AST and map specific node types (like a paragraph, heading, or embedded image) to specific React/Vue/Angular components.
Complex Filtering Mechanisms
AEM's filter argument allows for deep, complex queries using a nested object syntax.
Basic Filtering:
query {
productList(
filter: {
brand: {
_expressions: [{ value: "Adobe", match: EQUALS_CASE_INSENSITIVE }]
}
}
) {
items { title }
}
}Advanced Filtering: Multiple Expressions (AND/OR logic)
You can supply multiple expressions. By default, multiple expressions on the same field act as an OR. Multiple fields act as an AND.
query {
eventList(
filter: {
eventType: {
_expressions: [
{ value: "webinar", match: EQUALS }
{ value: "conference", match: EQUALS }
]
},
isPublic: {
_expressions: [{ value: true }]
}
}
) {
items { eventName date }
}
}Advanced Filtering: Date Ranges Querying by date requires specific match operators.
query {
articleList(
filter: {
publishDate: {
_expressions: [
{ value: "2026-01-01T00:00:00.000Z", match: GREATER_EQUAL }
{ value: "2026-12-31T23:59:59.999Z", match: LESS_EQUAL }
]
}
}
) {
items { title publishDate }
}
}Supported match operators include: EQUALS, EQUALS_NOT, EQUALS_CASE_INSENSITIVE, CONTAINS, CONTAINS_NOT, STARTS_WITH, ENDS_WITH, GREATER, GREATER_EQUAL, LESS, LESS_EQUAL.
Querying Content Fragment Variations
Content Fragments support Variations (e.g., "Master", "Mobile", "Summary"). You can explicitly request a specific variation in your query using the _variation argument.
query {
articleList(
_variation: "mobile",
filter: {
category: { _expressions: [{ value: "news" }] }
}
) {
items {
title
summary { plaintext }
}
}
}If the requested variation does not exist on a specific fragment, AEM will fall back to the "Master" variation data.
Cursor-based Pagination (Relay Standard)
For fetching a handful of items, you might use standard offset pagination. However, for large datasets or infinite scrolling implementations, cursor-based pagination (adhering to the GraphQL Relay specification) is strictly required for performance and data consistency.
query GetPaginatedArticles($cursor: String) {
articlePaginated(
first: 20,
after: $cursor,
sort: "publishDate DESC"
) {
edges {
node {
_path
title
}
cursor
}
pageInfo {
hasNextPage
endCursor
startCursor
}
}
}The frontend uses the endCursor from the previous response as the $cursor variable for the next request. This ensures stable pagination even if new articles are published while the user is scrolling.
Handling Deep and Inline References
One of AEM's most powerful capabilities is resolving references entirely server-side, preventing the frontend from making cascading N+1 API calls.
Nested Model References
When an Article Content Fragment contains a Fragment Reference field pointing to an Author fragment, you can request the Author's data directly in the same query.
query {
articleList {
items {
title
authorReference {
... on AuthorModel {
firstName
lastName
profileImage {
... on ImageRef {
_path
width
height
}
}
}
}
}
}
}Because references can point to multiple allowed models, you must use GraphQL inline fragments (... on ModelName) to specify which fields to extract based on the actual model type returned.
Inline References in Rich Text (_references)
AEM allows authors to drop references to other Content Fragments directly inside a multi-line text editor (Rich Text). To extract these inline references, you must query the _references object alongside the json node.
query {
articleByPath(_path: "/content/dam/my-tenant/test-article") {
item {
body {
json
}
_references {
... on ProductModel {
_path
price
sku
}
... on ImageRef {
_path
mimeType
}
}
}
}
}The frontend parses the json AST. When it encounters a node of type reference, it looks at the reference's path attribute, finds the corresponding object in the _references array, and renders the appropriate React component with that data.
Performance Warning: AEM strictly limits reference resolution depth to prevent catastrophic recursive queries (e.g., Article A references Article B which references Article A) from blowing up the JVM heap. Be extremely careful with cyclic references in your data modeling.
Dispatcher Configuration and Cache Invalidation
Deploying Persisted Queries without properly configuring the Dispatcher is a recipe for disaster. The Dispatcher must be explicitly told to allow and cache the GraphQL execution endpoints.
The VHost and Filters
In your Dispatcher configuration (e.g., dispatcher/src/conf.d/available_vhosts/my-tenant.vhost), ensure the execute paths are allowed. Note that variables in Persisted Queries are passed via URL parameters (semicolon delimited).
# Allow GraphQL Execute
<LocationMatch "^/graphql/execute\.json/.*$">
Require all granted
</LocationMatch>In your dispatcher/src/conf.d/filters/filters.any, you must explicitly permit the GET requests:
/0100 { /type "allow" /method "GET" /url "/graphql/execute.json/*" }Advanced Cache Invalidation Strategies
Caching JSON responses in the Dispatcher introduces a complex challenge: when a Content Fragment is published, how does the Dispatcher know which GraphQL queries to invalidate?
A simple query like get-homepage-products might depend on 50 different product fragments. AEM handles this via Dependency Tracking. When a persisted query is executed and cached, AEM tracks which JCR nodes contributed to that response. When an author publishes an update to one of those fragments, AEM sends a specialized invalidation request to the Dispatcher.
Furthermore, AEM as a Cloud Service automatically appends Cache-Control headers for persisted queries, leveraging Surrogate-Control keys for the Fastly CDN. It generates ETag headers based on the underlying Content Fragment modification dates. When a client sends an If-None-Match header, the Dispatcher or CDN can return a 304 Not Modified, saving immense bandwidth and parsing time on mobile devices.
Error Handling and Edge Cases
GraphQL errors are fundamentally different from REST errors. A GraphQL response often returns an HTTP 200 OK even if parts of the query failed. The errors are enclosed in an errors array alongside the data object.
{
"data": {
"articleList": {
"items": [
{
"title": "Valid Article",
"authorReference": null
}
]
}
},
"errors": [
{
"message": "Exception while fetching data (/articleList/items[0]/authorReference) : Cannot resolve reference /content/dam/missing-author",
"locations": [{"line": 6, "column": 7}],
"path": ["articleList", "items", 0, "authorReference"]
}
]
}What happens when a query references a deleted fragment?
If an Article references an Author, but the Author fragment was deleted or unpublished, the authorReference field will return null, and an error will be appended to the errors array.
Crucial Frontend Practice: Your frontend application must defensively check for null values on all references. Do not assume that because the parent fragment exists, the referenced fragments exist. A null reference check prevents complete page crashes.
Security: Rate Limiting and Introspection
Exposing a headless API to the internet requires significant security hardening.
Preventing GraphQL Introspection
GraphQL Introspection allows clients to query the schema itself to discover all available queries, types, and fields. While incredibly useful during development (it powers the GraphiQL IDE), it is a massive security risk in production, as it maps out your entire data structure for attackers.
In AEM, restrict introspection queries at the Dispatcher level. Introspection queries typically contain specific keywords like __schema or __type. You should block dynamic POSTs entirely anyway, which mitigates this, but if you must allow them, block introspection payloads.
Rate Limiting and Query Whitelisting
Because you are using Persisted Queries, you effectively have Query Whitelisting enabled by default—attackers cannot run arbitrary queries. However, they can still spam your persisted endpoints. Implement aggressive rate limiting at the CDN level (Fastly/Cloudflare) or WAF (Web Application Firewall) to prevent scraping of your headless data endpoints.
Authentication: Service Credentials vs Developer Tokens
If your headless application needs secure, internal data (e.g., pricing matrices, internal employee directories), the GraphQL endpoint cannot be public.
AEM as a Cloud Service (OAuth Server-to-Server)
AEMaaCS enforces highly secure OAuth 2.0 Server-to-Server (S2S) credentials.
- Generate Service Credentials in the Adobe Developer Console.
- Your Node.js/Next.js backend exchanges these credentials for a JWT Access Token.
- The backend attaches the token to the GraphQL request:
Authorization: Bearer <token>.
Never hardcode Developer Tokens (which expire every 24 hours) in your production application. Always use the automated S2S flow. Review the AEM Security guide for deep details on service users.
The AEM Headless SDK
Adobe provides official SDKs (both for Node.js and Vanilla JS/Browser environments) to abstract the complexities of formatting requests, handling variables, and managing authentication.
// Node.js example using @adobe/aem-headless-client-nodejs
import AEMHeadless from '@adobe/aem-headless-client-nodejs';
const aemHeadlessClient = new AEMHeadless({
serviceURL: 'https://publish-p12345-e6789.adobeaemcloud.com',
endpoint: 'my-tenant',
auth: process.env.AEM_S2S_TOKEN // Optional: For secure APIs
});
async function fetchProducts() {
try {
// The SDK automatically formats the GET request and URL parameters
const response = await aemHeadlessClient.runPersistedQuery('my-tenant/get-products', {
category: 'electronics'
});
return response.data.productList.items;
} catch (error) {
console.error("AEM Headless Error:", error);
}
}The SDK automatically handles GET request formatting for persisted queries, manages headers, and parses the JSON response seamlessly.
Real-World Architecture: AEM, Next.js, and ISR
In a modern enterprise stack, AEM acts as the pure content repository, while Next.js handles presentation. For peak performance, implement the Incremental Static Regeneration (ISR) pattern.
+--------------+ +---------------+ +---------------+
| | | | | |
| AEM Author +------->+ AEM Publish +------->+ Fastly CDN |
| (Creates CFs)| | (Persisted Qs)| | (AEM Managed) |
+--------------+ +-------+-------+ +-------+-------+
^ ^
| |
| |
+--------------+ +-------+-------+ +-------+-------+
| | | | | |
| User Browser |<-------+ Next.js App +------->+ Edge / Vercel |
| | | (React / ISR) | | Network |
+--------------+ +---------------+ +---------------+The Webhook Revalidation Flow:
- Next.js builds the initial pages statically at build time, pulling data via Persisted Queries from AEM.
- The Vercel edge caches the rendered HTML globally.
- When an author updates an article in AEM and hits publish, a custom workflow or OSGi event fires a Webhook to a Next.js API route (
/api/revalidate?secret=xyz). - Next.js receives the webhook, identifies the specific path, and calls
res.revalidate('/blog/article-name'). - Next.js refetches the Persisted Query from AEM (or the AEM CDN), generates the new HTML in the background, and seamlessly replaces the stale edge cache.
This architecture ensures users always experience millisecond response times. For deeper Next.js integration details, read the Next.js App Router for AEM Developers guide.
Content Fragment OpenAPI (REST) vs GraphQL
While GraphQL is incredible for fetching precise trees of data for frontend rendering, AEM also provides the Content Fragment OpenAPI (a RESTful API).
When should you use which?
- Use GraphQL for: Frontend delivery, fetching data to render pages, querying complex relationships, and mobile apps. It is optimized for Read operations.
- Use OpenAPI (REST) for: System-to-system integrations, mass content migrations, programmatic creation or updating of Content Fragments, and external PIM synchronizations. It is optimized for CRUD (Create, Read, Update, Delete) operations.
Read more in the AEM APIs and Integrations guide.
Cheat Sheet
Critical Endpoints
- Global Schema:
/content/cq:graphql/global/endpoint.json - Configured Schema:
/content/cq:graphql/<config>/endpoint.json - Execute Persisted (GET):
/graphql/execute.json/<config>/<query-name> - Save Persisted (PUT):
/graphql/persist.json/<config>/<query-name> - List Persisted (GET):
/graphql/persist.json/<config>
OSGi Configurations
- Endpoint mapping:
com.adobe.cq.graphql.core.internal.GraphQLConfigurationImpl - Logs:
com.adobe.cq.graphql(Set to DEBUG to view raw Oak queries executing).
Best Practices
- Aggressive Namespacing: Prefix Content Fragment Models (e.g.,
B2B_Article,B2C_Author) to prevent schema collisions in shared environments. - Rich Text Rendering: Always request the
jsonformat forMultiFormatStringfields if rendering in React/Vue. Never use the raw HTML output. - Cache Header Audits: Work closely with devops to ensure Dispatcher
filters.anyand vhost configurations are not accidentally stripping out theCache-ControlorETagheaders generated by AEM for the.jsonendpoints. - Defensive Frontend Coding: Always assume any GraphQL reference could return
nulland wrap component rendering logic in existence checks.
Do's & Don'ts
- DO use Persisted Queries for 100% of production traffic.
- DO test your queries and complex filters in the built-in GraphiQL IDE before committing them to application code.
- DO implement cursor-based pagination for lists expected to exceed 50 items.
- DON'T use the Global Endpoint. Always strictly scope queries to a tenant configuration path.
- DON'T create deep, cyclical references in your Content Fragment Models (e.g., Article -> Category -> Related Articles -> Category). AEM OSGi depth limits will trigger and crash your query response.
- DON'T send dynamic POST queries from a client browser. This is a severe performance liability and a massive security anti-pattern.
- DON'T blindly parse dates on the frontend. Use AEM's structured calendar outputs to manage timezone offsets correctly.
Mastering AEM's GraphQL API elevates your architecture from a fragile, dynamic integration to a robust, statically-edge-cacheable powerhouse. Adhere to persisted queries, respect the schema generation limits, and leverage the AEMHeadless clients to build the next generation of resilient headless experiences.
Discussion
Loading discussion…
Try a related tool
Subscribe to the Newsletter
Get the latest articles, tutorials, and tech insights delivered straight to your inbox. No spam, unsubscribe anytime.