Migrating from WordPress to Headless CMS Architecture

Migrating WordPress to headless breaks on three things: serialized PHP meta, runtime-interpreted post types, and HTML stuffed with shortcodes and inline styles. None of it survives contact with a strongly typed schema. This guide covers the normalization layer, schema mapping, extraction pipeline, and routing that get you to a zero-downtime cutover. When weighing Headless CMS Architecture & Platform Selection, favor schema rigidity and API predictability over the legacy plugin ecosystem. The guide sits under GraphQL vs REST API Tradeoffs because the delivery API you pick shapes the new frontend’s data layer from the first migrated page.

A phased WordPress migrationA twelve-week migration: schema mapping in weeks one and two, pipeline and asset relocation in weeks two to five, parallel running of old and new sites from week five, redirect testing in weeks eight to ten, and the DNS cutover in week ten followed by a monitoring period.Schema mappingPipeline + assetsParallel runningold + newRedirect testingMonitoring0 weeks2 weeks4 weeks6 weeks8 weeks10 weeks12 weekscutover
Running both sites in parallel before the cutover is what makes zero downtime realistic.

Where decoupling breaks

WordPress stores content in a relational but loosely typed schema. wp_posts is a catch-all for pages, posts, attachments, media, and revisions; wp_postmeta leans on serialized PHP arrays for field groups like ACF or plugin config. Direct JSON or SQL exports break because frontend frameworks expect strict, predictable types.

Failure scenario: A Next.js frontend renders migrated content.rendered via dangerouslySetInnerHTML. The payload carries unescaped shortcodes ([gallery ids="12,14"]), inline style attributes, and relative image paths (/wp-content/uploads/2023/05/image.jpg). Static generation fails: the HTML parser hits undefined DOM nodes and the asset resolver can’t find the media on the new CDN.

Fix: A pre-ingestion normalization layer between the legacy database and the target platform that strips WordPress artifacts, rewrites relative URLs to absolute CDN paths, sanitizes markup, and enforces content boundaries before anything reaches the destination API.

Schema mapping

WordPress interprets flat post types and dynamic meta at runtime; headless platforms need explicit, nested, validated schemas. The mapping preserves relationships while removing runtime ambiguity.

WordPress Entity Headless CMS Equivalent Migration Action
post_type Collection / Content Type Define explicit type with strict validation rules and required fields
post_meta Typed Fields / Components Flatten nested serialized arrays into structured JSON objects
taxonomies Relations / References Map to foreign-key relationships with slug indexing and hierarchical trees
revisions Version History Archive or discard; retain only the latest published state for migration

Enforcing types at ingestion keeps malformed payloads out of the headless environment. This Zod schema sanitizes and reshapes legacy WordPress payloads:

TypeScript
// schema-validator.ts
import { z } from 'zod';

export const WpPostSchema = z.object({
  id: z.number(),
  slug: z.string().regex(/^[a-z0-9-]+$/),
  title: z.string().min(1),
  content: z.string().transform((html) => {
    // Strip WP shortcodes and normalize relative URLs
    return html
      // Remove only known shortcodes; a generic /\[.*?\]/ would also delete legitimate bracketed text.
      .replace(/\[\/?(gallery|caption|embed|contact-form-7|vc_[a-z_]+)[^\]]*\]/g, '')
      .replace(/src="\/wp-content\//g, 'src="https://cdn.example.com/assets/')
      .replace(/style="[^"]*"/g, ''); // Remove inline styles for CSS-in-JS compatibility
  }),
  meta: z.record(z.unknown()).transform((meta) => ({
    acf_fields: meta._acf_fields || {},
    seo_title: meta._yoast_wpseo_title || '',
    canonical_url: meta._yoast_wpseo_canonical || '',
  })),
});

export type NormalizedPost = z.infer<typeof WpPostSchema>;

export function normalizeWpPayload(raw: unknown): NormalizedPost {
  return WpPostSchema.parse(raw);
}

Extraction and normalization pipeline

Run the pipeline in discrete, idempotent stages — extraction, asset relocation, URL rewriting, and transactional ingestion:

The extraction and normalization pipelinePosts, pages and media are extracted in batches from the WordPress REST API, assets are relocated and transcoded, URLs are rewritten to the CDN, content is normalized with Zod, and each batch is validated before transactional ingestion, with rejected batches going back to extraction.Batch extractwp/v2 posts, mediaRelocate assetsWebP / AVIFRewrite URLsto CDNNormalizeZodValidatetypesTransactionalingestpassreject batch
Each stage is idempotent, so a rejected batch can be rerun from the start without duplicates.

Run the pipeline in discrete, idempotent stages. Manual CSV exports and raw database dumps corrupt data; use the WordPress REST API or WP-CLI for programmatic extraction.

  1. Batch extraction. Paginate wp/v2/posts, wp/v2/pages, and wp/v2/media, respecting rate limits with exponential backoff on large datasets.
  2. Asset relocation. Download attachments, drop unnecessary EXIF, transcode to WebP/AVIF, and upload to dedicated object storage or CDN.
  3. URL rewriting. Keep an old-wp-content-to-CDN mapping table and apply it across all content and meta fields before ingestion.
  4. Headless ingestion. Push normalized payloads via the target REST or GraphQL API in transactional batches — a failed batch retries whole, never partially.

Deterministic routing

WordPress routing depends on rewrite rules and query params; headless needs file-system routing that matches SSG/ISR. Map legacy URLs to framework route handlers — in Next.js, generateStaticParams for known slugs plus a catch-all [[...slug]].tsx for fallback. Add a 404 handler that checks the CMS for draft content and returns the right status (404, 410, or 301).

For preview, wire webhook-driven revalidation: a publish triggers revalidatePath() or revalidateTag() so editors see updates without a full rebuild. See Next.js Data Fetching.

API strategy

The transport layer drives migration complexity and frontend performance — the GraphQL vs REST API Tradeoffs apply directly here.

  • REST: simpler for flat types but needs multiple round-trips for nested relationships (post, then author, then categories). Cache hard with Cache-Control and ETag validation.
  • GraphQL: better for nested models and component-driven UIs. Lock allowed operations with persisted queries and batch ingestion calls with DataLoader.

Either way, put an edge cache (Varnish, Cloudflare, Fastly) in front, key it by slug and content-type tags, and invalidate via webhook payloads.

Governance and DX metrics

A migration adds operational responsibilities: content lifecycle management, RBAC, and compliance auditing. Snapshot the headless database before each major batch for rollback. Track:

  • Build duration: SSG/SSR compile time — target sub-60s for large content sets.
  • Cache hit ratio: aim for >95% on public content.
  • Editor latency: publish-to-visible time — target <2s for ISR revalidation.
  • Error rate: hydration errors and 4xx/5xx during migration windows.

Treat the cutover as a data-engineering project, not a platform swap, and deployments stay predictable.

Preserving URLs and Search Rankings

A migration that changes URLs without redirects loses search traffic and breaks inbound links. Export every public URL from WordPress, including category archives, tag pages, paginated lists and attachment pages, and decide for each whether the new site keeps it, redirects it with a 301 or retires it with a 410. Serve the redirect map from the edge, not from application code, so it applies before any rendering. Before cutover, crawl the old site’s URL list against the new site and require every URL to return 200, 301 to a 200, or an intentional 410.

TypeScript
// middleware.ts: legacy URL redirects from a generated map
import { NextResponse, type NextRequest } from "next/server";
import redirects from "./generated/wp-redirects.json"; // { "/2023/05/old-slug/": "/articles/old-slug" }

const map = redirects as Record<string, string>;

export function middleware(req: NextRequest) {
  const path = req.nextUrl.pathname.endsWith("/") ? req.nextUrl.pathname : `${req.nextUrl.pathname}/`;
  const target = map[path];
  if (target) return NextResponse.redirect(new URL(target, req.url), 301);
  return NextResponse.next();
}
Outcome of the pre-cutover URL crawlResults of crawling all 18,400 legacy WordPress URLs against the new headless site before cutover, by status after fixes.Kept, 20011200 URLsRedirected, 3016700 URLsRetired, 410500 URLs
Every legacy URL resolved to an intentional status before DNS was switched.

Gotchas & Edge Cases

  • Shortcodes with content. Gallery and embed shortcodes carry meaning. Convert them to structured blocks during normalization instead of stripping them, or migrated posts lose images.
  • Page builders. Content built with visual page builders is stored as nested shortcodes or serialized data. Budget for manual rebuilding of key pages; automated conversion is rarely clean.
  • Authors and users. WordPress users are both authors and accounts. Migrate only the public author data into an author type, never passwords or emails.
  • Comments. Decide early whether comments move to a dedicated service, are archived as static content or are dropped, since the headless CMS will not hold them.

Worked Example

A publisher with 9,000 posts and 18,400 public URLs migrated over twelve weeks. The normalization layer converted three shortcode types into blocks and rewrote 41,000 image URLs. Both sites ran in parallel for five weeks, with editors publishing in the new CMS and a sync job mirroring new posts back to WordPress for safety. The URL crawl found 1,100 URLs without a decision, mostly tag pages, which were redirected to topic pages. After the DNS switch, organic traffic dipped by about four percent for two weeks and then returned to its previous level.

The largest unplanned effort was page-builder content: forty landing pages had to be rebuilt by hand from the new block library, which the team would budget explicitly next time.

Rollout Checklist

  • Map every post type, meta field and taxonomy to the new schema.
  • Build the idempotent pipeline and run it repeatedly against a staging CMS.
  • Relocate assets and rewrite every URL to the new CDN.
  • Generate a redirect map and serve it at the edge.
  • Crawl all legacy URLs against the new site before cutover.
  • Run both sites in parallel, then switch DNS and monitor traffic and errors.

Frequently Asked Questions

Can WordPress itself be used headless instead?

Yes, through its REST API or WPGraphQL, which avoids a content migration. It keeps WordPress’s data model and plugin maintenance, so it suits teams who want a new frontend but not a new editorial tool.

How long should both systems run in parallel?

Long enough for editors to work comfortably in the new CMS and for the URL crawl to pass, typically two to six weeks.

Do we need to migrate revisions?

Rarely. Migrate the latest published state and archive the WordPress database for reference.

What about SEO plugin data?

Migrate titles, descriptions and canonicals into the new SEO fields, as described in modeling SEO fields.

Should the new site launch with all old content?

Usually yes, to keep URLs and rankings, but it is a good moment to retire thin or outdated posts with 410s or redirects to better pages, decided with the editorial team.