Next.js Technical SEO Audit: Canonicals to Crawl Evidence

Quick overview

Audit a Next.js site for rendered metadata, canonical URLs, pagination, robots.txt, sitemap.xml and source failures using concrete release checks.

A Next.js technical SEO audit should verify what crawlers and visitors actually receive. Route names, configuration files, and a successful development session are useful inputs, but the evidence is the deployed response: status, headers, canonical, readable content, and discovery links.

This checklist is designed for an agency website or content library with public services, a blog, and private administration routes. It focuses on failure patterns that can hide behind a page that looks correct in the browser. Use representative URLs and repeat the checks after deployment, because hosting redirects and production data access can differ from local behavior.

Build a URL inventory with expected behavior

Start with a short matrix rather than crawling blindly. Include the homepage, a service detail, an article, the second page of the blog, an out-of-range page, a search variant, an archived article, an admin page, and the discovery endpoints.

For each, record the expected response and indexing policy. A published article intended for search should identify its own canonical URL. A search variant may be intentionally excluded from indexing. An archived article may remain accessible while absent from public discovery. These are editorial and product decisions that the implementation should express consistently.

The inventory should also include:

  • HTTP and HTTPS forms of the preferred hostname.

  • Apex and www domain behavior.

  • Trailing-slash variants where relevant.

  • Common tracking parameters.

  • Valid and invalid pagination values.

  • A genuinely missing content identifier.

  • A temporarily unavailable content source.

  • A public login route and an authenticated private route.

Verify metadata in the returned document

The Next.js metadata guide documents static and generated metadata plus streaming behavior. Inspect the rendered response for the relevant request rather than assuming every user agent sees metadata in the same position.

Check that the title and description describe the actual page and that the canonical points to the intended URL. Inspect Open Graph and article fields against the visible heading, publication date, and actual modified date. A default homepage description repeated on every article is easy to miss during visual review.

Use a crawler-like request as well as a normal browser request where the rendering policy differs. Keep the input user agent in the test record. A test that reads only a source module cannot establish that metadata is delivered correctly by the deployed app and hosting layer.

Canonicals must preserve distinct pagination

A canonical normaliser should remove irrelevant tracking parameters without merging content that is genuinely different. Page two of a blog index may contain different article links from page one; canonicalising both to the root index can obscure that distinction.

Validate numeric page values rather than accepting a partial parse of malformed input. Decide how a request beyond the final page is handled, and ensure the canonical, visible pagination, and displayed content agree. If the application clamps the page number, its metadata should use the same resolved value.

Google’s canonical documentation describes canonical signals and the need for consistency. Your application test should compare the chosen canonical with internal links and sitemap URLs rather than treating the tag as an isolated fix.

Robots.txt and noindex solve different problems

Robots.txt controls crawling behavior for compliant crawlers; it does not authenticate users or reliably remove a URL from search. Google’s robots introduction explains why access and indexing controls should not be confused.

For private pages, require real access control and an appropriate indexing policy. If the crawler cannot fetch a page, it may not observe the page’s noindex instruction. A public login page can be accessible while excluded from indexing, whereas confidential records should require an authorised session.

Inspect every user-agent group in the actual robots response. A named AI-crawler group can accidentally override the wildcard group’s intended private-path rules. Test group-specific behavior and confirm the published sitemap location uses the preferred hostname.

Sitemap.xml should describe intended canonical content

A sitemap is a discovery file. It should list the public canonical URLs the business wants considered for indexing, including approved articles and real service details. Do not include every generated archive URL merely because the content directory contains a file.

Check that the XML parses, URL values are correctly escaped, and there are no duplicate entries. Match article modification dates to substantive changes. Publishing a fresh lastmod value for unchanged content creates misleading freshness information rather than an editorial update.

A content-source outage needs an error signal. If fetching services fails, returning a successful sitemap that silently omits every service misrepresents the site. Design and test the unavailable-source path, including cache behavior, so a temporary failure does not become a long-lived discovery artifact.

Audit structured data against visible evidence

Article markup should describe the actual headline, author identity, dates, and canonical page. Breadcrumbs should reflect real navigation. Product or organisation data should contain supported facts, not invented ratings or placeholder values copied from an example.

Inspect the generated JSON and the rendered page together. Valid JSON alone does not establish that the claims are accurate or that the page is eligible for a particular search feature. Keep the publisher’s business details consistent across the site and confirm any claimed address or profile belongs to the organisation.

On-page questions can be useful for readers even when there is no promised rich-result outcome. Treat schema as a description of the content rather than a substitute for answering the question well.

Distinguish content absence from a failing dependency

An unavailable database is not evidence that the requested service no longer exists. A catch block that returns an empty array or undefined can turn a timeout into a false empty listing or missing-page response. That affects users, metadata, and discovery endpoints simultaneously.

Test the shared data-access boundary with a deliberate failure. Confirm that a genuine empty result remains valid and a failed request propagates an operational error that the route handles appropriately. Fixing only the sitemap can leave the same issue in service and portfolio pages.

Record the error without exposing secrets or full connection strings. The useful diagnostics are the operation, route, error category, and correlation information needed to investigate it. Public responses should remain understandable and avoid presenting internal implementation details.

Validate the release with a compact evidence set

Keep checks focused on behavior changed by the release. A practical evidence set includes:

  • Unique article slugs and valid frontmatter.

  • Metadata for a representative article and listing page.

  • Canonicals for valid and invalid pagination inputs.

  • Indexing controls for archives and private pages.

  • Robots rules for wildcard and named crawler groups.

  • A parseable sitemap with approved content only.

  • A source-failure test without silent content disappearance.

  • Deployed-hostname redirects and final page responses.

Run code checks appropriate to the implementation, then inspect the deployed output. If the hosting platform applies an earlier redirect, the application configuration may never receive that request. Record that boundary rather than claiming a repository edit changed the externally observed status.

Frequently asked questions

Does a perfect audit guarantee indexing?

No. It establishes that selected technical problems have been checked or corrected. Search systems still decide whether and how content appears.

Should every blog enter the sitemap?

Use an editorially defined indexing set. Accessible archives can remain outside discovery while being rewritten, consolidated, or assessed for relevance.

Is a browser screenshot enough?

It helps assess the visible page, but headers, robots, sitemap XML, metadata, and error behavior need direct inspection too.

Read the Next.js architecture guide, explore web development services, or request a technical audit with representative URLs and the behavior you expect.

MT

Written by

Moueen Togarvi

Founder & CEO at Voquarn Code, focused on product engineering, search growth, and practical AI systems.

About author
Turn the insight into action

Need a practical plan for your next digital project?

Tell us what you are building. We will help you clarify the scope, technical approach, and highest-value first step.

Discuss your project