October 5, 2026

How Did utm_source=chatgpt.com End Up in Disney+'s Canonical URL?

Sometimes a tiny HTML tag can reveal a much larger technical SEO issue.

An interesting example recently surfaced when someone searched Google for “Disney login” and noticed that the Disney+ login page had a canonical URL containing:

?utm_source=chatgpt.com

The canonical tag reportedly looked roughly like this:

<link
  rel="canonical"
  href="https://www.disneyplus.com/identity/login/enter-email?utm_source=chatgpt.com"
/>

At first glance, it looks like nothing more than a tracking parameter.

But seeing it inside a canonical URL raises a much more interesting technical SEO question.


First: what is a canonical URL?

A canonical tag tells search engines which URL should be treated as the preferred version of a page.

For example, all of these URLs may display exactly the same content:

example.com/product
example.com/product?utm_source=google
example.com/product?utm_source=newsletter
example.com/product?ref=homepage

Ideally, they would all point to:

https://example.com/product

as their canonical URL.

Effectively, the site is telling search engines:

These URL variations may exist, but this is the primary version of the page.


So how could utm_source=chatgpt.com get into the canonical?

One possible explanation is that the site's canonical URL is being dynamically generated from the current request URL.

A simplified implementation might look like this:

const canonical = window.location.href;

If someone arrives through:

https://example.com/login?utm_source=chatgpt.com

the application could accidentally output:

<link
  rel="canonical"
  href="https://example.com/login?utm_source=chatgpt.com"
/>

The interesting part here isn't ChatGPT itself.

The real issue is that tracking parameters appear to have leaked into canonical URL generation without proper URL normalization.


Why shouldn't tracking parameters normally appear in canonical URLs?

UTM parameters are designed to identify traffic sources and campaigns.

For example:

?utm_source=google
?utm_source=linkedin
?utm_source=chatgpt.com
?utm_campaign=summer-sale

They can be useful for analytics.

But they normally don't change the underlying content.

In other words:

/product

and:

/product?utm_source=linkedin

usually represent the same page.

If canonical URLs preserve those tracking parameters, a site could potentially generate different canonical URLs depending on the visitor's acquisition source:

/product?utm_source=google
/product?utm_source=linkedin
/product?utm_source=chatgpt.com
/product?utm_source=newsletter

If every request declares itself canonical, the purpose of canonicalization starts to break down.


Does Google always follow the canonical tag?

No.

This distinction matters.

Google generally treats rel="canonical" as a signal rather than an absolute directive.

Canonical selection can also be influenced by signals such as:

  • redirects,

  • internal linking,

  • sitemap URLs,

  • content similarity,

  • HTTPS,

  • consistency between URL signals.

So an incorrect canonical tag doesn't automatically mean a page will disappear from search or suffer an immediate ranking problem.

Google may still choose the cleaner URL itself.

But that doesn't make the implementation desirable.

At scale, inconsistent canonicalization can contribute to unnecessary crawling, duplicate URL discovery and conflicting indexing signals.


The bigger issue: URL normalization

The more interesting lesson here isn't really about canonical tags.

It's about URL normalization.

Applications need to distinguish between parameters that change the resource and parameters that only describe how a user reached it.

For example:

/product?id=123

may legitimately identify a specific resource.

But:

/product?id=123&utm_source=twitter

still represents the same product.

Tracking parameters can therefore be removed when generating the canonical URL.

For example:

const url = new URL(request.url);

[
  "utm_source",
  "utm_medium",
  "utm_campaign",
  "utm_term",
  "utm_content",
  "gclid",
  "fbclid"
].forEach(param => {
  url.searchParams.delete(param);
});

The normalized URL can then be used for the canonical tag.


Canonicalization isn't enough on its own

Clean URL management requires more than adding a canonical tag.

If a website generates many versions of the same URL, several signals should be checked together.

Canonical tags

Does every variation point to the correct preferred URL?

Redirects

Are obsolete or unnecessary URLs redirected when appropriate?

Internal links

Is the site itself continuously linking to parameterized URLs?

Sitemaps

Does the sitemap contain only canonical URLs?

404 and legacy URLs

What happens when users or bots request URLs that no longer exist?

All of these are parts of the same URL management problem.


AI referrals will make this increasingly visible

Historically, marketers mostly dealt with referral and campaign identifiers such as:

google
facebook
twitter
newsletter

Now websites are increasingly receiving traffic from sources such as:

chatgpt.com
perplexity.ai
copilot.microsoft.com
gemini.google.com

As AI-driven discovery grows, AI-related attribution parameters are likely to become more visible in analytics and URLs.

That makes good URL normalization even more important.

The problem isn't:

utm_source=chatgpt.com

The problem is treating that parameter as if it were part of the page's permanent identity.


Conclusion

A single utm_source=chatgpt.com parameter inside a Disney+ canonical tag may look trivial.

But it highlights a useful technical SEO principle:

Tracking information and URL identity should be treated as two different things.

Canonical URLs should remain clean and stable regardless of where a visitor came from.

And URL management shouldn't stop at canonical tags.

404s, legacy URLs, redirects, query parameters and canonical signals are all parts of the same URL lifecycle.

That's also where No404 focuses: detecting broken or outdated requests, matching them with relevant live URLs, and recovering traffic that would otherwise end at a dead page.

Detect → Match → Redirect → Recover.