Why HTTP, HTTPS, WWW and Non-WWW "4 Doors" matter
*Button to download a practical checklist at the bottom
Sometimes digital maturity is not about introducing another platform. Sometimes it is about testing the assumptions underneath the platforms you already have. So perhaps take five minutes today and knock on all four doors.

There is a wonderfully simple technical question I like to ask when looking at a website:
What is your website address?
Most people answer confidently.
site.com
Excellent. Case closed. Except, technically speaking, your website may have a number of different ways of being reached:
https://site.com
https://www.site.com
http://site.com
http://www.site.com
To a human being, these are obviously the same website. However, to the various browsers, servers, crawlers, analytics platforms, ad platforms and assorted machines involved in modern digital marketing, things are slightly less relaxed.
The HTTP and HTTPS versions use different protocols. The www and non www versions use different hostnames. Depending on how your DNS, SSL, CDN, redirects, server rules and website configuration have been set up, each of these requests can potentially behave differently.
Which is comforting.
Google itself identifies protocol variants, including HTTP and HTTPS versions of the same page, as situations where duplicate URLs can exist and canonicalisation becomes important.
So I tend to think of these as your website's four front doors. Ideally, three of them should politely direct everybody towards the fourth.
That is the theory.
The more entertaining question is whether anybody has actually checked. Because what appears to be a tiny technical detail can quietly influence a surprisingly large portion of your digital marketing operation. And, as is often the case with websites, everybody assumes someone else checked.
Four Doors, Hopefully One Destination
Let us assume your chosen website address is:
A sensible setup would mean that requests for:
http://site.com
http://www.site.com
https://site.com
all permanently redirect to:
https://www.site.com
More importantly, the same logic should apply to individual pages.
If someone requests:
http://site.com/services
they should eventually arrive at:
https://www.site.com/services
rather than being helpfully deposited on the homepage with no explanation.
This sounds obvious.
And when something in technology sounds obvious, there is usually a reasonable chance it was configured eight years ago by someone who has since left the company.
Websites evolve. Hosting changes. DNS records change. SSL certificates are added. CDNs arrive. Developers migrate platforms. Agencies change. Campaign landing pages appear. Tracking scripts multiply. Subdomains reproduce quietly in the background.
Redirect rules get added for very good reasons and then remain there long after anyone remembers what those reasons were.
Before long, what everyone believes to be one beautifully organised website is actually a collection of historical decisions held together by redirects and optimism.
Google and other WebCrawler’s and indexing platform recommend choosing a preferred canonical destination when a homepage is available through multiple URLs and permanently redirecting alternatives to it.
In other words:
If you know where you want people and machines to go, tell them. Do not turn basic website configuration into an escape room.
Does Google Really Think You Have Four Different Websites?
Not exactly.
That would be an unnecessarily dramatic interpretation.
Google is quite capable of recognising that:
http://site.com
and:
https://www.site.com
may represent the same organisation and substantially the same content. Search engines perform canonicalisation.
When Google encounters multiple URLs containing the same or very similar content, it attempts to group them together and select a representative version.
That selected version is called the canonical URL.
Google also generally prefers HTTPS versions of equivalent pages.
So the point is not:
"Google sees four websites and immediately has a nervous breakdown."
The better question is:
Why are we making Google, Bing, analytics platforms, advertising systems and AI crawlers work out something we could simply configure properly ourselves?
Machines are becoming very good at compensating for our mess. This should not be confused with permission to create more of it. And today, considerably more than Google is looking at your website. Bing crawls it. Advertising platforms inspect it. Social networks retrieve URLs to create previews. Analytics systems record page locations and hostnames. Security platforms analyse requests. AI tools retrieve information. OpenAI operates OAI-SearchBot for content that may appear in ChatGPT search experiences. Microsoft Bing increasingly supports not only traditional search, but also AI powered experiences where indexed web content can be used to support answers.
So the number of systems knocking on the front door of your website is increasing. It would therefore be mildly useful if all four doors behaved sensibly.
The Tiny Technical Detail That Quietly Touches Everything
The reason this matters is not merely SEO. Your website sits underneath a significant portion of your marketing technology stack. If the foundation is inconsistent, those inconsistencies can surface in some surprisingly unrelated places.
1. Crawling and Indexing
Let us start with the obvious one.
If several URL variants return the same content rather than redirecting towards a single preferred version, search crawlers may encounter several URLs representing essentially the same page.
This is not automatically disastrous. Google is clear that duplicate content is not inherently a spam violation. The internet contains duplication everywhere. The problem is less about punishment and more about unnecessary ambiguity. Google must decide which version to index. Signals may need to be consolidated. Crawling effort can be spent visiting versions you do not actually care about. Reporting can become messier.
All of this is manageable.
Of course, so is not creating the problem in the first place.
2. Canonicalisation
Canonicalisation is essentially your way of saying:
Of the several URLs that might show this content, this is the one I consider primary.
For example:
https://www.site.com/services
might include a canonical reference pointing back to itself.
That is useful.
But canonical tags should not become the technical equivalent of writing "please tidy this up" on top of poor architecture.
The strongest setup is one where everything agrees. Unwanted URL variants permanently redirect to the preferred version. The preferred page references itself canonically. Internal links point to it. The XML sitemap contains it. Campaign links use it. HTTPS is consistent. The server agrees. The CMS agrees. The analytics setup agrees.
Everybody has, for once, received the same memo.
3. Analytics and Reporting
This is where a small technical inconsistency starts wandering into rooms it was never invited into.
Google Analytics can record both hostname and full page location.
site.com
and:
www.site.com
are different hostnames.
That does not mean GA4 automatically treats them as four entirely separate properties. But it does mean inconsistent URL behaviour can appear in the data. Now imagine a reporting team attempting to answer:
Which landing pages are performing best?
Where are users entering?
Which campaigns drove conversions?
Which content is generating leads?
Why does one dashboard say 42,000 sessions while another says 39,800?
At this point somebody normally says:
"It's probably attribution."
A useful sentence because it sounds technical enough to temporarily end the conversation.
Sometimes it is attribution. Sometimes your basic URL configuration is just messy. And messy inputs have a remarkable ability to produce beautifully designed dashboards.
4. Paid Search
Suppose one paid search campaign links to:
http://site.com/product
Another links to:
https://site.com/product
A third links directly to:
https://www.site.com/product
All three may ultimately reach the correct page.
So technically, everything may "work".
Wonderful.
But each additional redirect or variation introduces another moving part between:
the person clicking an advert
and:
the page on which you expect to measure that person.
Paid media measurement already involves destination URLs, tracking parameters, click identifiers, cookies, consent, attribution windows and platform specific logic.
It has enough hobbies. There is no compelling reason to add inconsistent website URLs to the list. Paid campaigns should point directly at the preferred final destination wherever possible. Not eventually arrive there after a short scenic tour of your redirect architecture.
5. Paid and Organic Social
The same issue appears in social media.
One person copies:
https://site.com
Another uses:
https://www.site.com
An old post links to:
http://www.site.com
Someone finds a ten year old campaign URL.
An influencer uses another variation.
Someone shortens the URL because apparently five extra characters are now an obstacle to civilisation.
If your infrastructure is sound, these should all resolve correctly.
But again, the point is governance.
Why allow everyone involved in marketing to use a different representation of the same destination? A preferred URL convention is a very small thing to establish. And small pieces of governance tend to be considerably cheaper than large pieces of forensic analytics later.
6. Attribution
Attribution is already an area where otherwise reasonable adults can spend several hours discussing who deserves credit for a conversion.
There is little need to make it more exciting.
Digital attribution can rely on combinations of:
- referrers
- campaign parameters
- click identifiers
- cookies
- events
- sessions
- users
- devices
- channels
- conversion paths.
A correctly implemented redirect does not magically destroy attribution.
Redirects are normal. But long redirect chains, inconsistent tracking setups, lost query parameters or host specific configuration differences can certainly make the environment harder to reason about.
The principle is not particularly glamorous:
Reduce unnecessary changes between click and conversion.
Unfortunately "correctly configure the redirects" tends not to receive the same conference speaking slots as "AI powered predictive omnichannel orchestration".
But it is still rather useful.
7. Site Speed
Every redirect requires another request.
For example:
http://site.com
responds:
"Actually, please go over here."
The browser then requests:
https://www.site.com
This is not a catastrophe. The internet will survive. But unnecessary redirects are still unnecessary work. The more entertaining version is the redirect chain:
http://site.com
to:
https://site.com
to:
https://www.site.com
to:
https://www.site.com/en/
At which point the visitor has technically travelled more than some people do on holiday before reaching your homepage. Each step may introduce latency. Each is another interaction with infrastructure. Each is another potential failure point.
The sensible approach is therefore to send alternate URLs directly to their final destination wherever possible.
Very boring but very effective.
8. Research, Audience Creation and Analysis
This is perhaps less obvious.
Marketing teams increasingly combine data from numerous platforms:
Analytics.
CRM.
Social listening.
Search platforms.
Advertising systems.
Data warehouses.
AI tools.
Competitive research.
Audience platforms.
Web crawling.
Now imagine that the same organisation appears as:
site.com
www.site.com
http://site.com
and:
https://www.site.com
A human immediately understands these are related. A machine understands exactly what its rules and data model tell it. If you are joining, grouping or analysing data using URLs or domains as identifiers, inconsistency needs to be normalised somewhere.
So once again, a tiny technical inconsistency at source quietly creates work elsewhere. Data engineers have a technical term for this.
I believe it is: - "Why is this like this?"
9. AI Search and Answer Engines
And then we arrive at the newer part of the conversation. Websites are no longer being built purely for traditional search engines.
AI based discovery is becoming increasingly relevant. ChatGPT can search and cite information from websites. Microsoft uses indexed web information in Copilot experiences. Google increasingly blends generated answers with traditional search.
Other AI platforms retrieve and interpret web content in different ways.
The practical implication is important. Your website increasingly serves two audiences:
Humans
and:
Machines trying to help humans.
Those machines need predictable access to information.
If your various URL versions behave differently, return different status codes, apply different crawler rules, expose inconsistent canonical tags or trigger different security behaviour, you are creating ambiguity at the exact point where machines are trying to understand your content.
Again, this does not mean AI systems instantly fail because you forgot to redirect http://.
Modern systems are resilient. They are quite capable of working around many configuration issues.
But we seem to have developed an odd habit in technology where the increasing intelligence of machines is used to justify decreasing discipline from humans.
- AI can probably work it out.
- Google can probably work it out.
- Analytics can probably clean it up.
- Development can probably patch it.
- Data can probably normalise it.
This is all technically true.
It is also not a particularly impressive operating model.
The Ideal Setup Is Deeply Unexciting
Fortunately, the correct setup does not require a twelve month transformation roadmap.
It is spectacularly dull. Which, in infrastructure, is often an excellent sign.
Choose your preferred format.
For example:
Then make everything agree.
Redirects
Every unwanted protocol and hostname variation should permanently redirect directly to its equivalent preferred URL.
Not eventually.
Directly.
HTTPS
Use HTTPS consistently.
Make sure the certificate works properly.
A revolutionary concept.
Canonicals
Your indexable pages should use appropriate canonical references pointing towards the preferred version.
Internal Links
Your own website should link directly to final preferred URLs.
There is little point creating a redirect and then making your own navigation repeatedly trigger it.
That is effectively leaving yourself directions to the wrong address and relying on a road sign to correct you every time.
XML Sitemap
Your XML sitemap should contain URLs you actually want indexed.
Preferably the same URLs your canonical tags say you want indexed.
Again, consistency.
Marketing Links
Paid search.
Paid social.
Email.
Organic social.
Influencer campaigns.
QR codes.
Campaign landing pages.
All should use the preferred destination format.
Your official website address should not depend on which member of the marketing department copied it.
Analytics and Tracking
Check that tracking loads properly on the preferred destination.
Check that required campaign parameters survive redirects.
Check hostname reporting.
Check page locations.
Check that your analytics implementation has not quietly evolved differently across different variants of the domain.
Because nothing says "single source of truth" quite like three subtly different implementations of the same analytics script.
Robots and AI Crawlers
Review robots.txt.
Review llms.txt
Understand which crawlers you want to allow.
Understand which you do not.
Do not accidentally block a crawler you expect to surface your content.
Also do not throw the doors open to absolutely everything because somebody attended a webinar about Generative Engine Optimisation.
Intentional configuration remains fashionable.
CDN, Firewall and Bot Protection
Modern websites often sit behind CDNs, web application firewalls and bot management systems.
These are extremely useful.
They are also capable of blocking legitimate automated traffic if configured incorrectly.
So your ability to be discovered by search or AI platforms may depend not only on your CMS and SEO setup, but also on infrastructure and security rules.
Which means your SEO person may eventually need to speak to your infrastructure person.
Please ensure there is coffee available.
A Five Minute Test You Can Do Yourself
This is the bit I rather enjoy because it requires no enterprise strategy document.
Take your own domain and try all four versions:
https://yourdomain.com
https://www.yourdomain.com
http://yourdomain.com
http://www.yourdomain.com
Then see what happens.
Do all four end up at the exact same preferred address?
Do they all resolve securely?
Do any fail?
Does one behave differently?
Does one generate a certificate warning?
Do some pass through multiple redirects?
Then test a page deeper in the site.
For example:
http://yourdomain.com/services
Does it arrive at:
https://www.yourdomain.com/services
or does it dump you at the homepage and effectively say:
"I know you were looking for something specific, but here is the logo."
Then continue.
Check your canonical tag.
Check your XML sitemap.
Check internal links.
Check Google Search Console – This is really important and you can report on each “Door” independently on Search console.
Check analytics hostname reporting.
Check paid campaign URLs.
Check robots.txt.
Check whether the search and AI crawlers relevant to your business can actually access your content.
If everything works perfectly, excellent.
Close the tab and go back to discussing AI.
If it does not, you have discovered something rather useful from a test that took less time than most meetings spend discussing the agenda.
This Is Not Really an SEO Problem
There is a tendency in organisations to hear words such as:
- redirects
- HTTP
- HTTPS
- canonical tags
- DNS
- SSL
- crawlability
and immediately classify the conversation as:
SEO stuff.
Which is a convenient way for everybody else to leave. But websites are not merely an SEO asset. They are often the primary connection point between:
Media. Content. Data. CRM. Analytics. Lead generation. Customer experience. Research. Sales. Automation.
And increasingly, AI powered discovery.
So website architecture is not simply the concern of the person responsible for organic rankings. It is digital infrastructure. SEO may be the department that first notices the issue. That does not mean SEO is the only place affected by it.
Machines Are Very Good at Cleaning Up After Us
There is an interesting broader principle here.
Modern technology is increasingly capable of compensating for weak technical discipline. Google can choose a canonical. Analytics can filter hostnames. Developers can redirect URLs. Data engineers can normalise domains. AI systems can infer that two URLs refer to the same organisation.
Fantastic.
But there is a subtle trap in that.
The better machines become at compensating for disorder, the easier it becomes for humans to tolerate disorder.
We start accepting:
"The platform will figure it out."
rather than asking:
"Why are we making the platform figure it out?"
There is a significant difference between those two operating standards.
Good digital environments generally reduce ambiguity wherever possible.
One naming convention. One campaign taxonomy. One measurement standard. One customer identifier where possible. One definitive version of each web page.
Simple inputs make sophisticated systems perform better. It is not terribly exciting. But neither is foundation concrete. Most people still prefer their building to have it.
The Larger Lesson Behind Four Very Boring URLs
The reason I like the four doors example is precisely because it appears trivial.
Digital transformation discussions tend to focus on impressive things. Generative AI. Predictive modelling. Personalisation. Automation. CDPs. Machine learning. Agentic workflows. Advanced attribution.
Everyone wants to work on the impressive layer.
Nobody puts:
"Checked whether HTTP redirects properly"
on the first slide of the innovation strategy.
Yet sophisticated technology sits on top of basic technical foundations. A brilliant attribution model still depends on trustworthy tracking. A sophisticated AI system still depends on accessible source information. A beautifully designed performance dashboard still depends on clean data.
And millions in media spend can still ultimately send somebody through a series of badly configured redirects because nobody checked the plumbing.
Sometimes digital maturity is not about introducing another platform. Sometimes it is about testing the assumptions underneath the platforms you already have. So perhaps take five minutes today and knock on all four doors.
Hopefully, each one takes you to exactly the same place. If it does, excellent.
Nobody will congratulate you. Which is generally how you know infrastructure is working properly.
If it does not, however, that tiny technical detail may suddenly become considerably more interesting.
So for those who are still here – please see the link below to a basic checklist to which may be of help if you want to add this to a compliance file.
Onwards and upwards
(Downloads)
- Download
Steven_in_the_loop_four_doors_audit_checklist.xlsx
VND.OPENXMLFORMATS-OFFICEDOCUMENT.SPREADSHEETML.SHEET · 19.6 KB


