SEO guide · Chapter 5 · 14 min read
Technical SEO: make sure Google can find, read and store your site
Technical SEO is the work that lets search engines reach, render and index every important page of your website. This chapter explains crawling and indexing, robots.txt, sitemaps, canonical tags, redirects, hreflang and JavaScript in plain language, with a checklist.
By Lennox de Wolff, founder of Delta Six · updated on
This is chapter 5 of 13 of the Delta Six SEO guide.
Key takeaways
- A page can only rank after Google has crawled it, rendered it and stored it in the index.
- The Page indexing report in Google Search Console shows exactly which pages are in the index and why others are left out.
- The robots.txt file controls crawling. A noindex tag controls indexing. They do different jobs.
- An XML sitemap lists the pages you want indexed, and a canonical tag names the preferred version of similar pages.
- Use a 301 redirect for every URL that moves, and keep it in place for at least a year.
- Put your main content and links in the HTML itself so every crawler can read them, including AI crawlers.
The basics
What is technical SEO? The foundation under every ranking.
Technical SEO is everything that helps search engines access and understand your website as a whole. It deals with the structure and the code of the site, where on-page work deals with the content of a single page.
Think of a shop. Your content is the product on the shelves. Technical SEO is the building: the doors open, the lights work, the aisles are labeled. A great product behind a locked door sells nothing.
The good news for a small business: a site with a few dozen pages has few technical needs. Get the basics in this chapter right once, check them a few times a year, and spend the rest of your energy on content.
- AccessSearch engines can reach every important page.
- UnderstandingThe site sends clear signals about which URL is the right one for each piece of content.
- ExperiencePages are secure, work on phones and load quickly. Chapter 6 covers speed.
How search works
Crawling and indexing: how Google processes a page
Google's documentation 'How Search works' describes three stages: crawling, indexing and serving results. For modern sites there is a rendering step in between. Technical SEO removes the obstacles at each stage.
Crawling is done by Googlebot, the program that fetches pages. It discovers URLs by following links and by reading sitemaps. Google uses its smartphone crawler for indexing, so the mobile version of your page is the one that counts.
Indexing means Google analyzes the page and stores it in its database, the index. Google chooses which pages earn a place. Pages with little unique value, or pages that duplicate others, are often left out.
| Stage | What happens | What can go wrong |
|---|---|---|
| Discovery | Google learns a URL exists through links or a sitemap | The page has zero links pointing to it |
| Crawling | Googlebot requests the page from your server | Blocked by robots.txt, server errors, slow responses |
| Rendering | Google runs the code to see the page as a browser would | Content appears only after a click or a failed script |
| Indexing | Google stores the page and picks a canonical URL | noindex tag, duplicate content, thin content |
| Serving | The page appears for relevant searches | Stronger competing pages, weak relevance |
First check
Check which of your pages Google has indexed
Start every technical review with one question: are my important pages in the index? Two free methods give the answer within minutes.
The quick method is the site: search. Type site:yourdomain.com into Google. You see a sample of indexed pages. Add a word to check a section, such as site:yourdomain.com/services.
The precise method is Google Search Console, Google's free tool for site owners. Under Indexing, open the Pages report, officially called the Page indexing report. It shows the number of indexed pages and lists reasons for every page left out.
For a single page, paste its address into the URL Inspection tool at the top of Search Console. You see when Google last crawled it, which canonical URL Google chose, and whether the page is indexed.
- Step 1Verify your site in Search Console, ideally as a domain property through your DNS settings.
- Step 2Open Indexing, then Pages, and note the totals.
- Step 3Write down your 20 most important URLs and inspect any that are missing from the index.
- Step 4After a fix, use 'Request indexing' in URL Inspection and 'Validate fix' in the report.
The offer
Your complete website, built free first. You decide once you have seen it.
- 01Built free firstWe build your complete website up front. You pay nothing until you have seen it.
- 02See it in 15 minutesIn a 15-minute online showcase we walk through your new website together.
- 03Live within two weeksChoose us and your website is online within two weeks.
- 04Money-back guaranteeTwo months after launch you choose again. Stop after two improvement rounds and you get your money back.
- 05Everything includedOnline booking, an AI chatbot, a lead magnet, review follow-up and follow-up by email and WhatsApp.
- 06Made for your brandPremium 3D design in your colors, with your story and your customers up front.
The report
How to read the Page indexing report
The report can look alarming, since nearly every site has pages outside the index. Many of those are fine: redirected URLs, duplicates with a correct canonical tag and pages you excluded on purpose.
Focus on important pages that are missing. The table lists common statuses in the report and the action each one calls for.
| Status in the report | Meaning | Action |
|---|---|---|
| Page with redirect | The URL forwards to another URL | Fine when the redirect is intended |
| Alternate page with proper canonical tag | A duplicate that points to the main version | Fine, this is the tag doing its job |
| Excluded by 'noindex' tag | The page asks to stay out of the index | Remove the tag if the page should rank |
| Blocked by robots.txt | Googlebot is told to stay away | Remove the rule if the page matters |
| Soft 404 | The page looks empty or like an error page | Add real content or return a true 404 |
| Server error (5xx) | Your server failed to respond properly | Check hosting and server logs |
| Duplicate, Google chose different canonical than user | Google prefers another URL | Align links, sitemap and canonical tag |
| Crawled or discovered, awaiting indexing | Google knows the URL and has yet to index it | Improve content and add internal links |
robots.txt
robots.txt: tell crawlers where they may go
The robots.txt file is a small text file at the root of your site, at yourdomain.com/robots.txt. It tells crawlers which parts of the site they may request. Every well-behaved crawler reads it before anything else.
Google's documentation stresses one point: robots.txt manages crawling, and it is a poor tool for keeping a page out of Google. A blocked URL can still appear in results, shown as a bare link, when other sites link to it.
Most small business sites need only a few lines. Allow everything, block a handful of private or endless areas such as internal search results and shopping carts, and state where your sitemap lives.
- User-agent: *The rules below apply to all crawlers.
- Disallow: /cart/Crawlers skip every URL that starts with /cart/.
- Disallow:An empty Disallow line means the whole site is open.
- Disallow: /A single slash blocks the entire site. This one line has removed many sites from Google after a launch.
- Sitemap: https://yourdomain.com/sitemap.xmlTells crawlers where to find your sitemap.
Keep CSS and JavaScript files open for crawling. Google needs them to render your pages the way a visitor sees them.
Two tools
robots.txt or noindex: which one do you need?
A noindex rule is a line in the HTML head of a page, or in the HTTP response header, that asks search engines to leave the page out of the index. Google must be able to crawl the page to see that request.
This creates a classic mistake: blocking a page in robots.txt and adding noindex. Google then stays away, misses the noindex and may keep the URL in the index. Choose one tool per goal.
| Goal | Right tool | How |
|---|---|---|
| Keep a page out of search results | noindex | Meta robots tag with noindex, page stays crawlable |
| Stop crawlers wasting time on endless URLs | robots.txt | Disallow rule for the folder or pattern |
| Keep private content private | Login | Password protection on the server |
| Remove a page for good | 404 or 410 status | Delete the page, or redirect it to a fitting replacement |
| Merge duplicate versions | Canonical tag or redirect | Point every version to the preferred URL |
Sitemaps
The XML sitemap: your list of pages for Google
An XML sitemap is a file that lists the URLs you want search engines to index. XML is simply a structured text format. The file usually lives at yourdomain.com/sitemap.xml, and most website platforms generate it automatically.
A sitemap helps discovery. It is most valuable for new sites, large sites and pages with few links. It is a suggestion to Google and gives zero guarantee of indexing.
Google's sitemap documentation gives the limits: 50,000 URLs or 50 MB uncompressed per file. Google ignores the priority and changefreq values. It does use the lastmod date when that date is consistently accurate.
- IncludeOnly indexable pages that return status 200 and are their own canonical version.
- ExcludeRedirected URLs, noindex pages, error pages and duplicate versions.
- SubmitIn Search Console, open Sitemaps, enter the address and press Submit.
- ReferenceAdd the Sitemap line to your robots.txt so every crawler finds it.
- MonitorCheck the Sitemaps report for errors and for the number of discovered pages.
Structure
Site architecture: every page within three clicks
Site architecture is the way your pages are organized and linked. Google's crawler travels along links, so a page that sits far from the home page gets visited less often and receives less authority.
Aim for a flat structure. A visitor should reach any important page from the home page in three clicks or fewer. Main menu, category pages and contextual links in the text make that possible.
Add breadcrumbs, the small trail of links that shows the path to the current page. They help visitors orient themselves and give crawlers a clear picture of your hierarchy.
- Level 1Home page.
- Level 2Main sections such as services, shop, about, blog and contact.
- Level 3Individual service pages, categories and articles.
- Level 4Detail pages such as products, cases and sub-services.
Status codes
HTTP status codes every site owner should know
Every time a browser or crawler requests a URL, your server answers with a three-digit status code. The code tells Google how to treat the page. You can see it in URL Inspection or in free online header checkers.
| Code | Meaning | Effect on SEO |
|---|---|---|
| 200 | OK, the page exists | Page can be indexed |
| 301 | Moved permanently | Google indexes the new URL and passes signals on |
| 302 | Moved temporarily | Google usually keeps the old URL indexed |
| 404 | Page is missing | Page drops out of the index over time |
| 410 | Gone on purpose | Like 404, a clear signal of removal |
| 500 | Server error | Crawling slows down, pages drop if it persists |
| 503 | Temporarily unavailable | Correct code for short maintenance |
A 404 for a page that truly ended is healthy. The problems are 404s on pages that still have links or traffic, and soft 404s: empty or error-like pages that still return code 200.
Redirects
Redirects: move pages and keep your rankings
A redirect sends visitors and crawlers from one URL to another. Use a 301 redirect whenever a page moves for good: a new URL, a merged page, a new domain or a switch from http to https.
Google's documentation on site moves advises keeping redirects in place as long as possible, and at least one year. That gives Google time to transfer all signals and catches visitors who use old links.
Redirect each old URL to the most relevant new URL. Sending every old page to the home page confuses visitors, and Google may treat such redirects as soft 404s.
- Map before you moveList every old URL next to its new destination before a redesign or migration.
- Avoid chainsA to B to C wastes time. Point A directly to C. Googlebot follows up to 10 hops and then stops.
- Avoid loopsA to B and B back to A makes the page unreachable.
- Update internal linksChange links on your own site to the final URL so crawlers skip the detour.
- Test after launchCrawl the old URL list and confirm each one returns a 301 to a page with status 200.
Canonicals
The canonical tag: one preferred URL per page
The same content is often reachable at several addresses. Think of a product in two categories, a URL with tracking codes, or one page with a trailing slash and a twin that omits it. Google sees each address as a separate page.
A canonical tag is a line in the HTML head that names the preferred address: rel='canonical' with the full URL. Google then combines the signals of all versions on that one URL.
Google's documentation calls the canonical tag a strong hint, and Google makes the final choice. It also weighs redirects, internal links, the sitemap and https. So keep all those signals pointing at the same URL.
- Self-referenceGive every indexable page a canonical tag that points to itself.
- Use full URLsWrite https://yourdomain.com/page/ in full, including the protocol.
- One per pageTwo canonical tags on a page cancel each other out.
- Point to live pagesThe canonical URL returns status 200 and is free of noindex.
- Stay consistentLink internally to the canonical version and list only that version in the sitemap.
Duplicates
Where duplicate content comes from
Duplicate content means the same or nearly the same content at more than one URL. Google's guidance is calm about it: there is zero penalty for ordinary duplication. The cost is diluted signals and wasted crawling.
Most duplication is created by the website system itself. Check the usual sources below and fix each with a redirect or a canonical tag.
- http and httpsBoth versions load. Fix: 301 redirect everything to https.
- www and bare domainBoth versions load. Fix: pick one and redirect the other.
- Trailing slash and capitals/Page and /page/ both work. Fix: one format, redirect the rest.
- URL parametersSorting, filters and tracking codes create endless versions. Fix: canonical tag to the clean URL.
- Print and staging copiesTest sites and print versions get indexed. Fix: password protection for staging, canonical for print.
Security and mobile
HTTPS and mobile-first indexing
HTTPS encrypts the connection between visitor and website. Google confirmed it as a ranking signal in 2014, and browsers label plain http pages as insecure. Every page, image and script should load over https.
Mobile-first indexing means Google crawls and judges the mobile version of your site. Google completed that switch for the web as a whole. Whatever is missing on mobile is missing for Google.
- CertificateInstall a valid SSL certificate and renew it automatically.
- RedirectSend every http URL to its https twin with a 301.
- Mixed contentReplace http links to images and scripts inside https pages.
- Same content on mobileShow the same text, headings, links and structured data on phone and desktop.
- Responsive designUse one URL per page with a layout that adapts to the screen, as Google recommends.
Languages
Hreflang: the right language for the right visitor
Hreflang is a tag that tells Google which language and country version of a page to show. You need it only when the same content exists for several languages or regions, such as English for the US and for the UK.
Each version lists all versions, itself included. Google's documentation requires this to be mutual: when page A names page B, page B names page A. Missing return links are the most common hreflang error.
Use standard codes: a two-letter language code such as en or nl, optionally with a country code such as en-GB or nl-BE. Add x-default for the page that serves everyone else.
| Situation | Example values | Remark |
|---|---|---|
| English, all countries | en | Language only |
| English for the United Kingdom | en-GB | GB is the code, UK is invalid |
| Dutch for Belgium | nl-BE | Language first, then country |
| Fallback for all other visitors | x-default | Often the English or language-picker page |
You can place hreflang in the HTML head, in the HTTP header or in the XML sitemap. Pick one method and apply it to every language version.
JavaScript
JavaScript SEO: make sure your content is in the HTML
JavaScript is the programming language that makes pages interactive. Many modern sites use it to build the whole page in the browser. JavaScript SEO is making sure search engines still see that content.
Google can run JavaScript. Its documentation describes a second phase: Google first reads the raw HTML, then places the page in a queue for rendering with a current version of Chrome. Content in the raw HTML is processed soonest.
Google's crawler loads pages and skips clicking, scrolling and typing. Content that appears only after an action stays unseen. Links must be standard anchor elements with an href address to be followed.
The dependable solution is server-side rendering or static generation: the server sends complete HTML, and JavaScript adds interaction afterward. Google's documentation recommends these approaches over workarounds.
- Test 1Open 'View page source'. Is your main text visible in the code?
- Test 2Use URL Inspection, 'Test live URL', then view the rendered HTML and screenshot.
- Test 3Check that menu and content links are real anchor elements with an href.
- Test 4Confirm that titles, canonical tags and structured data are in the initial HTML.
AI crawlers
Technical basics for AI search
AI assistants such as ChatGPT, Claude and Perplexity use their own crawlers, with names like GPTBot, ClaudeBot and PerplexityBot. Your robots.txt decides whether they may read your site.
An analysis published by Vercel in December 2024 found that the major AI crawlers fetched JavaScript files and skipped executing them. So content that exists only after rendering may be invisible to those assistants.
The practical rule is the same as for Google, only stricter: put the content in the HTML. Chapter 12 of this guide covers optimization for AI search in depth.
Scale
Crawl budget: who needs to care?
Crawl budget is the number of URLs Google is willing and able to crawl on your site in a given period. Google's guide on the subject is aimed at very large sites, from roughly a million pages, or large sites that change daily.
A business site with a few hundred pages is crawled comfortably. Your attention belongs elsewhere, unless the site generates endless URLs through filters, calendars or internal search.
You can see Google's activity under Settings, Crawl stats in Search Console. Look at the share of requests that end in errors and at the average response time of your server.
Audit
Technical SEO checklist: a one-hour audit
A technical seo checklist is most useful as a routine you repeat every quarter and after every redesign. The steps below need only Search Console and your browser.
For a deeper look, a crawler tool visits every page the way Googlebot does and lists errors. Screaming Frog SEO Spider is a widely used one, with a free mode for small sites. Chapter 10 compares SEO tools.
- Minute 0 to 10Open the Page indexing report. Confirm your 20 key pages are indexed.
- Minute 10 to 20Read your robots.txt line by line and open your XML sitemap. Remove wrong rules and dead URLs.
- Minute 20 to 30Test http, https, www and bare domain. All should land on one version with a single 301.
- Minute 30 to 40Inspect five key URLs. Check the Google-selected canonical and the rendered HTML.
- Minute 40 to 50Review pages with status 404 and 5xx. Redirect or repair those with links or traffic.
- Minute 50 to 60Check the site on your phone: same content, readable text, working menu and forms.
Our approach
How Delta Six handles the technical side
Delta Six builds websites on its own platform with hand-written code. That gives us direct control over the HTML each page sends, the sitemap, canonical tags, redirects and https.
Monthly SEO and optimization for AI search are included as standard. You see your complete website in a 15-minute online showcase, and you pay only when you are happy.
Checklist
Work through this chapter. Point by point.
- Verify your site in Google Search Console and open the Page indexing report.
- Confirm that your 20 most important pages are indexed, using URL Inspection.
- Read your robots.txt and remove any rule that blocks pages, CSS or JavaScript you need indexed.
- Use noindex for pages that should stay out of search, and keep those pages crawlable.
- Submit an XML sitemap that lists only indexable pages with status 200, and add the Sitemap line to your robots.txt.
- Redirect http to https and choose either www or the bare domain, with a single 301.
- Give every indexable page a self-referencing canonical tag with the full URL.
- Redirect every moved or merged URL with a 301 to its closest replacement.
- Remove redirect chains and update internal links to the final URL.
- Repair or redirect 404 pages that still receive links or visitors.
- Make sure every important page is reachable within three clicks from the home page.
- Show the same content and links on mobile as on desktop.
- Confirm that main content and links are present in the page source.
- Add mutual hreflang tags if you serve several languages or countries.
- Repeat this audit every quarter and after every redesign or migration.
Frequently asked questions
Good to know.
What is technical SEO?
Technical SEO is optimizing the structure and code of a website so search engines can crawl, render and index it. It covers robots.txt, sitemaps, redirects, canonical tags, https, mobile setup and site speed.
What is the difference between crawling and indexing?
Crawling is Google fetching a page. Indexing is Google analyzing that page and storing it in its database. A page must be crawled and indexed before it can appear in search results.
Does robots.txt remove a page from Google?
It only stops crawling. A blocked URL can still show up as a bare link when other sites link to it. To keep a page out of results, use a noindex tag and leave the page crawlable.
Do I need an XML sitemap for a small website?
A small, well-linked site gets crawled fine on links alone. A sitemap still helps: it speeds up discovery of new pages and gives you a useful report in Search Console. Most platforms create one automatically.
What is a canonical tag?
A canonical tag is a line in the HTML head that names the preferred URL for a page. It tells Google which version to index when the same content is reachable at several addresses.
Should I use a 301 or a 302 redirect?
Use a 301 when a page moves permanently, so Google indexes the new URL and transfers signals. Use a 302 for a short, temporary move where the original URL will return.
When do I need hreflang?
You need hreflang when the same page exists in several languages or for several countries. Each version lists all versions, itself included, so Google shows visitors the one that fits them.
Is JavaScript bad for SEO?
JavaScript works well for SEO when the main content and links are present in the HTML the server sends. Use server-side rendering or static generation, and test the rendered page in URL Inspection.
How often should I do a technical SEO audit?
Run a basic audit every quarter and right after a redesign, migration or platform change. Check Search Console monthly for new indexing errors so problems are caught early.
Read next
Keep reading.
More about web design and SEO
Find what fits you.
Services
- Web design
- 3D website design
- AI SEO agency
- AI chatbot for website
- Conversion rate optimization
- Custom website design
- Ecommerce website design
- Landing page design
- Local SEO services
- Online booking website
- SEO company near me
- SEO services
- Small business website design
- Web designers near me
- Website developers
- Website redesign
- Website speed optimization
By industry
By city
Your next step
Open the doors for Google once, then let your content do the work.
Fill in your details in two minutes. We build your complete website and show it in a 15-minute showcase.