Skip to content
Lovable Field Guide
Start here

sitemap.xml and robots.txt on a Lovable site: verify, fix, and ship

Lovable generates sitemap.xml and robots.txt automatically, but not always up front. Verify your Lovable sitemap, kill invented URLs, and ship a correct robots.txt.

12 min read

Lovable generates both files for you. The sitemap is the one to watch, because Lovable can add URLs that don’t exist as your site grows. Its SEO docs say sitemap.xml and robots.txt are created automatically, and then hedge, in the same breath, that they are “not always generated up front.” That hedge is the entire reason this post exists. Encited’s free robots.txt analyzer checks the file against the major crawlers. Two files you assume are handled are the two files most likely to be missing, stale, or full of URLs that were never real.

Here is the order of operations: confirm the files exist at your live domain, run Lovable’s built-in review to create or repair them, then audit the sitemap’s contents by hand, because that last part is the one nothing automates well. The surrounding technical picture (rendering, meta tags, canonicals) is covered in the Lovable SEO guide.

Which stack you are on changes half of this

Settle this before you debug anything. On 13 May 2026 Lovable switched the default for new projects from React + Vite to TanStack Start, and the two stacks behave differently on exactly the things this post is about.

  • Project created before 13 May 2026: React + Vite, client-rendered. Lovable applies on-request pre-rendering on deployed public URLs, served only to verified crawlers. Humans still get the SPA.
  • Project created on or after 13 May 2026: TanStack Start on Cloudflare Workers, server-rendered by default.
  • Older project that wants the new stack: the upgrade runs from chat, costs credits, and is reversible from version history. Lovable’s docs warn that browser-only libraries can break server rendering in ways the upgrade’s checks do not catch, so test afterwards.

Do not settle it by reading a single doc page. Lovable’s own documentation contradicts itself here (the deployment/hosting/ownership page and llms.txt still describe projects as Vite + React) so check your project’s dependencies, or curl your live URL and see what comes back.

Next: prove the files exist at the published domain

Do not check this in the editor preview. Lovable publishes by snapshot (the live site does not change until you republish) so the preview and the public URL can disagree. Ask the live origin directly:

# Status line and content-type only. Filter rather than truncate: header
# order is server-dependent, and behind a CDN content-type is rarely in the
# first few lines.
curl -sI https://yourdomain.com/sitemap.xml | grep -iE '^(HTTP/|content-type)'
curl -sI https://yourdomain.com/robots.txt  | grep -iE '^(HTTP/|content-type)'

# Then look at the first few lines of the body.
curl -s https://yourdomain.com/sitemap.xml | head -n 5
curl -s https://yourdomain.com/robots.txt

You are looking for 200 plus a content type of application/xml (or text/xml) for the sitemap and text/plain for robots.txt. A 404 means the file is not in the deployed output.

The sneakier failure is a 200 whose body starts with <!doctype html>. That is the app’s catch-all route answering for /sitemap.xml and handing back index.html. Search Console reports it as a fetch failure or “not XML”; people describe it as a 404 because the practical effect is identical. The status code lies, so read the body.

If the files are missing, the documented fix is Lovable’s own tooling rather than hand-editing.

What Lovable’s built-in SEO review actually does here

As of September 2026, Lovable ships an SEO and AI search review. On sitemaps and robots.txt specifically, the docs say it surfaces missing or out-of-sync items and can create, update, or repair them in one click. Three concrete checks are documented:

  1. It validates the sitemap XML.
  2. It checks robots.txt for a proper Sitemap: directive.
  3. It checks URL consistency across both files.

That third one matters more than it sounds. A sitemap listing https://yourdomain.com/pricing while robots.txt advertises a sitemap at the old *.lovable.app host is a consistency problem the review is built to catch, and it is a very easy one to create by moving to a custom domain halfway through a project.

Findings are graded green (passing), blue for low impact, amber for medium, and red for high. Run it, take the one-click repair, then re-run the curl checks above, because the review is inspecting your project, and curl is inspecting what the internet receives. Those are different questions, and a green review with an empty Search Console index report means you asked the wrong one.

The failure that actually costs you: URLs that were never real

This is the part the one-click repair does not solve. Lovable users report sitemaps that list routes which were never built and omit real pages that do exist; Encited’s own write-up calls Lovable sitemaps “notorious for hallucinating URLs.” Lovable’s docs do not document this behavior, so verify rather than assume. The mechanism is easy to believe, though: a sitemap produced by prompting an AI to write one is plausible-looking rather than accurate: routes that were discussed but never built, blog post slugs invented from the pattern of the ones that exist, and real pages left out entirely. A static sitemap is also a snapshot, so any site with a database behind it drifts out of sync the moment someone publishes a row.

Both directions hurt, but the invented URLs hurt more, for a mechanical reason. Every fake URL in your sitemap is a fetch you explicitly asked a crawler to make. On a React + Vite Lovable project, an unmatched route returns the app shell with HTTP 200, so the crawler does not even get a clean 404: it gets a soft 404, indexes nothing, and learns that URLs from your sitemap are not worth much. On TanStack Start the router can answer an unmatched route with a real 404 status, but only if the route tree is set up to; confirm it with curl -sI https://yourdomain.com/this-does-not-exist and read the status line rather than assuming. Either way, repeat the mistake a few hundred times and you have spent your crawl allowance teaching Google to ignore your best signal.

AI crawlers punish it harder, because they have no equivalent of Google’s crawl scheduling. Vercel’s December 2024 log analysis found GPTBot hitting 404s on 34.82% of fetches, plus 14.36% redirects, and ClaudeBot on 34.16%, against 8.22% for Googlebot. Those bots do not retry and do not wait. Handing them a list padded with dead URLs is throwing away a third of an already-wasteful budget. (That measurement is nearly two years old now and no first-party re-measurement exists, so treat the exact percentages as directional.)

Generate the sitemap from something that knows the truth

The fix is to stop treating the sitemap as content and start treating it as build output. Two shapes work:

  • Static routes: generate from the route manifest at build time, so a route that does not exist cannot be listed.
  • Dynamically loaded content: generate at request time from the rows that are actually published: a server function or edge function that queries the content table and returns XML.

Whatever generates it, audit the output once by hand. This takes about a minute and finds nearly every version of the problem. First, pull every URL into a file. The --compressed flag covers a gzipped sitemap.xml.gz:

# Case A: a single sitemap file.
curl -sS --compressed https://yourdomain.com/sitemap.xml \
  | grep -oE '<loc>[^<]+' | sed 's/<loc>//' > urls.txt

If /sitemap.xml is a sitemap index, its <loc> values are child sitemaps rather than pages, so fetch one more level:

# Case B: a sitemap index. Expand the children first.
curl -sS --compressed https://yourdomain.com/sitemap.xml \
  | grep -oE '<loc>[^<]+' | sed 's/<loc>//' \
  | while read -r child; do
      curl -sS --compressed "$child" | grep -oE '<loc>[^<]+' | sed 's/<loc>//'
    done > urls.txt

Then check what each URL actually returns, showing only the failures:

while read -r u; do
  printf '%s %s\n' "$(curl -s -o /dev/null -w '%{http_code}' "$u")" "$u"
done < urls.txt | grep -v '^200 '

A 200 is not proof the page exists. On a client-rendered Lovable project the catch-all answers every invented URL with 200 and the app shell, which is exactly the set of URLs this audit exists to find, so the status loop alone will tell you everything is fine while the sitemap is still full of fiction. Capture the shell once and compare against it:

# What the site serves for a URL that definitely does not exist.
curl -sS https://yourdomain.com/this-does-not-exist > shell.html

while read -r u; do
  curl -sS "$u" | cmp -s - shell.html && echo "SOFT 404: $u"
done < urls.txt

If your shell is not byte-identical between requests (a build hash, a nonce, an injected timestamp) grep for a per-page marker instead: the route’s own <title> or <h1>. A response missing it is the shell wearing a 200.

Anything that returns a 301, 404 or 500, and anything that comes back as the shell, comes out of the sitemap. Then invert the test: list the pages you know exist and check each one appears in the file. Missing pages are the quieter half of this bug.

The hard limits Google actually enforces

Google’s sitemap documentation is short and worth obeying literally.

Rule Value Consequence of breaking it
Max URLs per sitemap 50,000 Split into multiple files and reference them from a sitemap index
Max file size 50 MB uncompressed Same fix: split and index
Encoding UTF-8 Parse failure
URL form Fully qualified and absolute Relative URLs are invalid
<priority> Ignored by Google Effort wasted
<changefreq> Ignored by Google Effort wasted
<lastmod> Used, but only when consistently and verifiably accurate Bumping it on every deploy destroys its credibility

That last row is the one CI pipelines get wrong. <lastmod> should reflect a meaningful change to the page’s content, not the timestamp of the build that happened to redeploy it. If every URL’s lastmod moves every time you push a CSS tweak, you have converted the one field Google reads into noise.

Submit the finished file through the Search Console Sitemaps report, the Search Console API, or a Sitemap: line in robots.txt. All three are documented; the robots.txt line is free and should be there regardless.

robots.txt: four rules people get wrong

Disallow is not noindex (debugging guide). Google’s documentation states it cannot index the content of a disallowed page but may still index the URL and show it in results without a snippet. Disallow governs crawling. noindex governs indexing. They also actively conflict: a URL you disallowed can never be fetched, so the noindex tag on it is never read. If you want a page out of the index, allow the crawl and serve noindex.

Never block JS, CSS, or your assets folder. Google needs those files to render the page, which is why Googlebot is allowed to crawl them by default. On a client-rendered app this is not a minor degradation: Disallow: /assets/ is the single most effective way to make Google genuinely unable to see your app. A blanket Disallow: / has a second casualty most people forget: the social unfurlers that build your link previews are crawlers too, so chat and social previews go blank alongside your rankings. If link previews are the symptom you arrived with, how crawlers fetch a Lovable page is the better starting point.

There is a size cap and a cache window. Google ignores anything past 500 kibibytes, and it generally caches robots.txt for up to 24 hours. So a bad rule you just deleted can keep biting for another day, and the file has to live at the top-level directory of the host.

Group matching is winner-takes-all. Only one group applies to a given crawler: the most specific match. The moment you give GPTBot its own group, GPTBot stops reading the User-agent: * group entirely, including every Disallow in it. This is how private paths quietly leak to exactly the bots you wrote a special group for.

A correct robots.txt for a Lovable site that welcomes AI crawlers

Because anything not disallowed is allowed, welcoming AI crawlers mostly means not blocking them. The minimal correct file is this:

# https://yourdomain.com/robots.txt

User-agent: *
Allow: /
Disallow: /admin/
Disallow: /account/

# Do not add /assets/ or your JS/CSS paths here: Google needs them to render.

Sitemap: https://yourdomain.com/sitemap.xml

Absolute URL on the Sitemap: line, and it must point at the host you actually want indexed. If you have moved to a custom domain, a Sitemap: line still aimed at *.lovable.app is a live contradiction between your two most-read files.

If you want the permissions written down explicitly (for a colleague, an auditor, or your own future self) one group with several User-agent lines is valid and keeps the rules in one place:

# Explicit opt-in. Note these bots now IGNORE the User-agent: * group,
# so every Disallow you care about must be repeated here.
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-User
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Applebot
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
User-agent: CCBot
User-agent: Bytespider
User-agent: Google-Extended
Allow: /
Disallow: /admin/
Disallow: /account/

Two honest caveats on that second file. Letting a bot in doesn’t mean it can read you. The Vercel/MERJ crawler study found the major AI crawlers don’t execute JavaScript (Applebot is the documented exception), so they get whatever is in your HTML response.

Which brings the stack question back. Lovable’s docs say older React + Vite projects get on-request pre-rendering for verified crawlers, and that TanStack Start projects (the default since 13 May 2026) server-render. Both statements are about the page shell. Here is what you will actually observe: Lovable’s generated code typically fetches dynamic data from Supabase or Lovable Cloud inside a React hook that runs after the component mounts, in the visitor’s browser. That query is not in the pre-rendered or cached HTML, so the products, listings, posts and profiles it returns are not in what a crawler receives. Test it against a page with dynamically loaded content (CMS posts, product pages, directory listings: anything that isn’t written into the page’s source code), never the homepage:

# A page with dynamically loaded content. The marketing homepage won't tell you anything.
curl -sS https://yourdomain.com/products | grep -c 'product-card'
# Compare with what you count in the browser. A gap is the whole problem.

Correct URLs pointing at pages that come back without their content is a well-organized way to get nothing indexed, and it is the whole subject of getting a Lovable site into AI answers. The second caveat is simpler: robots.txt is advisory. Well-behaved crawlers honor it; nothing enforces it.

What to do next

  1. Establish which stack the project is on, then run the curl checks against your live published domain: status, content type, first lines of body. Do it before changing anything, so you know what you are fixing.
  2. Run Lovable’s SEO and AI search review and accept the one-click create or repair for anything missing or out of sync.
  3. Pipe every <loc> through the status-code loop above, then through the shell comparison. Remove every non-200 and every soft 404. Add every real page that is absent.
  4. Replace any hand-written static sitemap on a site with dynamically loaded content with a generated one, then fetch it unauthenticated to confirm RLS did not empty it.
  5. Read your robots.txt line by line against the four rules: no asset blocking, no Disallow standing in for noindex, an absolute Sitemap: URL on the right host, and no per-bot group that silently drops your wildcard disallows.
  6. Submit the sitemap in Search Console and check back in a few days. If pages are still sitting at “Crawled – currently not indexed,” the sitemap was never the problem: start with why a Lovable site does not get indexed instead.

Found something wrong or out of date? Lovable changes fast and we'd rather fix a guide than let it rot.

Disclosure: the team behind this guide also builds Encited, mentioned above.

Frequently asked questions

Does Lovable create a sitemap.xml automatically?
Yes. Lovable's SEO documentation says sitemap.xml and robots.txt are generated automatically, but it also hedges that they are 'not always generated up front.' Treat them as probably-there rather than definitely-there, and confirm with a request to your live published domain before assuming.
Why does my Lovable sitemap.xml return a 404?
Two causes dominate. Either the file was never written into the deployed build output, or the single-page-app catch-all is answering /sitemap.xml with index.html. The second case returns HTTP 200 with an HTML body, which Google Search Console rejects as an invalid sitemap and which users usually describe as a 404. Check the content-type header as well as the status code.
Should I set priority and changefreq in my sitemap?
No. Google's sitemap documentation states plainly that it ignores both the priority and changefreq elements. Only lastmod is used, and only when it is consistently and verifiably accurate. A build pipeline that stamps today's date on every URL at every deploy trains Google to distrust the field entirely.
Does Disallow in robots.txt keep a page out of Google?
No. Google's robots.txt documentation says it cannot index the content of a disallowed page but may still index the URL and show it without a snippet. Disallow controls crawling; noindex controls indexing. The two also conflict: a page you have disallowed can never be crawled, so its noindex tag is never read.
Do I need to list AI crawlers in robots.txt to let them in?
No. Anything not disallowed is allowed by default, so a permissive file already welcomes GPTBot, ClaudeBot, PerplexityBot and the rest. Explicit per-bot groups are useful only as documentation or when you want different rules per bot, and they carry a trap: a crawler that matches its own group ignores the wildcard group completely.
Is Google-Extended an AI crawler I should allow?
Google-Extended is not a crawler. Google's crawler documentation states it has no separate HTTP request user agent string and that the robots.txt token is used only in a control capacity, governing Gemini training and grounding use of content Google already crawled. Allowing or blocking it only changes permissions.

Read next

  • Making a Lovable app crawlable

    Lovable crawlability comes down to links as much as rendering: real anchors, history routing, honest status codes, and a crawl you can run yourself.

  • Lovable site not showing up on Google

    A Lovable app not indexed on Google is usually not a rendering bug, but check anyway. Seven causes ranked by likelihood, with the curl test that proves which one.

  • Getting a Lovable app cited by ChatGPT and Perplexity

    Lovable AI search visibility depends on what is in your raw HTML. Which bots fetch what, why SSR alone does not settle it, and how to measure citations.

  • Lovable SEO: the complete 2026 guide

    Lovable SEO splits across two stacks, and SSR alone does not put your data in the HTML. Test both on your own routes, then close the gaps Lovable leaves to you.