Making a Lovable app crawlable
Lovable crawlability comes down to links as much as rendering: real anchors, history routing, honest status codes, and a crawl you can run yourself.
13 min read
Crawlability is a different problem from rendering, and fixing the second does not fix the first. A Lovable app can return perfect server-rendered HTML on every URL and still leave most of its pages undiscoverable, because a crawler finds pages by following href attributes. It never clicks anything. This post covers everything between “the HTML renders” and “Google can walk the site”: anchors, routing mode, status codes, redirects, pagination, orphans, and the robots.txt rules that quietly break all of it.
The short version: real <a href> for anything that is a page, the History API instead of hash routes, a genuine 404 status for URLs that do not exist, server redirects instead of window.location, paginated URLs behind your infinite scroll, an inbound link for every page in your sitemap, and never a Disallow on your JavaScript.
First, know which stack you are on
Lovable changed its default on 13 May 2026. Projects created after that date are TanStack Start on Cloudflare Workers, where SSR is on by default. Older projects are React + Vite, client-rendered, and Lovable applies on-request pre-rendering on deployed public URLs that is served only to verified crawlers: Google, Bing, social bots and AI engines. Humans, and third-party scanners like Screaming Frog, still get the SPA shell.
SSR being on is not the same fact as your content being in the HTML. SSR renders the component tree; it does not fetch your data for you. Lovable’s own description of the new stack is carefully worded: the server “runs the React tree, executes the loaders, and streams the result.” Loaders are the hinge. Data pulled in a route loader or a createServerFn is serialized into the HTML; data pulled in a component hook is not, because React’s documentation says effects “only run on the client. They don’t run during server rendering,” and TanStack Query’s SSR guide says un-prefetched queries “wont be server rendered, instead they will be fetched on the client after the application is interactive.” A page can be one hundred percent server-rendered and still ship a skeleton. SSR versus SPA for SEO works through the consequences.
Check rather than guess:
grep -E '"(@tanstack/react-start|@tanstack/react-router|react-router-dom)"' package.json
@tanstack/react-start means SSR. Only react-router-dom means you are on the pre-rendering path. The Lovable SEO guide covers the consequences of each, and Lovable’s pre-rendering behaviour gets its own treatment.
Then check whether your data actually lands in the HTML. Run it against a page with dynamically loaded content (anything that isn’t written into the page’s source code, like a listing or detail page), never the homepage, and grep for a phrase that isn’t in your source code, such as a product name:
curl -s -A 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)' \
https://yourapp.com/products > /tmp/page.html
grep -qiF 'Ceramic Pour-Over Kettle' /tmp/page.html \
&& echo 'server-rendered' || echo 'client-fetched: crawlers never see it'
The search phrase has to be one contiguous run of text; markup splits phrases across elements and a multi-word grep then misses content that is genuinely there. On the React + Vite stack this returns the shell whatever user agent you send, and what verified crawlers receive can’t be verified from outside.
Whichever stack you are on, and whichever way the grep comes out, everything below still applies. Every problem in this post is stack-independent: SSR decides what is inside the page you asked for, but it cannot invent a link to a page you never linked.
Buttons with onClick are not links, and the distinction is absolute
This is the most common crawlability defect in Lovable-generated code, and it comes from a reasonable place: the model reaches for a Button component because the design system has one, then wires navigation through the router’s imperative API:
// Generated code that works for humans and is invisible to crawlers.
import { useNavigate } from 'react-router-dom'
export function NavBar() {
const navigate = useNavigate()
return (
<Button onClick={() => navigate('/pricing')}>Pricing</Button>
)
}
A crawler collects URLs from href attributes. That is the entire discovery mechanism. Googlebot does render JavaScript, on an evergreen Chromium, but it never interacts with the page. It does not click, hover, scroll or submit. After rendering it reads the DOM and harvests href attributes, exactly as it would from static HTML, so an onClick handler is never invoked and the URL behind it is never learned.
For the non-rendering crawlers it is worse. The Vercel/MERJ AI crawler study (measured December 2024; no first-party re-measurement has been published since) reported that none of the major AI crawlers execute JavaScript: OpenAI’s OAI-SearchBot, ChatGPT-User and GPTBot, Anthropic’s ClaudeBot, Meta-ExternalAgent, Bytespider and PerplexityBot. Those agents never reach the DOM your button lives in.
The fix is to render an anchor and let the router intercept it.
// react-router-dom: Link renders a real <a href> and calls preventDefault
// on left-click, so humans get client-side nav and crawlers get a URL.
import { Link } from 'react-router-dom'
export function NavBar() {
return <Link to="/pricing" className={buttonStyles}>Pricing</Link>
}
On the TanStack Start stack, Link from @tanstack/react-router behaves the same way: it emits an <a href> with the resolved path. If your design system’s Button is doing the styling work, most shadcn/ui buttons accept an asChild prop, which lets the button render the child element instead of a <button>, so <Button asChild><Link to="/pricing">Pricing</Link></Button> keeps the styling and produces an anchor.
Find every offender in the repo:
# Imperative navigation that produces no href
rg -n "navigate\(['\"/]|useNavigate\(|router\.navigate|window\.location\s*=" src/
# Clickable elements that are not anchors
rg -n "<(div|span|Button|Card)[^>]*onClick" src/
Not every hit is a bug. A button is correct for an action: opening a dialog, submitting a form, toggling a filter, advancing a wizard step that has no URL of its own. A button is wrong whenever the destination is a page you would be happy to see in search results. The test is simple: if you would put the destination in your sitemap, it needs an href.
Hash routing cannot be fixed, only migrated away from
If your URLs look like https://example.com/#/pricing, nothing downstream of the browser can help you.
Two separate things are going on. The first is Google’s position: the AJAX-crawling scheme has been deprecated since 2015, and Google’s guidance for single-page apps is to “use the History API to load different content based on the URL in a SPA.” Fragment URLs will not work reliably with Googlebot.
The second is the one most write-ups miss, and it is decisive. Per RFC 3986 the fragment (everything from the # onward) is a client-side construct, stripped by the browser before it builds the HTTP request line. It is never transmitted. Your origin never sees it, your CDN never sees it, a Cloudflare Worker never sees it, and a prerender middleware never sees it. Every /#/... URL on your site is, to every piece of infrastructure you own, the same single request for /.
Hash routing therefore makes edge prerendering, per-route Open Graph tags, per-route canonicals and per-route status codes physically impossible. No vendor and no configuration gets around it, because the information is not in the request.
Migrating means swapping HashRouter for BrowserRouter (or using TanStack Router’s default history), then making sure your host serves the app shell for a cold request to /pricing instead of 404ing. On Lovable hosting that rewrite is already in place; if you exported and deployed elsewhere it is a _redirects line on Netlify, a rewrite rule on Vercel, or Cloudflare Pages’ not-found handling. Redirect any already-indexed hash URLs server-side, per the next-but-one section.
Soft 404s: the status line is the only thing that counts
Here is the mechanism, precisely. Your host has a catch-all rewrite so that direct requests to client-side routes work. A crawler requests /this-page-does-not-exist. The rewrite serves index.html with HTTP 200. The router then mounts a “Not found” component. The crawler recorded a 200, so it indexes the URL and later flags it as a soft 404. At scale, across hallucinated sitemap entries and stale external links, this burns crawl budget on an unbounded set of URLs that do not exist.
Your not-found component being beautiful is irrelevant. Google judges the status code.
Google documents exactly two remedies on its “Fix search-related JavaScript problems” page. The first is to redirect to a URL that really does return 404:
// Google's documented fix #1 for SPA soft 404s.
fetch(`https://api.kitten.club/cats/${id}`)
.then(res => res.json())
.then((cat) => {
if (!cat.exists) {
// redirect to page that gives a 404
window.location.href = '/not-found';
}
});
The second is to add a noindex robots meta tag when the record is missing:
// Google's documented fix #2 for SPA soft 404s.
fetch(`https://api.kitten.club/cats/${id}`)
.then(res => res.json())
.then((cat) => {
if (!cat.exists) {
const metaRobots = document.createElement('meta');
metaRobots.name = 'robots';
metaRobots.content = 'noindex';
document.head.appendChild(metaRobots);
}
});
Both remedies are client-side, which means both only work for crawlers that render. That is a real limitation, and it is why the better answer, when you have a server, is to emit the status from the server.
On the TanStack Start stack you can set the response status inside a route loader when the record is not found, and the crawler gets a genuine 404 with no JavaScript involved. That is the version that also works for the crawlers in the Vercel/MERJ measurement above, which reported no JavaScript execution from GPTBot or ClaudeBot. Verify either way:
curl -sI -o /dev/null -w '%{http_code}\n' https://yourapp.com/this-does-not-exist
# 200 => soft 404. 404 => correct.
Client-side redirects are a dead end for everything except Googlebot
A window.location.href = '/new-url' or a <Navigate to="/new" /> in a route component is a redirect only for agents that run JavaScript. Googlebot follows it, because it renders. Applebot renders too, and Bing says it is “generally able to process JavaScript” while conceding it is “difficult for bingbot to process JavaScript at scale”, so treat Bing as unreliable here rather than incapable. The link-preview bots and the AI crawlers are the clear-cut case: in the Vercel/MERJ measurement cited above, facebookexternalhit, Slackbot, GPTBot, ClaudeBot and PerplexityBot all receive the pre-redirect shell and stop there.
That matters more than it sounds, because non-rendering crawlers are already wasting most of their budget. In the Vercel/MERJ measurement, 34.82% of GPTBot fetches hit 404s and another 14.36% hit redirects; ClaudeBot recorded 34.16% 404s. Googlebot, by comparison, saw 8.22% 404s and 1.49% redirects. Adding client-side redirects to that mix compounds a problem those crawlers already have.
Use a server 301 or 308 at the edge for anything that must be followed by a non-JS agent: renamed routes, the apex-versus-www decision, trailing-slash normalization, retired URLs. Reserve client-side navigation for post-authentication app flow, where no crawler belongs anyway. If your canonical tags and your redirects disagree about which URL form is real, you have a second problem: canonical tags on Lovable covers that pairing.
Infinite scroll needs real paginated URLs, and rel=next/prev is not coming back
Crawlers do not scroll. Content revealed by an IntersectionObserver or a “Load more” button is unreachable, no matter how well it renders, unless there is a parallel path made of URLs.
Google stopped using rel="next" and rel="prev" in March 2019, and its current pagination guidance no longer mentions them. Adding them today does nothing. What works is unglamorous:
- Give every page state its own URL:
?page=2, or/blog/page/2. - Make that URL reachable from a real
<a href>in the HTML. - Return the correct slice at that URL on a cold, direct request, with no prior state.
- Give each paginated page a self-referencing canonical. Pointing pages 2..n at page 1 de-indexes them and takes the items that only appear there with them.
The pattern that keeps both audiences happy is a “Load more” control that is an anchor with a working href, intercepted for humans:
// The href is the truth; preventDefault is the enhancement.
<a
href={`/blog?page=${page + 1}`}
onClick={(e) => { e.preventDefault(); loadMore() }}
>
Load more
</a>
A crawler sees a link to /blog?page=2 and follows it. A human clicks and never leaves the page. Middle-click and “open in new tab” also start working, which is a nice side effect.
Orphan pages: in the sitemap, linked from nowhere
An orphan is a URL that exists and is valid but has no internal link pointing at it. Lovable auto-generates sitemap.xml and robots.txt, though its own docs hedge that they are “not always generated up front,” and the generated sitemap will happily list routes that nothing on the site links to.
A sitemap is a discovery hint. It says a URL exists. Internal links say a URL matters, and they are how authority moves around your site. Pages that appear only in a sitemap are the classic residents of “Crawled: currently not indexed,” and if that is where you are, the diagnosis in why a Lovable site isn’t indexed starts one level up from here.
Three things create orphans in Lovable projects: routes added by a chat prompt that never got a nav entry, detail pages reachable only through a search or filter UI, and (the one that surprises people) pages whose only inbound links are the onClick buttons from the first section. Fix the anchors and a chunk of the orphan list resolves itself. Where a nav link would be clutter, an HTML index page is the cheap answer: a /blog archive, a /directory listing, a footer block of category links. Anything that gives every indexable URL at least one href from a page that is itself crawlable.
The robots.txt rules that break crawling silently
Four facts from Google’s robots.txt documentation, each of which costs people real traffic:
| Rule | What it actually means |
|---|---|
Disallow is not noindex |
Google “can’t index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet.” |
A disallowed page’s noindex is unreadable |
If you blocked the URL, Google never fetches it, so it never sees the tag telling it to stay out. |
| Blocking JS/CSS breaks rendering | Google needs those files for rendering, “so Googlebot is allowed to crawl them.” A Disallow: /assets/ is the one reliable way to make Google genuinely unable to see your React app. |
| The file is capped and cached | Maximum 500 kibibytes; anything beyond is ignored. Google “generally caches the contents of robots.txt file for up to 24 hours,” so a fix is not instant. |
A working baseline:
User-agent: *
Allow: /
Disallow: /app/
Disallow: /api/
# Do NOT do this: it prevents Googlebot rendering your app:
# Disallow: /assets/
Sitemap: https://yourapp.com/sitemap.xml
Applebot is worth a specific mention here, because it is the AI-adjacent crawler that does render: Apple’s documentation says it “may render the content of your website within a browser,” and that if JavaScript is robots-blocked it “may not be able to render the content properly.” Blocking your bundle costs you Google and Apple at once. The full treatment lives in sitemap and robots.txt for Lovable.
Crawl it yourself: the check that settles every argument
Every section above reduces to one question: can a program that never runs JavaScript walk from your homepage to every page you care about? Answer it directly. This script fetches raw HTML, extracts anchors with a regex, and follows same-origin links breadth-first. It is deliberately dumb, because that is the point: it approximates what a non-rendering crawler can reach.
// scripts/crawl-raw.mjs: Node 20+, no dependencies.
// node scripts/crawl-raw.mjs https://yourapp.com
const start = new URL(process.argv[2]);
const UA = 'Mozilla/5.0 (compatible; RawCrawl/1.0)';
const seen = new Map(); // url -> status
const linkedFrom = new Map(); // url -> Set of parents
const queue = [start.href];
const norm = (href, base) => {
try {
const u = new URL(href, base);
if (u.origin !== start.origin) return null;
if (!/^https?:$/.test(u.protocol)) return null;
u.hash = ''; // fragments never reach a server
u.pathname = u.pathname.replace(/\/+$/, '') || '/';
return u.href;
} catch { return null; }
};
while (queue.length && seen.size < 500) {
const url = queue.shift();
if (seen.has(url)) continue;
const res = await fetch(url, { headers: { 'user-agent': UA }, redirect: 'manual' });
seen.set(url, res.status);
if (res.status >= 300 && res.status < 400) {
console.log(`${res.status} ${url} -> ${res.headers.get('location')}`);
continue;
}
if (!res.ok || !(res.headers.get('content-type') || '').includes('text/html')) {
console.log(`${res.status} ${url}`);
continue;
}
const html = await res.text();
const hrefs = [...html.matchAll(/<a\b[^>]*\shref=["']([^"']+)["']/gi)].map(m => m[1]);
console.log(`${res.status} ${url} (${hrefs.length} anchors in raw HTML)`);
for (const h of hrefs) {
const next = norm(h, url);
if (!next) continue;
if (!linkedFrom.has(next)) linkedFrom.set(next, new Set());
linkedFrom.get(next).add(url);
if (!seen.has(next)) queue.push(next);
}
}
// Orphan check: sitemap URLs the link graph never reached.
const sm = await fetch(new URL('/sitemap.xml', start), { headers: { 'user-agent': UA } });
if (sm.ok) {
const locs = [...(await sm.text()).matchAll(/<loc>\s*([^<\s]+)\s*<\/loc>/g)]
.map(m => norm(m[1], start))
.filter(Boolean);
const orphans = locs.filter(l => !linkedFrom.has(l));
console.log(`\nIn sitemap, never linked from a crawled page (${orphans.length}):`);
orphans.forEach(o => console.log(' ' + o));
}
Read the output for four things.
Anchor count on the homepage. Compare it against what the browser has after hydration: open devtools and run document.querySelectorAll('a[href]').length. A large gap is your button-navigation problem, quantified. If the raw count is near zero on an older React + Vite Lovable project, a Googlebot user agent won’t change the result, and what verified crawlers receive can’t be verified from outside. Use Search Console’s URL Inspection to see what Googlebot actually received on that stack.
Pages reached versus pages expected. Anything you know exists but that never appears in the crawl is unreachable by link. That is your orphan list, before you even get to the sitemap comparison.
Every 3xx line. Each is a redirect a non-rendering crawler has to follow, or a chain to collapse. Redirects that exist only client-side will not show up here at all, which is itself the finding.
The sitemap orphan list. Your homepage will appear there (nothing linked to it), and so will pages the crawl stopped short of at the 500-URL cap. Everything else on that list needs an inbound link or removal from the sitemap.
Then re-run the soft-404 status test from earlier, plus a crawler-UA fetch to confirm what a bot actually receives.
curl -s -A 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)' \
https://yourapp.com/pricing | grep -oE '<a [^>]*href="[^"]*"' | head -40
On the TanStack Start stack that is a fair sample of what any crawler gets. On React + Vite it is not (the UA string alone does not make you a verified crawler) so compare it against Search Console’s URL Inspection before you draw a conclusion from it.
The script is a snapshot from your laptop, which is enough to catch the obvious problems. Seeing how bots actually walked the site over time (which URLs Googlebot requested, which returned 404s, which it never reached) takes request logs. Cloudflare and Vercel both expose raw logs you can filter by user agent. Encited packages the same data as per-bot crawl logs, alongside a site audit that flags orphan pages, soft 404s and redirect chains. Its free SEO audit runs the same checks on a single URL.
What to do next
- Run the script against production today and save the output. It is your before-picture.
- Grep for
useNavigate(andonClickon non-anchor elements. Convert everything whose destination belongs in your sitemap. - Confirm a nonexistent URL returns 404, not 200. Fix it at the server if you have one, with Google’s documented client-side remedy if you do not.
- If you see
/#/anywhere in your URLs, stop and migrate routing first: no other fix on this list will hold. - Diff the sitemap against the crawl. Every orphan gets an inbound link or gets dropped.
- Re-run the script after the changes ship, and remember that on Lovable, publishing is a snapshot: edits do not reach the live site until you republish.
Frequently asked questions
- Does Googlebot click buttons to find pages?
- No. Googlebot renders JavaScript, but it does not click, hover, scroll or type. It discovers URLs from href attributes in the HTML it has after rendering. A div or button with an onClick handler that calls a router navigate() is a dead end for link discovery, even though it works perfectly for a human.
- Why can't a prerendering service fix hash routing?
- Everything after the # in a URL is a fragment, and per RFC 3986 the browser strips it before sending the request. No server, CDN, edge worker or prerender middleware ever learns which hash route was asked for. Moving from HashRouter to BrowserRouter comes first; nothing else works without it.
- My 404 page looks right but Search Console calls it a soft 404. Why?
- Because the HTTP status line said 200. A single-page app's catch-all rewrite serves index.html with a 200 for every unmatched URL, and Google judges by the status code, whatever the component renders. Google documents two fixes: redirect to a URL that genuinely returns 404, or inject a noindex robots meta tag when the record is missing.
- Is a sitemap enough to get a page crawled if nothing links to it?
- Sometimes, but it is a weak signal on its own. A sitemap tells Google a URL exists; internal links tell Google the URL matters and pass authority to it. Pages that appear only in the sitemap are orphans, and they are the ones that most often sit in Crawled - currently not indexed.
- Should I ever block /assets/ in robots.txt?
- No. Google's robots.txt documentation says it needs JavaScript and CSS files for rendering, so Googlebot is allowed to crawl them. Disallowing your bundle directory is the one reliable way to make Google genuinely unable to see a React app, and it also breaks Applebot, which renders too.
Read next
Keep going
-
Lovable SEO: the complete 2026 guide
Lovable SEO splits across two stacks, and SSR alone does not put your data in the HTML. Test both on your own routes, then close the gaps Lovable leaves to you.
-
Lovable site not showing up on Google
A Lovable app not indexed on Google is usually not a rendering bug, but check anyway. Seven causes ranked by likelihood, with the curl test that proves which one.
-
sitemap.xml and robots.txt on a Lovable site
Lovable generates sitemap.xml and robots.txt automatically, but not always up front. Verify your Lovable sitemap, kill invented URLs, and ship a correct robots.txt.
-
Canonical tags in Lovable: the rules and the traps
Google's canonical rules, then the SPA traps that ruin a Lovable canonical tag: one baked URL on every route, trailing-slash forks, and lovable.app vs your domain.