Getting a Lovable app cited by ChatGPT and Perplexity
Lovable AI search visibility depends on what is in your raw HTML. Which bots fetch what, why SSR alone does not settle it, and how to measure citations.
17 min read
If you want a Lovable app cited in ChatGPT, Claude or Perplexity, the first thing to settle is whether those engines can read the page at all. Every public log study of these agents points the same way: almost none of them execute JavaScript. Google does, which is why a client-rendered app can rank perfectly well in Google Search and be completely absent from AI answers, and why nothing in your Google-shaped analytics stack will tell you that is happening.
The short version: confirm a page with dynamically loaded content (CMS posts, product pages, directory listings: anything that isn’t written into the page’s source code) returns your actual content to a non-rendering fetch, allow the right bots in robots.txt, skip llms.txt, put structured data in the server response but do not expect citations from it, write pages whose answers survive being chopped into Markdown, and measure with server logs rather than a dashboard. For everything else that affects how a Lovable site ranks, see the complete Lovable SEO guide.
Two kinds of AI crawler, and only one of them cites you
Lumping every AI bot together is what makes this topic confusing. They split into two groups with different jobs, and the fix for each is different.
Corpus builders crawl broadly to assemble training data: GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider, CCBot. Cloudflare’s Radar data for July 2025 found training accounted for roughly 80% of AI bot crawling, with user-action and undeclared crawling combined under 5%. Whether you allow them is a policy decision about your content. Blocking them doesn’t remove you from answers, and allowing them doesn’t get you cited.
Retrieval agents fetch at answer time, and these are the ones that produce citations: OAI-SearchBot and ChatGPT-User for OpenAI, Claude-SearchBot and Claude-User for Anthropic, PerplexityBot and Perplexity-User for Perplexity. The distinction inside each pair matters too: the -User agents fire when a specific person’s question causes a specific page to be opened. A July 2026 investigation of ChatGPT’s retrieval stack (RESONEO, reported by Search Engine Land, covering 1,200 answers and around 26,900 pages) found pages that were actually opened got cited 74% of the time, against 7% for pages that were retrieved but not opened. Being fetched is most of the battle.
The practical consequence: if you only ever see GPTBot in your logs and never ChatGPT-User or OAI-SearchBot, you are in the training corpus and absent from the retrieval path. Those are different problems with different fixes.
What the log studies say: they read raw HTML and nothing else
Vercel and MERJ analyzed real request logs across Vercel’s network and published the result in December 2024. Their wording is unambiguous: “none of the major AI crawlers currently render JavaScript”, naming OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider and PerplexityBot. Encited’s free AI crawl checker shows what those bots can read on a given URL. They do fetch .js files (11.50% of ChatGPT’s requests and 23.84% of Claude’s) but fetching a script is not executing one, and log-based claims that “the bot downloaded my bundle so it must have rendered” are misreading that number.
The July 2026 ChatGPT retrieval study reached the same conclusion independently, about 19 months later, stating the system does not execute JavaScript and that client-side content is invisible to it. Two different methods, same answer.
| Agent | Runs JavaScript? | What it is for | Basis |
|---|---|---|---|
| Googlebot | Yes, evergreen Chromium | Search index, and the grounding for AI Overviews and AI Mode | Google docs |
| Bingbot | Sometimes, unreliably at scale | Bing index, and Copilot downstream of it | Bing’s own guidance |
| Applebot | Yes: the documented exception | Apple search and Apple Intelligence features | Apple docs |
| GPTBot | No | Training corpus | Vercel/MERJ logs |
| OAI-SearchBot | No | Surfacing sites in ChatGPT search | Vercel/MERJ logs |
| ChatGPT-User | No | User-initiated page opens | Vercel/MERJ logs |
| ClaudeBot | No | Training corpus | Vercel/MERJ logs |
| PerplexityBot | No | Surfacing and linking sites | Vercel/MERJ logs |
| Meta-ExternalAgent | No | Meta AI | Vercel/MERJ logs |
| Bytespider | No | ByteDance | Vercel/MERJ logs |
| Google-Extended | Never makes a request | robots.txt control token only | Google docs |
Two entries on that table deserve more than a row.
Applebot renders. Apple’s documentation says Applebot may render the content of your website within a browser, and that blocking JavaScript in robots.txt may stop it rendering properly. It is the one agent in this class where a client-rendered page has a chance.
Google-Extended is not a crawler. Google’s documentation states it has no separate HTTP request user agent string and that the robots.txt token is used in a control capacity: crawling is done by existing Google user agents. It governs Gemini training and grounding use. Asking whether Google-Extended runs JavaScript is a malformed question, and any bot list that reports latency or crawl counts for it is guessing.
Which Lovable stack you are on decides how this bites you
Lovable changed its default stack on 13 May 2026, and the shape of the problem differs on each side of that line. Neither side gets a free pass.
Projects created after 13 May 2026 run TanStack Start on Cloudflare Workers, where SSR is on by default. Read Lovable’s own description of what the server does closely, though: it runs the React tree, executes the loaders, and streams the result. Loaders are the hinge. Server-side rendering renders your components, and it only includes data that a route loader fetched. React’s useEffect documentation is explicit that effects “only run on the client. They don’t run during server rendering,” and TanStack Query’s SSR guide says queries that are not prefetched “wont be server rendered, instead they will be fetched on the client after the application is interactive.” A route that loads its content in a route loader (or a createServerFn the loader awaits) serializes that content into the HTML. A createServerFn called from a component after hydration is still a client fetch; what puts the data in the bytes is awaiting it during server rendering. A route that loads it in a component hook ships the loading skeleton: server-rendered perfectly, and empty of your data.
That gap costs Google a trip through the rendering queue. For the agents in the table above it costs you everything: if the log studies are right that they never execute JavaScript, content fetched in a component hook does not exist for them in any form. A page can be one hundred percent server-rendered and completely uncitable. “SSR is enabled” and “my content is in the HTML” are independent facts, and only the second one decides whether ChatGPT can quote you. The curl below settles it in about thirty seconds, per route.
And fetching in a hook is what Lovable writes unless you ask for something else. When we studied 82 Lovable apps on GitHub, only 4.4% of page routes loaded their data on the server. Nearly half fetched it in the browser after the page loaded, and most of the apps queried Supabase straight from the browser.
Put that next to the log studies and the chain is three links long: the retrieval agents do not execute JavaScript, a typical Lovable route fetches its data after hydration, so that data never reaches them. The markup is genuinely server-rendered (anyone can confirm that with View Source) and the dynamically loaded content behind it generally is not. One app in the study makes the cost concrete: its public blog post page has no loader: and no head:, and queries a blogs table in a useEffect with loading set to true. What a non-rendering agent receives for every post on that site is a spinner, with no title, description or og tags on it.
Older React + Vite projects are client-rendered, and Lovable applies on-request pre-rendering on deployed public URLs. Per Lovable’s docs, that HTML is served only to verified crawlers: Google, Bing, social preview bots, and AI engines including ChatGPT, Perplexity, Claude and Gemini. The AI crawlers are explicitly on that list, and the same docs say external pre-rendering services are unnecessary for projects Lovable hosts. Both are Lovable’s documented positions, but they are not the same kind of statement. The first is a behavior you can verify; the second is a vendor’s claim about its own product, so treat it as a starting assumption and check your own pages with dynamically loaded content rather than taking it on faith.
Check which one you are on rather than guessing:
grep -E '"(@tanstack/react-start|react-router-dom)"' package.json
Three situations still leave a real hole, and they are the ones worth your attention:
- You left Lovable hosting. Exported to GitHub and deployed to Vercel, Cloudflare Pages, Netlify or your own box, the verified-crawler pre-rendering does not come with you. A plain Vite build on someone else’s CDN is a client-rendered app again, and every non-rendering agent gets an empty root div.
- Verified-crawler lists drift. Serving bot-only HTML means maintaining an allowlist, and new agents ship constantly. An agent that is not on the list gets the shell, silently, with no error anywhere.
- Your content is fetched in a component hook. This one is stack-independent and the most common of the three. The HTML contains your layout plus whatever the route loader returned; a
useEffector an un-prefetcheduseQueryruns after hydration, and a non-rendering agent never gets that far. A pricing table, a listings grid, reviews: server-rendered page, missing content. Given how rarely Lovable apps use loaders, assume this applies to you until you’ve checked. It takes one curl per route to settle. SSR versus SPA rendering covers that gap in detail, and Lovable’s pre-rendering behavior covers the verified-crawler path specifically.
Prove it in one command
Everything above is theory until you fetch your own page the way a bot does. No JavaScript, no browser. Run it against a route whose content comes out of your database (products, listings, a blog index) and not against the homepage, where the words are usually written into the component and prove nothing about your data.
URL="https://yourapp.com/products"
BOT='Mozilla/5.0 (compatible; GPTBot/1.1; +https://openai.com/gptbot)'
# What a non-rendering AI agent receives. The search phrase must be one contiguous run of text
# that isn't in your source code: markup splits phrases across elements.
curl -sL -A "$BOT" "$URL" > /tmp/page.html
grep -qiF 'Ceramic Pour-Over Kettle' /tmp/page.html \
&& echo 'in the HTML' || echo 'client-fetched: non-rendering agents never see it'
# Size delta between a plain fetch and a crawler UA reveals bot-only prerendering.
curl -sL "$URL" | wc -c
curl -sL -A "$BOT" "$URL" | wc -c
# Are you accidentally blocking the agents that cite you?
curl -sL https://yourapp.com/robots.txt \
| grep -iE "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot"
A large byte delta between the two fetches means bot-only HTML is in play. Similar sizes with your product name present means real HTML for everyone. Similar sizes with it absent means non-rendering agents are getting nothing. On Lovable-hosted React + Vite, curl can’t tell you what verified crawlers receive. Check which stack you are on first, then read the three ways to misread this below.
robots.txt, per bot, without the cargo cult
Three rules first, because they cause most of the damage:
Disallowcontrols crawling only. A disallowed URL can still be indexed. Google’s docs say it may still index a disallowed URL and show it without a snippet, and it can never read anoindexon a page it was told not to fetch.- Never block your JavaScript or CSS. Googlebot and Applebot need them to render, so blocking
/assets/is the one reliable way to make a React app genuinely unreadable to the crawlers that do. - A blanket
Disallow: /underUser-agent: *also kills social link unfurlers, usually discovered by a marketing person on a Monday.
A working file for a Lovable site that wants AI visibility:
User-agent: *
Allow: /
Disallow: /app/
Disallow: /api/
# Retrieval agents: these are the ones that produce citations.
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
# Corpus builders: a content policy decision, not a visibility one.
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
Sitemap: https://yourapp.com/sitemap.xml
Two footnotes. Perplexity documents that Perplexity-User generally ignores robots.txt, on the grounds that a person initiated the fetch, so that line is advisory rather than binding. And the usual limits apply: the file must sit at the domain root, anything past 500 kibibytes is ignored, and Google caches it for up to 24 hours: a fix is not instant. Sitemaps and robots.txt on Lovable goes deeper.
llms.txt: the honest status
Skip it, or publish one for your own coding agents and expect nothing from search.
Ahrefs studied 137,210 domains in May 2026. About 28% published a valid llms.txt, and 97% of those files received zero requests that month. Of the 3% that saw any traffic, 96% of the requests were bots, and the composition is the tell: SEO audit tools were the largest single category and retrieval bots barely registered. Ahrefs called the files largely decoration.
Google representatives have said repeatedly that Search does not use it, with John Mueller reportedly comparing it to the meta keywords tag: ignored for a decade precisely because the site operator controls it. Those are secondhand reports of remarks at events rather than a written Google doc, so hold them loosely. Lovable’s position is documented, though: its docs state llms.txt is not required and that the built-in SEO and AI search review does not treat a missing file as a problem.
Structured data: real, but not the lever you were sold
The layered answer, because a single yes or no is wrong here.
Google reads it, including client-injected JSON-LD: Google’s docs say Search can understand structured data available in the DOM when it renders the page. The warning attached is that dynamically generated markup makes some crawls less frequent and less reliable, and the common SPA bug is a race: JSON-LD written in a useEffect after an async fetch gets snapshotted with "name": "" because the data had not landed.
AI crawlers see nothing. Client-injected JSON-LD does not exist in any form to an agent that never runs the effect.
ChatGPT may not see it even server-rendered. The July 2026 retrieval investigation found that pages are converted to Markdown for the model and that scripts, iframes and JSON-LD are stripped in that conversion, while image alt text survives. That is one study by one agency, and nobody has reproduced it or had it confirmed by OpenAI. It is still the only rigorous public look at the pipeline, and it contradicts the entire genre of posts promising schema boosts AI citations. The circulating statistics on that point (“2.5x more likely to appear in AI answers”, “61% of ChatGPT citations”) have no traceable study behind them.
So: put JSON-LD in the server response because it costs nothing and Google reads it reliably there. Do not tell a client it drives ChatGPT citations. And accept the implication: the visible text in your initial HTML is doing the work.
Write pages that survive being turned into Markdown
Once the HTML is real, AI visibility stops being an infrastructure problem and becomes an editing one. A few things follow from how these systems actually consume a page.
The opening lines under a heading carry enormous weight. The same July 2026 investigation found that in ChatGPT’s fast path the model sees roughly 200 characters after the H1, and that this snippet is frozen and query-independent: the same URL yields the same snippet whether the question was about pricing or history. The same study reports meta descriptions are ignored entirely for results from OpenAI’s own index. Lead every section with the answer, in plain sentences, immediately under the heading.
Make claims self-contained. An engine lifts a single passage out of your page. “This is roughly double the previous figure” is unusable out of context; “GPTBot hit a 404 on 34.82% of its fetches, against 8.22% for Googlebot” survives the trip. Name the subject, carry the number, carry the date.
Write headings that read like the question. Fan-out retrieval decomposes a prompt into subqueries, so a plain heading is a better anchor than a clever one.
Do not make an unsophisticated crawler work. Vercel’s data has AI crawlers wasting roughly a third of their fetches on 404s (34.82% for ChatGPT’s bot, 34.16% for Claude’s, against 8.22% for Googlebot) plus 14.36% redirects for ChatGPT’s. They do not retry and they do not wait. Server 301s rather than window.location redirects, real status codes rather than soft 404s, and no content behind infinite scroll or a click.
Keep alt text honest. It is one of the few non-body elements reported to survive the conversion to Markdown.
Measuring whether any of it worked
Two signals, and they diagnose different failures. Analytics, the number most people reach for first, is not one of them.
Server logs are the free, unfakeable one. Count agents by name and you learn whether you are being fetched at all:
grep -hoE "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User" \
/var/log/nginx/access.log* | sort | uniq -c | sort -rn
Read it this way: no retrieval-agent hits means a reachability problem: rendering, robots.txt, a WAF rule rejecting bot user agents, or a bot challenge on public pages. Hits but no citations means the content is being read and not chosen, which is a different job entirely. The catch on Lovable is that you probably do not have that log file. As of September 2026 Lovable’s docs describe backend and edge-function logs inside the editor but do not document access to raw HTTP request logs, so this check realistically only exists if you self-host or put something in front of the origin.
Analytics will undercount. The utm_source=chatgpt.com parameter does not appear on pages opened during ChatGPT’s thinking mode, so GA4 alone misses a chunk of the traffic. Do not report an AI referral number as if it were complete.
Visibility trackers report sample statistics. Profound, Peec, Otterly, Ahrefs Brand Radar and Semrush’s AI visibility product all run a prompt set on a schedule and parse the answers. Results are non-deterministic and sensitive to prompt phrasing, personalization and model version, so a “share of voice” number describes the prompt set you chose: fine for tracking your own trend over time. Encited’s AI visibility tracker works this way, sampling ChatGPT, Perplexity and Gemini answers for the prompts you choose.
One last reason not to optimize once and assume it generalizes: Ahrefs compared 540,000 query pairs in September 2025 (US only) and found only 13.7% citation overlap between Google’s AI Overviews and AI Mode, despite the answers being 86% semantically similar. AI Overviews moved to Gemini 3 in January 2026, so treat the exact figure as a snapshot; the structural point, that the two surfaces cite different sources, is the durable part. Measure each surface separately.
Doing this by hand means re-running curl against a handful of URLs and asking the answer engines the same questions every few weeks. Making it continuous takes two pieces: request logs that record which AI bots fetched which URL and what they received, and a tracker that samples answers for the prompts you care about. Encited does both (per-bot crawl logs that show whether GPTBot or ChatGPT-User got rendered HTML or the shell, and citation tracking across ChatGPT, Perplexity and Gemini) and it can pre-render pages whose data is fetched in a component hook for agents that don’t run JavaScript. Your CDN’s raw logs filtered by user agent cover the first piece if you’d rather assemble it yourself.
What to do next
- Check
package.jsonfor@tanstack/react-startfirst: the stack decides which test is valid. On the TanStack stack, run the GPTBot curl against three URLs with dynamically loaded content and grep for a phrase that isn’t in your source code, such as a product name. If you are still on Lovable-hosted React + Vite, curl can’t answer the question, and what verified crawlers receive can’t be verified from outside. Only once a real crawler confirms the content is missing is rendering the thing to fix. - Then find where the failing route actually fetches. Moving a fetch from a component hook into a route
loader(or into acreateServerFnthe loader awaits) is the cheapest fix available and it works on the TanStack stack today. A server function called from the component after hydration is still a client fetch and changes nothing. If the project is still Vite and has moved off Lovable hosting, you need build-time prerendering, SSR, or an edge prerender layer instead. - Grep robots.txt for the retrieval agents specifically. Decide the corpus builders separately, as a content policy question.
- Delete
llms.txtfrom your to-do list. Spend that hour rewriting the first 200 characters under each H1 instead. - Move any JSON-LD out of
useEffectand into the server response, then stop counting on it for AI citations. - Start recording which AI agents reach your origin, and re-read the numbers in a month. If retrieval agents are fetching you and still not citing you, the next fix is editorial. Start with the Lovable SEO guide.
Frequently asked questions
- Do AI crawlers execute JavaScript?
- Almost none of them do. Vercel's December 2024 log study across its network found that GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Meta-ExternalAgent, Bytespider and PerplexityBot all read raw HTML without running JavaScript, and a separate July 2026 investigation of ChatGPT's retrieval stack reached the same conclusion. Applebot is the documented exception: Apple states it may render pages in a browser.
- Do I need an llms.txt file for AI search visibility?
- There is no evidence it helps. Ahrefs studied 137,210 domains in May 2026 and found 97% of published llms.txt files received zero requests that month; most of the traffic the remaining files got came from SEO and audit tools rather than answer engines. Google representatives have said Search does not use it, and Lovable's own documentation says it is not required and a missing file is not treated as a problem.
- Does adding schema markup get my page cited by ChatGPT?
- There is no credible public evidence that it does. The most detailed public investigation of ChatGPT's retrieval stack, published in July 2026, reported that JSON-LD is stripped when pages are converted to Markdown for the model. Structured data is still well evidenced for Google rich results and worth server-rendering, but treat LLM citation claims about it as unproven.
- My Lovable app is TanStack Start with SSR. Does that make it visible to AI crawlers?
- Not on its own. SSR renders your component tree, not your data: React's documentation states that effects only run on the client and do not run during server rendering, and TanStack Query's SSR guide says unprefetched queries are fetched on the client after the app is interactive. What a route loader awaits during server rendering is in the HTML; what a component hook fetches after hydration is not, however thoroughly the page was server-rendered. In a study of 82 Lovable apps, only 4.4% of page routes loaded their data in a route loader, so assume yours is missing from the HTML until you have checked.
- Is Google-Extended an AI crawler I should worry about?
- No, because it never makes a request. Google's documentation states Google-Extended has no separate HTTP user agent string and that the robots.txt token is used in a control capacity only. It governs whether your content is used for Gemini training and grounding; the actual fetching is done by Google's existing crawlers.
- My Lovable site gets AI crawler hits but is never cited. What now?
- That is a content and authority problem. Crawler hits in your logs prove the HTML is reachable and parseable, so the remaining work is ordinary: cover the question directly, answer it in the first lines under the heading, make claims self-contained, and earn the kind of references that make an engine pick you over a competitor.
Read next
Keep going
-
What 82 Lovable TanStack apps actually server-render
Lovable TanStack SSR, measured: in 82 Lovable apps, only 4.4% of pages load their data on the server. The layout reaches the HTML; the data usually doesn't.
-
SSR vs SPA for SEO: what actually changes
SSR vs SPA SEO, settled: Googlebot renders your SPA, but unfurlers and AI crawlers read raw bytes, and SSR only helps if your data is fetched in a route loader.
-
Pre-rendering a Lovable app: when you need it
Lovable prerendering, honestly: SSR puts your page layout in the HTML but usually not your data. Why that happens, and a 30-second test for your own site.
-
Lovable SEO: the complete 2026 guide
Lovable SEO splits across two stacks, and SSR alone does not put your data in the HTML. Test both on your own routes, then close the gaps Lovable leaves to you.