More and more people ask an AI to recommend where to eat, where to stay or whom to hire. The AI answers with what it finds and understands on the web.
North team · Updated October 7, 2026 · 3 min read
Where assistants get their information
Assistant
Crawler that reads your site to answer
What else carries weight
ChatGPT with search
OAI-SearchBot
Pages that answer the question, reviews and mentions on other sites
Gemini and AI Overviews
Googlebot, the same one as Google Search
Your position on Google and your Business Profile
Perplexity
PerplexityBot
Sources it can cite: guides, forums, reviews
Claude
Claude-SearchBot
Clear pages and third-party sources
Training crawlers (GPTBot, ClaudeBot and the Google-Extended permission) are a separate decision: blocking them does not remove you from search-based answers. Blocking the search crawlers does.
Methods
Opening the site to AI search engines
The robots.txt says which crawler can come in. Search crawlers decide whether you appear in answers; training crawlers decide whether your text is used to train models.
robots.txttext
# Search: read and cite the site in answers
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Model training: a separate decision
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
Sitemap: https://casaaurora.mx/sitemap.xml
Recommendation. Also check the firewall or CDN: many block bots by default even when the robots.txt lets them in.
Seeing which sources AI cites
When an assistant doesn’t mention you, the answer shows where it got the information. That is where the work is.
Analysis of one questionlog
question: "boutique hotel with restaurant in Tulum"
assistant mentions Casa Aurora? sources cited
ChatGPT yes, 1st casaaurora.mx · TripAdvisor · guide from a travel magazine
Perplexity no Booking · Reddit thread about Tulum · travel blog
action: reply in the Reddit thread as the hotel, with facts and without selling
ask for TripAdvisor reviews · pitch the hotel to the magazine
Recommendation. Working on the sources the assistant already cites pays off more than publishing new pages nobody mentions.
Asking the customer
The most direct way to measure what AI brings in is to ask at booking time, as Tally does in the case below.
Question in the booking formtext
How did you hear about us?
( ) Google ( ) Instagram ( ) Recommendation ( ) ChatGPT or another AI
If it was an AI, what did you ask it? ______________________
What AI looks at
A website that says clearly what you do, where and for whom.
Structured information that machines can read.
Reviews and ratings on well-known sites.
Mentions in media, guides and directories in your industry.
Case analysis
Tally: when ChatGPT became its main source of customers
Tally · Belgium, 2025 and 2026
A form builder with a small team and no investors competes against giants. Its advantage: being where AI looks for answers.
What they did, step by step
Built in public and answered every message from its community.
Took part in forums and on Reddit by actually contributing, not advertising.
Added a question at sign-up: how did you hear about us and what did you ask the AI.
Used those answers to learn which AI questions brought in the most customers.
Published comparison pages against Typeform, Jotform and Google Forms.
No. 1source of new sign-ups: ChatGPT
+2,000sign-ups from AI tools, counting only the ones it could track
5 monthsahead of plan it reached USD 3M in annual revenue
The mechanism
Community and forumsReddit, real contributions
Comparison pagesagainst competitors
Sign-up surveywhat did you ask the AI?
Questions that bring the mostaccording to the survey
Cited by AIChatGPT, Claude, Gemini
Asking customers what they asked the AI turns a black box into a measurable data point.
Why it worked. AI assistants cite real conversations and clear pages. Tally was in both.
Analysis of a public case. North Marketing did not take part in this work. Figures self-reported by the company on its blog.
Assistants that search the web in real time, like ChatGPT with search or Perplexity, reflect a change when they crawl the page again. What comes from a model’s training takes months and can’t be controlled. That is why the work targets search.
Is the llms.txt file useful?
It is a proposal for summarizing a site for language models. No major search engine has confirmed using it to choose what to recommend, and Google said it does not use it. You can have one, but it does not replace robots.txt, the sitemap or clear pages.
Let’s start with a free diagnosis
30 minutes to see where your business is losing customers and what to do first. If you don’t see results in 30 days, we give your money back.