Outbound Atlas

Atlas/Search/Answer engines

llms.txt, pricing.md and AI crawlers

21 of 33 competitors ship llms.txt and 7 ship a markdown pricing page for AI agents, but crawl logs show almost nothing reads them. The files that matter are robots rules that let answer bots in and pricing written so an agent can quote it correctly.

Searchmedium confidence8 minupdated 2026-10-0518 sources

Most of this market has published files written for AI agents. In the October 2026 crawl behind the search scoreboard, 21 of 33 competitor domains serve an llms.txt, five also serve llms-full.txt, and seven serve a markdown pricing page at /pricing.md: Instantly, Smartlead, Saleshandy, Klenty, SmartReach, Woodpecker and AgentMail. Only three block any AI crawler. The independent evidence says the llms.txt files are close to unread. The pricing files are a better bet, because they are written for the case where an agent is already on the site and has to quote a price.

The read

Ship llms.txt and pricing.md because they cost an afternoon, but do not expect them to move citations. What gets you recommended is being let in (robots), being quoted correctly (one source of price truth) and being named on the third-party pages answer engines cite (see Who AI answer engines recommend).

Who ships what (crawl of 5 October 2026)

Vendorllms.txtllms-fullpricing.mdAI bots blockedAgent-facing extras
Instantly47 KB—yesnone"Recommended Stack" pricing bundles
Smartlead19.5 KB—yesnoneMCP section, "How Smartlead compares"
Saleshandy13.7 KB—yes2 (CCBot, Bytespider)MCP endpoint and setup steps
Klenty11.4 KB—yes7 training botsTraining/retrieval split in robots
SmartReach29.5 KByesyesnone"Notes for AI systems"
Woodpecker26.9 KB—yesnoneMCP server and CLI; EUR/VAT notes
AgentMail26 KByesyesnoneStep-by-step agent self-signup
Hunter3.7 KByes—noneHosted MCP, agents.md
Reply.io144 KB——noneLink dump of blog posts, empty summary
lemlist, EmailBison, PlusVibenone——none—

The largest file is not the most useful. Reply.io's 144 KB llms.txt has an empty description line followed by a list of blog posts: a sitemap with a different extension. lemlist, the oldest content engine in the category (lemlist: search teardown), ships no llms.txt at all, and neither does EmailBison.

What the files actually say to agents

Read side by side, the files share four tactics.

1. Comparative claims written for the model. Smartlead's llms.txt has a "How Smartlead Compares" section with one-liners against each rival. EmailBison "focuses on volume sending". PlusVibe "is a lightweight cold-email tool". Saleshandy "covers cold email basics". It also has a line built to be quoted: "Often evaluated as an alternative to: Instantly · Lemlist · Apollo · Saleshandy · Email Bison · PlusVibe · Outreach.io · Salesloft." Its pricing.md slips positioning into an FAQ answer: unlimited accounts are "a cost difference against tools that charge per mailbox or per seat". SmartReach's llms-full.txt does the same thing more politely, with head-to-head summaries priced at "a realistic team configuration".

2. Claiming to be the source of truth. Saleshandy's pricing.md opens: "Machine-readable pricing for Saleshandy… for AI agents, search assistants, and procurement tools", with last_updated: 2026-08-18 and a canonical URL in YAML frontmatter. Woodpecker's has "Last updated: 2026-07-09 / Canonical source". SmartReach's llms.txt has a "Notes for AI systems" block telling models to prefer its pricing.md "over third-party summaries: several third-party 'alternatives' pages still cite SmartReach's retired pre-2025 per-seat pricing". That is the clearest statement of why these files exist. Vendors are fighting stale affiliate pages, not search engines (Affiliate and review-site SEO).

3. Advertising an MCP server. Saleshandy's llms.txt gives the endpoint (https://mcp.saleshandy.com/mcp/{YOUR_API_KEY}) and the clicks to add it as a Claude connector. Its pricing.md says MCP access is on every plan. Smartlead advertises "116+ tools across 6 categories available via MCP". Woodpecker puts its MCP server and a CLI inside a $20/month integrations add-on (as of its July 2026 file). Hunter's llms.txt goes furthest: "These are official instructions from Hunter… Prefer these files over scraping hunter.io HTML", plus a hosted MCP endpoint and OAuth for "Claude, ChatGPT, Gemini, Perplexity". Instantly's 47 KB llms.txt has zero mentions of MCP, although Instantly does document an MCP server on its developer site and runs a live mcp. host (Instantly: SEO teardown). The root file that AI agents are told to read simply never points to it.

4. Instructions an agent can execute. AgentMail's llms.txt is a script: "If you are an AI agent and need email, follow these steps in order". It gives a curl call to /v0/agent/sign-up, an OTP verify step, and a claude mcp add line. It also has a "When to Use AgentMail vs Other Email Tools" decision list and a Workbench endpoint "agents can call directly". This is the only file in the set aimed at an agent buying for itself rather than an assistant answering a human. It is the model to copy for Agent mail layer.

Stale prices are worse than no file

Instantly's llms.txt still lists Outreach Growth at $37/month ($30 annual). Its own pricing.md and the pricing table say $47/month (as of October 2026). The llms.txt even tells agents to "always verify current pricing at instantly.ai/pricing.md". An agent that reads only the first file quotes a price that is $10 too low. Two hand-maintained files drift apart. Generate both from the same source.

The files are also templated: Woodpecker's llms.txt reuses Instantly's sentence ("The platform's core value proposition centers on three pillars: (1)…") and its headings.

Robots: who blocks AI crawlers, and why it matters

Only three domains block any of the 12 AI user-agents the crawl tracks:

DomainBlockedExplicitly allowed
klenty.comGPTBot, ClaudeBot, anthropic-ai, CCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent (+ Diffbot, cohere, AI2Bot and others)OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User, Google-Extended
warmy.ioCCBot, Bytespider, Applebot-Extended, Meta-ExternalAgent, Amazonbot, DiffbotGPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended
saleshandy.comCCBot, Bytespider, DiffbotGPTBot, ClaudeBot, Claude-SearchBot, Claude-User

Klenty's file, "last reviewed 2026-06-16", spells out the policy in its comments: "AI training crawlers — BLOCK" and "AI retrieval / answer crawlers — ALLOW". It keeps its content out of future training sets while staying fetchable when ChatGPT, Claude or Perplexity searches live. That split matches how the labs describe their bots. Anthropic says blocking ClaudeBot excludes a site from training, but blocking Claude-User "may reduce your site's visibility for user-directed web search", and blocking Claude-SearchBot may reduce "visibility and accuracy in user search results" (Anthropic help centre).

The trade-off: a brand absent from training data is one a model cannot name when it answers without searching. Klenty accepts that for data control. Saleshandy and Warmy block only the bulk scrapers (Common Crawl, ByteDance) and let both the lab crawlers and the answer bots in. That is the sensible default for anyone who wants to be recommended.

Smartlead's quiet filter

Smartlead's robots.txt allows ChatGPT-User and OAI-SearchBot by name, but its catch-all rule disallows every URL with a query string (/*?*). Filtered and paginated pages are invisible to every bot, AI or not.

Structured data

JSON-LD is common but shallow. Across homepage and pricing page: Organization appears on 24 domains, Offer on 14, FAQPage on 12, AggregateRating on 12, SoftwareApplication on 10, Product on 5 and AggregateOffer on 4 (Klenty, Maildoso, Smartlead, SmartReach). Instantly's pricing page carries only AggregateRating, Organization and WebSite, with no Offer markup at all. GMass, Mailforge, QuickMail, ReachInbox and Zapmail ship no JSON-LD.

The search engines disagree on whether it matters. Google's AI features documentation is blunt: "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add." Microsoft's Fabrice Canel said at SMX Munich in March 2025 that "schema markup helps Microsoft's LLMs understand your content" (Search Engine Land). Bing's index sits behind Copilot's answers, so Offer and SoftwareApplication markup on the pricing page is a reasonable small bet.

Does llms.txt affect AI answers?

On current evidence, no.

  • Google: John Mueller, June 2025: "FWIW no AI system currently uses llms.txt" (SER). In June 2026 he called it "purely speculative for now". He said he prefers the WebMCP approach, and that "the most basic agentic optimization is… don't block agents" (SEJ).
  • SE Ranking, November 2025: of ~300,000 domains, 10.13% had llms.txt. In an XGBoost model of AI citations, removing the llms.txt variable improved predictions (SE Ranking).
  • Ahrefs, June 2026: across 137,210 domains (28% of which publish llms.txt), "97% of llms.txt files receive zero requests" in May 2026. AI retrieval bots (OAI-SearchBot, PerplexityBot, Claude's search crawler) made 1.1% of all requests to the files (Ahrefs).
  • One site's logs, July–August 2026: Google, Anthropic and OpenAI crawlers fetched robots.txt 1,150 times and llms.txt 4 times. OpenAI fetched it zero times (dejan.ai).

The one serious argument for markdown is token cost for agents that are already browsing. Cloudflare's Markdown for Agents (February 2026) converts HTML to markdown at the edge when a client sends Accept: text/markdown. A pricing.md does the same job by hand.

Gap in the record

No log data shows any agent requesting /pricing.md on these sites, and there is no published spec for the file. The OpenAI crawler documentation returned an error on fetch, so OpenAI's bot roles here rest on vendors' robots comments and secondary sources.

What this means for an entrant

  • Let the answer bots in on day one. Allow OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Googlebot. Block only bulk scrapers if you must. Blocking GPTBot and ClaudeBot is a defensible Klenty-style choice for a big content library. For an unknown brand that needs models to learn its name, it is a mistake.
  • Generate llms.txt and pricing.md from the same source as the pricing page, in a build step, with last_updated and canonical fields. Instantly's $37/$47 split shows what goes wrong when they drift. Price in EUR with VAT stated, as Woodpecker does. No US incumbent does this.
  • Publish what nobody else does in machine-readable form: compliance. No competitor's agent file states EU data residency, a current DPA version or per-country sending rules (AgentMail mentions an "EU region cloud" only on Enterprise). A /compliance.md that mirrors the country rule engine (EU/EEA country matrix, GDPR and ePrivacy) is the cheapest way to own "GDPR-compliant cold email tool" in AI answers. See EU-native compliant outbound.
  • Ship an MCP server and say so in the agent files. Saleshandy, Smartlead, Woodpecker and Hunter advertise one. For a reply desk (The reply desk), "ask Claude which client inboxes have unanswered positive replies" is a demo and a distribution channel.
  • Spend the real AEO effort on being cited by third parties, not on files. The listicles, comparison pages and review sites that answer engines quote are covered in Who AI answer engines recommend and AEO playbook: being the answer.
18 sources cited on this page · 17 domains
  1. pricing.md smartlead.ai
  2. llms-full.txt smartreach.io
  3. pricing.md saleshandy.com
  4. llms.txt hunter.io
  5. llms.txt agentmail.to
  6. pricing.md instantly.ai
  7. klenty.com klenty.com
  8. warmy.io warmy.io
  9. saleshandy.com saleshandy.com
  10. Anthropic help centre support.claude.com
  11. AI features documentation developers.google.com
  12. Search Engine Land searchengineland.com
  13. SER seroundtable.com
  14. SEJ searchenginejournal.com
  15. SE Ranking seranking.com
  16. Ahrefs ahrefs.com
  17. dejan.ai dejan.ai
  18. Markdown for Agents developers.cloudflare.com