· Marc Price · ai-strategy  · 11 min read

Your Next B2B Buyer Is an AI, and Your Data Isn't Ready

51% of B2B software buyers now start vendor research in an AI chatbot, and 69% end up choosing a vendor they did not expect. The shortlist is being written somewhere your content cannot reach.

51% of B2B software buyers now start vendor research in an AI chatbot, and 69% end up choosing a vendor they did not expect. The shortlist is being written somewhere your content cannot reach.

TL;DR

Half of B2B software buyers now begin their research in an AI chatbot rather than a search engine, and more than two-thirds of them end up choosing a vendor they did not have in mind when they started. That shortlist is written by a model reading whatever it can reach, which for most £10M-£50M firms means the public website and third-party review sites, and not the gated case studies, the sales-deck PDFs or the CRM where the real proof lives. This post sets out what an AI-readable content architecture looks like: one canonical facts layer, proof with numbers on public pages, third-party corroboration, an audit of the “no robots” signs you did not know you had, and a single source of truth underneath so the AI, the website and the rep all say the same thing.


Who is actually reading your website now?

An AI, on behalf of a buyer who will never see your homepage.

G2’s The Answer Economy report, published in April 2026 from a survey of 1,076 B2B decision-makers, found that 51% of B2B software buyers now start research with an AI chatbot more often than with Google. A year earlier it was 29%. Once they start there, the outcome moves: 69% chose a different vendor from the one they initially planned, and 33% bought from a vendor they had not previously heard of, because of what the AI told them. 85% think more highly of vendors an AI recommends.

The sobering line in the same report is that 64% of buyers say they encounter inaccurate AI recommendations “often” or “very often”. They know the answers are patchy. They use them anyway, because the alternative is twenty tabs and a form-fill. Forrester’s State of Business Buying 2026 reaches the same place from the other direction: generative AI was the single most cited meaningful interaction type for researching purchases, and 36% of buyers said it made them more confident in their decision.

So the first conversation about your firm is now happening between a buyer and a model. We wrote about GEO last November when this was an emerging behaviour. It is now the default, and the question has changed from “how do we rank in AI search” to “what does the AI actually know about us, and where did it get it”.

Why can’t the AI see what you actually do?

Because the good stuff is locked in formats it cannot read, behind doors it cannot open.

Walk through a typical mid-market B2B firm’s evidence base. The case study with the real numbers is a PDF behind a form. The pricing logic is in a sales deck. The customer list is in the CRM, along with the win reasons, the integrations that actually get used and the sectors where retention is strongest. The website says “trusted by leading organisations” and shows six logos. When a buyer asks ChatGPT “who does X for a company like ours in the UK, and what do they cost”, the model has your six logos and an adjective, and it has the competitor’s public pricing page and a G2 profile with 140 reviews. It recommends the competitor. It is not wrong to.

Then there is the door itself. A July 2026 audit of the top 2,000 domains found that 19.1% block at least one AI crawler in robots.txt, and blocking is nearly all-or-nothing: 75% of blockers shut out three or more crawlers, with no distinction between training bots and the answer bots that cite sources. More telling, a further 8.9% of domains allowed crawlers in robots.txt but refused them at the network edge - a firewall rule, a bot-management product or a CDN default that nobody in marketing chose. That is a “no robots” sign in the shop window that the shop does not know it has hung.

Both problems have the same root: the content was built for a human on a browser, and the proof was built for a rep on a call. Neither was built to be queried.

Doesn’t schema markup fix this?

No. It helps machines classify what they already found. It does not make them find or trust it.

This matters because “add structured data” is the reflex answer, and the evidence has moved. Ahrefs tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026, matched against 4,000 control pages, and measured citations across Google AI Overviews, AI Mode and ChatGPT. The result: AI Overviews -4.6%, AI Mode +2.4%, ChatGPT +2.2%, none of it distinguishable from random variation. Keep your schema; the FAQ block at the bottom of this post has it. Just stop expecting it to do the heavy lifting.

What the models reward instead is unglamorous: facts stated plainly, on public pages, consistently across the web, corroborated by someone who is not you. G2 found 45% of buyers say a citation from a review site is the single most confidence-inspiring signal in an AI answer. Gartner’s 2026 survey of 645 B2B buyers found they used an average of seven information sources per purchase. The AI is triangulating. If your own site is the only place your claims exist, they weigh very little.

What does an AI-readable content architecture look like?

Five layers, built in this order. None of them is a new tool.

1. One canonical facts layer. A single public page, or a small set, that states in plain text what you do, for whom, in which sectors and geographies, how pricing works, what you integrate with, and how you differ from the two or three alternatives a buyer will name. Written for a reader who has ninety seconds, not for a designer. This is the page the model will quote. If it does not exist, the model writes its own version from fragments.

2. Proof as claims, not assets. Every case study gets an ungated public summary: the customer’s sector and size, the problem, what was done, the number. “A 200-person logistics firm cut quote turnaround from 3 days to 4 hours” is a claim a model can retrieve and repeat. A 12-page PDF behind a form is not. Keep the gated version for the humans who want depth; put the claim in the open.

3. Third-party corroboration. Review platforms, directories, partner listings and the trade press are where the model checks you. A profile with twelve honest reviews beats a website with none of them, and it is where 45% of buyers say they get their confidence. This is a quarterly operational task, not a campaign.

4. Open the door and check the window. Audit robots.txt and the edge. Fetch your own key pages with the user-agents of the major answer bots and confirm they get a 200, not a challenge page. Decide deliberately which crawlers you allow. If you want to be cited, the answer bots have to get in. Publish an llms.txt if you like, but treat it as a courtesy rather than a fix.

5. One source of truth underneath. This is the layer that keeps the other four honest. The facts page, the proof claims, the review responses and the rep’s talk track all have to carry the same numbers, which only happens when they draw from the same system. That is the argument we made in ten tools to one system: the unified data layer is not a RevOps nicety, it is what stops your public story drifting away from your CRM. When the customer count on the website, the number in the deck and the figure the AI quotes are all different, the buyer notices at exactly the wrong moment.

What about the humans?

They arrive later, and they arrive to check.

Gartner’s survey found 67% of B2B buyers prefer a rep-free experience and 70% prefer a completely digital, self-service journey, which sounds like the end of sales. But 69% of the same buyers turn to a sales rep to validate AI-generated insights, and they are split almost evenly on whom to distrust: 51% expect misleading information from GenAI, 49% from a rep. The buyer has read the AI’s summary of you before the first call. The rep’s job is to confirm it, correct it gently, and add the specifics the model could not reach.

That only works if the rep and the model are telling the same story. If the public facts layer says one thing and the rep says another, the buyer does not conclude the AI was wrong. They conclude you are inconsistent, and the CFO who has to sign off already has a shortlist that did not include you a week ago.

The Bottom Line

The B2B buying journey has acquired a new first participant, and it does not fill in forms. It reads what is public, weighs what is corroborated, and shortlists from that. For most £10M-£50M firms the honest position is that the most persuasive evidence they own is invisible to it.

Do this in the next quarter:

  1. Ask the three main chatbots who they would recommend for what you do, in your region. Write down what they say about you and where they got it.
  2. Fetch your key pages as an answer bot. Fix any door that is shut.
  3. Publish one canonical facts page and an ungated claim for every case study.
  4. Get your review profiles to a state you would be happy for a model to quote.
  5. Put the numbers on all of the above under one source of truth, so they stay the same.

None of this is clever. It is plumbing. But the shortlist is being written now, and it is being written from what is reachable.

If you would like to know what the AI currently says about your firm and where the gaps are, that is a two-hour piece of work we do regularly. Book a discovery call and we will run it for you.


Frequently Asked Questions

How many B2B buyers actually start their research with an AI chatbot?

51%, according to G2’s April 2026 survey of 1,076 B2B decision-makers - up from 29% a year earlier. 63% of them use ChatGPT as their primary tool. That is the first-touch point for a majority of software purchases, and it is a channel your analytics cannot see.

Does adding schema markup get us cited by AI?

Not on its own. Ahrefs tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026 against 4,000 control pages: ChatGPT citations moved +2.2%, Google AI Mode +2.4% and AI Overviews -4.6%, none of it distinguishable from noise. Schema helps machines classify content. Plain, answerable, consistent facts are what get cited.

Could our website be blocking AI crawlers without us knowing?

Yes, and it is common. A July 2026 audit of the top 2,000 domains found 19.1% block at least one AI crawler in robots.txt, and a further 8.9% allow crawlers in robots.txt but refuse them at the network edge - a WAF rule or CDN default nobody chose. Check both, not just robots.txt.

What should an AI be able to find about us in under a minute?

Six things, each on a public page, each in plain text: what you do and for whom, what it costs or how pricing works, where you operate, three proof points with numbers, how you compare to the obvious alternatives, and who is accountable for the claims. If any of those lives only in a PDF, a gated download or a CRM field, the AI will fill the gap from a competitor or a review site.

If buyers use AI to shortlist, do sales reps still matter?

More than the numbers suggest. Gartner’s 2026 survey of 645 B2B buyers found 69% turn to a sales rep to validate what the AI told them, even though 67% would prefer a rep-free experience. The rep’s job is now to confirm the story the buyer already read. If your reps and your public content disagree, you lose the deal at validation.


References


Marc Price is the founder of Aandai, a B2B automation and AI consultancy helping mid-market businesses achieve more with less. With 25+ years in B2B technology marketing and web development, Marc specialises in connecting legacy systems, eliminating manual processes, and implementing practical AI solutions that deliver measurable ROI. Aandai runs its own agentic stack on OpenClaw to automate parts of its consultancy delivery - including the research that informed this article.

Related Posts

View All Posts »