AI Commerce · Structured Data

Why Unstructured Product Data Makes Your B2B Catalog Invisible to AI Shopping Agents

AI-driven orders grew 15x year over year, but most product pages still aren’t built for AI agents to read. Here’s what a B2B catalog needs to fix.

Ceejay S Teku September 2, 2026
Catsy PIM and DAM structuring product data across multiple connected storefronts

Direct answer: Structured product data for AI shopping agents means AI agents read structured data, not page design. They parse JSON-LD and schema.org fields, not marketing copy or a spec sheet PDF. Most B2B product pages score around 66% machine-readable, the lowest of any page type Adobe measured, because exact dimensions, compliance data, and technical specs usually live in a document instead of a labeled field. Fixing that starts with an audit, not a full catalog rebuild.

What’s the Difference Between Structured Data, Schema Markup, and JSON-LD?

Structured data is the general concept: product facts organized into labeled fields instead of free text. Schema.org is the shared vocabulary that defines what those fields are called, like Product, Offer, and their related properties. JSON-LD is the specific format most AI crawlers expect that vocabulary to arrive in: a clean data block in a page’s code, separate from what a shopper sees. A field can exist in a Product Information Management (PIM) system without ever reaching an agent. It has to be structured, mapped to the schema.org vocabulary, and delivered as JSON-LD on the actual page before it counts as visible.

That data layer is a different problem from the AI shopping and checkout protocols this article covers later: MCP, ACP, and UCP. Those protocols move already-structured data between systems. None of them can expose a spec that was never structured in the first place.

Is AI Shopping Traffic Growing Faster Than Catalogs Can Adapt?

Yes. Agents are already placing real orders, and most catalogs haven’t caught up. Shopify President Harley Finkelstein told investors on the company’s Q4 2025 earnings call that orders from AI-powered search were up roughly 15 times since January 2025. He added the growth was on a small base, but accelerating fast.

That traffic also converts unusually well. Adobe analyzed over one trillion visits to U.S. retail sites and found that in March 2026, AI-referred traffic converted 42% better than traditional web traffic. That’s a record high, and a sharp reversal from a year earlier when the same traffic converted worse than average.

Checkout infrastructure for this traffic is live now, not experimental. Stripe and OpenAI co-developed the Agentic Commerce Protocol, which now powers Instant Checkout inside ChatGPT. Separately, Google introduced the Universal Commerce Protocol at NRF’s 2026 conference, an open standard that lets agents complete purchases directly inside Search and Gemini. Shopify, Etsy, Wayfair, Target, and Walmart helped build it.

These three facts describe one shift from three angles. Agents are placing real orders. That traffic converts better than the traffic catalogs spent a decade optimizing for. The biggest platforms in AI and commerce have already built the checkout rails to support it.

Why Aren’t Most Product Pages Built for a Machine to Read?

Most product pages aren’t built for a machine because they were built for a person to scan, not a script to parse. Adobe’s AI Content Visibility Checker scores how much of a page’s content is actually readable by machines. Across the U.S. retail sector, individual product pages scored an average of just 66%, the lowest of any page type Adobe measured. Homepages scored 75%. Category pages scored 74%.

That gap matters most on product pages, because that’s exactly where the comparison detail an AI agent needs actually lives, or fails to. A page can look complete to a person and still be unreadable to an agent. A spec table rendered as an image is invisible to a system parsing structured fields, not pixels. A dimension buried in marketing copy is invisible the same way.

Marketing language substitutes for data more often than teams realize. “Industrial-grade durability” tells a shopper something. It gives an AI agent nothing to compare, filter, or verify against a competing product.

The gap between sites is already wide, and it will only compound as more shopping activity routes through agents. Adobe found the best-performing U.S. retail sites scored 82.5% on average, against 54.2% for the lowest performers.

What Does “Machine-Readable” Actually Mean for a Product Record?

Machine-readable means a script can parse a product’s facts without guessing. It isn’t about writing better product copy. It’s about exposing the same facts in a format built for parsing, not persuading.

RequirementWhat It Means in Practice
JSON-LD, not just visible textA structured data block in the page’s code gives crawlers a clean, unambiguous version of a product’s name, price, availability, and attributes.
Schema.org’s Product typeProduct, Offer, and related properties are the vocabulary most AI crawlers expect. A field that never reaches that vocabulary is invisible to an agent, even if it’s technically in the PIM.
Exact values, not descriptive rangesAn agent comparing a dimension needs an actual number in a structured field, not “compact” or “fits most standard mounts.”
Live data, not a static exportA weekly feed can’t reflect a price change, a stock-out, or a spec correction in the window that matters. Agents checking availability need a live connection to the source PIM.
PIM and DAM interface showing product data enrichment across different product types and variants
Enrichment at the attribute level, across product types and variants, is what turns a PIM field into something schema.org’s Product vocabulary can actually use.

Where Do B2B Catalogs Have Specific AI-Visibility Blind Spots?

B2B catalogs have a worse version of this problem than consumer retail, because the data that matters most is also the data least likely to be structured. Most of the AI-visibility conversation so far has centered on consumer retail, not manufacturing or distribution.

Blind SpotWhy It Blocks Agent Visibility
Technical specs live in PDFs, not fieldsA spec sheet attached as a digital asset is readable by a person who opens it, not by an agent scanning structured attributes.
Exact dimensions and tolerancesA B2C shopper tolerates “approximately 12 inches.” A B2B buyer sourcing a replacement part needs an exact, comparable figure.
Compliance and certification dataCertifications often exist as a badge on a page or a line in a document, not a queryable attribute an agent can filter for.
Catalog scaleA distributor with tens of thousands of SKUs can’t patch data by hand. The fix has to happen at the attribute and hierarchy level.
B2B buyer relying on accurate, structured product data to make a purchasing decision
B2B buyers, human or AI, decide on the same thing: whether the data in front of them is complete enough to trust.

What’s a Practical Starting Checklist for AI-Ready Product Data?

Fixing this doesn’t require rebuilding a catalog from scratch. It requires being specific about what’s missing, then prioritizing the fields that actually block AI agent visibility.

ActionWhy It Matters
Audit a sample of product pages against a machine-readability frameworkEstablishes a real baseline instead of assuming pages are fine because they look fine to a person.
Identify specs that live only in a document, not a fieldStart with dimensions, materials, certifications, and compatibility data, the attributes most likely to appear in a comparison query.
Confirm each channel’s feed is live, not a periodic exportA PIM with real API connections keeps price, stock, and spec changes visible to an agent in near real time.
Check whether structured fields actually reach the page’s schema markupA field that never makes it into the page’s markup doesn’t help an agent, no matter how clean it is inside the source system. Catsy’s guidance on structured schema data applies to AI visibility the same way it applies to traditional SEO.
PIM and DAM syndication feed pushing product data to a storefront and B2B distributor channels
A live feed to each channel, not a periodic export, is what keeps a distributor listing and an AI agent’s view of your catalog in sync.

How Does This Fit Alongside MCP, ACP, and UCP?

This article covers the data layer: getting specs, dimensions, and attributes into a structured, machine-readable form in the first place. That’s earlier-stage work than any single protocol integration. Catsy’s guide to Shopify’s MCP integration covers what shopper-facing agents see once they’re already querying a live storefront. Catsy’s MCP integration with HubSpot covers a different downstream case again: generating marketing content from already-approved product data. Both depend on the groundwork this article describes. An agent connected through MCP, ACP, or UCP can only work with what’s already structured and exposed.

Diagram showing scattered, unstructured product data on one side and a streamlined PIM syndication pipeline on the other
The fix is structural, not protocol-specific: one validated source feeding every channel, instead of a separate patch for each new AI surface.

A PIM that validates attributes at the field level, then pushes them through live API connections, makes a catalog visible to every one of these channels at once. That beats fixing each new protocol separately as it launches. For manufacturers managing complex technical catalogs, that structural foundation matters more, not less, because the specs an agent needs most are exactly the ones most likely trapped in a document instead of a field. The same is true for distributors aggregating feeds from multiple suppliers, where inconsistent source data multiplies the problem across every feed.

Key Takeaways

01. AI-driven shopping is already producing real orders and converting better than traditional traffic.
02. Most product pages, especially in B2B catalogs, remain hard for agents to read even when they look complete to a person.
03. Machine-readability means structured JSON-LD using schema.org’s Product vocabulary, exact values, and live feeds.
04. B2B catalogs have specific blind spots: specs trapped in PDFs, imprecise dimensions, unstructured compliance data, and scale that blocks manual fixes.
05. Start with an honest visibility audit, not a full catalog rebuild, then prioritize the fields most likely to block agent comparison.
06. Structuring and exposing product data comes before any protocol integration like MCP, ACP, or UCP can actually help.
PIM for AI Agents

See How Catsy Structures Product Data for AI Visibility

If your pages look complete to a shopper but you’re not confident an agent could parse your specs, dimensions, and compliance data, the gap is usually in the data layer, not the page design. Catsy validates attributes at the field level, then pushes them through live API connections so structured data stays accurate across every channel. [BOOK-A-DEMO-URL TO FILL] Or browse Catsy’s PIM and DAM resource center for more on preparing your catalog for agentic commerce.

FAQs

Because AI agents parse structured data, like JSON-LD and schema markup, rather than visual layout. A page can look complete to a human shopper and still be unreadable to an agent if its specs live in an image, a PDF, or plain marketing copy instead of a structured field.

They overlap significantly. Both depend on clean schema markup and accurate structured fields. The difference comes down to stakes: a search engine might still rank a page with weak structured data using other signals, but an AI shopping agent comparing products often can’t include a product in its results at all if the data isn’t there in a usable format.

Because the data B2B buyers and AI agents need most, exact technical specifications, compliance certifications, and compatibility details, is exactly the data most likely to live in an unstructured document rather than a labeled field. B2C product data tends to be simpler and easier to structure by comparison.

Both matter, but a static export that’s accurate once a week can’t support an agent that needs current pricing, stock status, or spec accuracy at the moment of the query. A live API connection is what keeps structured data trustworthy enough for an agent to act on.

No, though they’re related. Structuring and exposing product data is the foundational layer. Protocol-specific work, like configuring an MCP connection, assumes that foundational data is already clean, complete, and structured. Skipping straight to protocol integration without fixing the underlying data just exposes the same gaps through a new channel.

SHARE