...
Shopify Product Data for AI Agents

AI Commerce Insights

How to Structure Your Shopify Product Data for AI Agents: The Complete Optimisation Guide

Shopify structures and syndicates product data for AI agents using Shopify Catalog, which automatically syncs real-time inventory, pricing, and attributes to external AI platforms.
AI CommerceAugust 30, 2026By NOIR & BLANCO

Your product data is invisible

Not because AI agents cannot see your Shopify store. They can. It is because the way your product information is structured does not match how those agents parse, interpret and recommend products.

When a shopper asks an assistant for “a lab grown diamond solitaire under one lakh, IGI certified, at least one carat”, the agent is not browsing your website the way a person does. It is looking for specific attributes: diamond origin, certification body, carat weight, metal, price ceiling.

If your product record does not state those attributes in a form a machine can read, your ring gets skipped. Even when it is exactly what the shopper wanted.

This is the quiet conversion problem most Shopify merchants have never audited. Your site may be beautiful. Your photography may be extraordinary. Your copy may be genuinely well written. None of that helps if the underlying data cannot answer a direct question.

Examples throughout are drawn from three categories we work in most: lab grown diamond jewellery, fashion, and lifestyle. Brand names are illustrative.

How AI agents actually evaluate a product

Step one, query parsing. The request is broken into attributes and constraints. “Lab grown diamond solitaire under one lakh, IGI certified, at least one carat” becomes category equals ring, origin equals lab grown, certification equals IGI, carat minimum equals 1.0, price maximum equals 100000.

Step two, data retrieval. The agent queries structured product data. It is reading fields, not admiring your copy.

Step three, attribute matching. Is the origin explicitly stated as lab grown? Is the certifying body named? Is carat weight a number it can compare against, or is it buried in a sentence?

Step four, ranking. Matching products are ranked on relevance, price, ratings and availability.

Step five, purchase facilitation. The agent needs exact price, available sizes, delivery estimate and returns terms.

Notice what plays no part in this. Lifestyle imagery. Brand film. The collection page you sequenced so carefully. Those matter enormously to human shoppers and contribute nothing to the machine evaluation.

Which fields carry the most weight

Tier one, critical. Without these an agent will not confidently recommend you at all.

  • Product title, descriptive and carrying key attributes
  • Product type and category, mapped to a standard taxonomy
  • Price, exact and current, with currency
  • Availability, distinguishing in stock, out of stock, preorder and made to order
  • Core attributes, the two or three a buyer would filter on

Tier two, important. These decide ranking between products that all match.

  • Product description, factual and structured
  • Metafields carrying category specific attributes
  • Reviews and ratings, exposed as structured data
  • Schema markup on the product template
  • Images with descriptive alt text

Tier three, supporting context.

  • Brand, SKU, barcode, GTIN
  • Dimensions and weight
  • Care instructions
  • Warranty and certification detail

The pattern we see in audits is a catalogue strong on tier three, because those feel like product page furniture, and weak on tier one and two, because those feel like admin.

Product titles: the highest weighted field you own

Most brands waste the title in one of two directions.

The poetic title. “The Aria.” “The Sunday Shirt.” “Object No. 4.” Beautiful on a collection page, meaningless to a machine. No category, no material, nothing matchable.

The keyword dump. “Lab Grown Diamond Ring Solitaire Engagement Women Gold Certified Real Diamond Wedding Gift Ring for Her Best Price.” Unreadable to humans and read as spam by the systems you are trying to impress.

The workable pattern is a documented formula applied consistently: brand, then primary category, then key attributes, then one differentiator.

Lab grown diamond jewellery

  • Solene Aria 1.5ct Lab Grown Diamond Solitaire Ring, 18k White Gold, IGI Certified
  • Solene Halo Lab Grown Diamond Stud Earrings, 1ct Total, 14k Rose Gold, VS Clarity
  • Solene Tennis Bracelet, 3ct Lab Grown Diamonds, 18k Yellow Gold, IGI Certified

Each answers the four questions a buyer in this category asks first: is it lab grown, how many carats, what metal, who certified it.

Fashion

  • Kora Relaxed Fit Cotton Poplin Shirt, 100 Percent Cotton, Unisex, Oversized
  • Kora High Rise Straight Leg Denim, Rigid Selvedge, 13oz, Ankle Length
  • Kora Handwoven Chanderi Silk Cotton Saree, Zari Border, 6.3 Metres

Lifestyle

  • Terra Stoneware Dinner Plate Set of Four, Handglazed, Microwave and Dishwasher Safe, 26cm
  • Terra Solid Sheesham Wood Side Table, 45cm, Natural Finish, Preassembled
  • Terra Pure Soy Wax Candle, 200g, 45 Hour Burn, Oud and Amber

The title checklist

  • Lead with brand where the brand carries recognition
  • Name the actual category in the words a buyer would use
  • Include the defining measurement: carat weight, GSM, dimensions, volume
  • Include material or the key attribute: 18k white gold, 100 percent linen, solid sheesham
  • Add one genuine differentiator, not an adjective
  • Keep under roughly one hundred characters, because longer titles get truncated
  • Never put variant information in the title. Colour, size and ring size belong in the variant
  • Remove marketing words. Premium, ultimate, exquisite and luxurious carry no information
  • Use industry standard terminology. Say solitaire, not statement stone

Consistency across the catalogue matters as much as any single title, because it lets a system infer structure from your naming.

Descriptions: structure for both readers

You do not have to choose between copy that moves a person and copy a machine can parse. You have to sequence them. Seven blocks.

1. What it is. One sentence, plain. “A 1.5 carat lab grown round brilliant solitaire set in 18k white gold, certified by IGI.”

2. Specifications. Bulleted, factual, with units.

  • Centre stone: 1.52ct lab grown round brilliant
  • Growth method: CVD
  • Colour grade: F
  • Clarity grade: VS1
  • Cut grade: Excellent
  • Certification: IGI, certificate supplied with order
  • Metal: 18k white gold, BIS hallmarked
  • Metal weight: 3.1g approximately
  • Setting: four prong
  • Band width: 1.9mm
  • Ring sizes: 6 to 22, Indian sizing

3. Primary use cases. Two or three sentences. “Designed as an engagement ring and worn daily. The low four prong setting sits close to the finger and clears most stacking bands, which makes it practical for everyday wear rather than occasion only.”

4. Features and benefits. Each feature tied to an outcome.

  • Four prong setting maximises light entry, so the stone reads brighter than a bezel at the same carat weight
  • Low profile reduces snagging on fabric during daily wear
  • Lab grown origin means an identical specification costs meaningfully less than a mined equivalent
  • BIS hallmarked gold, independently verified for purity

5. Who it is best for. One sentence. “Buyers who want a certified centre stone above one carat at a price a mined equivalent would not reach.”

6. What is included. Packaging, certificate, care kit, warranty card, resizing entitlement.

7. Care and terms. Cleaning, resizing window, warranty period, buyback or exchange terms.

Two rules govern all of it. Never put a fact only inside an image or video. And never let the description carry information that belongs in a structured field, because a description is text a system must interpret while a metafield is data it simply reads. The facts should appear in both.

One correction to the standard advice. Some guides say that if you must choose between optimising for humans or for agents, choose agents. That is wrong today. Your revenue this quarter comes from humans on your site. Structure so the facts are legible to machines and the opening and closing carry your voice. That serves both and costs nothing extra.

Metafields: where the real advantage sits

Standard Shopify fields cannot capture what makes your products distinguishable. Metafields can, and this is where most competitors will simply not do the work.

When an agent receives “lab grown solitaire, at least one carat, VS clarity or better, under one lakh”, metafields let it evaluate directly. Without them it parses prose and guesses, and guessing produces skipping.

One correction worth making: metafields are often described as invisible to shoppers. They should not be. Render them into the product page as a specification block. Humans want them too.

Lab grown diamond jewellery

  • diamond_origin, restricted to lab grown or natural
  • growth_method, CVD or HPHT
  • carat_total_weight, number
  • centre_stone_carat, number
  • stone_shape, controlled list covering round brilliant, oval, emerald, princess, pear, cushion, marquise
  • colour_grade, controlled list D through J
  • clarity_grade, controlled list VVS1 through SI2
  • cut_grade, controlled list
  • certification_body, IGI, GIA or SGL
  • certificate_supplied, boolean
  • metal_purity, 14k, 18k, platinum, sterling silver
  • metal_colour, yellow, white, rose
  • metal_weight_grams, number
  • bis_hallmarked, boolean
  • setting_type, prong, bezel, halo, pave, channel
  • number_of_stones, number
  • band_width_mm, number
  • available_ring_sizes, list
  • resizing_available, boolean and resizing_range
  • nickel_free, boolean
  • made_to_order, boolean and lead_time_days
  • warranty_months, number
  • buyback_or_exchange_policy, metaobject reference

In this category, certification and origin are not optional fields. They are the two facts that decide whether a shopper trusts the listing at all, and the two most frequently left implicit.

Fashion

  • fabric_composition, structured with percentages
  • fabric_weight_gsm, number
  • weave_or_knit, controlled list
  • fit_type, slim, regular, relaxed, oversized
  • garment_measurements_by_size, metaobject reference
  • model_height_cm and model_size_worn
  • sleeve_length, neckline, closure_type, lining, pockets
  • opacity, opaque, semi sheer, sheer
  • stretch, none, slight, moderate, high
  • care_instructions, metaobject reference
  • wash_type, machine, hand, dry clean only
  • shrinkage_expected, boolean
  • occasion and season, controlled lists
  • country_of_manufacture
  • handloom_or_handcrafted, boolean
  • certifications, GOTS, OEKO TEX, handloom mark

Sizing is decisive here. Indian apparel sizing is inconsistent across brands and shoppers know it. Garment measurements by size, as text in centimetres and inches, removes a real purchase barrier for humans and gives an agent something to reason with.

Lifestyle and home

  • material and secondary_material, controlled lists
  • dimensions_cm, structured length, width, height
  • weight_grams, number
  • capacity_ml or capacity_litres, number
  • set_quantity, number
  • microwave_safe, dishwasher_safe, oven_safe, food_safe, all boolean
  • oven_safe_max_temp_c, number
  • assembly_required, boolean and assembly_time_minutes
  • handmade, boolean
  • craft_technique, text
  • origin_cluster, for example Channapatna, Moradabad, Khurja
  • burn_time_hours and fragrance_notes, for candles
  • care_instructions, metaobject reference
  • warranty_months, number

For handcrafted goods the origin and technique fields do double duty. They are exactly the provenance an agent needs to justify a recommendation, and exactly the story a human buyer wants.

Getting metafields right

Use typed fields. Weight as a weight type. Dimensions as a dimension type. Carat as a number. Dishwasher safe as a boolean. Typed data can be filtered and compared. “approximately 3.1 grams” in a text field cannot.

Use controlled vocabularies. For any field with a fixed set of values, define the list and restrict entry to it. Three people typing clarity grades freely produces three formats and one unusable field.

Standardise units once. Centimetres or inches, grams or carats. Decide and never mix.

Complete beats broad. Twenty fields populated on every SKU beats sixty populated on some. Partial data teaches every system reading you that a blank means nothing rather than no.

Creating them in Shopify

Settings, then Custom Data, then Products. Create a namespace such as product_specs. Define each metafield with a clear name, the correct type, and a description explaining what belongs in it.

For population, work outside the admin. Export, edit in a spreadsheet against a written specification, reimport with the bulk editor or a migration tool. Editing two thousand products one at a time is how these projects die.

Then make sure the data reaches the outside world. Render metafields into the product template as visible text and include them in your schema markup. A metafield that lives only in the admin does nothing.

Schema markup: the output layer

{
“@context”: “https://schema.org/”,
“@type”: “Product”,
“name”: “Solene Aria 1.5ct Lab Grown Diamond Solitaire Ring, 18k White Gold”,
“image”: “https://example.com/aria_solitaire_18k_white_gold.jpg”,
“description”: “A 1.52 carat lab grown round brilliant solitaire in 18k white gold, IGI certified.”,
“brand”: { “@type”: “Brand”, “name”: “Solene” },
“sku”: “SOL_ARIA_150_18KW”,
“material”: “18k white gold”,
“additionalProperty”: [
{ “@type”: “PropertyValue”, “name”: “Diamond origin”, “value”: “Lab grown” },
{ “@type”: “PropertyValue”, “name”: “Carat weight”, “value”: “1.52” },
{ “@type”: “PropertyValue”, “name”: “Colour grade”, “value”: “F” },
{ “@type”: “PropertyValue”, “name”: “Clarity grade”, “value”: “VS1” },
{ “@type”: “PropertyValue”, “name”: “Certification”, “value”: “IGI” }
],
“offers”: {
“@type”: “Offer”,
“url”: “https://example.com/products/aria-solitaire”,
“priceCurrency”: “INR”,
“price”: “94500”,
“availability”: “https://schema.org/InStock”,
“itemCondition”: “https://schema.org/NewCondition”
},
“aggregateRating”: {
“@type”: “AggregateRating”,
“ratingValue”: “4.8”,
“reviewCount”: “213”
}
}

 

Most modern Shopify themes output basic product schema already. The question is whether it is complete and whether it survived your theme customisations. Validate with the Rich Results Test and the Schema validator after every significant theme change, because partial or broken markup is extremely common and nothing on the page looks wrong when it fails.

If you would rather not touch Liquid, several schema apps provide an interface. Whichever route you take, check the rendered page source rather than trusting the app’s description of itself.

Images and alt text

Weak alt text. “ring”, “product image”, “jewellery”

Strong alt text. “Solene Aria 1.5 carat lab grown diamond solitaire ring in 18k white gold, three quarter view showing four prong setting and 1.9mm band”

Fashion. “Kora relaxed fit cotton poplin shirt in ecru, front view on model, showing dropped shoulder and mother of pearl buttons”

Lifestyle. “Terra handglazed stoneware dinner plate in slate blue, overhead view, 26cm diameter with visible glaze variation”

File names count too. Use aria_solitaire_18k_white_gold_front.jpg rather than IMG_4471.jpg.

Include the shots that answer questions: scale reference, detail on the setting or the stitching or the glaze, the reverse, and the product in use. In jewellery specifically, a hand shot at true scale prevents a large share of returns and presale queries.

The implementation roadmap

Phase one, audit and specification, weeks one and two. Export the catalogue and score completeness field by field. Then write the specification: field names, types, units, permitted values, owners. Do this before anyone populates anything. This document is what stops the project decaying in six months.

Phase two, core data, weeks two and three. Rewrite titles against the formula. Restructure descriptions into the seven blocks. Fix alt text and file names. Work through highest revenue products first and finish them completely rather than half finishing everything.

Phase three, metafields, weeks three and four. Create the namespace and definitions. Populate the priority range through bulk import. Render the fields into the product template.

Phase four, markup and testing, week four. Validate schema, then test against real assistants. Fix what fails.

Phase five, ongoing. Gate publishing so no new product goes live without required fields. Score completeness monthly. Retest with assistants monthly. Update availability and price promptly, because a confidently wrong record produces confidently wrong recommendations that arrive back as returns.

Testing whether it worked

Validators tell you the markup is well formed. They do not tell you whether an agent can answer a customer’s question. Take twenty real presale questions from your support inbox and ask them of the major assistants, both with and without your brand name.

Jewellery. “Find me a one carat lab grown diamond solitaire under one lakh with IGI certification.” “Is the Solene Aria resizable and what sizes does it come in?”

Fashion. “I am 5 foot 9 with a 40 inch chest, which size should I take in the Kora poplin shirt?” “Find me an oversized cotton shirt under three thousand rupees that is not sheer.”

Lifestyle. “Are these stoneware plates microwave safe?” “Find me a handmade dinner set for four under five thousand rupees.”

Check three things. Does your brand appear at all. Is the answer correct. Where did it come from.

The failure modes are diagnostic. A wrong price points to caching or markup. A wrong sizing answer means the size data is still in an image. Absence on unbranded queries means the record is too thin to be selected. An answer sourced from a marketplace listing means someone else is currently speaking for your brand.

Common mistakes

Vague titles. “The Aria” communicates nothing matchable.

Marketing led descriptions. Emotional benefit in place of specification.

Missing or inconsistent metafields. Half a catalogue populated is close to useless.

Facts locked in images and PDFs. Size charts, specification grids, certification details and care instructions delivered as graphics are invisible. The single most common and most costly error.

No validated schema. Or schema that broke during a theme edit eighteen months ago and nobody noticed.

Generic or keyword stuffed alt text.

Inconsistent option naming. Colour, Color and Shade across one catalogue destroys comparability.

Duplicate products for colourways. Splits reviews, splits inventory logic.

Stale availability. In made to order categories, state lead time as a number of days.

Supplier descriptions copied verbatim. Identical to forty other listings, and no reason to select you.

The India layer

Certification and hallmarking are trust infrastructure. For lab grown diamond jewellery, IGI or SGL certification and BIS hallmarking are what make a listing credible to both a shopper and an agent. Put them in structured fields, not in a trust badge image.

Sizing conventions need translating. Indian ring sizing and apparel sizing both vary. State the standard you use and give the conversion.

Delivery and COD terms need numbers. Serviceability varies enormously by pincode. Whether COD is available, on which order values and in which pincodes, is a top three question for Indian shoppers and is frequently undocumented. If your checkout runs through a layer such as GoKwik, confirm what data that layer exposes.

Our view

Nothing on this list requires a bet on agentic commerce.

Complete typed attributes, a standard taxonomy mapping, consistent titles, text based specifications, proper alt text and valid schema markup each improve your organic search, your Shopping and Performance Max feed quality, your on site filtering, your presale support volume and your product page conversion rate. Those returns arrive on the traffic you already have.

The agentic benefit is a further return on the same work. When an assistant compares three lab grown solitaires and can answer every question about one of them, that is the one it recommends. Not because the brand is larger, but because the record was finished.

Most merchants will not finish theirs. That is the opportunity.

Frequently Asked Questions: Shopify Product Data for AI Agents

What does it cost to optimise product data for AI agents?
Mostly time rather than money if you do it internally. The cost drivers are catalogue size and how much information currently sits in images or in people’s heads rather than in fields. A hundred SKUs is a few focused weeks. Several thousand needs a phased plan and probably contracted data entry. Be sceptical of anyone quoting a guaranteed conversion uplift figure, because the returns show up across organic search, feed quality, support volume and conversion at the same time, which makes clean attribution difficult.

Should I rewrite descriptions for humans or for agents?
Both, in the same description. Open with two sentences of positioning in your voice, carry the facts in the middle in plain declarative sentences, close with terms and care. Advice that says to prioritise agents over humans is wrong for now, because your revenue this quarter comes from people on your site.

What is the difference between metafields and schema markup?
Metafields store structured attributes inside Shopify. Schema markup exposes information on the rendered page in a standard format external systems read. Metafields are the database, schema is the publication. A metafield that is never rendered or marked up does nothing outside your admin.

Do metafields help human shoppers?
Yes, if you render them. Output them as a specification block. Shoppers in considered categories such as jewellery and furniture actively want that detail, and it reduces presale queries.

How often should product data be updated?
Price and availability immediately. New products correct from day one, enforced as a publishing rule. Everything else, a quarterly completeness review plus a monthly assistant test to catch drift.

Can I use AI to help populate this?
For rewriting titles and descriptions to a specification, yes, and it saves real time. For generating attribute values, be careful. An assistant will confidently invent a clarity grade or a fabric composition it cannot know. Facts must come from your own specifications or your supplier, and anything generated must be verified before publishing.

What if I have thousands of SKUs?
Start with the products driving most of your revenue and finish them completely, then work outward. A catalogue fully correct on two hundred SKUs beats one half correct on two thousand. Use export, spreadsheet and reimport rather than the admin.

How do I know whether it is working?
Segment traffic referred by assistants as best you can, and run the twenty question assistant test monthly. The products that never surface are telling you their records are still too thin.

Should I optimise differently for different assistants?
No. Structure to universal standards and test across all of them. Testing across several surfaces gaps that one alone would not, but the underlying work is the same.

What if my products have unusual attributes with no standard field?
Create custom metafields. A handloom saree with a named weave cluster, a candle with a documented burn time, a ring with a specific growth method all need fields of their own. Your genuinely distinctive attributes are exactly the ones that most need documenting.

NOIR & BLANCO builds and scales Shopify storefronts for premium D2C brands. If you want an audit of how complete your catalogue data actually is, [get in touch].

Privacy Preference Center