How do you get an AI model to recommend your product?
Not by rewriting your homepage. The lever is what other pages already say about you, and there's finally a way to measure it.
An AI model recommends what it already associates with your category, and that association comes mostly from what other pages say about you, not from your own site's copy. Latent Space ran six prompt variations across seven frontier models over 161 product categories on September 7, 2026, and found every model agreeing on one clear winner in just 28 of them. The other 133 are still open. Start by running your own category through several models and logging exactly what comes back.
What did the Latent Space AEO tracker actually measure?
Latent Space's AEO tracker, published September 7, 2026, ran six differently-phrased prompts through seven frontier models across 161 product and tool categories, from coding agents to voice platforms. Each answer was scored with weights on first-choice picks, alternative mentions, and plain mentions, with negative weights when a model actively warned against a product. The headline number: only 28 of the 161 categories produced a single primary choice every model agreed on. The other 133, more than 80% of the categories tracked, are contested ground where different models pick different products for the same question.
Why does one model recommend Claude Code and another recommend Cursor for the same question?
Because each model favors the tool nearest to its own lab first. The tracker found Claude Code dominant in Fable and Opus's answers, Grok favoring Cursor, and SWE-1.7 leading with Devin. That split is not random noise, it's the same mechanism that decides everything else a model recommends: an answer reflects whichever name the model already associates most strongly with a category, and a lab's own product sits closer to that association than a stranger's does. House bias is just the most visible case of a rule that applies everywhere else too.
So what actually decides who wins?
About 60% of ChatGPT-style answers get generated purely from the model's training memory, with no live web lookup at all (Digital Bloom, 2025 AI Citation & LLM Visibility Report). For most questions, a model either already has an opinion of you baked in from training or it doesn't, and there's no page you can publish today that edits data the model was already trained on.
Which means the input that actually shapes that opinion is everything other people wrote about you before the training cutoff, plus whatever a live retrieval pass turns up after it. Muck Rack put a number on the split: 84 to 85% of AI citations point to earned, third-party sources rather than brand-owned websites (Muck Rack, 2025; a separate 5W Releases analysis distributed via PR Newswire puts the same figure at 85.5%). A domain mentioned across four or more independent platforms is about 2.8 times more likely to show up in a ChatGPT answer than one mentioned on fewer, and ChatGPT names a brand about 3.2 times more often than it actually links to it, because naming reflects association, not clickable citation (Digital Bloom, 2025).
A correlation study backs the same read from a different angle. Brand mentions correlate with AI-search visibility at about 3 times the strength of raw backlinks (Search Engine Land, citing an Ahrefs analysis of 75,000 brands, 2025), and a separate keyword-level breakdown put the actual coefficients at 0.664 for mentions versus 0.218 for backlinks (Keyword.com, 2025). Backlinks still help you rank on Google. They are a weaker signal for whether a model recommends you.
Put together, the mechanism is this: a model's opinion of you is built out of how often, how consistently, and how credibly other people and other pages already talk about you. It is not built out of anything your own page asserts about itself.
A founder named Cheonkyu hit exactly this gap in April 2026, before there was a tracker to name it. He typed his own category into ChatGPT and asked for the best tool for it. "It recommended 5 competitors. Mine wasn't on the list," he wrote on Indie Hackers on April 16, 2026. His product scored 12 out of 100 on a visibility tracker he later built. Over the next two weeks he wrote 13 blog posts, posted in one Reddit thread, and submitted to six directories, all aimed at getting mentioned in more places rather than rewriting his own site. The score moved to 32 out of 100, and Perplexity started recognizing the product in 8 of 10 probes. One product, small sample, but it's the mechanism showing up in real life: mentions moved the number, copy never touched it.
Does the source-count spread (5 vs. 15) actually change your strategy?
The tracker logged a real spread in how many sources each model consults per answer: a median of 5 for Astra, 9 for Sol, 11 for Opus, and 15 for Fable. Latent Space doesn't explain why, and we don't have a controlled test proving the causal link, so take this next part as our own read, not a cited finding: a model pulling a median of 15 sources per answer is casting a wide net, and winning its recommendation likely rewards being mentioned across many independent pages rather than owning one authoritative one. A model settling on 5 is leaning harder on whatever it already trusts most, probably training memory or a small set of high-authority sources, which rewards being the canonical answer somewhere specific over being scattered thinly everywhere.
The practical version: if you can only do one thing, getting into 10 mediocre-but-real mentions probably moves a high-source-count model faster than it moves a low-source-count one, and the reverse may hold for a single strong placement on a page a low-source-count model already trusts. We haven't tested this split ourselves. Treat it as a hypothesis worth checking against your own logs, not a rule to build a whole strategy on.
The test to run before you spend an hour on anything else
- Write the question the way a buyer actually asks it, not your product name. "Best tool for X" beats "tell me about [product]."
- Run 5-6 phrasings of that question across at least 3-4 models from different labs, not six variations on one model. The tracker's own house-bias finding means one model tells you almost nothing about another. Save the raw text of every answer, not a summary.
- Score each answer the way the tracker did: first choice, alternative, mention, absent, or actively warned against. Note how many sources the model claims to have used, if it says.
- If you're absent, ask the model to name its sources, or what it based the answer on. That list is your actual target-mention list, not a wish list you made up beforehand.
- Check whether those sources, or the obvious equivalents, comparison roundups, the relevant subreddit, other people's blogs in your category, already mention you. If they don't, that's the real gap. It is very rarely "our homepage needs better copy."
- Close the gap by getting mentioned, not by rewriting your own page. Get included in an existing comparison page, answer real questions in the relevant community with your name and a real number attached, pitch your data to whoever already writes the roundup.
- Re-run the same prompts monthly. 133 of 161 categories in the tracker had no agreed winner, which means most categories move. Track the delta over time, not one snapshot.
What doesn't work
- Rewriting your homepage copy to sound more like "the best." Close to 60% of answers never touch your page at all (Digital Bloom, 2025). A sharper sentence on a page the model isn't reading changes nothing.
- Adding an llms.txt file. Google has said directly it doesn't use it for search, and nothing in the tracker's data points to models weighting it for recommendation either. It's an unconfirmed hypothesis, not a lever.
- Schema markup as the fix. Basic Article and FAQ schema is fine hygiene, but there's no proven correlation between schema and getting recommended. Machine-readable is not the same thing as machine-recommended.
- One prompt, once, on one model. Models change their answer under light paraphrasing, which is exactly why the tracker ran six phrasings per model instead of one. A single ChatGPT query that skips you tells you almost nothing about whether you're actually absent from the category.
- Serving markdown to crawlers and calling it done. It matters for whether an agent can read your docs at all, though only 3 of 7 coding agents even request markdown content negotiation as of February 2026 (Checkly, 2026). That's a crawlability fix, not a recommendation one. A page a bot can parse perfectly and that nobody else ever mentions still doesn't get recommended.
- Buying a visibility-tracking subscription and stopping there. A dashboard that tells you your score is 12 out of 100 hasn't moved it. The work is still getting mentioned somewhere new.
Where this could be wrong
Three honest limits. First, the tracker excluded Gemini, GLM, and DeepSeek entirely for errors and rate limits, so whatever governs recommendation on three major model families sits outside this data, and a strategy tuned only to the models the tracker covered may miss how a large slice of real queries actually get answered. Second, one model was measurably more stable under paraphrasing than the rest, which is good news if you're optimizing for that model and a caution everywhere else: a model that flips its answer on a slightly different phrasing can flip back just as fast, so a win today is not a lock. Third, the source-count-spread read two sections up is our own inference, not a cited finding, and it's the single piece of this page most likely to be wrong.
Related: Is an AI search visibility tool worth building in 2026?, on the funded dashboard layer that measures this same gap without fixing it. Can you trust AI benchmark leaderboards in 2026? for the same lesson applied to model claims instead of product recommendations: verify the number, don't take the surface reading. And Best AI model for building a startup in 2026 if the model doing the recommending is also the one you're deciding whether to build on.
Frequently asked questions
How do you get an AI model to recommend your product?
Mostly by getting mentioned by other people and other pages, not by editing your own site. Roughly 84 to 85% of AI citations trace back to earned, third-party sources rather than brand-owned websites (Muck Rack, 2025), so the fastest lever is getting into the comparison pages, forum threads, and roundups a model already pulls from. Find those by running your own category through several models first and logging what it names.
What did Latent Space's AEO tracker actually find?
Running six prompt variations across seven models over 161 product categories on September 7, 2026, it found a single agreed-on primary choice in only 28 categories. The other 133 are contested, with different models favoring different products for the identical question, and each lab shows a clear bias toward its own family's tools: Claude Code for Fable and Opus, Cursor for Grok, Devin for SWE-1.7.
Why do AI models lean on third-party mentions instead of a company's own website?
Because about 60% of ChatGPT-style answers are generated straight from training memory with no live web lookup at all (Digital Bloom, 2025), so a model's opinion of you is mostly set by what it absorbed from other sources during training, not by what your site says today. When a model does look something up live, it still weighs independent mentions heavily: a domain mentioned across 4+ platforms is about 2.8x more likely to appear in an answer (Digital Bloom, 2025).
Does adding an llms.txt file or schema markup get you recommended?
Not on its own. Google has confirmed it doesn't use llms.txt for search, and there's no proven correlation between schema markup and getting cited or recommended. Both are hygiene at best, helping a bot parse a page it already decided to read. Neither changes whether a model associates you with the category in the first place.
Why were Gemini, GLM, and DeepSeek left out of the AEO tracker?
Latent Space excluded them for errors and rate limits during the run, not for a substantive finding about their recommendation behavior. That's a real gap: the tracker's picture of who gets recommended covers Astra, Sol, Opus, Fable and a couple of others, not the full model landscape people actually query.
How many sources does an AI model check before recommending something?
It varies a lot by model. The tracker logged a median of 5 sources for Astra, 9 for Sol, 11 for Opus, and 15 for Fable per answer, a real 3x spread across models answering the same 161 categories.
What's the fastest way to test whether a model already recommends you?
Ask the exact question a buyer would ask, phrased 5-6 different ways, across at least 3-4 models from different labs, and log the raw text of every answer. One query on one model tells you almost nothing, since some models change their pick under light paraphrasing.
Is a single ChatGPT query a reliable way to check if you're recommended?
No. The tracker used six prompt variations specifically because a single phrasing isn't stable across models, and it found one model notably more resistant to flipping its answer under paraphrasing than the rest, implying the others are less stable by comparison. Treat one query as a data point, not a verdict.
The free pack: 100 AI ideas actually worth building, each with the receipts and a clear verdict. No fake MRR screenshots.