AI News AI资讯 5h ago Updated 1h ago 更新于 1小时前 46

The Frontier AEO Tracker: What Astra Chooses (and every other frontier model, and what you can do about it) 前沿AEO追踪器:Astra的选择(以及所有其他前沿模型的选择,以及你可以做什么)

Latent Space Frontier built an AEO tracker analyzing 6 prompt variations across 7 frontier AI models covering 161 categories, revealing how AI agents recommend products and tools Strong self-bias observed: models tend to recommend competitors from the same ecosystem (Claude Code recommended by Claude, Codex by Sol/Astra, Cursor by Grok, Devin by SWE-1.7) 28 out of 161 categories showed universal dominance across all surveyed frontier models, while many others remain competitive battlegrounds New Latent Space发布AEO追踪器,测试7个前沿模型在161个类别中的推荐表现,发现模型存在明显"软偏见"(如Claude推荐Claude Code、GPT推荐Claude) Opus→Fable和Sol→Astra迭代显示Anthropic在数据/RL优先级上发生显著变化,Astra比Sol更"自信"且搜索来源更少(中位数5 vs 9) 28个类别存在全模型一致的首选推荐,其余多为"接近竞争"状态,构成AEO核心战场 验证了Ora和Vercel提出的AEO实践(如markdown content-negotiation)真实有效,失败会阻止模型读取内容 因API限制,Gemini/GL

62
Hot 热度
72
Quality 质量
65
Impact 影响力

Analysis 深度分析

TL;DR

  • Latent Space Frontier built an AEO tracker analyzing 6 prompt variations across 7 frontier AI models covering 161 categories, revealing how AI agents recommend products and tools
  • Strong self-bias observed: models tend to recommend competitors from the same ecosystem (Claude Code recommended by Claude, Codex by Sol/Astra, Cursor by Grok, Devin by SWE-1.7)
  • 28 out of 161 categories showed universal dominance across all surveyed frontier models, while many others remain competitive battlegrounds
  • Newer model generations (Astra, Fable) show increased confidence and efficiency — searching fewer sources and changing recommendations less when questions are paraphrased
  • AEO practices like markdown content-negotiation were validated as real and impactful; failures in these practices actively discourage models from reading content

Why It Matters

This research provides the first large-scale empirical look at how AI agents influence product discovery and recommendation — a critical factor as AI agents become primary interfaces for users. For AI practitioners and companies, understanding AEO dynamics is now essential for visibility in agent-driven search, and the observed model biases reveal both opportunities and risks in how products get recommended.

Technical Details

  • Methodology: Extended AmplifyingAI's framework to run 6 prompt variations across 7 frontier models (Claude Opus/Fable, Gemini Sol/Astra, Grok, Muse, SWE-1.7) covering 161 categories spanning coding agents, AI podcasts, sandboxes, managed databases, ASR models, angel investors, corporate spend, and payroll software
  • Scoring system: A proprietary AEO score weighting first choices, alternative choices, and mentions, with negative weights applied to mild and strong anti-recommendations
  • Data transparency: Every prompt and answer pair is publicly inspectable; contamination checks were performed and found negative
  • Source extraction: Top cited sources influencing agent recommendations were extracted, along with analysis of top recommendation failures
  • Limitations: Gemini/Antigravity, GLM/Zcode, and DeepSeek/DeepCode were excluded due to rate limits and errors in the first run

Industry Insight

  • AEO is emerging as a legitimate and high-value optimization discipline — as models become more confident and less random in their recommendations, the ROI of proper AEO practices (like markdown content-negotiation) will only increase
  • Companies should monitor "soft biases" where models favor same-ecosystem products and invest in AEO strategies that can break through these recommendation loops, especially in the 133 non-universal categories that remain competitive
  • The shift toward fewer source lookups and higher confidence in newer model generations means first-mover advantage in AEO will be critical — once a product establishes dominance in a category, it becomes increasingly difficult for competitors to displace it

TL;DR

  • Latent Space发布AEO追踪器,测试7个前沿模型在161个类别中的推荐表现,发现模型存在明显"软偏见"(如Claude推荐Claude Code、GPT推荐Claude)
  • Opus→Fable和Sol→Astra迭代显示Anthropic在数据/RL优先级上发生显著变化,Astra比Sol更"自信"且搜索来源更少(中位数5 vs 9)
  • 28个类别存在全模型一致的首选推荐,其余多为"接近竞争"状态,构成AEO核心战场
  • 验证了Ora和Vercel提出的AEO实践(如markdown content-negotiation)真实有效,失败会阻止模型读取内容
  • 因API限制,Gemini/GLM/DeepSeek等模型未能纳入首批分析

为什么值得看

本文为AI从业者提供了首个系统性的Agent Engine Optimization(AEO)基准测试框架,揭示了模型推荐行为中的偏见模式和迭代趋势。对希望优化产品被AI Agent推荐策略的创业公司和开发者具有直接参考价值。

技术解析

  • 测试架构:6种提示变体 × 7个模型 × 161个类别,覆盖编程Agent、AI播客、托管数据库、ASR模型甚至天使投资人等非常规类别
  • 评分体系: proprietary AEO Score,加权首选项、备选、提及,同时对轻度/强反推荐给予负分
  • 模型对比:Sol中位数搜索9个来源,Astra仅5个;Opus中位数11个,Fable达15个,显示Anthropic在迭代中调整了搜索策略
  • 偏见验证:各模型倾向推荐自家生态产品(Claude→Claude Code、Codex→Cursor、Muse→Muse Code、SWE-1.7→Devin),但GPT推荐Claude被视为"值得称赞的无偏见"
  • 可解释性:所有提示-答案对公开可查,同时提取影响推荐的关键来源和失败案例分析

行业启示

  • AEO正成为AI Agent时代的关键增长渠道,模型推荐行为的"软偏见"为产品定位提供明确策略方向
  • 模型迭代中搜索效率与推荐稳定性的权衡(如Astra的"自信"模式)将影响AEO价值,减少选择随机性可提升推荐转化率
  • 现有AEO实践(markdown content-negotiation等)已被验证有效,但DeepSeek/Gemini等主流模型缺失限制了行业基准的完整性,需推动API可访问性改善

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Agent Agent Evaluation 评测 Research 科学研究 Benchmark 基准测试