Outrank Competitors With an AI Visibility Benchmark

Winning in AI answers starts with knowing where competing brands are getting cited, mentioned, and reinforced across the web. This guide shows how to benchmark AI visibility, spot the gaps that matter most, and prioritize actions that can improve your competitive position.

Seonix team·September 29, 2026·22 min read
outrank competitors - Marketing team reviewing competitor visibility data on a dashboard

To outrank competitors in AI answers, you need a benchmark that measures who gets cited, mentioned, and associated with buyer-relevant topics across search and AI engines. A brand can lose the shortlist before a visitor ever reaches its site. ChatGPT, Google AI Overviews, Perplexity, Gemini, Copilot, and Claude often shape the first round of comparison. That shift makes AI visibility tracking for businesses a practical revenue issue, not a vanity report.

Buyer research changed fast once answer engines became a normal starting point. A prospect can ask five comparison prompts in two minutes, see three recurring brands, and never click beyond the answer block. In national markets, that pattern matters because one strong competitor can dominate broad category prompts, local-intent prompts, and buyer-stage prompts at the same time. AI Mode and answer engines compress the shortlist earlier than classic organic search alone.

This article shows how to benchmark AI Share of Voice, how to measure citation and mention gaps, and how to decide what to fix first. You will see a sampling method, a prioritization framework, an ethical authority map for Reddit, Quora, YouTube, TikTok, and media sources, plus a 90-day rollout plan. You will also learn what proof a service provider should show before claiming it can help you outrank competitors.

Why do competitor AI mentions change buyer shortlists before a website visit?

Competitor AI mentions change shortlists early because answer engines reduce the number of brands a buyer reviews. If one brand appears in 6 of 10 comparison prompts and your brand appears in 1 of 10, the buyer often treats that repeated brand as safer before checking any website. AI brand mentions versus traditional brand monitoring is the key difference here: classic monitoring tracks where your name appears, while AI benchmarking tracks whether engines surface your brand during decision-making.

Traditional brand monitoring usually counts press mentions, social mentions, or backlink alerts. AI visibility benchmarking goes further and asks four harder questions. Which brands are named in answers? Which sources are cited to support them? Which topics does the engine connect to each brand? Which brands get reinforced across multiple engines and prompt types? Those four checks tell you whether a buyer sees your brand as established in the category and whether you can outrank competitors in buyer research.

Shortlist pressure starts before the click

A simple example shows the gap. Imagine a B2B software firm appears on page one for several commercial terms, yet ChatGPT and Perplexity keep naming three other vendors in prompts like “best tools for mid-sized teams” and “top options for compliance reporting.” Search traffic may still look decent; however, assisted conversions soften because buyers arrive with another shortlist. The website did not fail first. The off-site proof and extractable content did.

Tip: If the same two or three brands recur across engines, treat that pattern as a shortlist signal, not a random mention. Repetition across 20 to 30 prompts usually matters more than one prominent answer on a single day.

Decision visibility matters more than awareness alone

Buyer-readiness signals make the effect stronger. If a competitor appears with pricing context, implementation detail, third-party validation, and category-fit language, the answer feels safer. A brand that only appears in broad awareness prompts still loses commercial intent prompts. That is why the benchmark must separate awareness visibility from decision visibility instead of rolling both into one score. If you want to outrank competitors, you need proof in the prompts that shape buying decisions.

How to Benchmark Your AI Visibility Against Competitors

Benchmarking AI visibility means measuring answer presence, citation support, topic association, and buyer-stage coverage across a defined prompt set. The most useful benchmark compares your brand with three to five real competitors over at least 30 prompts per engine. That sample size is large enough to reveal repeat patterns and small enough to run monthly without wasting time.

A marketer reviewing a visibility dashboard on a laptop with charts and comparison metrics.

Start with prompt-level visibility tracking across AI engines, not broad impressions. Use prompts that mirror the buying path: category discovery, comparisons, alternatives, use-case questions, objections, pricing, implementation, integrations, and trust checks. Then run each prompt in ChatGPT, Google AI Overviews, Perplexity, Gemini, Copilot, and, if relevant, Claude. Record whether your brand appears, where it appears, whether the answer cites sources, and what themes the answer attaches to your brand. This is one of the clearest ways to outrank competitors with evidence instead of guesswork.

AI visibility benchmark areas to compare against competitors
Benchmark AreaWhat to MeasureWhy It MattersPriority if WeakRecommended Next Action
CitationsCited sources per promptSupports answer trustHighPublish source-worthy pages
Unlinked mentionsBrand named without citationShows entity recognitionMediumGrow corroborating mentions
Topic coverageAssociated buyer topicsShapes category fitHighBuild answer-first content
CorroborationThird-party support countReduces trust gapHighEarn media and expert references
Channel presenceForum, video, social, media spreadImproves discovery breadthMediumMap ethical off-site channels

Use a repeatable sampling method

A useful benchmark needs a stable method. Split prompts into five groups of six prompts each: informational, commercial, comparison, problem-solution, and buyer-objection. Weight the groups by search demand and revenue value, not by volume alone. For example, six implementation prompts may drive more pipeline than six broad educational prompts, so they should carry more weight in your reporting.

Record the date, engine, prompt, answer position, brand mentions, citations, linked sources, and brand sentiment. Additionally, add one more field: buyer stage. That field shows whether your visibility is concentrated at the top of the funnel or spread into high-intent prompts. A brand that appears in 40% of awareness prompts but only 5% of comparison prompts does not have a strong competitive position. Therefore, it will struggle to outrank competitors where revenue starts.

Good to know: A 30-prompt benchmark across six engines creates 180 answer observations per reporting cycle. That gives you enough data to spot direction without pretending AI outputs are perfectly fixed.

Calculate AI Share of Voice the right way

AI Share of Voice is the percentage of tracked prompts where a brand appears in the answer set. Keep the formula simple: brand appearances divided by total prompts run within an engine or prompt group. If your brand appears in 18 of 60 prompts in ChatGPT, your ChatGPT AI Share of Voice is 30%. If a competitor appears in 33 of 60, the gap is 15 percentage points.

Do not stop there. Split AI Share of Voice into named mentions, cited mentions, and buyer-ready mentions. Named mentions tell you entity recall. Cited mentions show support. Buyer-ready mentions show whether the answer presents your brand with commercial context such as use cases, implementation, or proof. That split avoids a common mistake where teams celebrate mentions that never influence a buying decision. It also shows what must change if you plan to outrank competitors in commercial prompts.

A reliable dashboard usually includes at least these eight metrics:

  • AI Share of Voice by engine
  • AI Share of Voice by prompt group
  • Brand citation rate
  • Unlinked mention rate
  • Associated topic count
  • Third-party corroboration count
  • Buyer-ready answer rate
  • Trend change since last cycle

You can see a broader explanation of how to track visibility with repeatable guardrails if you want the reporting layer in more detail.

Check engine-specific differences

Each engine behaves differently, so an average score can hide a weak channel. Google AI Overviews often draws from strong web pages and search-visible sources. Perplexity tends to foreground citations more clearly. ChatGPT visibility can be shaped by widely corroborated web entities, structured pages, and recurring mentions across trusted sources. Copilot often mirrors broader search evidence, while Gemini can reward well-structured explanatory content.

Suppose your brand appears in 35% of Perplexity prompts, 28% of Google AI Overviews prompts, and 8% of ChatGPT prompts. That pattern suggests you may already have publishable source material, but weaker entity reinforcement across the web. The fix is not one generic “do more content” task. Instead, add corroboration, strengthen answer-first pages, and improve how your brand is discussed beyond your own domain. Those are the practical moves that help outrank competitors.

How do you measure whether competitors appear more often in AI answers than your brand?

You measure competitor visibility by running the same prompt set across the same engines on the same day, then comparing appearance rate, citation rate, and buyer-stage coverage. The cleanest comparison uses a controlled sample of at least three competitor brands and one primary brand. That method shows whether competitors appear more often, and it also shows why.

Build a competitor panel that reflects actual buyer alternatives. Include direct category rivals, one adjacent substitute, and one high-authority brand that often gets recommended even if it is not your closest operational match. That mix reveals whether you lose on relevance, authority, or both. If you benchmark against the wrong set, the report becomes flattering but useless.

Set the comparison frame first

The strongest measurement workflow looks like this:

  1. Define 30 to 50 prompts across awareness, comparison, use case, objection, and purchase intent.
  2. Run the full prompt set in each engine within a short time window, ideally 24 to 48 hours.
  3. Log every named brand, source citation, associated topic, and commercial cue.
  4. Score each answer for buyer readiness on a simple 0 to 2 scale.
  5. Repeat monthly, then compare both share and direction.

A 0 to 2 buyer-readiness score is practical. Score 0 if the brand is absent, 1 if the brand is named without proof or fit detail, and 2 if the brand is named with supporting context that could move a shortlist. That extra layer helps you avoid false positives and focus on what will outrank competitors in real buying journeys.

Rule of thumb: if a competitor leads your brand by 10 percentage points or more in high-intent prompts, treat that gap as a pipeline issue first and a visibility issue second.

Use examples and proof, not broad claims

Consider a worked example. Your team tracks 36 prompts across ChatGPT, Google AI Overviews, and Perplexity, for 108 total observations. Your brand appears in 27 answers, with 9 cited mentions and 6 buyer-ready mentions. Competitor A appears in 49 answers, with 22 cited mentions and 18 buyer-ready mentions. Competitor B appears in 34 answers, with 11 cited mentions and 10 buyer-ready mentions. The largest gap is not total appearances alone. Rather, Competitor A turns almost half of its appearances into buyer-ready answers.

That distinction matters during procurement. A service provider that promises to help you outrank competitors should be able to show engine-by-engine screenshots, prompt lists, scoring logic, and before-and-after deltas. Claims without prompt evidence are weak because AI answer outputs vary, and broad rank claims often hide cherry-picked prompts.

How to find citation and mention gaps your competitors already own

Citation and mention gaps show where competing brands have proof that answer engines can reuse. A citation gap exists when sources are cited for a competitor but not for you. A mention gap exists when engines recognize a competitor as relevant even without linking to a source. Both gaps matter, yet citation gaps usually deserve higher priority because they are easier to trace to a concrete page, source type, or missing topic.

Start by sorting every competitor appearance into three buckets: cited first-party source, cited third-party source, and unlinked mention. First-party citations come from the brand’s own site. Third-party citations come from media, forums, reviews, research summaries, or expert commentary. Unlinked mentions suggest the engine associates the brand with the topic even when no visible citation appears. That can happen when entity recognition is strong across the web.

Next, compare source types. If a competitor keeps getting cited from documentation pages, use-case explainers, glossary pages, and integration pages, the pattern points to a content structure advantage. If a competitor appears through roundups, trade press, podcasts, and forum threads, the pattern points to broader corroboration. In practice, most brands need both if they want to outrank competitors.

You can explore the wider idea of getting your brand mentioned in AI answers and proving it if you want the broader visibility playbook. For benchmarking, stay focused on comparative gaps.

Which gaps matter most for higher rankings against competitors?

The gap that matters most depends on prompt type, but there is a clear order for most commercial categories. Citation gaps usually come first, because they tie directly to sourceable proof. Associated topic gaps come next, because a brand cannot win prompts for topics it is never linked to. Third-party corroboration gaps follow closely, especially in expert or high-consideration categories. Unlinked mention gaps matter too, but they often improve after the other three strengthen.

A useful priority order looks like this:

  1. Citation gaps on high-intent prompts
  2. Associated topic gaps on category and use-case prompts
  3. Third-party corroboration gaps across independent sources
  4. Channel presence gaps on forums, video, social, and media
  5. Unlinked mention gaps across lower-intent prompts

Imagine your brand sells a service with long evaluation cycles. Competitors appear in “best options,” “alternatives,” “pricing,” and “implementation” prompts with cited comparison pages and media references. Your brand appears only in broad “what is” prompts. The urgent issue is not awareness. Instead, the urgent issue is missing proof in buyer-stage prompts if your goal is to outrank competitors.

Watch out: Teams often chase unlinked mentions because they feel broad and exciting. In most commercial benchmarks, one missing cited source on a high-intent prompt matters more than ten weak mentions on low-intent prompts.

Associated topics deserve careful review. If answer engines connect a competitor with “enterprise security,” “fast onboarding,” or “best for agencies,” and they never connect your brand with those topics, your content and off-site proof are too thin in those areas. The fix may include new pages, stronger schema markup, better on-page extraction structure, and more external reinforcement from relevant sources. As a result, you get a clearer path to outrank competitors by topic, not just by brand mention count.

How to document proof with a prompt pack

A prompt pack is a dated record of the exact queries, engines, screenshots, and scoring notes used in the benchmark. It sounds simple, but it changes decision quality. Without a prompt pack, teams argue from memory. With one, you can compare like for like over 30, 60, and 90 days.

A solid proof pack includes the prompt text, engine used, date, screenshot, named brands, cited sources, associated topics, buyer-readiness score, and follow-up action. Keep one pack per cycle. Then create a trend sheet that rolls up the score changes. If ChatGPT visibility rises from 8% to 18% and cited mentions rise from 2 to 7 across the same prompts, you have usable proof of movement even if not every individual answer changes in your favor. That proof matters when you need to show how you will outrank competitors over time.

The brands that win AI discovery usually do not win because they publish more. They win because more sources consistently make the same claim about their relevance.

That logic also explains why answer-first content matters. Pages built for extraction give engines clearer material to reuse. Definitions, short sections, concrete comparisons, tables, FAQs, and crisp evidence blocks often perform better than long pages that bury the answer. If your content is strong but hard to extract, the benchmark may reveal a structure problem rather than a topic problem. Fixing that structure can help outrank competitors without chasing low-value noise.

How can forum, social, video, and media visibility be used ethically to support AI discovery?

Forum, social, video, and media visibility can support AI discovery when they provide real corroboration rather than fake buzz. Ethical use means showing up where buyers already discuss the problem, adding useful information, and avoiding planted posts or deceptive incentives. Reddit, Quora, YouTube, TikTok, podcasts, and trade media can all strengthen discovery if the contribution is authentic and relevant.

A small team planning content across forum, video, and social channels at a desk.

Answer engines often synthesize from public discussion and widely repeated explanations. Reddit threads can surface practical comparisons and objections. Quora answers can capture use-case framing. YouTube videos can reinforce demos, tutorials, and product fit. TikTok can spread simple category education and customer language, though its commercial value varies by audience. Media mentions can provide independent validation that supports entity trust. Together, these signals can help outrank competitors if the information stays useful and specific.

The ethical test is plain: would the content still deserve to exist if no AI engine ever used it? If the answer is no, the tactic is weak. Forum spam, fake testimonials, paid seeding without disclosure, and scripted “community” posts create short-lived signals at best and reputational risk at worst.

Build an authority map by source type

An authority map groups source types by the role they play in buyer trust. For many B2B categories, the map has five layers: your site, expert third-party coverage, community discussion, video proof, and social reinforcement. For consumer categories, video and community discussion may carry more weight earlier in the decision path. The map should fit the buying motion, not a generic checklist.

Here is a practical national-level source map for many service and software categories:

  • Your own site: comparison pages, use cases, FAQs, documentation, pricing explanations
  • Media: trade publications, interviews, contributed commentary, product roundups
  • Forums: Reddit, Quora, niche communities, professional discussion boards
  • Video: YouTube tutorials, webinar clips, product walkthroughs, expert reviews
  • Social: LinkedIn posts, short-form video clips, industry threads, customer commentary

Suppose a competitor is repeatedly cited for onboarding speed. Check whether that message appears on its site, in YouTube demos, in community discussions, and in media summaries. If the same claim appears in four source types, the engine has repeated corroboration. Therefore, your authority map should help you spot where your message is thin or absent if you want to outrank competitors.

Tip: One honest, information-rich Reddit thread can be more useful than ten shallow social posts. Depth and specificity travel farther than volume in AI extraction.

Set governance rules before participating

Governance rules keep authority-building useful and safe. Teams need clear boundaries for who can post, what disclosure is required, what claims need proof, and how customer stories can be used. A lightweight policy beats ad hoc posting every time.

Use rules like these:

  • Disclose affiliation when a team member speaks for the brand.
  • Do not post scripted praise or ask others to do it.
  • Answer the question first, then mention the brand only if relevant.
  • Never copy the same response across multiple threads.
  • Use customer examples only with permission and without exaggeration.
  • Link only when the linked page directly solves the question.

Those rules sound basic, yet they solve a common problem. Many brands damage trust by treating Reddit or Quora as link drops. Engines and people both detect that pattern quickly.

A related technical issue also matters. If your owned content cannot be crawled or extracted well, off-site visibility will not convert into strong citations. Check robots.txt, llms.txt where used, indexability, page rendering, and structured data. GPTBot, Bingbot, and traditional search crawlers need accessible pages. Google Search Console plus Bing Webmaster Tools can reveal indexing or crawl issues that suppress source visibility. If your pages are missing from answer engines despite strong content, review indexing and crawlability problems that block AI answer visibility.

Generative Engine Optimization is the discipline that aligns these pieces: extractable content, entity clarity, corroborating sources, and prompt-level measurement. If you want the baseline definition and checks, this explanation of Generative Engine Optimization covers the core ideas. In practice, those checks help outrank competitors because they connect visibility with pages engines can reuse.

How to prioritize actions that improve competitive visibility without gaming AI answers

The fastest way to improve competitive visibility is to fix the weakest high-intent gap first, then support it with technical access and off-site corroboration. Most teams should prioritize in this order: technical crawlability, answer-first page coverage, third-party proof, and ethical channel participation. That order works because answer engines cannot cite pages they cannot access, cannot trust claims they cannot corroborate, and cannot recommend brands that lack buyer-stage evidence.

A project roadmap on a whiteboard with tasks, timelines, and priorities for a 90-day plan.

Use a simple framework with four columns: technical, content, authority, and measurement. Score each action by impact, effort, and proof speed. Impact means how strongly the action can affect high-intent prompts. Effort means team time across SEO, content, design, and subject experts. Proof speed means how quickly you can measure movement in a 30- to 90-day cycle. This makes it easier to choose the work most likely to outrank competitors first.

Actions that often deserve first priority include:

  • Fixing noindex, crawl, rendering, or blocked-resource issues
  • Publishing buyer-stage comparison and use-case pages
  • Rewriting key pages in an answer-first structure
  • Adding schema markup for clear page meaning
  • Expanding FAQ sections that match real prompts
  • Earning third-party references that validate category fit
  • Creating video explainers that support recurring objections

A 90-day plan to outrank competitors with proof

A 90-day plan makes the benchmark operational. Month 1 should establish the baseline, repair access issues, and publish the first buyer-stage pages. Month 2 should expand topic coverage, strengthen corroboration, and launch channel-specific authority work. Month 3 should test prompt movement, refine pages that now get partial visibility, and scale what improved cited mentions.

Here is a practical monthly deliverable view:

  • Month 1: 30 to 50 prompt benchmark, engine screenshots, crawlability audit, priority page list, 3 to 5 answer-first content updates
  • Month 2: 4 to 8 new buyer-stage pages, schema review, source outreach plan, 2 to 4 media or expert contribution targets, forum governance rollout
  • Month 3: repeat benchmark, gap delta report, prompt proof pack, content refreshes, video or community expansion based on source wins

Example: a team starts with 36 prompts across three engines and finds 12 high-intent gaps. In Month 1 it fixes indexing on 6 pages and rewrites 4 comparison pages. In Month 2 it publishes 5 use-case pages and secures 3 third-party mentions. In Month 3 it reruns the same prompts and finds buyer-ready mentions moved from 6 to 14 out of 108 observations, which is a clearer gain than raw mention count alone. Moreover, that kind of gain is easier to defend in a plan to outrank competitors.

Team roles that support outrank competitors goals

Execution usually needs four roles, even in a lean team. One person owns measurement and reporting. Another owns content structure and publishing. A third handles technical fixes. A fourth manages off-site authority work or subject-matter input. Sometimes one person covers two roles, but the work still needs all four functions.

A provider that claims it can help you outrank competitors should explain who does what, how often benchmarks are rerun, which prompts are tracked, how evidence is captured, and what gets published each month. If the answer is vague, the delivery model is vague too. Commercial buyers should ask for a sample report, a sample prompt pack, and a sample action plan before signing anything.

Automation can reduce the manual load a lot here. A platform that combines research, writing, optimization, publishing, and tracking removes handoffs that often slow teams down by weeks. Seonix is built around that operating model, and an automated workflow from URL to published, tracked SEO content is often what turns a benchmark into shipped pages instead of a stalled spreadsheet. As a result, teams can outrank competitors faster because execution keeps moving.

What evidence should a service provider show before claiming it can help a brand outrank competitors?

A credible provider should show benchmark methodology, prompt-level proof, sample deliverables, and a realistic reporting cadence. Claims without examples are weak because AI outputs are fluid and category contexts vary. The provider does not need to promise exact answer placement, but it should show how it measures movement and what actions it controls.

Proof that supports outrank competitors plans

Ask for five kinds of proof. First, request a sample benchmark with prompts, engines, and scoring logic. Second, ask for before-and-after screenshots from the same prompt set over time. Third, ask what monthly deliverables are included, such as page creation, updates, technical checks, or authority work. Fourth, ask how they separate brand mentions from cited mentions. Fifth, ask how they handle ethics on forums and social platforms.

A strong buyer-readiness scorecard for providers includes these criteria:

  • Prompt sampling method is explicit
  • Engine coverage includes at least three major answer environments
  • Reporting shows citations, mentions, and associated topics separately
  • Technical readiness checks include robots.txt, schema markup, indexing, and rendering
  • Off-site work has disclosure rules and no fake UGC tactics
  • Publishing workflow is clear and repeatable
  • Proof packs use screenshots and dated logs

How to judge delivery and reporting quality

If a provider only talks about one engine, one vanity score, or generic “AI optimization,” the offer is too thin for procurement. Buyer teams need evidence that the service can connect measurement to execution. That means seeing what gets done in Month 1, Month 2, and Month 3, not just hearing theory.

Many teams make the same mistake: they pay for a benchmark that never becomes published improvements. The useful providers connect measurement, content changes, technical fixes, and proof into one operating loop. Over a few reporting cycles, that loop gives you a much better chance to outrank competitors than a one-off audit ever will.

Conclusion

To outrank competitors in AI answers, you need more than brand monitoring and more than a few extra articles. You need a benchmark that shows AI Share of Voice by engine, citation gaps on high-intent prompts, associated topic gaps, third-party corroboration strength, and channel presence across the places buyers already trust. Once those gaps are visible, the next actions become clear: fix crawlability, publish answer-first pages, add proof that engines can reuse, and build ethical reinforcement across forums, video, social, and media.

The practical advantage of this approach is focus. Instead of guessing why a competitor gets named more often, you can see the source patterns, prompt patterns, and buyer-stage patterns behind the gap. Consequently, AI search visibility and brand mentions become an operating system for organic growth rather than a vague branding exercise. Teams that use this process consistently are far more likely to outrank competitors because they improve the right pages, the right proof, and the right channels in the right order.

FAQ

Here are the most common questions buyers ask when they want to compare brands and outrank competitors in AI answers.

How do you measure whether competitors appear more often in AI answers than your brand?

Use the same prompt set across the same engines within a short time window, then compare appearance rate, citation rate, and buyer-readiness score. A useful baseline is 30 to 50 prompts across at least three engines, repeated monthly with screenshots and dated logs. This shows where you can outrank competitors and where proof is still missing.

Which gaps matter most: citations, unlinked mentions, associated topics, or third-party corroboration?

Citation gaps usually matter most on commercial prompts because they show missing proof that engines can cite directly. Associated topic gaps come next, then third-party corroboration gaps, while unlinked mentions usually become stronger after the first three areas improve. In that order, teams can outrank competitors with more control.

How can forum, social, video, and media visibility be used ethically to support AI discovery?

Use those channels to add honest, useful information where buyers already ask questions. Disclose affiliation, avoid scripted praise, answer the question first, and treat each contribution as content that should help a person even if no AI engine ever reuses it. Done well, this can help outrank competitors without gaming the system.

What evidence should a service provider show before claiming it can help a brand outrank competitors?

A provider should show prompt lists, engine coverage, scoring logic, before-and-after screenshots, monthly deliverables, and governance rules for off-site participation. It should also separate cited mentions from simple brand mentions and explain how technical fixes and publishing work connect to the benchmark. That evidence shows whether the provider can truly help outrank competitors.

Early movement can appear within one reporting cycle if the issue is technical access or missing answer-first structure on existing pages. Broader gains usually need 60 to 90 days because stronger citations, topic associations, and third-party corroboration take time to spread across the web. Meanwhile, teams that stick with the same benchmark can see whether those changes help outrank competitors.

If you want to turn this benchmark into an action plan, start with a free SEO analysis. Seonix can map the gaps, publish the right content to your site, and keep tracking what changes across search and AI answers.

Related articles

Outrank Competitors With an AI Visibility Benchmark