ChatGPT & Perplexity Optimisation | Retrieve · Rerank · Cite

Ten blue links became one answer.
Only a few sources get named in it.

ChatGPT and Perplexity do not rank pages. They retrieve passages, weigh how well each claim is corroborated, and cite the handful they build the answer from. We work on the parts of your site and your off-site footprint that decide whether you are in that handful, then measure it on both platforms separately.

  • Fixed prompt set, baselined before any work ships
  • ChatGPT and Perplexity measured separately
  • Entity and schema work shipped, not filed
  • 10+ years of engineering heritage
Get a free citation audit →

Prompt set · citation baseline · yours to keep either way

Prompt Fan-out Retrieve Rerank CiteEvery stage is a place a brand quietly drops out of the answer.

What we work on

Six things decide whether a model can quote you.

They run in sequence and they depend on each other. A passage cannot be cited if a crawler never retrieved it, and a retrieved page is no use if the claim inside it falls apart the moment it is lifted out. We work through them in the order the dependencies run.

01

Retrieval and access

  • Crawler access decided per user agent rather than inherited from a template
  • Server-rendered HTML, since most AI crawlers do not execute your JavaScript
  • Canonical and status signals clean, so one URL carries each claim
  • Pages fast and stable enough to survive a live fetch mid-answer

02

Passage architecture

  • One question per section, answered completely inside that section
  • Claims carrying their own subject, figure and date so they survive extraction
  • Heading hierarchy that matches how the question is actually asked
  • Tables and definition lists wherever a model needs a discrete value

03

Entity clarity

  • One canonical name, description and fact set across every property you own
  • sameAs links pointed at the records that already describe your business
  • Founders, locations, services and credentials stated identically everywhere
  • Disambiguation where your name collides with another company

04

Structured data

  • Organization, Service, Product, Article and FAQPage matched to each template
  • Author and credential markup on anything asserting a factual claim
  • Validated against Google's requirements, not just valid JSON syntax
  • Kept in parity with the visible text, since a contradiction is worse than nothing

05

Citation strategy off-site

  • Coverage in the publications each platform demonstrably retrieves from
  • Presence in the documentation, directories and Q&A threads Perplexity favours
  • Third-party facts corrected at the source rather than argued on your own site
  • Comparison and alternatives pages where the buying question is a comparison

06

Visibility measurement

  • A prompt set built from real buying questions rather than keyword volume
  • ChatGPT and Perplexity sampled separately, repeatedly, on a schedule
  • Named citations kept separate from unlinked paraphrase
  • Competitor share of the same answers tracked alongside yours

None of this replaces the technical layer or the quality of what you publish, and we do not present it as one. Retrieval runs on pages a crawler can read. Citation runs on claims a model can lift without breaking them. Get both right and you are a candidate for every question your content genuinely answers.

How citations happen

A model does not rank your page. It lifts a passage out of it.

When someone asks ChatGPT or Perplexity a question, the prompt is expanded into several searches, a candidate set of pages comes back, and the passages inside those pages are scored for how completely they answer that specific question. The answer is then written from the few passages that survive, and the citations point at those.

This is why two pages holding the same information perform differently. One states its claim in a single self-contained paragraph, with the subject, the figure and the date all present. The other spreads the same claim across an introduction, a bullet list and a conclusion, so nothing can be lifted out of it without losing its meaning. The first gets quoted and named. The second gets paraphrased without attribution, if it is used at all.

Passage-level retrievalSelf-contained claimsNamed citationSampled per platform

Scope of work

What a ChatGPT and Perplexity engagement covers.

The complete scope, grouped the way the dependencies run rather than alphabetically. Not every site needs every item, and the audit says plainly which ones yours does.

ChatGPT optimisation

Everything that decides whether OpenAI's search surface holds you in its candidate set and is willing to name you in the answer.

  • Crawler access set deliberately per OpenAI user agent
  • Candidate set presence across your core buying questions
  • Brand coverage in the publications it retrieves from
  • Answer share against the competitors in the same answers

Perplexity optimisation

Perplexity retrieves at query time and names more sources per answer, which moves where the leverage actually sits.

  • PerplexityBot access and live fetch performance
  • Community corroborationin threads, docs and Q&A
  • Source count per answer and your position within it
  • Follow-up questions it generates around your topic

Entity SEO and Knowledge Graph

Making the identification of your business unambiguous before asking a model to trust anything it says about you.

  • Entity consistency across every property you control
  • sameAs and Wikidata links to records that already describe you
  • Knowledge Graph presence for the Google-side surfaces
  • Disambiguation where your name collides with another brand

Semantic search and content architecture

Writing for a retrieval step that scores passages rather than pages, without making the page worse for the person reading it.

  • Passage structure and one question per section
  • Query fan-out coverage across how a question gets asked
  • Topic clusters that answer a subject completely
  • Internal linking that keeps related claims reachable

Structured data

Telling machines what a page is in terms that leave nothing to be inferred, and keeping the markup honest.

  • Schema types matched to what each template genuinely is
  • Author and credentials on anything asserting a fact
  • Article and FAQPage markup validated, not merely well formed
  • Parity between the markup and the visible text

Measurement and reporting

Non-deterministic systems need sampling rather than screenshots, and a baseline recorded before anything changes.

  • Prompt set built from real buying questions
  • Repeated sampling on a schedule, per platform
  • Named citations separated from unlinked mentions
  • Baseline recorded in week one, before any work ships

Tooling

Not a performance claim, just the software behind the work, so you know what is running against your site. Citation data is checked by hand against live answers before it reaches a report, because both platforms change what they surface without announcing it.

ChatGPTPerplexityGoogle Search ConsoleScreaming FrogSchema validationGA4SemrushSEOCrawl AI

Two platforms

ChatGPT and Perplexity do not trust the same things.

Treating them as one audit misses where each one pulls its confidence from. The site work is shared. The off-site work and the measurement are not.

01

ChatGPT

  • Retrieves from a narrower candidate set and cites fewer sources per answer
  • Observed citations skew toward established publishers and recognisable brands
  • Separate user agents handle training, search indexing and live user fetches
  • One feature rarely moves it, steady trade and tier-one coverage does

02

Perplexity

  • Retrieves aggressively at query time and names more sources per answer
  • Community discussion, documentation and Q&A threads appear far more often
  • Generated follow-up questions widen the surface one topic can be cited in
  • A brand that comes up naturally in active threads feeds what it will cite

03

What both share

  • Neither can cite a passage it never retrieved, so the technical layer stays first
  • Both reward one self-contained claim over the same claim spread thin
  • Both are shaped by what third parties say about you, not only your own site
  • Neither is stable, so both need sampling rather than a single check

Google AI Overviews and Gemini sit on a different stack again, closer to the classic index and the Knowledge Graph, and that work is covered on our GEO service page rather than duplicated here. The overlap between platforms is the site. The difference is the off-site footprint and the measurement.

Our process

Baseline, audit, correct, restructure, re-measure.

Five stages, run in that order. The order carries the work: restructuring content before the baseline exists produces changes nobody can attribute to anything.

01

Prompt set and citation baseline

A prompt list built from how your buyers actually ask, run in both ChatGPT and Perplexity and sampled repeatedly, because one run of a non-deterministic system is an anecdote. The output records who is cited today, how often you appear, whether you are named or only paraphrased, and which page was used.

02

Retrieval and access audit

Whether a crawler can reach and read the pages carrying your claims. Access reviewed per AI user agent, server-rendered content confirmed, canonical and status signals checked, and the render each platform actually receives compared against the one you see. Nothing further matters until a page can be retrieved.

03

Entity and structured data correction

Name, description, facts and credentials made consistent across every property, sameAs links pointed at the records that already describe you, and schema corrected so it agrees with the page instead of contradicting it. This is the slow, compounding part, and it is where most sites have the largest unclaimed gap.

04

Content architecture

Existing pages restructured so each claim sits in one self-contained passage, and new pages built around the questions the baseline showed you absent from. Written for a person first and structured so a model can lift it without breaking it, which are the same goal more often than not.

05

Off-site corroboration and re-measurement

Coverage pursued and third-party facts corrected where each platform demonstrably pulls from, then the same prompt set re-run against the same baseline. A change that shipped cleanly but moved no citation rate is an open item rather than a deliverable.

Want to see which answers you are already in?

The prompt set and the citation baseline are yours to keep, whether you act on them with us, with your own team, or on your own schedule.

Get a free citation audit →

Track record

The track record predates the company.

The team behind Wegile DGTL has worked together for over a decade, running the complete marketing engine at its parent company, a software development company, and delivering campaigns for its clients before that. The results below were earned by this team across those years, on accounts where the technical, content and paid work were run together.

Home Services · California · Landscape & Christmas Lighting · Landscaping · Turf Installation · Irrigation · Outdoor Lighting

Elevated Seasons: from a wasted budget to market visibility.

8.76x

ROMI in 2 years

11.26x

ROAS (ad spend only)

2.28x

Lower cost per lead

3.82x

Lower customer acquisition cost

Elevated Seasons came to us after a previous agency spent their budget and delivered almost nothing. Sound familiar? We rebuilt the brand, launched Google Search and Performance Max, and ran the SEO strategy that took their revenue from a near standstill to consistent growth, still climbing two years in.

Brand repositioningGoogle Search adsPerformance MaxSEO strategy2yr ongoing partnership

eCommerce · Hair & Beauty · SEO + Meta Ads + Google Shopping

Hairbarnyc: technical foundation first, then content and spend on top of it.

320%

Organic revenue increase

38%

Customer acquisition cost reduction

8mo

To hit results

SEO, Meta Ads and Google Shopping run as one coordinated budget. Over eight months that combination drove a 320% increase in organic revenue and cut customer acquisition cost by 38%, with the crawl, indexation and product template issues resolved before any of the content or spend work was scaled.

Technical SEOIndexation cleanupProduct template fixesGoogle ShoppingMeta Ads
"Working with them transformed our business. Their creative testing, Meta retargeting and dedicated tech support boosted our product sales and appointments, helping us scale revenue and grow our brand. Their strategic insight and execution have been truly exceptional."

★★★★★ 5.0 · Beny, Hairbarnyc

Digital Products & eCommerce · Paid Social Launch

Deliciously Fit, with Chris Powell.

2,000

Books sold

<3mo

Time to sell out the run

A digital recipe book with hundreds of high-protein recipes built for weight-loss and GLP-1 audiences. Our team ran the paid social strategy behind the launch, testing hooks and formats against a cold audience, selling 2,000 copies in under three months.

Paid social launchCreative testingCold audience acquisition2,000 copies sold
Chris Powell, Deliciously Fit
"I've been working with them for nearly 10 years, and they've been an incredible partner every step of the way. From developing my fitness app and nonprofit app, to helping maintain and grow my website, their team has consistently delivered with professionalism and precision. They're responsive, reliable, and deeply knowledgeable, always guiding projects with care and expertise. I trust them fully and plan to continue working with them for years to come."

★★★★★ 5.0 · Chris Powell, Deliciously Fit

ChatGPT & Perplexity questions

Straight answers. No sales fluff.

The questions founders, marketing leads and in-house SEO teams ask before commissioning AI visibility work, answered plainly and in full, including where the honest answer is that something does not work yet.

Both run a retrieval step before they write anything. The prompt is expanded into several searches, a candidate set of pages comes back from an index, the passages inside those pages are scored for how completely they answer the specific question, and the answer is composed from the few that survive. Citations point at the passages that were actually used.

The unit of competition is therefore a passage rather than a page. A self-contained paragraph that answers one question completely, carrying its own subject, figure and date, can be lifted into an answer. The same information spread across an introduction, a bullet list and a conclusion often cannot be lifted without losing its meaning. Ranking well in Google correlates with being in the candidate set, but it does not decide the citation.

The on-site work overlaps heavily. The retrieval surfaces do not. ChatGPT search draws on a narrower candidate set and cites fewer sources per answer, and observed citation patterns skew toward established publishers and recognisable brand coverage. Perplexity retrieves more aggressively at query time, names more sources per answer, and surfaces community discussion, documentation and Q&A threads far more often.

Same site work in both cases, different off-site emphasis, and separate measurement, because a gain on one platform does not imply a gain on the other. We track them as two lists rather than one score.

Yes, and more than the framing of AI SEO as a separate discipline suggests. Nothing gets cited that could not be retrieved, and retrieval still runs on crawlable, server-rendered, indexed pages. Most AI crawlers do not execute JavaScript the way Google's rendering service does, so content that appears only after hydration is often invisible to them.

The technical layer is a prerequisite rather than an alternative. What is genuinely different is the emphasis: entity clarity, passage-level self-containment and off-site corroboration carry more weight here than they do for a conventional ranking.

Entity SEO means making the identification of your business unambiguous. One consistent name, one canonical description, and the same founders, locations, services and credentials stated the same way everywhere they appear, with explicit sameAs links from your schema to the records that already describe you.

Google's Knowledge Graph is one such structured record, and it feeds Google's own AI surfaces directly. ChatGPT and Perplexity do not query it, but they are shaped by many of the same underlying sources that populate it, which is why Wikidata entries, company registries, review platforms and consistent press coverage keep turning up as leverage. Where two companies share a name, that ambiguity is usually the whole problem.

Not directly. A language model reads the text of a page rather than its JSON-LD. What schema does is remove ambiguity for the systems upstream of the model: it states plainly what a page is, who wrote it, what an entity is called, what a product costs and where a business operates, which improves how cleanly a page is parsed, categorised and pulled into a candidate set. It also drives the Google surfaces where structured records are read directly.

We treat it as accuracy infrastructure rather than a citation lever, and we keep it consistent with the visible text, because markup that contradicts the page is worse than no markup at all.

With a fixed prompt set and repeated sampling. These systems are non-deterministic, so one query proves very little: the same prompt run five times can cite five different sets of sources. We build a list of prompts matching how your buyers actually ask, run each of them on a schedule in both ChatGPT and Perplexity, and record how often you appear, whether you are named or merely paraphrased, which URL was cited, and who else is in the answer.

The figure that matters is the change in that rate over time, not any single screenshot. If you are assessing agencies for this work, ask for the prompt list and the baseline. Anyone doing it seriously has both, and can show which content changes moved which answers.

It depends which crawler, and the distinction is worth understanding before the decision is made. Different user agents do different jobs. Some collect training data, some build the search index an assistant retrieves from at answer time, and some fetch a page live because a user asked about it.

Blocking the training crawler while allowing the search crawler is a coherent position. Blocking the search crawler removes you from the citations you are trying to win. Most sites we audit have never made this choice deliberately, they inherited a robots.txt file from a template or a plugin.

It is a proposal for a plain-text file that summarises a site for language models, and the major assistants have not committed to reading it. We will add one on request, since the cost is close to zero, but we do not present it as a lever and we would not build a strategy around it.

Separating the parts of AI search work that demonstrably move citations from the parts that are still speculative is a large share of what an audit is for.

Faster than a Google ranking in some cases and slower in others. Perplexity retrieves at query time, so a corrected, well-structured page can begin appearing within weeks of being crawled. ChatGPT's search index refreshes on its own schedule and its citation set is more conservative, so movement there tends to track off-site coverage and takes longer.

Entity and authority work is the slow part, measured in months, because it depends on third-party sources updating their own records. We record the baseline in week one so the change is visible rather than argued about.

One last thing

The audit costs nothing, and the prompt set is yours either way.

A fixed list of the questions your buyers actually ask, run in both ChatGPT and Perplexity, with a record of who gets cited today, where you appear, and whether you are named or paraphrased. Act on it with us, hand it to your own team, or keep it on file. The document is yours regardless.

Get your free citation audit

Prompt set · citation baseline · no obligation