The short answer
When people ask how much “weight” an AI answer gives to a company website, the news, WeChat posts, or creator content, they often imagine a fixed media mix: perhaps one percentage for official pages and another for social discussion. Public research does not establish such a mix.
A source can occupy at least three positions in the process. It may be retrieved as candidate material. It may be used while the system forms an answer. Or it may be shown to the reader as a citation. Those events overlap, but they are not interchangeable. A visible link shows what the system chose to display. Its absence does not reveal the full set of material the system considered.
The original retrieval-augmented generation paper makes the underlying distinction clear. Its model generates text while conditioning on retrieved documents, rather than relying only on parameters learned during training. That is a useful conceptual model, not a blueprint for every consumer AI search product today. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
The practical GEO question is what proposition needs support, which source can responsibly support it, and what can be observed in a particular answer.
Three layers that are easy to confuse
Take a buyer’s question: “Which parts of an AI workflow can this provider handle, and who approves external sends?” A useful answer needs a current description of the service and a clear account of the human decision point. The first question might also lead someone to ask for independent evidence of delivery.
The source set depends on how the question is pursued. OpenAI says ChatGPT Search may rewrite a question into one or more targeted searches, while Google says its AI features may search across subtopics and data sources. Neither description reveals every candidate that was considered and discarded. OpenAI's search guide Google's AI Search guide
The company’s own website is normally the right place for the first two claims. It is the source that can state what the company currently offers and what it is prepared to stand behind. A dated, named service page with a clear owner is more useful than a broad promise because the buyer can return to it when the answer is challenged.
An independent report has a different job. It can corroborate that a product launch happened, describe an external event, or bring an outside perspective. It cannot replace the provider’s current service terms merely because it is independent. A four-year-old press story may be excellent evidence about history and poor evidence about the service boundary today.
A post from a practitioner, a public account, or a creator can add operational texture: where a workflow was difficult, which wording confused users, or how people experienced a tool. It is weak evidence for a provider’s formal commitment unless the author has direct authority and the statement is attributable. It may also be inaccessible to some products or regions. No responsible GEO claim should assume that every assistant can read every post on a messaging or social platform.
This is why channel labels are too coarse. “News,” “social,” and “official site” describe where a page lives. They do not establish whether the claim is current, primary, attributable, independent, or entailed by the words on the page.
Retrieval, use, and citation are different observations
Availability tells you only that a page could be considered. A retrieved passage may never affect the wording. A citation tells you which link the product displayed; it does not give you access to the source's internal weight. The strongest check is smaller and more demanding: open the cited page and see whether it supports the particular sentence beside the link.
That check still leaves the rest of the answer to inspect. In the ALCE benchmark, citation evaluation separates correctness from completeness. A response can cite something accurately while leaving other important statements unsupported. Enabling Large Language Models to Generate Text with Citations
For example, an answer might cite a news story for context while drawing a product detail from an official page, then attach only one link to a paragraph. Counting displayed domains would miss that division of labour.
What the recent studies do and do not show
The academic GEO literature has started to measure source composition in answers. It is valuable evidence, provided the result is not stretched into a universal algorithm story.
The original GEO paper examined ways of changing web content to improve visibility in generative-engine responses. Its contribution was to make answer visibility an empirical research problem. It did not disclose a platform’s internal channel percentages, and its optimization results should not be read as proof that a source is inherently authoritative. GEO: Generative Engine Optimization
A 2026 preprint analyzed 602 controlled prompts and 21,143 valid search-layer citations. In its displayed citations, official, news, and vertical sources made up 34.22%, 31.17%, and 22.13% for ChatGPT; 46.35%, 18.99%, and 22.00% for Google AI Overview; and 44.07%, 16.07%, and 18.99% for Perplexity. These are sample compositions, not the products' internal source weights. “Official” includes several kinds of institutional pages, not only company websites. From Citation Selection to Citation Absorption
The study also builds a score from observable answer and page features to estimate how much a cited page was absorbed into an answer. That score is a proxy, not a causal trace of model attention. The sample is largely US and English oriented, and it does not include every page that was considered and rejected. It cannot tell a Chinese business what percentage of an answer will come from its site.
Another 2026 preprint asked 614 Chinese-language questions three times across the web and app interfaces of Doubao, DeepSeek, Tencent Yuanbao, and Qwen. Source sets differed even between interfaces of the same product. Its observed outputs come from several consumer categories; they cannot set a rule for every enterprise query or disclose a system's ranking function. What Do Chinese-Language Generative Search Engines Cite and Surface?
Together, these studies support a narrower claim. The source mix visible in generated answers varies with the question, language, interface, and product. It does not support putting a permanent “40% official site” label on GEO.
The official site matters because it carries accountability
The company website has a distinct role. It can hold the organization accountable for a current statement, even though no AI system is obliged to prefer it.
For the buyer question above, a good official page can name the service, define the work it will and will not do, identify the human decision point, and link to the governing evidence where that evidence is public. If the company changes the service, it can update the statement at the source. A buyer, journalist, or model reviewer can see what is being claimed and when it changed.
That makes the site the anchor for first-party facts. It also gives outside coverage something precise to confirm, dispute, or contextualize. Without that anchor, an answer assembled from announcements and reposts can become a collage of stale language with no authoritative place to resolve a contradiction.
This does not make every official claim persuasive. A vague page can be less helpful than a careful independent explanation. An official site also cannot establish independent praise, market adoption, or a customer outcome by assertion alone. Those propositions need appropriate sources.
Repetition is not corroboration
One press release can be republished by ten outlets. A creator may quote that release, and several public accounts may quote the creator. A citation sampler could show many domains while the underlying evidence still comes from one organization.
For each important claim, trace the line back to its origin:
- Who first made it?
- Who has authority to update or correct it?
- Does the later page add independently reported facts, or only repeat the earlier text?
- Is the cited passage actually sufficient for the sentence it is attached to?
The answers usually matter more than raw mention volume. They also keep a team from mistaking citation frequency for influence. A repeated claim can occupy many visible citations and still introduce no new corroboration.
A useful GEO review starts with one answer
Choose one buyer question with a real decision behind it. Before asking an AI product, write down what a careful answer would need. For the workflow question above, the company's current services page can establish its service scope and approval boundary. A dated outside report can establish an independently observed event. An attributable first-hand account can add experience, with its limitations visible. An undated repost is a poor substitute for any of these.
Ask the target AI system the exact question. Preserve the date, answer, citations, and product interface, then read each cited passage. Mark separately which claims were answered, which were supported, which were merely asserted, and which sources were repeated. Repeat the same questions after a website update to see whether the answer reflects the current statement.
OpenAI's search help says responses that use web search may include citations, and warns that cited material can be incomplete, outdated, or incorrect. Inspect the source beside each material claim. Searching the web with ChatGPT
Where Open GEO fits
Open GEO Console is useful when a team needs to turn that review into a website decision. Starting with a public site and target buyer questions, it checks the site's technical foundation, public content, question coverage, and citation evidence. It organizes the findings into an issue list, page recommendations, and remediation priorities.
It does not measure a platform’s hidden source percentages, prove that an uncited page was used, or promise that an AI product will cite the site. Those are outside what public answers can establish. The useful output is a reviewable set of questions: does the company’s own site carry the current facts it needs to own, where is corroboration required, and which pages leave a buyer unable to verify a material claim?
See the Open GEO project overview, the AI delivery services, or contact us with a public URL and a buyer question that matters to the business.
Sources
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Aggarwal et al., GEO: Generative Engine Optimization
- Zhang et al., From Citation Selection to Citation Absorption
- Zhen et al., What Do Chinese-Language Generative Search Engines Cite and Surface?
- Gao et al., Enabling Large Language Models to Generate Text with Citations
- OpenAI, Searching the web with ChatGPT
- Google, AI features and your website