How do LLMs decide which sources to cite?
Short answer: LLMs cite sources that are (1) accessible to their crawlers, (2) structured with schema.org, (3) consistent across the open web as an entity, and (4) original — they contain data or claims not found elsewhere. Backlinks and domain authority help but matter less than primary data.
Large language models choose which sources to cite through a multi-step process that combines their training data with live retrieval.
For training-time citation (what the model "knows"), the model is more likely to remember and repeat content that: (1) appears on multiple high-authority domains (cross-cited); (2) has clear authorship and publication date; (3) contains unique, specific data not summarized elsewhere; (4) is structured in clean Q&A or list format. Wikipedia, major news outlets, and government sites dominate this layer.
For retrieval-time citation (when the model searches the web live, as in ChatGPT Search or Gemini), the model ranks results using a combination of: (1) crawler access (can the bot even see your page?); (2) semantic similarity to the user's query; (3) source authority (backlinks, entity recognition, brand mentions); (4) content freshness; (5) structured data (schema.org makes extraction easy).
The single biggest mistake Indonesian businesses make is assuming that ranking on Google = being cited by AI. They are related but not the same. A page can rank #1 on Google and be completely absent from AI answers because the AI crawler was blocked, the schema was missing, or the content was a generic summary rather than primary data.
ShortcutSistem's Omni Audit specifically tests for these four conditions on every scan.