AI assistants pick sources through retrieval-augmented generation, pulling from indexed web content that matches the query and passes internal relevance filters. The systems favor pages with structured data, clear authorship signals, and factual density they can verify against other sources. Exact algorithms remain proprietary, but observable patterns reveal what consistently gets cited.

A potential client asks an AI assistant who sells waterfront condos in their area. The machine names three agents, and you are not one of them, even though you have closed more waterfront deals than those three combined. Nothing it can read tells it so. With 82% of Americans now using AI for housing market information (Realtor.com survey, October 2025), that silence has a price.

This guide covers how AI assistants decide which sources to cite. Its companion piece covers the other half of the equation: where ChatGPT gets its information about you in the first place. Together they explain why some agents get named and most stay invisible.

How AI Assistants Pick Sources: The Short Version

Most major AI assistants now use some form of retrieval-augmented generation. The model does not simply recall what it learned during training. It searches, retrieves content in real time, and synthesizes an answer from what it finds.

This means your website can be read today and cited tomorrow. It also means the system is choosing between your content and everyone else's every time someone asks a question you could answer.

Training data cutoff dates vary by model and affect source freshness. A model trained on older data may not know about your latest market report unless it retrieves it through search. Real-time retrieval is what makes current content discoverable.

The retrieval step works like search. The system queries an index, receives candidate pages, and decides which ones to trust. What happens next depends on signals it can observe.

The Signals That Decide Who Gets Cited

None of the companies disclose their exact ranking algorithms. But you can see what gets cited and work backward, and four signals show up consistently.

Authority. Pages from domains with established backlink profiles and consistent publishing histories appear more often than pages from new or thin sites. The machine is looking for evidence that other sources trust you.

Recency. A current-quarter market report will outrank a report from several years ago for questions about current prices. Timestamps are signals, and answer engines read them.

Structure. Structured data markup improves content discoverability. FAQ schema lets the system extract your answers directly. Article schema with clear authorship lets it attribute the source. AI Overviews pull directly from structured content when it is available. Pages without structured data force the machine to guess, and it usually guesses someone else.

Factual density. Consider two pages: one says "the market is strong," the other states specific median prices with quarter-over-quarter comparisons and named source attribution. The second gives the system something to cite. The first gives it nothing. Length alone does not win: a 500-word page with four concrete facts can outperform a 3,000-word page with none.

What does not get cited: pages with heavy scripts that block crawlers, pages behind login walls, pages with no clear author attribution, and pages that make claims without specifics.

What Our Own Tracking Shows

These patterns are not theoretical. Filtrs runs a repeating panel of the questions buyers and sellers actually ask AI, and logs every answer with its date and source list.

One observation from that log: a client's county valuation page, built with specific sale data and full schema markup, was cited by ChatGPT in 3 of 3 daily samples on the question "how much is my house worth in Bergen County NJ," holding the position 1 citation in the August 7, 2026 sweep. The pages that follow every rule in this guide are the ones the engines keep choosing. The full case study, including the losses and flat stretches, is public.

One client's results are observations, not a promise. But they are dated, logged, and repeatable measurements, which is more than most claims about AI visibility can say.

The Authority Signals AI Models Seem to Favor

Domain authority is one signal, but it is not the only one. Authorship matters too. When your name appears consistently across your website, your brokerage profile, and third-party mentions, the model can connect those dots.

Citation networks play a role. If other websites link to your content, that is evidence of trust. If your market report gets referenced by a local news outlet, your content may carry more weight the next time someone asks about that market.

Consistency is its own signal. An agent who publishes a monthly market report for three years has a different profile than an agent who published once and never again. Publishing patterns are observable, and the machine observes them.

None of this guarantees a citation. Exact ranking algorithms remain proprietary and undisclosed. But these signals appear in what gets cited, and their absence appears in what does not.

Common Misconceptions About AI Source Selection

Paying for ads does not improve AI citations. The retrieval system pulls from an index. Advertising placement does not change what it finds or how it ranks sources.

Longer content does not always win. Dense, specific content outperforms padded content, and the difference is measurable.

Social media presence does not directly feed AI retrieval in most cases. Instagram posts are not indexed the way website pages are. The relationship between platforms like Instagram and Facebook and AI search is more limited than most agents assume.

Practical Takeaways for Real Estate Agents

Publish specific content on your own domain. Market reports with real numbers. Neighborhood guides with named streets and landmarks. Building profiles with concrete details. The system needs something to cite.

Use structured data markup. Article schema with author attribution. FAQ schema for question-answer content. Machine-readable structure reduces ambiguity, and ambiguity is what keeps good agents invisible.

Build consistency over time. A single post may not move the needle. A year of monthly content on the same market builds a pattern the model can recognize.

Make your expertise retrievable. If you have closed more waterfront deals than anyone in your market, that fact needs to exist on a page the machine can read. Credentials that live only in your memory do not get cited.

Check what the machines currently say about you. Ask an AI assistant who sells condos in your area, and see whether your name appears. If it does not, now you know what to fix.

Frequently Asked Questions

Do AI assistants prefer certain websites over others?

AI assistants favor websites with clear authority signals, consistent publishing, and structured data markup. Domain authority matters, but so does content specificity and factual density. A newer site with well-structured, evidence-backed content can outperform an established site with thin content.

Can I pay to have my content cited by AI tools?

No. AI retrieval systems pull from indexed content based on relevance and authority signals, not advertising spend. Paying for ads does not change what AI assistants retrieve. The only way to improve citation rates is to improve the content itself.

How often do AI assistants update their source information?

Retrieval-augmented systems can access current web content in real time. Training data has fixed cutoff dates that vary by model, but retrieval extends that window. Content published today can be retrieved and cited tomorrow if it is indexed and accessible.

Why do AI tools sometimes cite outdated information?

AI tools may cite older content when it ranks higher on authority signals or when newer content lacks the structured data and specificity the system needs. Publishing fresh, well-structured content with current data is the fix.

How long does it take to get cited by AI?

Plan on citations typically following 4 to 8 weeks of held search rankings, because the engines cite pages that have already proven themselves. It can move faster: our first documented client saw a first genuine citation 12 days after his first posts published, logged July 14, 2026. That was observed, not typical, and anyone promising citations in week one is guessing.

See what AI can currently find about you with a free AI visibility report at filtrs.io