"What sources does ChatGPT use for health questions?" It's the right question for any healthcare marketer, and the honest answer is five. AI doesn't pull from one place. It builds each answer from a portfolio of five source types, and most brands optimize for exactly one of them.
That gap is the whole opportunity. Generative engine optimization (GEO) is the practice of shaping how AI assistants like ChatGPT, Gemini, Perplexity, and Claude name and cite your brand. Same goal as SEO, different rules for who gets surfaced. Once you know which source types models actually pull from, you can shape how AI answers healthcare questions instead of hoping for the best.
We analyzed 27,812 AI answers to 6,953 healthcare prompts across four major assistants in the US (Yolando Research, June 2026), capturing 675,425 citations spread across 33,141 domains. Five source types did most of the work. Here's the breakdown, and what you can do about each one.
What are the five AI citation sources for healthcare answers?
AI cites a portfolio of five source types. Ranked by share of individual citations, with how much control you have and the tactic that fits each:
# | Source type | Citation share | How much can you control it? | Primary tactic |
|---|---|---|---|---|
1 | Provider sites (your own domain) | ~61% | Full control, under-used | Deepen homepage, service, and condition pages so models can quote them |
2 | Social platforms (Reddit above all, forums, video) | ~12% | Low: steer, don't own | Earn honest recommendations in the threads patients already read |
3 | Aggregators, directories, and news | ~12% | Low: influence only | Keep every listing accurate, complete, and consistent |
4 | Government and academic (CDC, NIH, journals) | ~9% | None: align only | Mirror the guidance these bodies publish; never contradict it |
5 | Reference sites (Wikipedia, Wikidata) | ~6% | Partial | Fix the factual entity data about your organization |
Those percentages are shares of the 675,425 individual citations in the dataset. Provider sites are the layer you fully own. The other four you can move but can't control. Match your tactic to the control model and you stop pouring effort into the wrong one.
Why do provider sites earn ~61% of healthcare citations?
Provider sites earn the biggest slice because they answer the question the patient actually asked. Someone asks an AI assistant about a treatment, a condition, or where to get care, and the model reaches for pages that describe exactly that. Those pages usually live on your domain. It's the layer you fully control, and the one most healthcare brands neglect.
Here's the detail that should change where you spend. For the median cited brand, a single page drove about 32% of its citations. That page was the homepage 43% of the time, and a service or condition page 12% of the time. A blog post was the top-cited page only 5% of the time. AI reads your money pages, not your blog, yet most brands invest the opposite way, churning out posts the model barely reads while their core pages stay thin.
So build outward from the pages that already carry the load. Your service and condition pages should name the procedure, who it is for, the eligibility criteria, the cost signal, and the location, rather than leading with a booking form. A fertility clinic's IVF page that states cycle costs and who qualifies gives the model something to quote; a consultation form doesn't. Your condition and treatment explainers should define the topic in the first sentence, then answer the obvious follow-up a patient would ask. And your clinician bios need credentials, specialties, and the conditions each provider treats, so the model can match a provider to a query.
Blogs still have a job. They answer long-tail questions like "what does recovery from X look like." But they are a supporting layer, not the main event. Deepen the homepage and service pages first. Our breakdown of the pages AI actually cites in healthcare shows why that order matters.
How do you handle social and aggregators (~12% each)?
Social platforms and directories each account for roughly 12% of healthcare citations. Together that is nearly a quarter of the pool: too big to ignore, too public to control. With both, your relationship to the model is influence, not ownership.
Reddit sits at the center of the social slice. In our dataset it was the single most-cited domain at about 12% of all citations, the number-one source in every healthcare category we tracked, and present in roughly one of every three answers. That's enormous reach with almost no control: tens of thousands of independent threads you can't edit. You don't publish your way to the top of a Reddit thread. You steer it. Show up in your real voice, answer questions, correct misinformation, and be the practice patients mention because the experience was worth mentioning.
Directories are more mechanical, which makes them easier to work with. Keep every listing accurate: correct name, address, phone, services, hours, and specialties, matched across every directory that carries you. When one listing shows a service and another leaves it off, you hand the model conflicting signals. It's the same discipline as keeping a Google Business Profile current, extended to everywhere a model might read about you. After Reddit, each vertical leans on its own gatekeeper directory or authority, like Psychology Today, Yelp, or Healthgrades, so accuracy on the ones that matter in your category is table stakes. Which directories matter depends on the category, and a mental health practice has a very different list to keep clean than a multi-site surgical group.
How do government, academic, and reference sources fit in?
Government, academic, and reference sources together account for about 15% of healthcare citations. Your job with all three is alignment, not ownership.
Government and academic sources, meaning the CDC, NIH, and peer-reviewed journals, earn roughly 9% and carry outsized authority on clinical questions. Cite the same consensus these bodies publish and models get more comfortable naming you alongside them. Contradict it and the model routes around you. In some categories a specific authority like the CDC becomes the vertical's gatekeeper, so deference to published guidance isn't optional.
Reference sites, chiefly Wikipedia and Wikidata, earn about 6% but anchor more than that number suggests. Every major LLM is trained on Wikipedia content, and it's "almost always the largest source of training data" in these systems, according to the Wikimedia Foundation. Treat Wikipedia and Wikidata as your brand's fact sheet for AI. If your founding date, locations, leadership, or specialties are wrong there, the model repeats the error.
How do you build a healthcare AI citation portfolio?
Treat the five source types as a portfolio, not a single lever. A brand that pours everything into its website and ignores directories, or chases Reddit while its service pages stay thin, leaves citations on the table. If nearly all your citations trace to one source type, you're one algorithm change away from disappearing.
Start where you have the most control and the most upside at the same time:
Deepen your money pages. Provider content is roughly 61% of the opportunity and the layer you fully own. Fix the homepage and the service and condition pages first.
Clean up your directory listings so name, services, and hours match everywhere a model reads them.
Steer the social conversation by showing up honestly on Reddit and community forums where patients already are.
Align with authoritative guidance from the CDC, NIH, and the gatekeeper body in your vertical.
Fix your reference-data facts on Wikipedia and Wikidata so the model repeats the truth about you.
Each lever compounds the others. This is where GEO parts ways with the old SEO playbook: you win by owning the framework and the specific questions patients ask, not by brute force on one keyword. For how models pick and cite sources, see our guide to generative engine optimization.
The behavior underneath all of this isn't new; the channel is. Most Americans already draw health information from several types of sources rather than one. And Gartner projects traditional search engine volume will fall 25% by 2026 as people shift to AI assistants. The patients are the same. The answer layer is different, and it runs on these five sources.
See where your brand stands across all five source types
Our healthcare AI-search platform tracks how AI assistants cite your brand across every source type covered here. AI Visibility monitors your presence across ChatGPT, Gemini, Perplexity, and Claude daily, showing you exactly which prompts surface your brand and where competitors show up instead. When it finds a gap, the platform's recommendations and Marketing Studio tell you what to fix and generate on-brand content ready to publish.
See how Yolando tracks and grows your healthcare AI visibility. Book a demo.





