Methodology
The Observatory is built entirely from open bibliographic data. This page explains where that data comes from, how each figure on the site is computed, and how to read it well, including what each method can and cannot tell you.
The data
The corpus draws on two open sources. OpenAlex is the backbone, providing works, author and institution records, abstracts, and the pre-assigned concepts used for topic classification. PubMed (via the NCBI E-utilities) adds biomedical recall and records not indexed in OpenAlex. Coverage runs from 2015 to 2025.
After deduplication and false-positive removal, the corpus contains 20,777 articles published 2015–2025.
Because both sources index English-language journals most fully and the search terms are English, the corpus represents the indexed, English-language literature. This is worth keeping in mind when reading the geographic figures, which lean towards English-publishing countries. The sources are also primarily biomedical, so research on PMOS that lives in health services, nursing, social science and qualitative literatures is under-represented. That matters most for questions about care delivery, equity and lived experience, where the absence of papers here reflects our sources as much as the field.
Topic classification
Each article is tagged into a 15-category taxonomy derived from topic modelling of the corpus and validated against the natural clusters in the literature. The categories span metabolic health, fertility, genetics, endocrinology, mental health, inflammation, epidemiology, nutrition, pharmacology, lifestyle, gut microbiome, dermatology, diagnostics, animal models, and molecular mechanisms.
Classification is automated across three tiers. Most articles are classified using OpenAlex's pre-assigned concepts, with a keyword fallback and an AI pass for the remainder.
Classification pipeline
Classification is automated rather than expert hand-coding, so an individual tag is best read as a strong signal rather than a final verdict. At the aggregate level the taxonomy holds up well against the natural structure of the literature.
How topics are counted
An article can belong to up to three topics. Counting it once per tag would inflate multi-topic papers, so instead each article is split equally across its tags: a paper tagged under two topics contributes 0.5 to each. Topic totals therefore sum to the number of articles, not the number of tags.
This supports three views of the same data: raw topic mentions, weighted research share (the default on the site), and primary focus (each article's single highest-confidence tag).
Worked example: 3 papers
| Paper | Tags assigned | Weight each |
|---|---|---|
| Paper A | Metabolic HealthEndocrinology | 0.50 |
| Paper B | Mental Health | 1.00 |
| Paper C | GeneticsMolecular MechanismsInflammation | 0.33 |
Overview bar (weighted)
Sums to 3 = number of articles ✓
Explorer bar (raw mentions)
Sums to 6 = number of tag assignments
On the Overview, five charts use weighted research share: Top Research Topics, Topic Momentum, Research Approach, Citation Impact by Topic, and Funded % by Topic. In the Explorer, the Country × Topic heatmap also uses weighted counts, to avoid inflating multiple cells for papers tagged under more than one topic. All other Explorer figures, including the topic bars, country bars, top journals, and top institutions, use raw article mentions, so the numbers correspond directly to the article count shown in the results list.
How the Explorer works
All panels draw from the same filtered set of articles, and the results list shows the total count matching your selection. Topic bars, country bars, heatmap, trend chart, journals, and institutions are all breakdowns of that same set.
Topics use raw article mentions, so a paper tagged under three topics counts once in each and the bars can sum to more than the total article count.
Countries show the top publishing countries by default. When a country filter is active, the panel switches to co-authoring countries, meaning all countries sharing at least one paper with the selected country. Country is based on author affiliation at the time of publication.
The Country × Topic heatmap uses weighted counts per cell and offers two views: raw weighted share, and a Specialisation index (RCA) that normalises by country publication count and topic size to reveal comparative focus. RCA is the default.
Journals and institutions show raw article counts for the filtered set. When a country filter is active, only institutions with that country code are included.
Research approach: underlying mechanisms, symptom, bridging
PMOS is a multisystem condition encompassing endocrine, metabolic, reproductive, psychological, and dermatological features, yet research has largely developed across specialist domains: fertility, dermatology, mental health, endocrinology, nutrition, lifestyle. The underlying mechanisms / symptom-specific / bridging split was designed to make this fragmentation visible. It asks whether research is investigating why PMOS happens, treating one of its features within a specialist field, or connecting the two. This dimension is not captured by standard bibliometric schemes, which describe research phase rather than how the field relates to the condition as a whole.
It is applied by mapping each of the 15 topic categories to one bucket. The bucket is assigned at the topic level. A paper's weighted contribution (see above) flows into whichever bucket its topic belongs to, so a paper spanning two topics in different buckets is split proportionally between them. The mapping is published so the judgement behind it stays open to scrutiny.
A view based on the NIH translational continuum (basic, translational, clinical, implementation) is envisaged for a future iteration, to complement this framing for readers who prefer the standard research-phase taxonomy.
How the 15 topics are assigned to each bucket
Investigates why PMOS happens
Focuses on one symptom within a specialist domain
Connects cause to clinical care
Geographic concentration
Concentration is summarised three ways: a Gini index over per-country publications (0 means every country contributes equally, 1 means a single country produces everything), the combined share of the top five countries, and the collaboration rate (the share of papers with authors from two or more countries).
A country here is the author's affiliation, so these figures describe where research is produced rather than where patients are or where the need is greatest. That is the right lens for a question about the research community, and a limit to bear in mind for questions about populations.
Topic momentum
Momentum compares the share of the field each topic held in 2015 to 2017 against 2023 to 2025, expressed as a percentage change. It shows where the field's attention is shifting, not which topics matter more.
A large percentage change on a small base is still small in absolute terms, so each bar shows the topic's current share of the field alongside its change, letting a real shift be read in context.
Evidence pyramid
The evidence pyramid groups articles by study design, from the strongest tiers (meta-analyses and systematic reviews) down to case reports and commentary. Study design is taken from PubMed's Publication Type field, a fixed vocabulary maintained by the US National Library of Medicine.
A single paper often carries several Publication Type labels. A meta-analysis, for example, is usually also tagged as a review. To avoid double-counting, each article is counted only once, under the strongest design it qualifies for, so the tiers add up to the total covered.
Coverage is partial because PubMed only adds a specific design label when an article clearly fits one of its defined study-type categories. Everything else is filed under the generic “Journal Article” type, which carries no design signal, and lands in “Primary Research (unspecified design)”. That bucket is therefore large by construction and does not indicate weak research. Retracted articles are excluded from this chart.
Research versus disease burden
Each dot represents a country, positioned by how much PMOS research it produces relative to its reproductive-age female population (X axis, log scale, 2019 to 2025) against its recorded disease burden from GBD 2021 (Y axis, DALYs per 100,000 women). Colour shows World Bank income group. The companion ranking uses a gap score (burden percentile minus research percentile) to identify which countries are most under-researched relative to their recorded burden. Sources: GBD 2021 (IHME) via Yao et al. (PLOS ONE 2025) and World Bank reproductive-age population estimates.
Two limitations bear directly on how the chart should be read. First, GBD captures PMOS burden through disability only: PMOS does not directly cause death, so the DALY figures here consist entirely of Years Lived with Disability (YLD) with no Years of Life Lost (YLL) component. Additionally, the metabolic, cardiovascular, and mental health consequences of PMOS are attributed to those disease categories in GBD rather than back to PMOS, so the Y axis understates true burden for every country shown. Second, recorded burden reflects diagnostic capacity: countries with weaker health systems detect and code fewer diagnoses, so a low Y position in lower-income countries often signals invisible burden rather than absent burden. The chart is most reliably read for high-income countries where surveillance is strong; comparisons involving lower-income countries should be treated as indicative rather than precise.
Funding
Funding is enriched from three sources: OpenAlex funders and awards, PubMed grant lists, and Crossref funder records by DOI. About a quarter of papers carry an acknowledged funder.
We read this as a floor rather than a true funded rate, since acknowledgement varies by journal, country and sector. No source provides reliable monetary amounts, so the figure tracks whether funding is acknowledged, not how much money the field receives.
The funder chart counts distinct papers per funder, not amounts spent. A funder making many small grants will rank higher than one making fewer, larger ones. Funder names are lightly normalised to merge known name variants. Similarly named agencies from different countries are kept separate.
Comorbidities
Comorbidity figures come from a word-boundary regex pass over each article's title and abstract, covering 18 PMOS-adjacent conditions across metabolic, psychiatric, reproductive, oncological and other groups.
This records co-mention, whether a condition appears alongside PMOS in the text, which is a measure of how often the field connects the two rather than of dedicated research on that pairing.
Updates and scope
This release covers 2015 to 2025, indexed in English. As the field transitions to the PMOS name, new literature will appear under different terminology, making regular updates important to maintain full coverage. The longer ambition is for this to become a living map of PMOS science, one that goes beyond bibliometric counts to surface what the research actually says, where knowledge is consolidating, and where the field has yet to look. That means deeper text analysis, richer navigation, and eventually plain-language summaries that make findings accessible to patients and advocates, not just researchers.