When people search social media for GLP-1 advice, can video length, creator credentials, or popularity help them judge what to trust?
Study: Assessing the quality and reliability of health information on GLP-1 receptor agonists on TikTok and Bilibili: a cross-sectional analysis of the public health infodemic. Image Credit: CalypsoArt / Shutterstock
In a recent study published in the journal Therapeutic Advances in Endocrinology and Metabolism, a group of researchers compared the quality, reliability, and completeness of health information about glucagon-like peptide-1 (GLP-1) receptor agonists on Douyin (TikTok China) and Bilibili and examined factors associated with information quality.
Background
How reliable is the health information people encounter while scrolling through short videos about weight-loss medications? GLP-1 receptor agonists, initially developed for type 2 diabetes, have gained widespread attention because of their weight-loss effects.
As interest has grown, social media has become an increasingly important source of information about these medications. The World Health Organization describes an “infodemic” as an excess of accurate and inaccurate information that makes trustworthy sources difficult to identify.
Research on Western platforms has reported generally low-quality information, while China has a distinct digital environment shaped by different platforms, regulations, and market conditions.
About the study
The cross-sectional study examined publicly available videos on Douyin, the Chinese domestic version of TikTok, and Bilibili. The study followed the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines for cross-sectional studies. The keyword “GLP1” was searched in the mobile applications after account information and search history were cleared to reduce the influence of personalized recommendations.
All 106 videos returned by Douyin were collected, whereas the first 200 videos returned by Bilibili’s default search ranking were selected. The platforms used different search-result mechanisms, and Bilibili did not provide a total result count or popularity-based sorting option. Selecting only the first 200 Bilibili results could have introduced sampling bias.
Videos were excluded if they were irrelevant, duplicated, lacked substantive health education, had no sound, were posted within the previous 7 days, or were published before 2020. Videos identified as promoting or selling counterfeit GLP-1 products were also excluded from quality assessment. Metadata included publisher type, publication year, duration, likes, collections, comments, and retweets.
Video quality was assessed using the modified DISCERN tool, Global Quality Score (GQS), and JAMA benchmarks. Content completeness was assessed across six educational dimensions. Two medically trained researchers independently scored the videos, with disagreements resolved by a senior researcher. The researchers compared scores across platforms and examined associations between video characteristics, engagement, and quality.
Study results
After applying the predefined exclusion criteria, 56 Douyin videos and 72 Bilibili videos remained, producing a final sample of 128 videos. The characteristics of the two groups differed substantially. Bilibili videos were generally longer: 66.7% exceeded 300 seconds compared with 16.1% of Douyin videos.
The mix of publishers differed too. Doctors accounted for 66.1% of Douyin videos, whereas personal users produced 58.3% of Bilibili videos. While Douyin videos were shorter, they generated much higher median numbers of likes, comments, and retweets.
Bilibili videos demonstrated greater average content completeness across all six assessed dimensions: definition and mechanism, indications, risk factors and contraindications, usage and assessment, side effects and management, and outcomes and efficacy.
An analysis by publication year found no statistically significant trend toward greater completeness in the sampled videos on either platform from 2020 through 2025. Few of the sampled videos are from earlier years, limiting what this comparison can show about changes over time.
Videos produced by doctors and organizations appeared to score higher in the authors’ platform-specific heatmaps, but the pooled results did not show a consistent quality advantage for doctors over personal users across the modified DISCERN, GQS, and JAMA criteria. Only three organizational videos were included.
Differences in quality scores among publisher types were not statistically significant. Engagement differed substantially: videos from doctors and news media had higher median numbers of likes, comments, and retweets in the pooled sample, which combined platforms with markedly different engagement levels.
These findings showed that engagement did not align with researcher-rated information quality.
Bilibili videos had higher mean modified DISCERN scores and JAMA benchmark scores than Douyin videos. The study did not detect a difference in GQS scores between platforms. The results indicated higher scores for Bilibili on checklists assessing reliability, transparency, authorship, sources, and potential conflicts of interest. The scoring did not independently verify the factual accuracy of every claim made in the videos.
Video duration was strongly associated with quality scores. Mean modified DISCERN, GQS, and JAMA scores increased across the three duration categories and were highest among videos longer than 300 seconds. On Bilibili, duration was the only measured variable positively associated with all three quality scores; likes and comments showed no clear association with quality. On Douyin, likes, collections, retweets, and duration showed positive correlations with modified DISCERN and GQS scores.
Five videos identified as promoting or selling counterfeit GLP-1 products were excluded: four from Douyin and one from Bilibili. The videos promoted products using the GLP-1 name, such as herbal teas, slimming products, and unbranded liquids. The researchers did not report laboratory testing of those products.
Conclusions
The study found substantial differences in the measured quality and completeness of GLP-1 receptor agonist information between the sampled videos from Douyin and Bilibili. The sampled Bilibili videos were longer and scored higher on two measures of reliability and transparency, whereas neither platform consistently provided high-quality medical information.
Longer videos were associated with higher quality scores, while engagement metrics did not consistently indicate information quality. The identification of videos promoting counterfeit GLP-1 products represented an additional public health concern.
Because the study captured a single point in time and used different sampling methods for the two platforms, its findings cannot establish that either platform makes videos more reliable. It also did not assess viewers’ understanding or subsequent health behavior.
The authors called for stronger platform moderation, greater participation by healthcare professionals, and digital health-literacy initiatives. They also recommended longitudinal research and evaluation of interventions designed to improve the quality of online health information.
Further research is needed to clarify how these factors affect the quality of health information.
Journal reference:
- Peng, G., Wang, C., Zhang, H.-W., Xu, T., & Di, J.-Z. (2026). Assessing the quality and reliability of health information on GLP-1 receptor agonists on TikTok and Bilibili: A cross-sectional analysis of the public health infodemic. Therapeutic Advances in Endocrinology and Metabolism, 17. DOI: 10.1177/20420188261467887, https://journals.sagepub.com/doi/10.1177/20420188261467887
