This research develops a hybrid Hawkes–Transformer framework to forecast user engagement and to analyze competing information cascades on social media. The empirical setting is Sina Weibo. The dataset comprises 29,316 posts nested within 3,662 trending topics collected over a two-week period in 2024. The study links interpretable point-process estimates to a Transformer model by using Hawkes-estimated parameters as structured priors and time-varying inputs, with the goal of improving predictive accuracy while retaining interpretability.
The framework addresses two interrelated problems: (1) how different engagement behaviors—retweets, comments, and likes—interact over time to shape diffusion within a post, and (2) how multiple concurrent posts on the same topic compete for limited attention, producing parallel cascades that may fragment or reinforce information flow. The data and supporting information are reported in the paper and supplementary files.
The first component of the hybrid approach is a multivariate Hawkes point process that is retweet-centric. This estimation captures temporal dependencies and cross-behavioral effects among engagement types. Specifically, the Hawkes model quantifies self-excitation—how prior retweets increase the short-term hazard of future retweets—and cross-excitation or suppression between behaviors such as comments and likes.
Using the Hawkes formulation allows the authors to estimate interpretable diffusion parameters that reflect how each engagement type contributes to the intensity of subsequent events. These estimated parameters provide a compact, theory-grounded summary of intra-post dynamics that can be examined directly for signs of amplification or diversion of attention.
The second component is a Transformer-based deep learning model used for forecasting future engagement. Hawkes-estimated parameters are integrated into the Transformer as structured priors and as time-varying features. This design purposefully combines the theoretical transparency of point-process modeling with the sequence-learning and predictive capacity of a Transformer.
By embedding Hawkes-derived inputs, the Transformer can leverage both domain-informed parameter estimates and flexible attention mechanisms to model complex temporal patterns. The hybrid pipeline therefore aims to improve forecast performance relative to either method alone while maintaining a link to interpretable diffusion mechanisms.
Key empirical results highlight distinct roles for different engagement behaviors. The analysis indicates that retweets strongly amplify diffusion through self-excitation: past retweets increase the likelihood of subsequent retweets, creating cascading growth in visibility. By contrast, comments often act to suppress diffusion when they accumulate or cluster, interpreted as diverting user attention away from resharing toward discussion—this suppression effect can reduce the intensity of retweet-driven spread.
Likes are included in the multivariate framework as an engagement signal that interacts with other behaviors; the study treats likes, comments, and retweets as interdependent processes but emphasizes the retweet-centric nature of the Hawkes estimation. The findings stress that treating engagement behaviors in isolation risks missing these interdependencies and may mischaracterize how attention is allocated on platforms.
Beyond intra-post dynamics, the study examines competition among parallel cascades—multiple posts grouped under the same trending topic. The empirical evidence reported shows that concurrent posts generally compete for scarce attention and tend to fragment engagement rather than synergistically reinforce diffusion across posts.
This fragmentation implies that algorithmic aggregations (for example, trending lists or hashtag pages) produce environments in which multiple content streams vie for visibility. The competitive dynamics uncovered have implications for platform curation, content strategy, and interventions aimed at steering attention during high-stakes events such as crisis communication or product launches.
The hybrid Hawkes–Transformer approach contributes methodologically by demonstrating how interpretable, theory-driven parameters can be embedded into deep-learning architectures as priors and structured inputs. Empirically, the combination enables more accurate forecasting of engagement while offering direct measures of behavioral interactions (e.g., self-excitation of retweets, suppressive influence of comments).
Conceptually, the work highlights that attention on algorithmically mediated platforms is both scarce and contestable: early-stage engagement matters for subsequent visibility, comments and moderation choices can alter diffusion trajectories, and timing strategies must consider cross-post competition within topics. The authors argue that integrating statistical rigor (point-process modeling) with machine learning foresight (Transformer prediction) offers a pathway for transparent, actionable forecasting of digital attention dynamics.
All relevant data are reported within the paper and its supporting information files. The study received support from the National Natural Science Foundation of China (grant number 72472069) awarded to Y.G.; the funder had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. The article was received December 2, 2025, accepted July 8, 2026, and published July 24, 2026 in PLOS ONE. The authors declared no competing interests.