Figures
Abstract
This study examines how Chinese small and medium-sized cities gain visibility and sustain place-based meaning within Douyin’s platformized short-video ecology. Focusing on Shaoguan, Guangdong Province, it analyzes 30 high-interaction Douyin videos published between May 2023 and May 2025 and applies crisp-set Qualitative Comparative Analysis (csQCA) to identify configurations associated with high dissemination performance. The findings suggest that city-related hashtags function as a necessary platform-indexing condition in the sampled archive, while emotional resonance and narrative structure strengthen specific high-impact configurations. Two communicative logics are particularly salient: non-local creators can generate rapid visibility through immersive, fragmented, and affective presentations, whereas locally embedded storytelling can reinforce cultural resilience and place identity. The study contributes to digital urban communication by integrating the Heuristic-Systematic Model with debates on platformization, algorithmic visibility, and digital place branding.
Citation: Li F, Peng X (2026) Chinese small and medium-sized city image communication on Douyin: Algorithmic heuristics and cultural resilience in short-video ecologies. PLoS One 21(7): e0354468. https://doi.org/10.1371/journal.pone.0354468
Editor: Cheong Kim, Dong-A University College of Business Administration, KOREA, REPUBLIC OF
Received: April 23, 2026; Accepted: July 8, 2026; Published: July 23, 2026
Copyright: © 2026 Li, Peng. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All relevant data are available from Zenodo at https://doi.org/10.5281/zenodo.17084333.
Funding: This work was supported by a doctoral start-up project (Grant No. R20058) and a university-level humanities and social sciences project (Grant No. C20121) at Guangdong Ocean University, China. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
1. Introduction
China’s ongoing urbanization has heightened the strategic role of small and medium-sized cities (SMCs) in regional development, cultural continuity, and public communication. At the same time, short-video platforms such as Douyin have transformed city-image communication by embedding urban representation within recommendation systems, creator economies, hashtag publics, and platform-specific attention metrics. Under network media logic, visibility is no longer produced only by institutional messaging; it is co-produced by users, interfaces, algorithms, and circulating cultural cues [1–3]. This makes the study of SMCs on Douyin timely because peripheral and non-metropolitan cities must compete for attention in an environment where discoverability and recognizability are increasingly platform-governed.
Existing research has provided valuable accounts of city branding, multimodal urban representation, and short-video engagement, including recent work on city-image promotion short videos and Chinese urban imaginaries [4–7]. However, two gaps remain. First, many studies treat the city as a destination brand or a communication object, but pay less attention to how platform affordances structure the conditions under which small and medium-sized cities become visible. Second, single-factor explanations cannot fully explain why different combinations of hashtags, creator identity, emotion, narrative, and visual design may lead to similar dissemination outcomes.
This study addresses these gaps by examining Shaoguan, a prefecture-level city in northern Guangdong Province, as a theory-informed case of SMC communication. Shaoguan is analytically useful because it possesses recognizable cultural and ecological resources, including Danxia Mountain, Hakka culture, Nanhua Temple, and local foodways, but does not enjoy the automatic symbolic legibility of China’s major metropolitan centers. The case therefore enables us to ask how local cultural resources are translated into platform-recognizable signals and how such signals interact with deeper narratives of cultural resilience.
The article makes three contributions. Theoretically, it connects the Heuristic-Systematic Model (HSM) with platformization and digital place-branding scholarship to conceptualize a cue-narrative mechanism of Douyin-based city communication. Methodologically, it uses csQCA to model conjunctural causality and equifinality among 30 high-interaction videos. Empirically, it identifies how algorithmic heuristics and cultural narratives combine to produce high dissemination performance, thereby clarifying how SMCs can pursue short-term visibility without reducing place identity to purely traffic-oriented signs.
2. Theoretical background
2.1 City media image and place branding
Communication has long shaped how urban identities are imagined. Foundational urban semiotics highlights the role of visual and symbolic elements--landmarks, architecture, and spatial legibility--in structuring public imaginaries of the city [8]. Building on this tradition, place-branding scholarship stresses culture, tourism, civic identity, and stakeholder meaning-making as core dimensions of city brands [4,5,9,10]. Tourism and audiovisual communication research similarly shows that place images are constructed through organizational narratives, promotional mediation, and affective destination imaginaries [11–12]. With the diffusion of short-video platforms in China, city branding is increasingly driven by user-generated content and continuous micro-storytelling rather than top-down campaigns: affect-laden cues and everyday scenes travel quickly, shaping audience evaluations and shareability in digital streams [13–14].
2.2 Evolution of city image communication
Classical models framed communication as a linear process--‘who says what in which channel to whom with what effect’ [15]--and municipal promotions long relied on broadcast-style videos and print. In today’s China, platformized short-video communication reconfigures production and circulation: algorithmic feeds couple with participatory authorship, shifting from networked publics to interest-clustered publics and from institutional media logic to network media logic [1,3,13,16,17]. Platformization further means that cultural content, creator practices, and datafied governance are reorganized around platform infrastructures and metrics [2]. Chinese evidence suggests that multimodal videos articulate city history/modernity pairings and help municipal actors and creators co-construct attention, with specific content features linked to dissemination outcomes [6–7]. These developments motivate our configurational approach.
2.3 Cognitive psychology and the Heuristic–Systematic Model
The Heuristic-Systematic Model (HSM) explains how audiences process persuasive information via low-effort heuristics (e.g., source/format cues) and high-effort systematic processing (content-driven meaning-making) [18–19]. In digital environments, technological affordances trigger recognizable heuristics, while credibility judgments frequently rely on cognitive shortcuts [20–21]. Research on algorithmic visibility further shows that creators and influencers negotiate platform cues, attention economies, and visibility labor rather than merely publishing content into a neutral media space [22–23]. Applied to Douyin, we conceptualize hashtags and cover cues as algorithmic heuristics orienting attention, and emotional resonance and narrative structure as systematic cues sustaining identification. This cue-narrative alignment underpins our case design and variable selection.
2.4 Research propositions
Based on the preceding literature, three theoretical propositions guide the analysis.
- Proposition 1: In platformized short-video environments, city-related hashtags and visual packaging function as algorithmic heuristics that increase the likelihood of initial visibility for small and medium-sized cities.
- Proposition 2: High dissemination is more likely when heuristic cues are combined with systematic cultural resources, especially emotional resonance and narrative structure, because such combinations link rapid recognizability with deeper place meaning.
- Proposition 3: Creator positionality shapes how local cultural resources are translated into platform visibility: non-local creators can generate novelty and comparative attention, whereas local creators can strengthen embedded knowledge and cultural resilience.
3. Methods
To improve transparency, this section is organized around two linked components: the research procedure and sample, and the research tools used to analyze the archive. The first component explains why Shaoguan and why a bounded set of high-interaction Douyin videos is appropriate for the research question. The second component explains how HSM and csQCA are combined to identify configurational pathways rather than isolated net effects.
3.1 Research procedure and sample
The study follows a bounded case-archive design. Shaoguan was selected because it is a Chinese small and medium-sized city with recognizable cultural and ecological resources but comparatively limited national media visibility. The sample consists of 30 high-interaction Douyin videos about Shaoguan published between May 2023 and May 2025. The archive was constructed through purposive, cross-account keyword retrieval using ‘Shaoguan,’ ‘Shaoguan tourism,’ ‘Danxia Mountain,’ and ‘Hakka culture.’ This sampling strategy is not intended to represent all Shaoguan-related videos on Douyin statistically; rather, it captures videos that had already entered the platform’s attention circuit and is therefore appropriate for examining how SMC visibility is configured once dissemination has occurred.
3.2 Research tools
Two research tools structure the analysis. First, HSM provides the conceptual distinction between heuristic cues and systematic meaning-making. In China’s short-video ecology, hashtags, covers, creator badges, and duration operate as algorithmic heuristics that orient attention and retrieval, whereas emotional resonance, narrative coherence, and camera language support more systematic interpretation. Second, csQCA provides a set-theoretic tool for identifying necessity, sufficiency, conjunctural causality, and equifinality in a medium-N archive. We implemented the analysis with fsQCA (Windows v4.1) and report necessity, sufficiency, calibration, PRI-consistency, and robustness checks following established QCA guidance [24–25].
3.3 Sample documentation and representativeness
To reduce personalization bias, searches were conducted across new accounts, institutional accounts, and UGC accounts. The final archive represents a theoretically relevant subset of Shaoguan videos that achieved relatively high interaction, not a probability sample of all Douyin content. Its representativeness is therefore analytical rather than statistical: the cases reflect the visible segment of Shaoguan’s platform communication and allow comparison across creator identity, hashtag use, narrative form, affective cues, and visual design. This scope also defines the limits of generalization. The findings can inform comparable SMC communication on short-video platforms, but they should not be read as population-level estimates of all city-image videos. The selected cases and case identifiers are listed in Table 1.
3.4 Research conditions and outcome
This study employs the Heuristic–Systematic Model (HSM) to specify how multimodal cues in Douyin videos shape digital urban communication for Shaoguan, Guangdong. In HSM, heuristic processing relies on low-effort cues, whereas systematic processing engages high-effort evaluation of message content [18–19]. In China’s short-video ecology, we further treat platform-visible signals (e.g., hashtags, covers, creator badges, duration) as algorithmic heuristics and content-level elements (e.g., emotional resonance, narrative structure, camera language) as supports for systematic processing. This framing is consistent with research on Douyin/TikTok’s parallel platformization and Chinese short-video narratives [3,13]. We therefore model eight conditions: four heuristic (video duration, hashtags, cover visual design, creator attributes) and four systematic (content theme, emotional resonance, narrative structure, camera language).
Coding follows QCA best practices for necessity/sufficiency analysis in medium-N designs [24–25]. Two trained coders independently coded all cases; disagreements were resolved by discussion. By jointly operationalizing heuristic and systematic conditions, we tailor HSM to Chinese urban-communication contexts on Douyin rather than a generic “Asian” frame.
3.4.1 Condition 1: Video duration.
Shorter clips reduce cognitive effort and favor heuristic uptake, whereas longer clips allow deeper meaning‑making—consistent with HSM’s motivation/ability account [18–19]. Empirically, duration is a salient content feature associated with engagement composition in short‑video feeds [26].
We code duration as a crisp set relative to the one‑minute norm of short‑form video: videos with runtime ≤ 60 seconds are coded 0 (Short Duration), and > 60 seconds coded 1 (Non‑Short Duration). This boundary aligns with common platform practices while leaving content‑level effects to be captured by other conditions.
3.4.2 Condition 2: Hashtags.
Hashtags act as algorithmic and cultural signposts that index topics, guide recommendation, and provide quick semantic cues, thereby functioning as heuristics in Chinese short‑video contexts [3,13].
We record whether a video contains at least one city‑related or campaign/trending hashtag (e.g., #Shaoguan, #DanxiaMountain). Presence = 1; absence = 0. We also note tag specificity in Table 2 but retain a binary indicator for csQCA.
3.4.3 Condition 3: Cover visual design.
Covers (thumbnails) are first‑impression cues. Salience, iconicity, and compositional clarity prime attention and expectations; these design factors shape downstream judgments and are well‑documented in visual/branding research [19,27].
We identify whether the cover features iconic/place‑specific elements (e.g., Danxia Mountain silhouette, landmark typography) with clear, high‑contrast composition. Presence of iconic/clear cover = 1; otherwise = 0. Examples and counter‑examples are shown in Table 2.
3.4.4 Condition 4: Creator profile.
Creator profile conditions authenticity and novelty: local creators embed dialects and customs that foster identification, while non-local creators provide outsider perspectives that may generate short-term visibility--patterns observed in Douyin/TikTok scholarship and influencer video research [3,13,28].
We code creator profile using self‑descriptions and geotags: Local creator = 0; Non‑local creator = 1. Ambiguous cases (e.g., travel vloggers without declared origin) are adjudicated via profile/history and documented in Table 2.
3.4.5 Condition 5: Content theme.
Themes structure cognitive processing: information‑rich or composite themes stimulate systematic processing; Chinese multimodal work shows city videos braid modernity and heritage to widen appeal [6,19].
We operationalize this condition by distinguishing single-focus themes from composite themes based on whether one domain dominates (e.g., only scenery) or multiple domains co‑occur (e.g., scenery + cuisine + history). Following the existing calibration used downstream, Single theme = 1; Composite theme = 0 (see Table 2).
3.4.6 Condition 6: Emotional resonance.
Affective cues (awe, pride, nostalgia) increase sharing propensity and sustain identification; narrative‑emotion research further shows that emotional flow supports persuasive outcomes [14,29].
We code Emotion = 1 when videos contain explicit affective displays or cues (e.g., affect‑laden narration, expressive voice/music, on‑screen reactions) that foreground joy/pride/awe/nostalgia; otherwise 0. Coding anchors and examples are specified in Table 2.
3.4.7 Condition 7: Narrative structure.
Fragmented vignettes align with mobile scanning (rapid, low‑effort decoding), whereas coherent storylines support identification and distinctiveness—consistent with narrative/encoding theory [16,30–33].
To preserve consistency with downstream calibration, we code Fragmented narrative = 1 (e.g., montages, slogan‑driven cuts without an integrating arc); Coherent/structured narrative = 0. Decision rules and examples are detailed in Table 2.
3.4.8 Condition 8: Camera language.
Visual technique affects informational yield and engagement in city‑video contexts [7,34]. Distinctive techniques (e.g., drones, rapid edits) can signal production value; ordinary styles can index authenticity.
In line with the analysis that follows, we code Ordinary camera style = 1 (e.g., static/handheld with minimal effects) and Distinctive/professional techniques = 0 (e.g., aerials, hyperlapse). See Table 2 for illustrations.
3.4.9 Outcome: Digital content influence.
Short‑video dissemination is tightly coupled with user engagement on the platform. A substantial body of work on Douyin/short‑video research treats likes, comments, saves (favorites), and shares as core, observable indicators of a video’s influence and audience participation [35–36].
In industry practice, Qingbo Intelligence’s Douyin leaderboard operationalizes communication performance with the Douyin Communication Influence Index (DCI, V1.0), defined as Publishing Index (10%) + Engagement Index (76%) + Reach Index (14%); within the Engagement Index, likes (17%)/ comments (37%)/ shares (46%) are further weighted, underscoring the primacy of engagement in dissemination assessment [37]. This emphasis aligns with platform‑side mechanisms: Douyin’s personalized recommendation system explicitly considers behavioral signals such as likes, comments, shares/forwards, favorites, and watch time to predict interest and amplify exposure [38]. Empirical studies also observe that early accumulation of likes, comments, saves, and shares increases the likelihood of subsequent recommendations and broader reach, producing viral‑like diffusion patterns [39].
To ensure replicability using publicly accessible Douyin video-level data, this study focuses on engagement and constructs an author-defined index, DCI*, as a tractable proxy for a single video’s dissemination influence. Because the archive consists of already visible videos, the low-outcome category refers to relatively lower dissemination performance within the sampled high-interaction archive, rather than to low visibility on Douyin as a whole.
The weighting gives greater emphasis to immediately visible engagement signals, especially likes and comments, while retaining saves and shares as indicators of more deliberate engagement. This author-defined index is not intended to replicate Qingbo’s official DCI, but to provide a transparent and reproducible proxy based on publicly visible video-level data. Platform disclosures and algorithm-audit evidence jointly support the premise that engagement signals are central to recommendation-driven exposure [38–39].
3.5 Measurement and coding reliability
We binarized all conditions using a 0.50 membership threshold, consistent with set-theoretic logic and reporting conventions in csQCA [24–25]. Thresholds were anchored in the operational definitions reported in Table 2; cases at or above the stated boundary were coded 1 and the remaining cases were coded 0. Two trained coders independently coded the full sample using an HSM-informed codebook after pilot calibration. Intercoder reliability reached Cohen’s κ = 0.85, indicating substantial agreement [40]. Disagreements were resolved through adjudication and minor clarification of coding rules, and the reconciled dataset was used for downstream QCA.
3.6 Ethics statement
This study did not involve prospective recruitment, interventions, interviews, surveys, experiments, or access to private human-participant records. It analyzed publicly available Douyin short-video content, public account self-presentations/geotags used only to determine creator profile, hashtags, and publicly visible video-level engagement metrics. No private communications, restricted-access information, sensitive personal data, or non-public identifiers were collected. Informed consent was not obtained because the study did not recruit or interact with human participants, and no attempt was made to identify individuals beyond public account names or pseudonyms already displayed on the platform
4. QCA data analysis for small and medium-sized city image dissemination
4.1 Construction of the truth table
Following best practices for crisp-set Qualitative Comparative Analysis (csQCA), we calibrated eight conditions—video duration, hashtag use, cover visual design, creator profile, content theme, emotional resonance, narrative structure, and camera language—and one outcome, digital content influence, for 30 high-interaction Douyin videos about Shaoguan (May 2023–May 2025). Each case received binary membership (1 = presence, 0 = absence) according to the calibration rules in Table 2, and the resulting truth table (Table 3) maps condition configurations to the outcome, enabling the identification of causal pathways to effective city-image dissemination in Chinese small and medium-sized cities [24,25].
4.2 Necessity analysis of single conditions
The necessity analysis indicates that Hashtag is the only condition exceeding the conventional 0.90 consistency threshold. In the sampled archive, all high-dissemination cases contain city-related hashtags, whereas no high-dissemination case occurs without them. Given the high prevalence of hashtags in the sample, this finding should be interpreted as a necessary platform-indexing condition rather than as a standalone driver of dissemination performance. Full necessity results are reported in Table 4.
Table 5 and Fig 1 further confirm that no high-dissemination case lacks city-related hashtags, supporting the interpretation of hashtag use as a necessary platform-indexing condition in this sample.
4.3 Sufficiency analysis of condition configurations
To uncover configurational pathways leading to high dissemination performance of Shaoguan’s city image on Douyin, we analyzed the calibrated dataset with fsQCA 4.1 and constructed a truth table for cross-case comparison. Following best practice for small-N designs, we set frequency cutoff = 1 to retain rare but potentially informative combinations, and consistency cutoff = .80, a widely used benchmark in QCA applications [25,41]. Because Hashtag was established as necessary in Section 4.2, we fixed Hashtag = 1 during minimization; the remaining seven conditions (duration, cover, creator profile, theme, emotion, narrative, camera) were treated as dichotomous.
We estimated the complex, parsimonious, and intermediate solutions. The complex solution preserves all observed configurations without minimization, reflecting the multidimensionality of short-video communication. The parsimonious solution yields highly simplified expressions but risks overlooking context. We adopt the intermediate solution—which incorporates theory-informed directional expectations while excluding implausible counterfactuals—as it best balances empirical fit and theoretical plausibility [24,42,43]. The intermediate solution returns seven sufficient pathways (see Table 6), with solution coverage = 0.60 and solution consistency = 1.00.
To guard against spurious findings, we report PRI-consistency (Proportional Reduction in Inconsistency), which gauges the extent to which each configuration avoids simultaneously covering negative cases [25]. All retained configurations achieve PRI ≥ 0.80, comfortably above the conventional 0.65 threshold for sufficiency claims in configurational research [44].
Substantive interpretation. Two observations stand out.
First, the pathway with the largest unique coverage (Duration * Hashtag * Visuals * Profile * ~ Theme * Emotion * Narrative * Camera Language) indicates an immersive diffusion pattern. In this configuration, longer duration, iconic covers, non-local creators, composite themes, affective resonance, and fragmented storytelling combine to make Shaoguan legible and engaging within mobile feeds. This result supports the view that platform visibility is produced by bundles of cues rather than by any single content feature [3,7,17].
Second, smaller-coverage pathways show equifinality. Coherent narration can substitute for distinctive cinematography in some single-theme videos, while affective and visual cues can compensate for weaker narrative integration in others. The findings therefore point to multiple viable cue-narrative combinations for SMC city-image dissemination.
From an HSM perspective, fixing Hashtag = 1 in all sufficient paths confirms the heuristic role of platform indexing, while emotion and narrative structure explain how meaning and identification are sustained. The results thus connect algorithmic visibility with cultural interpretation rather than treating dissemination as a purely technical outcome.
4.4 Robustness inspection
To assess the reliability of the intermediate solution (solution coverage = 0.60, solution consistency = 1.00), we implemented multiple robustness checks recommended for small-N csQCA. First, we raised the consistency cutoff from 0.80 to 0.90 [25]. The re-estimated truth table preserved the core pathways, indicating that our findings are not sensitive to moderate calibration tightening. Second, we conducted a random case-deletion test by re-running the model on a 25-case subsample; the pathway with the largest unique coverage (Path 7) remained, while several very low-coverage paths disappeared as expected under reduced diversity. Third, we excluded the four cases with Emotion = 0 (13.3% of N = 30); Path 7 again persisted, underscoring emotional resonance as a stable contributor to high dissemination performance.
Consistent with best practice for datasets of this size, we set frequency cutoff = 1 to retain rare but potentially informative configurations; raising it to 2 would unduly prune empirically meaningful combinations in a 30-case design [24,25,41]. To avoid spurious sufficiency claims driven by simultaneous subset relations, we report PRI-consistency in addition to raw consistency [25]. As shown in Table 6, all retained configurations achieve PRI ≥ 0.80, comfortably above the commonly used 0.65 benchmark in configurational research [44].
Given eight conditions and 30 cases, limited diversity is expected—many logically possible configurations remain unobserved [24,25]. In the intermediate solution, we therefore borrowed a subset of logical remainders under theory-guided directional expectations (DEs) [42,43]: configurations with HASHTAG = 1 and EMOTION = 1 were assigned positive directional expectations, consistent with the HSM view of heuristic (hashtags) and systematic (emotion) synergy. Sensitivity tests that relaxed DEs for Narrative or Camera left the core pathway intact, indicating that our conclusions are not an artifact of a particular remainder-borrowing scheme. (See Table 7.)
Taken together--threshold elevation, random resampling, targeted case exclusion, PRI reporting, and transparent treatment of logical remainders--these checks support the relative stability of the identified pathways across calibration changes, sample variation, and alternative counterfactual assumptions.
5. The salient high-impact pathway: Immersive diffusion
Path 7 (Duration * Hashtag * Visuals * Profile * ~ Theme * Emotion * Narrative * Camera Language)—the pathway with the largest unique coverage—features longer videos (>60s), Shaoguan-related hashtags, iconic covers, non-local creators, composite themes, affective resonance, fragmented storytelling, and ordinary camera language. We interpret this as an “immersive diffusion” pattern (see Fig 2): longer duration enables multi-scene montage that compensates for non-distinctive cinematography; composite themes raise cognitive variety yet are held together by emotion and bite-sized narrative beats. This is consistent with dual-process accounts in which heuristic cues (e.g., tags, thumbnails) capture attention, and systematic elements (emotion, narrative structure) sustain meaning [18,19].
Two illustrative Douyin videos embody this immersive diffusion. One by user Los Angeles Yingzheng showcases Danxia Mountain and local cuisine through fragmented yet humor-laden sequences, conveying Shaoguan’s natural and cultural appeal. Another by Xu Ge Yuan interweaves scenes of regional culinary heritage, Danxia’s landscapes, and Shaoguan’s historic East Street into a sensory collage that stimulates sensory-rich viewer engagement. These narratives, aided by emotional resonance, leverage algorithmic visibility to enhance dissemination potential. In both, hashtags foreground topical indexing while emotion and fragmentation sustain scrolling-context engagement--an arrangement aligned with recent work on content logics in Douyin/TikTok’s parallel platformization [3].
5.1 The “Traffic Attraction–Cultural Resilience” dual-helix model
Within China’s platformized short-video ecology, city-image communication hinges on a dual logic: short-term attention capture (algorithmic exposure via heuristic cues) and long-term cultural resilience (identity-rooted meaning via systematic processing). Applying csQCA to Shaoguan shows how small and medium-sized cities can balance fleeting visibility with enduring identity. This dual-helix model, illustrated in Fig 3, reconciles creator-driven ephemeral attention with deeper narrative rootedness. It offers a potentially transferable framework for comparable small and medium-sized cities in platformized short-video environments.
5.2 Traffic attraction: Algorithm-driven dissemination and emotional engagement (see Fig 4)
Content embedded with event/locale-specific hashtags plus affective cues aligns with Douyin’s recommendation logic. Tags (e.g., #ShaoguanTravelGuide) act as algorithmic heuristics that index topics and audiences, while affective cues sustain viewer engagement and facilitate place-based identification. Empirical research documents tags’ privileged role within Douyin/TikTok’s platformization, and feature-level analyses of city-image videos show how content attributes map onto engagement bundles [3,7,13]. These dynamics turn ephemeral attention into durable city-image communication through word-of-mouth and networked sharing.
Case evidence suggests that decomposing a city image into recognizable cognitive units—iconic landmarks and everyday practices—accelerates reach. Yet surface imagery alone risks low retention; layering emotion and micro-narratives mitigates this risk and strengthens identification.
5.3 Cultural resilience: Deep narratives and cognitive engagement (See Fig 5)
Beyond traffic, cultural resilience depends on layered storytelling and meaning reconstruction. Non-local creators often act as “cultural decoders,” reframing urban narratives through outsider lenses; local creators anchor embedded knowledge and coherent narratives that reinforce place identity. Multimodal work on Chinese promotional videos similarly shows that deep narrative structures, sometimes with professional camera language, sustain long-tail engagement and place-based identification [6,7]. For SMCs, resilience is not only a creative task but also institutional curation: e.g., building a “narrative resource bank” to systematize local knowledge into reusable modules for short-video storytelling.
6. Discussion
The findings provide empirical grounding for the three theoretical propositions developed in this study, while also refining them in configurational terms. Shaoguan’s high-dissemination Douyin videos are not driven by a single universal formula. Instead, visibility emerges from configurations that combine algorithmic heuristics with affective, narrative, and positional resources. Proposition 1 is refined by the finding that hashtag presence operates as a necessary platform-indexing condition in the sampled archive, although its moderate coverage suggests that it should be interpreted as enabling rather than sufficient on its own. Visual packaging works less as an isolated driver than as a platform-addressability cue that becomes effective within broader content configurations. Proposition 2 receives clearer configurational support: high dissemination performance is more likely when rapid recognizability is coupled with emotional resonance and culturally meaningful narrative structures. Proposition 3 is also specified: non-local creators can generate novelty, contrast, and external attention, whereas local creators can strengthen embedded knowledge, cultural continuity, and place-based interpretation. Taken together, these findings show that short-video visibility for small and medium-sized cities depends not merely on attention capture, but on the coupling of traffic-oriented cues with culturally grounded meaning-making.
This also explains why Shaoguan’s visibility on Douyin should not be understood simply as the result of algorithmic amplification. Hashtags and visual cues help local symbols enter searchable and recommendable attention flows, but they do not automatically produce sustained communication effects. Dissemination becomes stronger when platform cues are paired with emotional resonance, creator positioning, and narrative structures that help viewers recognize Shaoguan as both a destination and a culturally meaningful place. In this sense, platform visibility is not only a technical outcome of indexing and recommendation, but also a cultural process through which place meanings are selected, packaged, circulated, and reinterpreted.
6.1 Theoretical implications
The study contributes to theory in three ways. First, it extends the heuristic-systematic model from a general model of persuasive processing to a platform-specific account of city-image communication. In short-video ecologies, heuristic cues such as hashtags, covers, visual symbols, and recognizable scenes help content gain initial attention, while systematic resources such as affective resonance, narrative organization, and cultural explanation sustain deeper place recognition. This extension shows that HSM can be used not only to explain individual information processing, but also to conceptualize how platform-native visibility is organized through the interaction between quick cues and culturally meaningful content.
Second, the study connects digital place branding with platformization research by demonstrating that city brands are not simply authored by municipal institutions. Rather, they are co-produced through creator labor, platform indexing, user engagement, and algorithmic visibility [2,4,5]. For small and medium-sized cities, this is especially important because their visibility often depends less on large-scale official campaigns than on dispersed creator practices and platform-mediated circulation. The case of Shaoguan shows that city image is increasingly shaped through a hybrid process in which official cultural resources, local memories, outsider curiosity, and platform logics are intertwined.
Third, the configurational findings move beyond linear explanations by showing that different cue-narrative-position bundles can produce similar dissemination outcomes. This is especially important for resource-constrained small and medium-sized cities because it suggests that city-image visibility does not depend on one optimal communication formula. Instead, different combinations of platform heuristics, creator positionality, emotional resonance, and narrative organization can generate comparable levels of digital content influence. The broader theoretical implication is that short-video city communication should be understood through a configurational logic: what matters is not whether a single factor is present, but how multiple cues and cultural resources are assembled into a recognizable and circulable form.
These implications are captured by the Traffic Attraction-Cultural Resilience dual-helix model proposed in this study. The traffic-attraction dimension refers to the platform-facing work of making the city searchable, recognizable, and recommendable. The cultural-resilience dimension refers to the place-facing work of preserving, renewing, and narratively organizing local cultural meaning. The two dimensions are not mutually exclusive. Effective city communication on short-video platforms requires their interaction: cities must enter attention flows through platform-native cues, but they must also sustain place identity through affective and narrative depth.
6.2 Practical implications
The findings offer practical implications for municipal communicators, cultural institutions, and local creators in China.
- (1) Standardize a hashtag taxonomy that combines official cultural identifiers (e.g., #ShaoguanHakka) with topical/trending descriptors to improve algorithmic addressability [3,13].
- (2) Co-produce with creator diversity: pair non-local creators (novelty/contrast) with local creators (embedded knowledge) to avoid echo chambers and widen reach [3].
- (3) Modularize narrative resources: curate a narrative resource bank (heritage episodes, symbolic landmarks, everyday practices) and couple emotion + micro-story beats to convert exposure into recognition [6–7].
- (4) Align design with attention ecology: use iconic covers and clear visual cues to capture attention during scrolling, then sustain attention with coherent story arcs consistent with the dual-helix logic.
For small and medium-sized cities, the practical lesson is therefore not simply to imitate the traffic strategies of major cities or popular tourist destinations. Instead, they need to build a communication system that links platform readability with local cultural depth. A successful short-video strategy should make local culture easy to find, easy to recognize, and easy to circulate, while also ensuring that the city is not reduced to a set of superficial visual symbols.
6.3 Limitations and future research
Several limitations should be acknowledged. First, DCI* is an author-defined composite indicator based on publicly visible engagement metrics. Although it is grounded in platform-disclosure logic and short-video research, alternative weighting schemes may produce different sensitivity patterns. Future research could compare different outcome functions and weighting strategies to test the stability of the results.
Second, the study focuses on one city and one platform. Shaoguan provides a useful case for examining the platform visibility of a non-metropolitan Chinese city, but the findings cannot be automatically generalized to all small and medium-sized cities. Future studies could replicate the model across multiple Chinese cities, different regional cultures, and different types of urban identity.
Third, the archive captures videos that had already achieved relatively high interaction. Therefore, the analysis explains configured visibility within an existing attention circuit rather than predicting virality for all city-related videos. Future research could include lower-engagement cases or longitudinal data to examine how videos move from initial exposure to wider dissemination.
Fourth, this study uses crisp-set QCA to identify interpretable configurations. This approach is appropriate for clarifying the presence or absence of key conditions, but it may compress fine-grained differences among cases. Future research could compare crisp-set results with fuzzy-set calibrations, alternative thresholds, or mixed-method designs to examine the sensitivity of the Traffic Attraction-Cultural Resilience dual-helix model.
Future research can extend the analysis in four directions: replicating the model across multiple Chinese small and medium-sized cities and subcultures; comparing Douyin with platforms such as Bilibili, Xiaohongshu, or WeChat Channels; connecting online engagement with offline indicators such as tourist footfall, bookings, or cultural-event participation; and testing alternative calibration thresholds and outcome functions to examine the sensitivity of the dual-helix model.
6.4 Conclusion
This article examined how Chinese small and medium-sized cities can achieve city-image visibility on Douyin. By integrating the heuristic-systematic model with crisp-set QCA, implemented with fsQCA 4.1 for Windows, it identified hashtag presence as a necessary platform-indexing condition and uncovered seven sufficient pathways to high digital content influence. Returning to the three theoretical propositions, the findings show that platform heuristics are enabling rather than sufficient conditions: hashtags and visual packaging help local cultural symbols enter algorithmic attention flows, but high dissemination performance depends on how these cues are configured with emotional resonance, narrative organization, and creator positionality.
The most salient pathway, immersive diffusion, combines algorithmic readability with affective and fragmented storytelling, suggesting that platform visibility is produced through both immediate recognizability and culturally meaningful engagement. The study’s broader contribution is the Traffic Attraction-Cultural Resilience dual-helix model, which explains how short-video city communication can balance immediate platform traffic with longer-term cultural meaning. For small and medium-sized cities, the central challenge is not simply to be seen, but to be represented in ways that preserve and renew place identity beyond momentary exposure.
Code availability
No custom programming scripts, macros, or author-generated software were used. The fsQCA analyses were performed using fsQCA 4.1 for Windows. The author-defined DCI* outcome score was calculated using transparent Excel formulas embedded in the deposited dataset: DCI* = 0.40 × Likes + 0.30 × Comments + 0.20 × Saves (Favorites) + 0.10 × Shares.
References
- 1. Klinger U, Svensson J. The emergence of network media logic in political communication: a theoretical approach. New Media Soc. 2014;17(8):1241–57.
- 2. Poell T, Nieborg DB, van Dijck J. Platformisation. Internet Policy Rev. 2019;8(4):1–13.
- 3. Kaye DBV, Chen X, Zeng J. The co-evolution of two Chinese mobile short video apps: parallel platformization of Douyin and TikTok. Mobile Media Commun. 2020;9(2):229–53.
- 4. Sevin HE. Understanding cities through city brands: city branding as a social and semantic network. Cities. 2014;38:47–56.
- 5. Zenker S, Braun E. Questioning a ‘one size fits all’ city brand: developing a branded house strategy for place brand management. J Place Manag Dev. 2017;10(3):270–87.
- 6. Wang Y, Feng WD. History, modernity, and city branding in China: a multimodal critical discourse analysis of Xi’an’s promotional videos on social media. Soc Semiot. 2023;33(2):402–25.
- 7. He J, Yang Z, Zhu J. A framework for visualizing and describing city image promotion short video data based on microcube model. PLoS One. 2025;20(4):e0317883. pmid:40208870
- 8.
Lynch K. The image of the city. Cambridge (MA): MIT Press; 1960.
- 9. Anholt S. The Anholt-GMI city brands index: how the world sees the world’s cities. Place Brand Public Dipl. 2006;2(1):18–31.
- 10. Kavaratzis M, Hatch MJ. The dynamics of place brands: an identity-based approach to place branding theory. Mark Theory. 2013;13(1):69–86.
- 11.
Elías-Zambrano R, Jiménez-Marín G. Reflections on organizational communication, advertising, and audiovisual communication from a multidisciplinary perspective. Madrid: Fragua; 2021.
- 12. Jiménez-Marín G, Correia P, Medina IG. Análisis del impacto turístico de la organización de bodas en la zona del Caribe. J Tour Dev. 2021;37:89–109.
- 13. Chen C, Kaye DBV, Zeng J. #PositiveEnergy Douyin: constructing ‘playful patriotism’ in a Chinese short video application. Chin J Commun. 2021;14(1):97–117.
- 14. Berger J, Milkman KL. What makes online content viral? J Mark Res. 2012;49(2):192–205.
- 15.
Lasswell HD. The structure and function of communication in society. In: Bryson L, editor. The communication of ideas. New York: Harper & Brothers; 1948. p. 37–51.
- 16.
Jenkins H. Convergence culture: where old and new media collide. New York: New York University Press; 2006.
- 17. Gerbaudo P. TikTok and the algorithmic transformation of social media publics: from social networks to social interest clusters. New Media Soc. 2024;28(3):1019–36.
- 18. Chaiken S. Heuristic versus systematic information processing and the use of source versus message cues in persuasion. J Pers Soc Psychol. 1980;39(5):752–66.
- 19.
Eagly AH, Chaiken S. The psychology of attitudes. Fort Worth (TX): Harcourt Brace Jovanovich; 1993.
- 20.
Sundar SS. The MAIN model: a heuristic approach to understanding technology effects on credibility. In: Metzger MJ, Flanagin AJ, editors. Digital media, youth, and credibility. Cambridge (MA): MIT Press; 2008. p. 73–100.
- 21. Metzger MJ, Flanagin AJ. Credibility and trust of information in online environments: the use of cognitive heuristics. J Pragmat. 2013;59:210–20.
- 22. Cotter K. Playing the visibility game: how digital influencers and algorithms negotiate influence on Instagram. New Media Soc. 2018;21(4):895–913.
- 23. Abidin C. Mapping internet celebrity on TikTok: exploring attention economies and visibility labours. Cult Sci J. 2021;12(1):77–103.
- 24.
Ragin CC. Redesigning social inquiry: fuzzy sets and beyond. Chicago: University of Chicago Press; 2008.
- 25.
Schneider CQ, Wagemann C. Set-theoretic methods for the social sciences: a guide to qualitative comparative analysis. Cambridge: Cambridge University Press; 2012.
- 26. Lu S, Yu M, Wang H. What matters for short videos’ user engagement: a multiblock model with variable screening. Expert Syst Appl. 2023;218:119542.
- 27. Ghorbani M, Westermann A. Exploring the role of packaging in the formation of brand images: a mixed methods investigation of consumer perspectives. JPBM. 2024;34(2):186–202.
- 28. Yang J, Zhang J, Zhang Y. Engagement that sells: influencer video advertising on TikTok. Mark Sci. 2025;44(2):247–67.
- 29. Nabi RL, Green MC. The role of a narrative’s emotional flow in promoting persuasive outcomes. Media Psychol. 2014;18(2):137–62.
- 30. Bruner J. The narrative construction of reality. Crit Inq. 1991;18(1):1–21.
- 31.
Hall S. Encoding/decoding. In: Hall S, Hobson D, Lowe A, Willis P, editors. Culture, media, language: working papers in cultural studies, 1972-79. London: Hutchinson; 1980. p. 128–38.
- 32.
Ryan M. Narrative across media: the languages of storytelling. Lincoln (NE): University of Nebraska Press; 2004.
- 33.
Couldry N, Hepp A. The mediated construction of reality. Cambridge: Polity Press; 2017.
- 34. Hartmann MC, Purves RS. Seeing through a new lens: exploring the potential of city walking tour videos for urban analytics. Int J Digit Earth. 2023;16(1):2555–73.
- 35. Lu Y, (Cindy) Shen C. Unpacking multimodal fact-checking: features and engagement of fact-checking videos on Chinese TikTok (Douyin). Soc Media Soc. 2023;9(1):20563051221150406.
- 36. Wang C, Li Z. Unraveling the relationship between audience engagement and audiovisual characteristics of automotive green advertising on Chinese TikTok (Douyin). PLoS One. 2024;19(4):e0299496. pmid:38573890
- 37. Qingbo Intelligence. 抖音号传播力指数 DCI (V1.0) 指标说明 [Douyin communication influence index DCI (V1.0) indicator description] [Internet]; 2025 [cited 2025 Sep 19]. Available from: https://www.gsdata.cn/site/usage-16
- 38. Douyin. “抖音”算法及模型备案公示说明 [Algorithm and model filing transparency note] [Internet]. [cited 2025 Sep 19]. Available from: https://lf3-cdn-tos.draftstatic.com/obj/ies-hotsoon-draft/douyin_agreement/70c3d13a-73cf-403a-8ecf-a16f70887c21.html
- 39. Shi W, Li J. New digital divide shaped by algorithm? Evidence from agent-based testing on Douyin’s health-related video recommendation. Commun Res. 2024;51(7):867–90.
- 40. Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–74. pmid:843571
- 41. Thomann E, Maggetti M. Designing research with qualitative comparative analysis (QCA): approaches, challenges, and tools. Sociol Methods Res. 2017;49(2):356–86.
- 42. Fiss PC. Building better causal theories: a fuzzy set approach to typologies in organization research. Acad Manage J. 2011;54(2):393–420.
- 43.
Rihoux B, Ragin CC, editors. Configurational comparative methods: qualitative comparative analysis (QCA) and related techniques. Thousand Oaks (CA): SAGE; 2009.
- 44. Greckhamer T, Furnari S, Fiss PC, Aguilera RV. Studying configurations with qualitative comparative analysis: best practices in strategy and organization research. Strateg Organ. 2018;16(4):482–95.