Fig 1.
Illustration of the synonym relationship types.
Fig 2.
Proposed method main steps and workflow.
Fig 3.
Text graph for the tweet replies text.
This graph has 193 vertices (34 vertices are singletons), 572 edges, and 45 connected components. 112 edges represent direct synonym relationships, and 460 represent indirect synonym relationships. The size of each vertex reflects its frequency in the original text. Thick and thin edges indicate direct and indirect synonym relationships, respectively.
Table 1.
High contributor singleton vertices in the text graph and their frequencies.
Table 2.
Summary of the qualities of the 22 main communities in the text graph in Fig 3.
Table 3.
Synonym relationships and clustering coefficients (CC) of four communities in the text graph.
The numbers represent relationship strengths (a strength of 1 is assigned between direct synonyms and 0.5 is assigned between indirect synonyms).
Table 4.
Keywords identified by three human extractors.
Table 5.
Performance comparison of the proposed method with TextRank and YAKE.
Table 6.
Statistical and text graph data of each dataset.
Number of words and number of tokens denote the number of words in the dataset before and after preprocessing respectively. Direct edges and indirect edges represent the number of direct and indirect synonym relationships between words in the text graph respectively.
Table 7.
Comparison of multiple keyword extraction methods.
Table 8.
Statistical and text graph data of each abstract in the HULTH dataset.
Number of words and number of tokens denote the number of words in the dataset before and after preprocessing respectively. Direct edges and indirect edges represent the number of direct and indirect synonym relationships between words in the text graph respectively.
Table 9.
Comparison of multiple keyword extraction methods against human annotators.