Fig 1.
Different Albanian dialects in western Balkans.
Table 1.
The textual statistics for DialectCorpus.
Fig 2.
List of assumptions, proposed solutions, and related sub-chapters.
Fig 3.
The visual description of the methodology used for using Twitter to collect a multi-dialectal corpus of Albanian.
Fig 4.
Displayed above is a bipartite-directed graph illustrating the relationships between standard Twitter users and hub Twitter users.
Accompanying this graphical representation a tabular dataset is shown, which provides a structured depiction of the connections present within the aforementioned bipartite-directed graph.
Table 2.
The summary of the network parameters.
Table 3.
Classification report.
Fig 5.
The overall language filtering workflow.
Table 4.
The F-1 score for each training algorithm, two dataset modes, and two TF-IDF modeling levels.
Fig 6.
The classification reports for experiments with Twitterer-level models.
Fig 7.
The classification reports for experiments with tweet-level classification models.
Fig 8.
The learning curve for SVC trained on Twitterer-level TF-IDF modeling.
Fig 9.
The learning curve for SVC trained on individual tweet-level TF-IDF modeling.
Fig 10.
The class prediction error analysis for Tweeterer-level model.
Fig 11.
The class prediction error analysis for Tweet-level models.
Fig 12.
The most archetypal features.
Table 5.
The accuracy rates and macro F1 scores of four human subjects that were asked to annotate 300 tweets and 30 users with dialectal labels.
The last row indicates the average values computed on all four annotators.
Fig 13.
A sum of dialect identification confusion matrices computed on all annotators included in the human evaluation study.
Table 6.
Sentences that have been labeled with ground-truth labels for the dialect identification task, along with the labels assigned by humans, are provided along with their corresponding English translations for improved understanding.
Table 7.
Two representative tweet sets extracted from the corpus: Each corresponding to the Twitterer from KS and the MK dialects/regions.
Fig 14.
Depiction of waterfall SHAP visualizations for a pair of exemplary observations derived from Table 7 in a comparative manner for MK dialect.
Fig 15.
Depiction of waterfall SHAP visualizations for a pair of exemplary observations derived from Table 7 in a comparative manner for KS dialect.
Fig 16.
Depiction of waterfall SHAP visualizations for a pair of exemplary observations derived from Table 7 in a comparative manner for AL dialect.