Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

Different Albanian dialects in western Balkans.

More »

Fig 1 Expand

Table 1.

The textual statistics for DialectCorpus.

More »

Table 1 Expand

Fig 2.

List of assumptions, proposed solutions, and related sub-chapters.

More »

Fig 2 Expand

Fig 3.

The visual description of the methodology used for using Twitter to collect a multi-dialectal corpus of Albanian.

More »

Fig 3 Expand

Fig 4.

Displayed above is a bipartite-directed graph illustrating the relationships between standard Twitter users and hub Twitter users.

Accompanying this graphical representation a tabular dataset is shown, which provides a structured depiction of the connections present within the aforementioned bipartite-directed graph.

More »

Fig 4 Expand

Table 2.

The summary of the network parameters.

More »

Table 2 Expand

Table 3.

Classification report.

More »

Table 3 Expand

Fig 5.

The overall language filtering workflow.

More »

Fig 5 Expand

Table 4.

The F-1 score for each training algorithm, two dataset modes, and two TF-IDF modeling levels.

More »

Table 4 Expand

Fig 6.

The classification reports for experiments with Twitterer-level models.

More »

Fig 6 Expand

Fig 7.

The classification reports for experiments with tweet-level classification models.

More »

Fig 7 Expand

Fig 8.

The learning curve for SVC trained on Twitterer-level TF-IDF modeling.

More »

Fig 8 Expand

Fig 9.

The learning curve for SVC trained on individual tweet-level TF-IDF modeling.

More »

Fig 9 Expand

Fig 10.

The class prediction error analysis for Tweeterer-level model.

More »

Fig 10 Expand

Fig 11.

The class prediction error analysis for Tweet-level models.

More »

Fig 11 Expand

Fig 12.

The most archetypal features.

More »

Fig 12 Expand

Table 5.

The accuracy rates and macro F1 scores of four human subjects that were asked to annotate 300 tweets and 30 users with dialectal labels.

The last row indicates the average values computed on all four annotators.

More »

Table 5 Expand

Fig 13.

A sum of dialect identification confusion matrices computed on all annotators included in the human evaluation study.

More »

Fig 13 Expand

Table 6.

Sentences that have been labeled with ground-truth labels for the dialect identification task, along with the labels assigned by humans, are provided along with their corresponding English translations for improved understanding.

More »

Table 6 Expand

Table 7.

Two representative tweet sets extracted from the corpus: Each corresponding to the Twitterer from KS and the MK dialects/regions.

More »

Table 7 Expand

Fig 14.

Depiction of waterfall SHAP visualizations for a pair of exemplary observations derived from Table 7 in a comparative manner for MK dialect.

More »

Fig 14 Expand

Fig 15.

Depiction of waterfall SHAP visualizations for a pair of exemplary observations derived from Table 7 in a comparative manner for KS dialect.

More »

Fig 15 Expand

Fig 16.

Depiction of waterfall SHAP visualizations for a pair of exemplary observations derived from Table 7 in a comparative manner for AL dialect.

More »

Fig 16 Expand