Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Table 1.

Subset of the SOC classification hierarchy.

More »

Table 1 Expand

Fig 1.

The distribution of yearly income for the users in our dataset.

The red dotted line represents the mean.

More »

Fig 1 Expand

Table 2.

Description of the user level features.

More »

Table 2 Expand

Table 3.

Prediction of income with our groups of features.

Pearson correlation (left columns) and Mean Average Error (right columns) between income and our models on 10 fold cross-validation using three different regression methods: Linear regression (LR), Support Vector Machines with RBF kernel (SVM) and Gaussian Processes (GP) and sets of features described in the User Features section.

More »

Table 3 Expand

Fig 2.

Mean income with confidence intervals for psycho-demographic groups.

All group mean differences are statistically significant (Mann-Whitney test, p < .001).

More »

Fig 2 Expand

Fig 3.

Linear and non-linear (GP) fit for Profile features.

Variation of income as a function of user profile features. Linear fit in red, non-linear Gaussian Process fit in black. Brackets show the GP lengthscales—the lower the value, the more important the feature is for prediction.

More »

Fig 3 Expand

Fig 4.

Linear and non-linear (GP) fit for emotions and sentiments.

Variation of income as a function of user emotion and sentiment scores. Linear fit in red, non-linear Gaussian Process fit in black. Brackets show the GP lengthscales—the lower the value, the more important the feature is for prediction.

More »

Fig 4 Expand

Fig 5.

Linear and non-linear (GP) fit for shallow textual features.

Variation of income as a function of user shallow textual features. Linear fit in red, non-linear Gaussian Process fit in black. Brackets show the GP lengthscales—the lower the value, the more important the feature is for prediction.

More »

Fig 5 Expand

Table 4.

Topics, represented by top 15 words, sorted by their ARD lengthscale.

Most predictive topics for income. Topic labels are manually added. Lower lengthscales (l) denote more predictive topics.

More »

Table 4 Expand

Fig 6.

Linear and non-linear (GP) fit for topics.

Variation of income as a function of user topic usage. Linear fit in red, non-linear Gaussian Process fit in black.

More »

Fig 6 Expand