Table 1.
Subset of the SOC classification hierarchy.
Fig 1.
The distribution of yearly income for the users in our dataset.
The red dotted line represents the mean.
Table 2.
Description of the user level features.
Table 3.
Prediction of income with our groups of features.
Pearson correlation (left columns) and Mean Average Error (right columns) between income and our models on 10 fold cross-validation using three different regression methods: Linear regression (LR), Support Vector Machines with RBF kernel (SVM) and Gaussian Processes (GP) and sets of features described in the User Features section.
Fig 2.
Mean income with confidence intervals for psycho-demographic groups.
All group mean differences are statistically significant (Mann-Whitney test, p < .001).
Fig 3.
Linear and non-linear (GP) fit for Profile features.
Variation of income as a function of user profile features. Linear fit in red, non-linear Gaussian Process fit in black. Brackets show the GP lengthscales—the lower the value, the more important the feature is for prediction.
Fig 4.
Linear and non-linear (GP) fit for emotions and sentiments.
Variation of income as a function of user emotion and sentiment scores. Linear fit in red, non-linear Gaussian Process fit in black. Brackets show the GP lengthscales—the lower the value, the more important the feature is for prediction.
Fig 5.
Linear and non-linear (GP) fit for shallow textual features.
Variation of income as a function of user shallow textual features. Linear fit in red, non-linear Gaussian Process fit in black. Brackets show the GP lengthscales—the lower the value, the more important the feature is for prediction.
Table 4.
Topics, represented by top 15 words, sorted by their ARD lengthscale.
Most predictive topics for income. Topic labels are manually added. Lower lengthscales (l) denote more predictive topics.
Fig 6.
Linear and non-linear (GP) fit for topics.
Variation of income as a function of user topic usage. Linear fit in red, non-linear Gaussian Process fit in black.