Table 1.
Likelihood ratio test compares models 1 and 2 only. Models 1 and 3 could not be compared due to different numbers of observations. Differences in AIC and BIC highlight different model aspects. For both AIC and BIC the lower scores are reflective of better model fit. BIC highlights tags as the driver of prediction in the model. AIC identifies the improved predictive power when using keywords with tags. LRT also identifies this relationship. Note that model fit improves dramatically when applied to data with greater degree of heterogeneity. All of San Francisco (model 4) and all of New York City (model 5) represent greater data heterogeneity than our original pilot study consisting of Chinese restaurants in San Francisco (models 1–3). This is seen when applied to datasets representing the entirety of San Francisco and the entirety of New York City.
Table 2.
An analysis of correlation between tags and keywords was used to decide which terms would be included in the model.
A correlation cutoff of .05 was used for inclusion in the model unless the authors strongly believed the keyword would be useful in the model despite low correlation (e.g., high quality, food poisoning, and employees, selected for relation to food quality, foodborne illness, and employee behavior). A liberal cut off point was used to include as many predictors as possible. Correlation is specific to pilot study training data which excluded all but Chinese restaurants.
Table 3.
Principal components analysis dimensions were set as covariates in a logistic regression model to show the predictive effect of each dimension on the outcome of receiving a health score <80.
The confidence intervals show that as keywords add weight to dimensions those dimensions are associated with corresponding increase or decrease in odds of low health score within the stated confidence interval.
Fig 1.
Receiver Operator Curve using validation dataset created using Yelp data compiled from 220 San Francisco restaurants.
Fig 2.
Positive Predictive Value using validation dataset created using Yelp data compiled from 220 San Francisco restaurants.
Fig 3.
Receiver Operator Curve created using Yelp data compiled from 1,543 San Francisco restaurants including all cuisine types.
Fig 4.
Receiver Operator Curve created using Yelp data compiled from 745 New York City restaurants including all cuisine types.
Table 4.
(Significant Predictors in San Francisco Model).
Odds Ratio and 95% Confidence Interval of Odds Ratio are listed above for predictive keywords and tag in the San Francisco model. Table is limited to significant predictors. Additional terms that were highly predictive but not identified as significant due to collinearity are not listed in this table.
Table 5.
(Significant Predictors in New York City Model).
Odds Ratio and 95% Confidence Interval of Odds Ratio are listed above for predictive keywords in the New York City model. Table is limited to significant predictors. Additional terms that were highly predictive but not identified as significant due to collinearity are not listed in this table.
Table 6.
Training Data refers to the first part of the dataset used for model creation.
Validation Data refers to the second part of the dataset used for validation purposes. Simulated data is a 200 restaurant random sampling and analysis repeated over 10,000 iterations using the complete dataset from the pilot study. “Prevalence” refers to the prevalence of restaurants with low health code rating in the specific dataset.
Fig 5.
Plot of observed and predicted prevalence of low health code rating over a two year period from the beginning of 2013 to the end of 2014 using validation dataset for first year and for second year.
Blue lines reflect predicted counts and red lines reflect the observed counts of restaurants with health code rating <80 (substandard).
Fig 6.
Plot of observed and predicted prevalence of low health code rating over a fifteen month period from the beginning of 2014 to the end of 2015 using sample from all restaurants in New York City.
Blue lines reflect predicted counts and red lines reflect the observed counts of restaurants with health code rating >14 (substandard).