Fig 1.
Conceptual diagram of the uncanny valley in the voice field.
Adapted from “The Uncanny Valley,” by M. Mori, 1970. Conceptual diagram of the theoretical graph presented in the original uncanny valley theory. X-axis corresponds to similarity between robots and humans and y-axis corresponds to familiarity of the robots. Recorded voice may represent the valley part and own voice the highest point after the valley. Sense of one’s self instead of similarity was used in the present study.
Fig 2.
Schematic of the task. After the presentation of stimuli, participants chose which of the stimuli sounded more like own-voice by button press.
Fig 3.
Individual results of pairwise comparison. The bar represents the similarity to own voice, rightmost represents the most own-voice like and leftmost represents the least own-voice like rating. The numbers on the top-half of the bar represents the result of the second session and the ones on bottom-half of the bar are the results of the third session. The numbers are for types of conditions: 1) Recorded voice, 2) Step filtered voice, 3) Bandpass filtered voice, 4) Lowpass filtered voice, 5) Adjusted voice.
Fig 4.
Consistency of own-voice rating across trials. The consistency of the most and the least own-voice like rating is presented. Blue represents the number of participants who rated both the most and least own voice-like voice consistently, orange represents the number of participants who rated only the least own voice-like sound consistently, yellow represents the number of participants who rated only the most voice-like sound consistently, and gray represents the number of participants who rated both the most and the least own voice-like sound inconsistently.
Fig 5.
Individual results of pairwise comparison. The bar represents the similarity to own voice, rightmost as the most own-voice like and leftmost as the least own-voice like rating. There were two non-lip synchronization sessions conducted and the results are presented as the numbers on the top bar. The numbers on the bottom bar represent the results of the lip synchronization session. The numbers are the types of conditions: 1) Recorded voice, 2) Step-filtered voice, 3) Bandpass filtered voice, 4) Lowpass filtered voice, 5) Adjusted voice.
Fig 6.
Participant own voice rating consistency across days. The consistency of own voice-like rating across participants is charted. Blue represents the number of participants who rated both the most and least own voice-like voice consistently, orange represents least choice consistency only, yellow represent most choice consistency only, and gray represents inconsistency for both the most and least own voice-like voice.
Fig 7.
Schematic of the experimental task. After the presentation of stimulus, the participant rated the stimulus in terms of the presented feature from one to nine by moving a cursor.
Fig 8.
Results of voice features scoring. The X-axis represents sense of oneself, y-axis represents familiarity for A and eeriness for B. Each individual score is plotted as green dots. The dotted line shows the Pearson’s correlation and the solid represents the cubic equation.