-
TV-watching partner robot: Analysis of User's Experience
Authors:
Donghuo Zeng,
Jianming Wu,
Gen Hattori,
Yasuhiro Takishima
Abstract:
Watching TV not only provides news information but also gives an opportunity for different generations to communicate. With the proliferation of smartphones, PC, and the Internet, increase the opportunities for communication in front of the television is also likely to diminish. This has led to some problems further from face-to-face such as a lack of self-control and insufficient development of c…
▽ More
Watching TV not only provides news information but also gives an opportunity for different generations to communicate. With the proliferation of smartphones, PC, and the Internet, increase the opportunities for communication in front of the television is also likely to diminish. This has led to some problems further from face-to-face such as a lack of self-control and insufficient development of communication skills. This paper proposes a TV-watching companion robot with open-domain chat ability. The robot contains two modes: TV-watching mode and conversation mode. In TV-watching mode, the robot first extracts keywords from the TV program and then generates the disclosure utterances based on the extracted keywords as if enjoying the TV program. In the conversation mode, the robot generates question utterances with keywords in the same way and then employs a topics-based dialog management method consisting of multiple dialog engines for rich conversations related to the TV program. We conduct the initial experiments and the result shows that all participants from the three groups enjoyed talking with the robot, and the question about their interests in the robot was rated 6.5/7-levels. This indicates that the proposed conversational features of TV-watching Companion Robot have the potential to make our daily lives more enjoyable. Under the analysis of the initial experiments, we achieve further experiments with more participants by dividing them into two groups: a control group without a robot and an intervention group with a robot. The results show that people prefer to talk to robots because the robot will bring more enjoyable, relaxed, and interesting.
△ Less
Submitted 2 March, 2023; v1 submitted 28 February, 2023;
originally announced February 2023.
-
Topic-switch adapted Japanese Dialogue System based on PLATO-2
Authors:
Donghuo Zeng,
Jianming Wu,
Yanan Wang,
Kazunori Matsumoto,
Gen Hattori,
Kazushi Ikeda
Abstract:
Large-scale open-domain dialogue systems such as PLATO-2 have achieved state-of-the-art scores in both English and Chinese. However, little work explores whether such dialogue systems also work well in the Japanese language. In this work, we create a large-scale Japanese dialogue dataset, Dialogue-Graph, which contains 1.656 million dialogue data in a tree structure from News, TV subtitles, and Wi…
▽ More
Large-scale open-domain dialogue systems such as PLATO-2 have achieved state-of-the-art scores in both English and Chinese. However, little work explores whether such dialogue systems also work well in the Japanese language. In this work, we create a large-scale Japanese dialogue dataset, Dialogue-Graph, which contains 1.656 million dialogue data in a tree structure from News, TV subtitles, and Wikipedia corpus. Then, we train PLATO-2 using Dialogue-Graph to build a large-scale Japanese dialogue system, PLATO-JDS. In addition, to improve the PLATO-JDS in the topic switch issue, we introduce a topic-switch algorithm composed of a topic discriminator to switch to a new topic when user input differs from the previous topic. We evaluate the user experience by using our model with respect to four metrics, namely, coherence, informativeness, engagingness, and humanness. As a result, our proposed PLATO-JDS achieves an average score of 1.500 for the human evaluation with human-bot chat strategy, which is close to the maximum score of 2.000 and suggests the high-quality dialogue generation capability of PLATO-2 in Japanese. Furthermore, our proposed topic-switch algorithm achieves an average score of 1.767 and outperforms PLATO-JDS by 0.267, indicating its effectiveness in improving the user experience of our system.
△ Less
Submitted 22 February, 2023;
originally announced February 2023.
-
Learning Explicit and Implicit Latent Common Spaces for Audio-Visual Cross-Modal Retrieval
Authors:
Donghuo Zeng,
Jianming Wu,
Gen Hattori,
Yi Yu,
Rong Xu
Abstract:
Learning common subspace is prevalent way in cross-modal retrieval to solve the problem of data from different modalities having inconsistent distributions and representations that cannot be directly compared. Previous cross-modal retrieval methods focus on projecting the cross-modal data into a common space by learning the correlation between them to bridge the modality gap. However, the rich sem…
▽ More
Learning common subspace is prevalent way in cross-modal retrieval to solve the problem of data from different modalities having inconsistent distributions and representations that cannot be directly compared. Previous cross-modal retrieval methods focus on projecting the cross-modal data into a common space by learning the correlation between them to bridge the modality gap. However, the rich semantic information in the video and the heterogeneous nature of audio-visual data leads to more serious heterogeneous gaps intuitively, which may lead to the loss of key semantic content of video with single clue by the previous methods when eliminating the modality gap, while the semantics of the categories may undermine the properties of the original features. In this work, we aim to learn effective audio-visual representations to support audio-visual cross-modal retrieval (AVCMR). We propose a novel model that maps audio-visual modalities into two distinct shared latent subspaces: explicit and implicit shared spaces. In particular, the explicit shared space is used to optimize pairwise correlations, where learned representations across modalities capture the commonalities of audio-visual pairs and reduce the modality gap. The implicit shared space is used to preserve the distinctive features between modalities by maintaining the discrimination of audio/video patterns from different semantic categories. Finally, the fusion of the features learned from the two latent subspaces is used for the similarity computation of the AVCMR task. The comprehensive experimental results on two audio-visual datasets demonstrate that our proposed model for using two different latent subspaces for audio-visual cross-modal learning is effective and significantly outperforms the state-of-the-art cross-modal models that learn features from a single subspace.
△ Less
Submitted 26 October, 2021;
originally announced October 2021.
-
SHECS: A Local Smart Hands-free Elderly Care Support System on Smart AR Glasses with AI Technology
Authors:
Donghuo Zeng,
Jianming Wu,
Bo Yang,
Tomohiro Obara,
Akeri Okawa,
Nobuko Iino,
Gen Hattori,
Ryoichi Kawada,
Yasuhiro Takishima
Abstract:
Some elderly care homes attempt to remedy the shortage of skilled caregivers and provide long-term care for the elderly residents, by enhancing the management of the care support system with the aid of smart devices such as mobile phones and tablets. Since mobile phones and tablets lack the flexibility required for laborious elderly care work, smart AR glasses have already been considered. Althoug…
▽ More
Some elderly care homes attempt to remedy the shortage of skilled caregivers and provide long-term care for the elderly residents, by enhancing the management of the care support system with the aid of smart devices such as mobile phones and tablets. Since mobile phones and tablets lack the flexibility required for laborious elderly care work, smart AR glasses have already been considered. Although lightweight smart AR devices with a transparent display are more convenient and responsive in an elderly care workplace, fetching data from the server through the Internet results in network congestion not to mention the limited display area. To devise portable smart AR devices that operate smoothly, we first present a no keep alive Internet required smart hands-free elderly care support system that employs smart glasses with facial recognition and text-to-speech synthesis technologies. Our support system utilizes automatic lightweight facial recognition to identify residents, and information about each resident in question can be obtained hands free link with a local database. Moreover, a resident information can be displayed on just a portion of the AR smart glasses on the spot. Due to the limited size of the display area, it cannot show all the necessary information. We exploit synthesized voices in the system to read out the elderly care related information. By using the support system, caregivers can gain an understanding of each resident condition immediately, instead of having to devote considerable time in advance in obtaining the complete information of all elderly residents. Our lightweight facial recognition model achieved high accuracy with fewer model parameters than current state-of-the-art methods. The validation rate of our facial recognition system was 99.3% or higher with the false accept rate of 0.001, and caregivers rated the acceptability at 3.6 (5 levels) or higher.
△ Less
Submitted 26 October, 2021;
originally announced October 2021.
-
A non-ordinary peridynamics implementation for anisotropic materials
Authors:
Gabriel Hattori,
Jon Trevelyan,
William M. Coombs
Abstract:
Peridynamics (PD) represents a new approach for modelling fracture mechanics, where a continuum domain is modelled through particles connected via physical bonds. This formulation allows us to model crack initiation, propagation, branching and coalescence without special assumptions. Up to date, anisotropic materials were modelled in the PD framework as different isotropic materials (for instance,…
▽ More
Peridynamics (PD) represents a new approach for modelling fracture mechanics, where a continuum domain is modelled through particles connected via physical bonds. This formulation allows us to model crack initiation, propagation, branching and coalescence without special assumptions. Up to date, anisotropic materials were modelled in the PD framework as different isotropic materials (for instance, fibre and matrix of a composite laminate), where the stiffness of the bond depends on its orientation. A non-ordinary state-based formulation will enable the modelling of generally anisotropic materials, where the material properties are directly embedded in the formulation. Other material models include rocks, concrete and biomaterials such as bones. In this paper, we implemented this model and validated it for anisotropic composite materials. A composite damage criterion has been employed to model the crack propagation behaviour. Several numerical examples have been used to validate the approach, and compared to other benchmark solution from the finite element method (FEM) and experimental results when available.
△ Less
Submitted 18 October, 2017;
originally announced October 2017.