Conversational agents in the form of virtual agents or social robots are rapidly becoming wide-spread. Humans use non-verbal behaviors to signal their intent, emotions and attitudes in human-human interactions. Conversational agents therefore need this ability as well in order to make an interaction pleasant and efficient. An important part of non-verbal communication is gesticulation: gestures communicate a large share of non-verbal content. Previous systems for gesture production were typically rule-based and could not represent the range of human gestures. Recently the gesture generation field has shifted to data-driven approaches. We follow this line of research by extending the state-of-the-art deep-learning based model. Our model leverages representation learning to enhance speech-gesture mapping. We provide analysis of different representations for the input (speech) and the output (motion) of the network by both objective and subjective evaluations. We also analyze the importance of smoothing of the produced motion and emphasize how challenging it is to evaluate gesture quality. In the future we plan to enrich input signal by taking semantic context (text transcription) as well, make the model probabilistic and evaluate our system on the social robot NAO.
My talk will be divided into two parts. In the first part, I will analyze Nesterov's accelerated gradient method from a dynamical systems point of view. More precisely, I will derive the accelerated gradient method by discretizing an ordinary differential equation with a semi-implicit Euler integration scheme. I will analyze both the ordinary differential equation and the discretization for obtaining insights into the phenomenon of acceleration. In particular, geometric properties of the dynamics, such as asymptotic stability, time-reversibility, and phase-space volume contraction are shown to be preserved through the discretization. In the second part, I will show that these geometric properties are enough for characterizing the convergence rate. The results therefore provide criteria that are easily verifiable for the accelerated convergence of any momentum-based optimization algorithm. The results also yield guidance for the design of new optimization algorithms. The talk will focus on unconstrained optimization problems with smooth and strongly-convex objective functions, even though the analysis potentially generalizes to non-convex or non-Euclidean settings, or when the decision variables are constrained to a smooth manifold.
Organizers: Sebastian Trimpe
Fingertip skin friction plays a critical role during object manipulation. We will describe a simple and reliable method to estimate the fingertip static coefficient of friction (CF) continuously and quickly during object manipulation, and we will describe a global expression of the CF as a function of the normal force and fingertip moisture. Then we will show how skin hydration modifies the skin deformation dynamics during grip-like contacts. Certain motor behaviours observed during object manipulation could be explained by the effects of skin hydration. Then the biomechanics of the partial slip phenomenon will be described, and we will examine how this partial slip phenomenon is related to the subjective perception of fingertip slip.
Future cities and infrastructure systems will evolve into complex conglomerates where autonomous aerial, aquatic and ground-based robots will coexist with people and cooperate in symbiosis. To create this human-robot ecosystem, robots will need to respond more flexibly, robustly and efficiently than they do today. They will need to be designed with the ability to move across terrain boundaries and physically interact with infrastructure elements to perform sensing and intervention tasks. Taking inspiration from nature, aerial robotic systems can integrate multi-functional morphology, new materials, energy-efficient locomotion principles and advanced perception abilities that will allow them to successfully operate and cooperate in complex and dynamic environments. This talk will describe the scientific fundamentals, design principles and technologies for the development of biologically inspired flying robots with adaptive morphology that can perform monitoring and manufacturing tasks for future infrastructure and building systems. Examples will include flying robots with perching capabilities and origami-based landing systems, drones for aerial construction and repair, and combustion-based jet thrusters for aerial-aquatic vehicles.
Organizers: Metin Sitti
Today’s robots have motor abilities and sensors that exceed those of humans in many ways: They move more accurately and faster; their sensors see more and at a higher precision and in contrast to humans they can accurately measure even the smallest forces and torques. Robot hands with three, four, or five fingers are commercially available, and, so are advanced dexterous arms. Indeed, modern motion-planning methods have rendered grasp trajectory generation a largely solved problem. Still, no robot to date matches the manipulation skills of industrial assembly workers despite that manipulation of mechanical objects remains essential for the industrial assembly of complex products. So, why are current robots still so bad at manipulation and humans so good?
Organizers: Katherine J. Kuchenbecker
Human shape estimation is an important task for video editing, animation and fashion industry. Predicting 3D human body shape from natural images, however, is highly challenging due to factors such as variation in human bodies, clothing and viewpoint. Prior methods addressing this problem typically attempt to fit parametric body models with certain priors on pose and shape. In this work we argue for an alternative representation and propose BodyNet, a neural network for direct inference of volumetric body shape from a single image. BodyNet is an end-to-end trainable network that benefits from (i) a volumetric 3D loss, (ii) a multi-view re-projection loss, and (iii) intermediate supervision of 2D pose, 2D body part segmentation, and 3D pose. Each of them results in performance improvement as demonstrated by our experiments. To evaluate the method, we fit the SMPL model to our network output and show state-of-the-art results on the SURREAL and Unite the People datasets, outperforming recent approaches. Besides achieving state-of-the-art performance, our method also enables volumetric body-part segmentation.
For many service robots, reactivity to changes in their surroundings is a must. However, developing software suitable for dynamic environments is difficult. Existing robotic middleware allows engineers to design behavior graphs by organizing communication between components. But because these graphs are structurally inflexible, they hardly support the development of complex reactive behavior. To address this limitation, we propose Playful, a software platform that applies reactive programming to the specification of robotic behavior. The front-end of Playful is a scripting language which is simple (only five keywords), yet results in the runtime coordinated activation and deactivation of an arbitrary number of higher-level sensory-motor couplings. When using Playful, developers describe actions of various levels of abstraction via behaviors trees. During runtime an underlying engine applies a mixture of logical constructs to obtain the desired behavior. These constructs include conditional ruling, dynamic prioritization based on resources management and finite state machines. Playful has been successfully used to program an upper-torso humanoid manipulator to perform lively interaction with any human approaching it.
Human footsteps can provide a unique behavioural pattern for robust biometric systems. Traditionally, security systems have been based on passwords or security access cards. Biometric recognition deals with the design of security systems for automatic identification or verification of a human subject (client) based on physical and behavioural characteristics. In this talk, I will present spatio-temporal raw and processed footstep data representations designed and evaluated on deep machine learning models based on a two-stream resnet architecture, by using the SFootBD database the largest footstep database to date with more than 120 people and almost 20,000 footstep signals. Our models deliver an artificial intelligence capable of effectively differentiating the fine-grained variability of footsteps between legitimate users (clients) and impostor users of the biometric system. We provide experimental results in 3 critical data-driven security scenarios, according to the amount of footstep data available for model training: at airports security checkpoints (smallest training set), workspace environments (medium training set) and home environments (largest training set). In these scenarios we report state-of-the-art footstep recognition rates.
Organizers: Dimitrios Tzionas
Animals are widespread in nature and the analysis of their shape and motion is of importance in many fields and industries. Modeling 3D animal shape, however, is difficult because the 3D scanning methods used to capture human shape are not applicable to wild animals or natural settings. In our previous SMAL model, we learn animal shape from toys figurines, but toys are limited in number and realism, and not every animal is sufficiently popular for there to be realistic toys depicting it. What is available in large quantities are images and videos of animals from nature photographs, animal documentaries, and webcams. In this talk I will present our recent work for capturing the detailed 3D shape of animals from images alone. Our method extracts significantly more 3D shape detail than previous work and is able to model new species using only a few video frames. Additionally, we extract realistic texture map from images for capturing both animal shape and appearance.
In academic and policy circles, there has been considerable interest in the impact of “big data” on firm performance. We examine the question of how the amount of data impacts the accuracy of Machine Learned models of weekly retail product forecasts using a proprietary data set obtained from Amazon. We examine the accuracy of forecasts in two relevant dimensions: the number of products (N), and the number of time periods for which a product is available for sale (T). Theory suggests diminishing returns to larger N and T, with relative forecast errors diminishing at rate 1/sqrt(N) + 1/sqrt(T) . Empirical results indicate gains in forecast improvement in the T dimension; as more and more data is available for a particular product, demand forecasts for that product improve over time, though with diminishing returns to scale. In contrast, we find an essentially flat N effect across the various lines of merchandise: with a few exceptions, expansion in the number of retail products within a category does not appear associated with increases in forecast performance. We do find that the firm’s overall forecast performance, controlling for N and T effects across product lines, has improved over time, suggesting gradual improvements in forecasting from the introduction of new models and improved technology.
My plan is to present the motivation behind Deep GPs as well as some of the current approximate inference schemes available with their limitations. Then, I will explain how Deep GPs fit into the BayesOpt framework and the specific problems they could potentially solve.
Organizers: Philipp Hennig
Political science is integrating computational methods like machine learning into its own toolbox. At the same time the awareness rises that the utilization of machine learning algorithms in our daily life is a highly political issue. These two trends - the integration of computational methods into political science and the political analysis of the digital revolution - form the ground for a new transdisciplinary approach: political data science. Interestingly, there is a rich tradition of crossing the borders of the disciplines, as can be seen in the works of Paul Werbos and Herbert Simon (both political scientists). Building on this tradition and integrating ideas from deep learning and Hegel's philosophy of logic a new perspective on causality might arise.
Organizers: Philipp Geiger
We present a novel probabilistic integrator for ordinary differential equations (ODEs) which allows for uncertainty quantification of the numerical error . In particular, we randomise the time steps and build a probability measure on the deterministic solution, which collapses to the true solution of the ODE with the same rate of convergence as the underlying deterministic scheme. The intrinsic nature of the random perturbation guarantees that our probabilistic integrator conserves some geometric properties of the deterministic method it is built on, such as the conservation of first integrals or the symplecticity of the flow. Finally, we present a procedure to incorporate our probabilistic solver into the frame of Bayesian inference inverse problems, showing how inaccurate posterior concentrations given by deterministic methods can be corrected by a probabilistic interpretation of the numerical solution.
Organizers: Hans Kersting
In this talk, I'd like to discuss the intertwining importance and connections of three principles of data science in the title. They will be demonstrated in the context of two collaborative projects in neuroscience and genomics, respectively. The first project in neuroscience uses transfer learning to integrate fitted convolutional neural networks (CNNs) on ImageNet with regression methods to provide predictive and stable characterizations of neurons from the challenging primary visual cortex V4. The second project proposes iterative random forests (iRF) as a stablized RF to seek predictable and interpretable high-order interactions among biomolecules.
Organizers: Michel Besserve