Empty presentation slide with a bullet list

No bullets, please!

This post has a simple mission: stop making template-based “title + bullet point”-style slides for your presentation. Yes, most presentation software suggests this as a default, but that doesn’t make it any better. Instead, think about your slides as complementary to what you are saying. Auditory–visual perception Human perception is inherently multimodal, meaning that we continuously experience the world with all our senses. In a presentation, that means the audience is processing your voice, your slides, your gestures, and the room itself as part of one combined experience. ...

April 3, 2026 · 3 min · 586 words · ARJ
Stones

Octaves aren't Rhythmic

I see that the concept of “tempo octave” is being used by some researchers in the music information retrieval (MIR) community. This is a confusing term from a musical perspective. Here I explain why this is a bad idea. Octaves An octave is a core term in (Western) music theory related to describing intervals, relationships between two notes (and tones!) with a frequency ratio of 2:1. Here is an example of an octave: ...

January 10, 2026 · 2 min · 351 words · ARJ

The History and Future of AI

Due to MishMash, I am nowadays lecturing on AI, music, and creativity several times a week. I usually include a brief overview of machine learning history, mainly to explain that ChatGPT didn’t come out of nowhere but was the result of decades of research. To check that my story holds and to get a few more critical years and names in place. This blog post summarizes the brief history of AI to date. ...

January 3, 2026 · 10 min · 2092 words · ARJ

New publication: Exploring music-related micromotion

I am happy to announce the publication of a new anthology that I have contributed a chapter to: Jensenius, A. R. (2017). Exploring music-related micromotion. In C. Wöllner (Ed.), Body, Sound and Space in Music and Beyond: Multimodal Explorations (pp. 29–48). Routledge. The chapter does not have an abstract, but the opening paragraph summarizes the content quite well: ...

April 13, 2017 · 1 min · 135 words · ARJ

New PhD Thesis: Kristian Nymoen

I am happy to announce that fourMs researcher Kristian Nymoen has successfully defended his PhD dissertation, and that the dissertation is now available in the DUO archive. I have had the pleasure of co-supervising Kristian’s project, and also to work closely with him on several of the papers included in the dissertation (and a few others). Reference K. Nymoen. Methods and Technologies for Analysing Links Between Musical Sound and Body Motion. PhD thesis, University of Oslo, 2013. Abstract There are strong indications that musical sound and body motion are related. For instance, musical sound is often the result of body motion in the form of sound-producing actions, and muscial sound may lead to body motion such as dance. The research presented in this dissertation is focused on technologies and methods of studying lower-level features of motion, and how people relate motion to sound. Two experiments on so-called sound-tracing, meaning representation of perceptual sound features through body motion, have been carried out and analysed quantitatively. The motion of a number of participants has been recorded using stateof- the-art motion capture technologies. In order to determine the quality of the data that has been recorded, these technologies themselves are also a subject of research in this thesis. A toolbox for storing and streaming music-related data is presented. This toolbox allows synchronised recording of motion capture data from several systems, independently of systemspecific characteristics like data types or sampling rates. The thesis presents evaluations of four motion tracking systems used in research on musicrelated body motion. They include the Xsens motion capture suit, optical infrared marker-based systems from NaturalPoint and Qualisys, as well as the inertial sensors of an iPod Touch. These systems cover a range of motion tracking technologies, from state-of-the-art to low-cost and ubiquitous mobile devices. Weaknesses and strengths of the various systems are pointed out, with a focus on applications for music performance and analysis of music-related motion. The process of extracting features from motion data is discussed in the thesis, along with motion features used in analysis of sound-tracing experiments, including time-varying features and global features. Features for realtime use are also discussed related to the development of a new motion-based musical instrument: The SoundSaber. Finally, four papers on sound-tracing experiments present results and methods of analysing people’s bodily responses to short sound objects. These papers cover two experiments, presenting various analytical approaches. In the first experiment participants moved a rod in the air to mimic the sound qualities in the motion of the rod. In the second experiment the participants held two handles and a different selection of sound stimuli was used. In both experiments optical infrared marker-based motion capture technology was used to record the motion. The links between sound and motion were analysed using four approaches. (1) A pattern recognition classifier was trained to classify sound-tracings, and the performance of the classifier was analysed to search for similarity in motion patterns exhibited by participants. (2) Spearman’s p correlation was applied to analyse the correlation between individual sound and motion features. (3) Canonical correlation analysis was applied in order to analyse correlations between combinations of sound features and motion features in the sound-tracing experiments. (4) Traditional statistical tests were applied to compare sound-tracing strategies between a variety of sounds and participants differing in levels of musical training. Since the individual analysis methods provide different perspectives on the links between sound and motion, the use of several methods of analysis is recommended to obtain a broad understanding of how sound may evoke bodily responses. ...

February 20, 2013 · 5 min · 917 words · ARJ

Music is not only sound

After working with music-related movements for some years, and thereby arguing that movement is an integral part of music, I tend to react when people use “music” as a synonym for either “score” or “sound”. I certainly agree that sound is an important part of music, and that scores (if they exist) are related to both musical sound and music in general. But I do not agree that music is sound. To me, sound is one (and an important one) component of music, but not the only one. From the perspective of embodied music cognition, music is truly multimodal, meaning that all our senses and modalities are involved in performance and perception. This is not to mention all the cultural and contextual elements involved in our experience of music. ...

October 25, 2010 · 1 min · 172 words · ARJ

Thought Conduit

Synchronisation is a core issue when carrying out research on multimodal sensing/acting and multimedia. My take on this has been through the work on GDIF, and we are currently implementing a GDIF/SDIF recorder/player using FTM for Max/MSP (see our ICMC2008 paper for more on this). I just came across a software called Thought Conduit{.external .text} which promises synchronisation of audio, video, annotations and even OSC-streams. This sounds very exciting and I hope to be able to test this in practice at some point.

February 9, 2009 · 1 min · 83 words · ARJ

Apple tries to patent multimodal sensing

AppleInsider reports on a set of patents for multimodal sensing (i.e. using two or more senses at the same time). Multimodal sensing has been a hot research topic in human-computer interaction for several years, based on the knowledge that human perception and cognition is fundamentally multimodal. If we want computers to respond more efficiently to human communication they will also have to use more than one modality in their sensing and communication. That said, I am not sure that everyone will be comfortable leaving the webcam on at all times to allow for computer vision techniques on everything happening in front of the screen (as the picture below depicts)… ...

September 9, 2008 · 1 min · 109 words · ARJ

Spatial and Temporal Resolution in Multimodal Perception

Stefania Serafin just held a great lecture on multimodal perception and sonic interaction design at the SMC summer school. It is fascinating with the differences in spatiotemporal resolution between the different senses: Spatial resolution eye — highest: very high spatial acuity (foveal acuity ≈ 1 arcmin), excellent for fine detail. tactile — medium: fingertip resolution ≈ 1–2 mm, good for textures and small features. ear — lowest: coarse spatial localization (degree-level accuracy), poor for fine spatial detail. Temporal resolution ear — highest: excellent temporal sensitivity (sub-ms to ms scale), great for timing and rapid changes. tactile — medium: temporal resolution on the order of a few ms to tens of ms. eye — lowest: relatively slow temporal processing (temporal integration ~50–200 ms; critical flicker fusion ≈ 50–60 Hz). Stefania showed a number of examples of the application of this knowledge in interaction design, one example being the tactile floor they are building at McGill.

June 11, 2008 · 1 min · 154 words · ARJ

EMMA: Extensible MultiModal Annotation markup language

Strange that I didn’t see this before. Apparently, W3C has made a draft for multimodal annotation called EMMA: Extensible MultiModal Annotation markup language. The abstract of the document reads: The W3C Multimodal Interaction working group aims to develop specifications to enable access to the Web using multimodal interaction. This document is part of a set of specifications for multimodal systems, and provides details of an XML markup language for containing and annotating the interpretation of user input. Examples of interpretation of user input are a transcription into words of a raw signal, for instance derived from speech, pen or keystroke input, a set of attribute/value pairs describing their meaning, or a set of attribute/value pairs describing a gesture. The interpretation of the user’s input is expected to be generated by signal interpretation processes, such as speech and ink recognition, semantic interpreters, and other types of processors for use by components that act on the user’s inputs such as interaction managers. ...

March 14, 2007 · 1 min · 201 words · ARJ