IRCAM Forum Workshops 2025 – ACIDS

From 26 to 28th of March, we (the sound design master, second semester) had the incredible opportunity to visit IRCAM (Institut de Recherche et Coordination Acoustique/Musique) in Paris as part of a student excursion. For anyone passionate about sound, music technology, and AI, IRCAM is like stepping into new fields of research, discussion and seeing prototypes in action. One of my personal highlights was learning about the ACIDS team (Artificial Creatiive Intelligence and Data Science) and their research projects—RAVE (Real-time Audio Variational autoEncoder) and AFTER (Audio Features Transfer and Exploration in Real-time

ACIDS – Team

The ACIDS team is a multidisciplinary group of researchers working at the intersection of machine learning, sound synthesis, and real-time audio processing. Their name stands for Audio, Communication, Information, Data, and Sound, reflecting their broad focus on computational audio research. During our visit, they gave us an inside look at their latest developments, including demonstrations from the IRCAM Forum Workshop (March 26–28, 2025), where they showcased some of their most exciting advancements. Beside their really good and catchy (also a bit funny) presentation I want to showcase two projects.

RAVE (Real-Time Neural Audio Synthesis)

One of the most impressive projects we explored was RAVE (Real-time Audio Variational autoEncoder), a deep learning model for high-quality audio synthesis and transformation. Unlike traditional digital signal processing, RAVE uses a latent space representation of sound, allowing for intuitive and expressive real-time manipulation.

Overall architecture of the proposed approach. Blocks in blue are the only ones optimized,
while blocks in grey are fixed or frozen operations.

Key Innovations

  1. Two-Stage Training:
    • Stage 1: Learns compact latent representations using a spectral loss.
    • Stage 2: Fine-tunes the decoder with adversarial training for ultra-realistic audio.
  2. Blazing Speed:
    • Runs 20× faster than real-time on a laptop CPU, thanks to a multi-band decomposition technique.
  3. Precision Control:
    • Post-training latent space analysis balances reconstruction quality vs. compactness.
    • Enables timbre transfer and signal compression (2048:1 ratio).

Performance

  • Outperforms NSynth and SING in audio quality (MOS: 3.01 vs. 2.68/1.15) with fewer parameters (17.6M).
  • Handles polyphonic music and speech, unlike many restricted models.

You can explore RAVE’s code and research on their GitHub repository and learn more about its applications on the IRCAM website.

AFTER

While many AI audio tools focus on raw sound generation, what sets AFTER (Audio Foundation Transformer) apart is its sophisticated control mechanisms—a priority highlighted in recent research from the ACIDS team. As their paper states:

“Deep generative models now synthesize high-quality audio signals, shifting the critical challenge from audio quality to control capabilities. While text-to-music generation is popular, explicit control and example-based style transfer better capture the intents of artists.”

How AFTER Achieves Precision

The team’s breakthrough lies in separating local and global audio information:

  • Global (timbre/style): Captured from a reference sound (e.g., a vintage synth’s character).
  • Local (structure): Controlled via MIDI, text prompts, or another audio’s rhythm/melody.

This is enabled by a diffusion autoencoder that builds two disentangled representation spaces, enforced through:

  1. Adversarial training to prevent overlap between timbre and structure.
  2. A two-stage training strategy for stability.
Detailed overview of our method. Input signal(s) are passed to structure and timbre encoders, which provides
semantic encodings that are further disentangled through confusion maximization. These are used to condition a latent
diffusion model to generate the output signal. Input signals are identical during training and but distinct at inference.

Why Musicians Care

In tests, AFTER outperformed existing models in:

  • One-shot timbre transfer (e.g., making a piano piece sound like a harp).
  • MIDI-to-audio generation with precise stylistic control.
  • Full “cover version” generation—transforming a classical piece into jazz while preserving its melody.

Check out AFTER’s progress on GitHub and stay updated via IRCAM’s research page.

References

Caillon, Antoine, and Philippe Esling. “RAVE: A Variational Autoencoder for Fast and High-Quality Neural Audio Synthesis.” arXiv preprint arXiv:2111.05011 (2021). https://arxiv.org/abs/2111.05011.

Demerle, Nils, Philippe Esling, Guillaume Doras, and David Genova. “Combining Audio Control and Style Transfer Using Latent Diffusion.” 

Prototyping III: Image Extender – Image sonification tool for immersive perception of sounds from images and new creation possibilities

Research on sonification of images / video material and different approaches – focus on RGB

The paper by Kopecek and Ošlejšek presents a system that enables visually impaired users to perceive color images through sound using a semantic color model. Each primary color (such as red, green, or blue) is assigned a unique sound, and colors in an image are approximated by the two closest primary colors. These are represented through two simultaneous tones, with volume indicating the proportion of each color. Users can explore images by selecting pixels or regions using input devices like a touchscreen or mouse. The system calculates the average color of the selected area and plays the corresponding sounds. Distinct audio cues indicate image boundaries, and sounds can be either synthetic or instrument-based, with timbre and pitch helping to differentiate them. Users can customize colors and sounds for a more personalized experience. This approach allows for dynamic, efficient exploration of images and supports navigation via annotated SVG formats.

image seperation by Kopecek and Ošlejšek

The review by Sarkar, Bakshi, and Sa offers an overview of various image sonification methods designed to help visually impaired users interpret visual scenes through sound. It covers techniques such as raster scanning, query-based, and path-based approaches, where visual data like pixel intensity and position are mapped to auditory cues. Systems like vOICe and NAVI use high and low-frequency tones to represent image regions vertically. The paper emphasizes the importance of transfer functions, which link image properties to sound attributes such as pitch, volume, and frequency. Different rendering methods—like audification, earcons, and parameter mapping—are discussed in relation to human auditory perception. Special attention is given to color sonification, including the semantic color model introduced by Kopecek and Ošlejšek, which improves usability through clearly distinguishable tones. The paper also explores applications in fields such as medical imaging, algorithm visualization, and network analysis, and briefly touches on sound-to-image conversions.

Principles of the image-to-sound mapping

Matta, Rudolph, and Kumar propose the theoretical system “Auditory Eyes,” which converts visual data into auditory and tactile signals to support blind users. The system comprises three main components: an image encoder that uses edge detection and triangulation to estimate object location and distance; a mapper that translates features like motion, brightness, and proximity into corresponding sound and vibration cues; and output generators that produce sound using tools like Csound and tactile feedback via vibrations. Motion is represented using effects like Doppler shift and interaural time difference, while spatial positioning is conveyed through head-related transfer functions. Brightness is mapped to pitch, and edges are conveyed through tone duration. The authors emphasize that combining auditory and tactile information can create a richer and more intuitive understanding of the environment, making the system potentially very useful for real-world navigation and object recognition.

References

Kopecek, Ivan, and Radek Ošlejšek. 2008. “Hybrid Approach to Sonification of Color Images.” In Third 2008 International Conference on Convergence and Hybrid Information Technology, 721–726. IEEE. https://doi.org/10.1109/ICCIT.2008.152.

Sarkar, Rajib, Sambit Bakshi, and Pankaj K Sa. 2012. “Review on Image Sonification: A Non-visual Scene Representation.” In 1st International Conference on Recent Advances in Information Technology (RAIT-2012), 1–5. IEEE. https://doi.org/10.1109/RAIT.2012.6194495.

Matta, Suresh, Heiko Rudolph, and Dinesh K Kumar. 2005. “Auditory Eyes: Representing Visual Information in Sound and Tactile Cues.” In Proceedings of the 13th European Signal Processing Conference (EUSIPCO 2005), 1–5. Antalya, Turkey. https://www.researchgate.net/publication/241256962.

02.02: Mehr als nur Einführung

Nach Absolvierung der ersten beiden “Lessons” des ausgewählten Kurses halte ich es jetzt mal für angebracht im ersten echten Blogpost ein kleines Zwischenfazit zu ziehen und meine Key-Learnings festzuhalten.

Insgesamt haben mich die beiden großen Eingangs-Tutorials nun eine gute Woche gekostet, wodurch ich auch zum Schluss gekommen, nicht den gesamten Kurs zu absolvieren, da mich damit wohl im vierten Semester noch nicht fertig wäre, sondern nur mehr ausgewählte Lessons zu machen und danach hier darüber zu berichten. Aber jetzt medias in res.

Lesson 1: Track Mattes, Masken und erste Paths

Im ersten, kürzeren Tutorial, ging es um die echten Basics von After Effects. Gottseidank bedeutete dies für mich, dass nicht allzu viel neues auf mich zu kam – anderenfalls müsste man auch meinen bisherigen Studienerfolg in Frage stellen. Zu den Key-Learnings zählt aber definitiv die Animation von Masken-Paths, von der ich nichts wusste, da diese meiner Meinung nach auch relativ versteckt ist. Damit ist es aber z.B. sehr einfach gelungen die im zweiten Shot sichtbaren Balken zu animieren, die, einen Schritt weitergedacht, im Grunde für Balkendiagramme genauso verwendet werden können. Außerdem gab es auch bereits eine Einführung in den Graph Editor und Dinge wie Easy Ease um schnell coolere Ergebnisse zu erzielen, die richtige Keyframe Action begann aber erst in Lesson 2.

Für mich außerdem interessant war die Verwendung von Blend-Modes mithilfe derer man im Grunde aus allen Bildern, die irgendeine Textur haben, diese z.B. auf den Hintergrund übertragen kann. Ich bin zwar noch nicht schlau daraus geworden was genau welches Modus jetzt macht, aber sich einfach durchzuklicken, bis einem etwas gefällt kostet auch nicht zu viel Zeit. Das Ergebnis des gesamten Tutorials ist nett, stellt aber wirklich nur eine grobe Einführung dar.

Lesson 2: Ich glaub´ ich kann jetzt alles

So einfach und anfängerfreundlich die erste Lektion auch wahr, richtige Motion Graphic Action gab es erste in Lesson 2, dafür dort aber richtig. Gefühlt kann ich schon nach dieser Einheit die Welt zerreissen und es ist irre wieviele verschiedene Techniken in gerade einmal 2,5h Einheit verpackt sind (auch wenn man beim Nachmachen sicher ein paar Tage dafür benötigt). Die wichtigen Learnings sind kaum an einer Hand abzählbar, deswegen hier ein Best-Of:

  • Der einfachste Weg um Text oder Zahlen zu erstellen, die man dann als Shapes auch manipulieren kann ist nicht durch Shapes selbst, sondern durch das Text Tool. Mit nur einem Klick (create shapes from text) kann man den dann nämlich umwandeln.
  • Über Null-Objekte auf welche man andere parented kann man super schnell eine ganze Gruppe kontrollieren ohne noch einmal zu precomposen.
  • Key-Frame Manipulation im Graph Editor funktioniert im Grunde immer gleich. Zieht man den die Geschwindigkeitskurve zu Beginn auf volle Kanne und dann immer lamgsamer, hat man in 90% der Fälle einen super smoothen look. Falls nicht, einfach umkehren, eins von den beiden ists immer.
  • Der einfachste Weg Objekte über einen definierten Pfad zu bewegen ist diesen mit dem Pen-Tool zu ziehen und dann mit “create nulls from path” direkt ein sich darauf bewegendes Null zu erzeugen, das auch gleich automatisch einen Progress Slider hat. Dieses dann einfach parenten und fertig.
  • Bounce Animations sind brutal scheisse und aufwendig wenn man sie mit einzelnen Keyframes macht, gibt´s da vielleicht etwas einfacheres???? Hilfe!!!!!

Das ganze Video, das sich meiner Meinung nach wirklich sehen lassen kann sieht nun so aus und beinhaltet meiner Meinung nach wirklich alle essenziellen Techniken für smoothe Motion Graphics:

Fazit

Da ich mit diesen beiden Einstiegskursen bereits das Gefühl habe gut dabei zu sein, werde ich mich in Zukunft mehr auf einzelne brauchbare Tutorials beschränken statt den gesamten Kurs zu machen, und bald davon berichten.

Plane of Emergence – Music between Machines (IRCAM)

When I arrived at the presentation of Plane of Emergence at IRCAM, the setup looked surprisingly simple at first. On the floor, inside a black marked rectangle, were two small cube-like devices, standing quietly next to each other. A big screen behind them showed a live camera view of the scene. I noticed a line connecting the two cubes on the projection, showing exactly how far they were apart. This was made possible by a motion-tracking camera mounted above, constantly measuring their positions.

The artist explained that these devices were not normal speakers or instruments, but autonomous machines. They were able to listen, react, and transform musical patterns based on how close or far they were from each other. There was no conductor or composer telling them what to play — everything emerged from their interaction alone.

While listening, I could feel how the soundscape was always shifting. Sometimes you could recognize small repetitive patterns, like a rhythm or a melody fragment. But just when you thought something stable was forming, it suddenly dissolved into something new. The artist described this as a balance between “territorialization” — when the devices settle into stable patterns — and “deterritorialization” — when they break free and surprise you with unexpected variations. It felt like watching two creatures communicating and constantly changing their language.

The idea behind it is inspired by the philosopher Deleuze and his concept of the plane of immanence — a space where things don’t follow strict rules but constantly create themselves from within. I liked that you could really hear this concept, it wasn’t just theory.

Technically, the system is based on a previous project called Spatially Distributed Instruments, where the machines not only send sounds but also “listen” to each other without noticeable delay. The sound you hear is not pre-composed, it is created in real-time from their relationship in space.

Unfortunately, as the artist mentioned, only two of the planned interaction methods were working that day. But even with these limitations, it was fascinating to see (and hear) how rich and alive the system already was.

For me, it was less like watching a performance and more like observing a small ecosystem made of sound and technology.

IRCAM Link: https://forum.ircam.fr/article/detail/plane-of-emergence/

A Journey Through Sound – (IRCAM)

When I entered the installation called ‘Hearing From Within A Crossfade by Lewis Wolstanholme’ at IRCAM, I was immediately surrounded by a very special atmosphere. Sounds were floating through the space — soft, detailed, and constantly changing. It didn’t feel like listening to a normal piece of music. Instead, it felt like the sounds were alive, moving gently around me and transforming into something new all the time.

What made this experience so fascinating was how the sounds seemed to blend into each other without clear breaks. One texture slowly became another, sometimes so smoothly that I barely noticed the change. I later found out that this was made possible by a special technique called Joint Time-Frequency Scattering Transform. This method allows sounds to be transformed and combined in a very natural way, almost like they were breathing.

The installation was created by Christopher Mitcheltree together with IRCAM. He used this technique to make sounds not only change over time but also move through space. Depending on how a sound behaved — for example, how much it was vibrating or how high or low it was — it appeared at different places in the room. This made the whole space feel like part of the music.

For me, it felt like I wasn’t just listening, but actually walking inside a sound. It was a very inspiring and calming experience, and I stayed much longer than I had planned.

IRCAM – Link: https://forum.ircam.fr/article/hearing-from-within-a-crossfade/

02.01: Wer ist das echte Beast?

Der Sinn dieser Blog-Post-Serie im zweiten Semester ist einfach wie eingänglich: Man soll etwas lernen… aber was? Rückblickend auf die ersten zehn Blogposts aus dem vergangenen Semester war das für mich recht klar: Ich möchte meine Datenvisualisierungen auf das nächste Level bringen. Dazu braucht es aber zweierlei, mit dem ich mich in den nächsten neun Blogposts befassen werde.

Einerseits muss die Basis einer jeden guten Datenvisualisierung wissenschaftlich evident gemacht werden. Dazu möchte ich mich durch verschiedenste Fachliteratur kämpfen, um herauszufinden warum manche Visualisierungen funktionieren und manche schlicht nicht. Basisliteratur soll das umfassende (und mir glücklicherweise vom besten Major-Leiter des Instituts zur Verfügung gestellte) Werk “Show Me The Numbers” werden, jedoch möchte ich einen Blogpost auch weiterer Fachliteratur zum Thema widmen.

Um mit diesem Wissen dann aber auch ordnungsgerecht umgehen zu können, muss natürlich eins her: After Effects. Wie bereits in meiner Themenvorstellung vor ein paar Wochen angekündigt möchte ich dazu einen online Kurs belegen. Welcher das jedoch sein wird, war in den letzten Tagen Kern angeregter Diskussionen zwischen mir und meinen drei anderen Persönlichkeiten (sowie meinen Studienkollegen natürlich). Dabei gibt es zwei große Überlegungen: Einerseits gibt es Kurse speziell für Datenvisualisierungen, diese sind aber eher nieschig, oft bereits älter und nicht von renommierten Lehrern oder Bildungshäusern, hätten aber natürlich den Vorteil genauer auf meine Bedürfnisse einzugehen. Andererseits gäbe es aber natürlich auch allgemeinere After Effects oder im speziellen Motion Design Kurse, deren Skills sicher gut auf Datenvisualisierungen umzumünzen wären.

Gerade in Hinblick auf andere Kurse in diesem Semester – ich denke da vor allem an Green Utopia sowie Moya – habe ich entschieden, dass es definitiv mehr Sinn machen würde einen allgemeineren Motion Design Kurs zu belegen, da ich die Fähigkeiten definitiv noch an der ein oder anderen Stelle brauchen werde. Stellt sich also nur die Frage welchen…

Die Auswahl dahingehend ist ja bekanntlich riesig, nicht nur auf Lernplattformen wie Udemy oder Skillshare, sondern auch bei privaten Anbietern. Die beiden Goldstandards in Sachen Motion Design scheinen dabei einerseits der Motion Beast Course und andererseits Design Breakthrough by Ben Marriott zu sein. Beide dieser Kurse warten aber mit horrenden Preisen auf (350 bzw 500 Euro), die ich ehrlicherweise im Moment nicht gewillt bin zu zahlen, auch wenn man mit Investitionen in die eigene Bildung ja eigentlich nie was falsch machen kann. (Ich muss ja auch was essen…) Daher habe ich mich in verschiedenen Foren auf die Suche nach einem preiswerten Ersatz gemacht und bin dabei auf die Seite https://www.learnto.day/aftereffects gestoßen, die im Endeffekt eine Zusammenstellung aller guten kostenfreien Ressourcen und Kurse zum Thema Motion Design darstellt und in einem großen Curriculum einen volleinheitlichen Kurs nachbilden soll. Da laut diverser Forenmitglieder darin mindestens gleich viel, wenn nicht sogar mehr gutes Wissen vorhanden sein soll, als in vielen pay-per-view Kursen habe ich mich schlussendlich dafür entschieden dieses Curriculum auch genau so durchzuarbeiten.

Zusammenfassend lässt sich also sagen, dass das restliche Semester DesRes für mich klar strukturiert ist. Nachfolgend wird es hier mit Ausschnitten und meinen Key-Learnings aus dem After Effects Curriculum weitergehen, da ich dieses Wissen wohl schon eher früher als später auch in anderen Kurse brauchen werde. Danach folgt die Fachliteratur, um am Ende mit all dem gelernten auch ein ansprechendes Werkstück gestalten zu können.

From Idea to Concept – The Final Project

After deciding to focus on 3D audio, I began shaping the project into a concrete concept. The result? Echoes of Addiction – Breaking Boundaries Through 3D Audio, a 3D audio concept album that explores addiction and dependency through immersive sound design and innovative production techniques.

The idea behind the album is to take listeners on an emotional journey, sonically representing different phases of addiction – from denial and isolation to dependency and recovery. Through spatial sound placement, dynamic movement, and immersive textures, I aim to create an auditory experience that goes beyond traditional music production. The goal is to make the emotional states of addiction tangible through sound.

Sound Design as a Narrative Tool

To achieve this, I will experiment with various sound design techniques that enhance the themes of the album. Some examples may be:

  • Ambisonic and Field Recordings: Capturing real-world environments and spatial depth to create immersive atmospheres that reflect emotional states. For example, the feeling of isolation could be represented by a vast, empty soundscape with distant, echoing voices.
  • Granular Synthesis: Breaking down recorded sounds (e.g., vocals, guitars, ambient noises) into tiny grains, which can be stretched, scattered, and manipulated across the 3D space to create a fragmented, distorted reality—symbolizing confusion or addiction-induced dissociation.
  • Spectral Processing: Transforming recognizable sounds into ghostly, abstract versions of themselves to reflect loss of control and mental instability. For instance, vocals could be decomposed and restructured into eerie, flickering echoes, representing the overwhelming thoughts in an addict’s mind.
  • Modulation and Dissonance: Using amplitude and frequency modulation to create tension, while the interplay between dissonant and consonant harmonies mirrors the chaos and clarity within the addiction cycle.
  • 3D Audio Movement: Sounds will not remain static but move dynamically around the listener, reinforcing the psychological aspects of addiction. A whisper circling the head, for example, could represent obsessive thoughts or withdrawal-induced paranoia.

By combining these techniques, I want to push the boundaries of what a concept album can achieve. The sound design will not only shape the individual songs but also serve as a connective tissue, creating transitions and atmospheric interludes that guide the listener through the emotional arc of the album.

Bringing the Project to Life

From a technical standpoint, I will work with both Ambisonics and Dolby Atmos to create a fully immersive experience. The challenge will be to ensure that the album translates well across different formats – from 3D speaker setups to binaural headphone mixes and even standard stereo. This will allow as many people as possible to experience the album’s message and emotional depth.

The production process will follow a structured timeline:

  1. Research & Concept Development – Deep diving into 3D audio techniques and defining the sound design approach.
  2. Songwriting & Pre-Production – Refining compositions while experimenting with 3D audio integration.
  3. Recording & Sound Design – Capturing performances, field recordings, and designing immersive sonic textures.
  4. Mixing & Finalization – Creating the 3D audio mix, testing binaural and stereo compatibility, and refining the album’s overall sonic identity.

This project is more than just an academic exercise; it is an artistic statement that merges my passion for music, immersive sound, and storytelling. Now that the concept is finalized, the next step is bringing Echoes of Addiction to life.

A Sudden Change of Direction – A New Focus on 3D Audio

Sometimes, things change unexpectedly – and that’s exactly what happened with my master’s project. I had already explored different ideas, including building my own studio monitors or a binaural microphone. But then, my focus suddenly shifted.

For some time now, I’ve been deeply fascinated by 3D audio and how sound can be experienced not just in stereo but in a fully immersive space. The more I explored the topic, the more I realized that this was something I truly wanted to dive into. However, making the final decision wasn’t easy.

About two weeks before my project expose was due, I had an inspiring conversation with Mr Sontacchi. His words gave me the confidence I needed and reinforced my belief that my growing interest in 3D audio was not just a temporary fascination but a real opportunity for my master’s project. This conversation gave me the final push to fully commit to this idea.

Then, I had an epiphany: Why not connect this passion with something I already love? My band, Flavor Amp, is currently working on a concept album about addiction and dependency. Since we want the music to be deeply emotional and immersive, integrating 3D audio into the production seemed like the perfect way to amplify the storytelling aspect of our songs.

Instead of focusing on hardware development, I decided to explore how 3D audio combined with Sound Design can enhance emotional storytelling in music. I want to experiment with different techniques, from Ambisonics to Dolby Atmos, and find out how spatial sound design can strengthen the narrative of our album.

This decision was a turning point for me. It wasn’t just about choosing a project; it was about following a passion that connects my studies, my creative interests, and my band’s music. The next step? Developing a concrete plan and shaping the project into something truly meaningful.

Narrowing Down My Project Ideas

At first I have to say that it was really hard for me to find a topic because I had not enough time to think about it and it’s a hard decision. But there are two main topics which I’m really interested in:

  • Developing and building my own microphone – “Kunstkopf Mikrofon” – like the Neumann KU100
  • Developing and building my own studio monitors

At the moment, I’m focusing more on the studio monitors, that’s the reason why I chose this topic for my second blog post. But I have to say that I’m not 100% sure which topic I will choose.

The “studio-monitor-project” combines my passion for sound quality with a desire to better understand the technical and artistic nuances of audio production. While I am new to the technical aspects of speaker design, I am excited to explore this field and gain knowledge in this area.

Core Purpose

The core purpose of my project is to create a pair of high-quality, custom studio monitors that I can use in my personal work. This project allows me to explore the intersection of acoustics, design, and audio technology.

Goals

  • Technical Learning: Build a foundational understanding of speaker design, including acoustic principles, driver selection, cabinet construction, measurements, …
  • Practical Application: Design and assemble functional monitors using ready-made drivers, amplifiers, and other components
  • Innovative Elements: Integrate an innovative aspect into the design. I’m still considering ideas such as using sustainable materials, exploring transaural speaker technology, or creating a modular design for easy upgrades and customization
  • Aesthetic and Functional Design: Create a design that balances professional audio quality with aesthetic appeal
  • Documentation and Sharing: Document my process and share it with people, which are interested in this topic

Relevance

This project holds personal significance for me as it helps me to widen my horizon and to dive into a (for me) unknown topic, which I’m really interested in. On a broader scale, it explores how accessible and customizable high-quality audio solutions can be created by individuals. Adding an innovative aspect also makes this project suitable for this FH-project, such as sustainability and evolving audio technologies.

Reference Works

I have drawn inspiration primarily from:

  • A project report from the Institute of Electronic Music and Acoustics (IEM), which provided valuable insights into transaural audio
  • Discussions and shared experiences in DIY speaker forums, where enthusiasts offer practical advice and troubleshooting tips

Planned Work Techniques

To structure my approach, I’ve broken the project into distinct phases:

  1. Research Phase: Deepen my understanding of acoustic principles and explore potential innovative aspects like sustainability or modularity
  2. Concept Phase: Create a mind map of possible designs, features, and materials
  3. Design Phase: Finalize the cabinet shape and dimensions, select drivers, and identify suitable amplifier modules
  4. Prototyping Phase: Build a prototype, test the sound, and refine the design
  5. Final Construction and Testing: Build the final version, measuring, measuring, measuring…

Open Questions

Several challenges remain, such as:

  • Deciding which project I’m going to choose

In terms of the “studio-monitor-project”:

  • Deciding which innovative aspect to prioritize and how to implement it effectively
  • Ensuring compatibility between all components
  • Balancing budget constraints with the desire for high-quality components

Over the next few weeks, I plan to:

  1. Decide for a project
  2. Research…

Prototyping II: Image Extender – Image sonification tool for immersive perception of sounds from images and new creation possibilities

Expanded research on sonification of images / video material and different approaches:

Yeo and Berger (2005) write in “A Framework for Designing Image Sonification Methods” about the challenge of mapping static, time-independent data like images into the time-dependent auditory domain. They introduce two main concepts: scanning and probing. Scanning follows a fixed, pre-determined order of sonification, whereas probing allows for arbitrary, user-controlled exploration. The paper also discusses the importance of pointers and paths in defining how data is mapped to sound. Several sonification techniques are analyzed, including inverse spectrogram mapping and the method of raster scanning (which already was explained in the Prototyping I – Blog entry), with examples illustrating their effectiveness. The authors suggest that combining scanning and probing offers a more comprehensive approach to image sonification, allowing for both global context and local feature exploration. Future work includes extending the framework to model human image perception for more intuitive sonification methods.

Sharma et al. (2017) explore action recognition in still images using Natural Language Processing (NLP) techniques in “Action Recognition in Still Images Using Word Embeddings from Natural Language Descriptions.” Rather than training visual action detectors, they propose detecting prominent objects in an image and inferring actions based on object relationships. The Object-Verb-Object (OVO) triplet model predicts verbs using object co-occurrence, while word2vec captures semantic relationships between objects and actions. Experimental results show that this approach reliably detects actions without computationally intensive visual action detectors. The authors highlight the potential of this method in resource-constrained environments, such as mobile devices, and suggest future work incorporating spatial relationships and global scene context.

Iovino et al. (1997) discuss developments in Modalys, a physical modeling synthesizer based on modal synthesis, in “Recent Work Around Modalys and Modal Synthesis.” Modalys allows users to create virtual instruments by defining physical structures (objects), their interactions (connections), and control parameters (controllers). The authors explore the musical possibilities of Modalys, emphasizing its flexibility and the challenges of controlling complex synthesis parameters. They propose applications such as virtual instrument construction, simulation of instrumental gestures, and convergence of signal and physical modeling synthesis. The paper also introduces single-point objects, which allow for spectral control of sound, bridging the gap between signal synthesis and physical modeling. Real-time control and expressivity are emphasized, with future work focused on integrating Modalys with real-time platforms.

McGee et al. (2012) describe Voice of Sisyphus, a multimedia installation that sonifies a black-and-white image using raster scanning and frequency domain filtering in “Voice of Sisyphus: An Image Sonification Multimedia Installation.” Unlike traditional spectrograph-based sonification methods, this project focuses on probing different image regions to create a dynamic audio-visual composition. Custom software enables real-time manipulation of image regions, polyphonic sound generation, and spatialization. The installation cycles through eight phrases, each with distinct visual and auditory characteristics, creating a continuous, evolving experience. The authors discuss balancing visual and auditory aesthetics, noting that visually coherent images often produce noisy sounds, while abstract images yield clearer tones. The project draws inspiration from early experiments in image sonification and aims to create a synchronized audio-visual experience engaging viewers on multiple levels.

Software Interface for Voice of Sisyphus (McGee et al., 2012)

Roodaki et al. (2017) introduce SonifEye, a system that uses physical modeling sound synthesis to convey visual information in high-precision tasks, in “SonifEye: Sonification of Visual Information Using Physical Modeling Sound Synthesis.” They propose three sonification mechanisms: touch, pressure, and angle of approach, each mapped to sounds generated by physical models (e.g., tapping on a wooden plate or plucking a string). The system aims to reduce cognitive load and avoid alarm fatigue by using intuitive, natural sounds. Two experiments compare the effectiveness of visual, auditory, and combined feedback in high-precision tasks. Results show that auditory feedback alone can improve task performance, particularly in scenarios where visual feedback may be distracting. The authors suggest applications in medical procedures and other fields requiring precise manual tasks.

Dubus and Bresin review mapping strategies for the sonification of physical quantities in “A Systematic Review of Mapping Strategies for the Sonification of Physical Quantities.” Their study analyzes 179 publications to identify trends and best practices in sonification. The authors find that pitch is the most commonly used auditory dimension, while spatial auditory mapping is primarily applied to kinematic data. They also highlight the lack of standardized evaluation methods for sonification efficiency. The paper proposes a mapping-based framework for characterizing sonification and suggests future work in refining mapping strategies to enhance usability.

References

Yeo, Woon Seung, and Jonathan Berger. 2005. “A Framework for Designing Image Sonification Methods.” In Proceedings of ICAD 05-Eleventh Meeting of the International Conference on Auditory Display, Limerick, Ireland, July 6-9, 2005.

Sharma, Karan, Arun CS Kumar, and Suchendra M. Bhandarkar. 2017. “Action Recognition in Still Images Using Word Embeddings from Natural Language Descriptions.” In 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), 978-1-5090-4941-7/17. DOI: 10.1109/WACVW.2017.17.

Iovino, Francisco, Rene Causse, and Richard Dudas. 1997. “Recent Work Around Modalys and Modal Synthesis.” In Proceedings of the International Computer Music Conference (ICMC).

McGee, Ryan, Joshua Dickinson, and George Legrady. 2012. “Voice of Sisyphus: An Image Sonification Multimedia Installation.” In Proceedings of the 18th International Conference on Auditory Display (ICAD-2012), Atlanta, USA, June 18–22, 2012.

Roodaki, Hessam, Navid Navab, Abouzar Eslami, Christopher Stapleton, and Nassir Navab. 2017. “SonifEye: Sonification of Visual Information Using Physical Modeling Sound Synthesis.” IEEE Transactions on Visualization and Computer Graphics 23, no. 11: 2366–2371. DOI: 10.1109/TVCG.2017.2734320.

Dubus, Gaël, and Roberto Bresin. 2013. “A Systematic Review of Mapping Strategies for the Sonification of Physical Quantities.” PLoS ONE 8(12): e82491. DOI: 10.1371/journal.pone.0082491.