Rachel Thomson
This is a question that we explored with Jonathan Chaim Reus who is part of the Making & Using Social Archives network at the University of Sussex and a vocal artist. Jonathan uses his voice as a tool for exploring and animating sound archives. Put simply he creates real-time generative AI models from the source material of a sound archive and then, in live performances, he interacts with these models using his voice.
In his solo performances, the archives that he engages with are chosen for his own proximity to the sources: the first most intimate layer was an archive of sound poetry performances created by his teacher, the vocal artist Jaap Blonk. By first engaging in the careful curation of the dataset, and then using his own voice to interact with the synthetic vocalisations of his teacher, Jonathan engages in a form of study through interaction. Jonathan only has access to this archive of material because of his intimate relationship with Blonk. There is an ethical and artistic alignment that allows this access, in the form of an intergenerational relationship of teaching, transmission and collaboration.
The other archives Jonathan vocalises ‘with’ represent different forms of distance from him as a performer. These archives are all primarily vocal. They include a student choir that he is connected to as a teaching mentor, multiple collections of animal vocalisations that he has recorded himself during specific field recording trips, and a collection of voices from datasets in the public domain – representing the kind of computationally standardised archive that forms the basis of the automated voices we hear as part of our everyday interactions with machines.
Finally, he performs with an archive of voice recordings that Jonathan collected from his own social media feed following the October 7th atrocities in Israel/Palestine – an archive of vocal distress. Each archive has its own unique affordances, posing distinct ethical questions in terms of how Jonathan as performer is able to access and interact with the material, resynthesized through the mediating layer of a machine learning system. Questions that arise include those of data ownership, access and consent, but also questions that concern the intention of the usage, the nature of the performance and the audience, the way that the material is received and narrated, and how it combines in the contingencies of the event.
For the past year Jonathan and I have been engaged in a rich conversation about reanimation as a way of interacting with archival materials, and how data-driven machine learning might bring about embodied encounters with the traces of individuals held within vocal archives. He recognised connections between his own artistic work with machine learning and datasets, and the reanimating data project, and has used the term in his writings to connect his work with voice datasets to widen explorations of the potential to work creatively with archives. We have shared interests in the ethics of working with archives, balancing the potential of recontextualization and combination of materials through collage with a desire and commitment to recognise the traditions within which we work including ideas of sovereignty, homage and integrity. The concept of collage itself is further complicated by the synthetic, but not sourceless, voices of machine learning systems.
Jonathan alerted me to the important work arising from open source and creative tech communities, such as: the RAIL license generator that circumscribes how data sets might be used for training ML models; his own vocal values core principles which seek to imagine a conceptualisation of voice data care principles that go beyond binaries of yes and no in the conceptualisation of consent and imagining both individual and collective rights associated with voices and the possibility of aligning vocal commons with creators’ values; the CC4r ‘collective conditions for reuse’ which understand authorship to be ‘part of a collective cultural effort… situated in social and historical conditions’; and the CARE principles of Indigenous Data Governance that seek to work the tensions between supporting open data and protecting rights and interests. In response, I offered him examples of ethical innovation associated with data reuse arising from the social science community including Moore et al.’s notion of careful risk and inventive ethics and the writings of Libby Bishop chronicling the evolution of thinking about the ethics of data re-use and sharing, as well as my own and Liam Berriman’s attempts to work collaboratively with research participants to imagine an archive prospectively – including the potential for reusing the data – and leaving messages for future users.
On 12th June 2026 we hosted Johnathan at the University of Sussex for a workshop entitled Tracing Archives with Voice, Memory and Machine Learning. The performance itself took place in the beautiful meeting room on the University campus. A round auditorium with walls clad in differently coloured glass windows. As the performance evolved a wall of sound was created that was shocking in its volume and intensity – thrilling yet also distressing. One member of the audience described arriving late to the building and having the sensation of a “rumbling space ship about to lift off from the ground”. In a post-performance presentation and discussion Jonathan explained his performative method – sharing a diagrammatic presentation of how his voice interacts with multiple synthetic voice models trained on the distinct archives, and how his self-made performance software enables him to build a relationship between the data sets and his voice in an embodied way. We understood that in creating the models in relation to the archives he is building a musical instrument – like a trumpet – into which his voice is the force that animates the underlying traces of the archives. The performance could best be described as a form of call and response between the singer and the archive as mediated by the model’s idiosyncratic ways of creating new sound material resembling that of the archive itself, a kind of impressionistic duplication of the voices within.

The performance gave rise to an animated discussion. The audience wanted to know how it worked, how the model was trained. Jonathan explained how he had used RAVE (Realtime Audio Variational autoEncoder) an AI tool that creates high-quality, real-time audio synthesis and timbre transfer. We asked him what it felt like to interact with the models and the archives in this way. We wanted to know if he was singing the archive or asking it questions. He told us that there were dead-zones in the model. Spaces of non-response – especially in the animal voice archives. Jonathan explained that the better you know an archive the better you become at bringing it out through your voice, but that there are always surprises. Was it like using an automated voice to say what you wanted to say – a kind of voice cloning? Or, might it be more a case of using your voice to tease out meaning from the archive through the language model? Can this relationship be two way? Can the archive sing you? Jonathan told us that engaging in this way is an interesting experience, a mediation. But it is also demanding, exhausting.
I think it is fair to say that the audience was left in some kind of state of surprise and shock -not only as a result of the quality and volume of the sound, but also the very idea of what they had encountered and some sense of transgression of crossing over borders of possibility and direction in the relationship between the archive, the model and the explorer. I certainly was left with questions about how this might be used in my own work. Could I create an archive of sound from a set of interviews and train a model to enable the sound archive to be interrogated? Are there ways of doing this other than singing? Can the archive speak sentences as well as expressing sound? It there the potential for a singular voice as well as a composite sound? Certainly, this event illustrated Moore et al.’s argument that data reuse is at the cutting-edge of innovation posing new and surprising research questions. It also enables us to understand AI (in this case LLMs) as partners in enquiry rather than as necessarily extractive tools through which data is stolen. It may well be possible to use these tools to study and mediate on a body of material, using experimental approaches to tease out meaning and to pay respects with two-way relations of learning.
