Can you Sing an Archive?

Rachel Thomson

This is a question that we explored with Jonathan Chaim Reus who  is part of the Making & Using Social Archives network at the University of Sussex and a vocal artist. Jonathan uses his voice as a tool for exploring and animating sound archives. Put simply he creates real-time generative AI models from the source material of a sound archive and then, in live performances, he interacts with these models using his voice.

In his solo performances, the archives that he engages with are chosen for his own proximity to the sources: the first most intimate layer was an archive of sound poetry performances created by his teacher, the vocal artist Jaap Blonk. By first engaging in the careful curation of the dataset, and then using his own voice to interact with the synthetic vocalisations of his teacher, Jonathan engages in a form of study through interaction. Jonathan only has access to this archive of material because of his intimate relationship with Blonk. There is an ethical and artistic alignment that allows this access, in the form of an intergenerational relationship of teaching, transmission and collaboration. 

The other archives Jonathan vocalises ‘with’ represent different forms of distance from him as a performer. These archives are all primarily vocal. They include a student choir that he is connected to as a teaching mentor, multiple collections of animal vocalisations that he has recorded himself during specific field recording trips, and a collection of voices from datasets in the public domain – representing the kind of computationally standardised archive that forms the basis of the automated voices we hear as part of our everyday interactions with machines. 

Finally, he performs with an archive of voice recordings that Jonathan collected from his own social media feed following the October 7th atrocities in Israel/Palestine – an archive of vocal distress. Each archive has its own unique affordances, posing distinct ethical questions in terms of how Jonathan as performer is able to access and interact with the material, resynthesized through the mediating layer of a machine learning system. Questions that arise include those of data ownership, access and consent, but also questions that concern the intention of the usage, the nature of the performance and the audience, the way that the material is received and narrated, and how it combines in the contingencies of the event.

For the past year Jonathan and I have been engaged in a rich conversation about reanimation as a way of interacting with archival materials, and how data-driven machine learning might bring about embodied encounters with the traces of individuals held within vocal archives. He recognised connections between his own artistic work with machine learning and datasets, and the reanimating data project, and has used the term in his writings to  connect his work with voice datasets to widen explorations of the potential to work creatively with archives. We have shared interests in the ethics of working with archives, balancing the potential of recontextualization and combination of materials through collage with a desire and commitment to recognise the traditions within which we work including ideas of sovereignty, homage and integrity. The concept of collage itself is further complicated by the synthetic, but not sourceless, voices of machine learning systems. 

Jonathan alerted me to the important work arising from open source and creative tech communities, such as: the RAIL license generator that circumscribes how data sets might be used for training ML models; his own vocal values core principles which seek to imagine a conceptualisation of voice data care principles that go beyond binaries of yes and no in the conceptualisation of consent and imagining both individual and collective rights associated with voices and the possibility of aligning vocal commons with creators’ values; the CC4r ‘collective conditions for reuse’ which understand authorship to be ‘part of a collective cultural effort… situated in social and historical conditions’; and the CARE principles of Indigenous Data Governance that seek to work the tensions between supporting open data and protecting rights and interests.  In response, I offered him examples of ethical innovation associated with data reuse arising from the social science community including Moore et al.’s notion of careful risk and inventive ethics and the writings of Libby Bishop chronicling the evolution of thinking about the ethics of data re-use and sharing, as well as  my own and Liam Berriman’s attempts to work collaboratively with research participants to imagine an archive prospectively – including the potential for reusing the data – and leaving messages for future users.

On 12th June 2026 we hosted Johnathan at the University of Sussex for a workshop entitled Tracing Archives with Voice, Memory and Machine Learning. The performance itself took place in the beautiful meeting room on the University campus. A round auditorium with walls clad in differently coloured glass windows. As the performance evolved a wall of sound was created that was shocking in its volume and intensity – thrilling yet also distressing. One member of the audience described arriving late to the building and having the sensation of a “rumbling space ship about to lift off from the ground”. In a post-performance presentation and discussion Jonathan explained his performative method – sharing a diagrammatic presentation of how his voice interacts with multiple synthetic voice models trained on the distinct archives, and how his self-made performance software enables him to build a relationship between the data sets and his voice in an embodied way. We understood that in creating the models in relation to the archives he is building a musical instrument – like a trumpet – into which his voice is the force that animates the underlying traces of the archives. The performance could best be described as a form of call and response between the singer and the archive as mediated by the model’s idiosyncratic ways of creating new sound material resembling that of the archive itself, a kind of impressionistic duplication of the voices within.

The performance gave rise to an animated discussion.  The audience wanted to know how it worked, how the model was trained. Jonathan explained how he had used RAVE (Realtime Audio Variational autoEncoder) an AI tool that creates high-quality, real-time audio synthesis and timbre transfer. We asked him what it felt like to interact with the models and the archives in this way. We wanted to know if he was singing the archive or asking it questions. He told us that there were dead-zones in the model. Spaces of non-response – especially in the animal voice archives. Jonathan explained that the better you know an archive the better you become at bringing it out through your voice, but that there are always surprises.  Was it like using an automated voice to say what you wanted to say – a kind of voice cloning? Or, might it be more a case of using your voice to tease out meaning from the archive through the language model? Can this relationship be two way? Can the archive sing you? Jonathan told us that engaging in this way is an interesting experience, a mediation. But it is also demanding, exhausting. 

I think it is fair to say that the audience was left in some kind of state of surprise and shock -not only as a result of the quality and volume of the sound, but also the very idea of what they had encountered and some sense of transgression of crossing over borders of possibility and direction in the relationship between the archive, the model and the explorer. I certainly was left with questions about how this might be used in my own work. Could I create an archive of sound from a set of interviews and train a model to enable the sound archive to be interrogated? Are there ways of doing this other than singing?  Can the archive speak sentences as well as expressing sound? It there the potential for a singular voice as well as a composite sound? Certainly, this event illustrated Moore et al.’s argument that data reuse is at the cutting-edge of innovation posing new and surprising research questions. It also enables us to understand AI (in this case LLMs) as partners in enquiry rather than as necessarily extractive tools through which data is stolen. It may well be possible to use these tools to study and mediate on a body of material, using experimental approaches to tease out meaning and to pay respects with two-way relations of learning.

Anonymity in the archive

Rosie Gahnstrom

Most of my last year has been spent on my sofa, watching daytime TV and preparing the Women, Risk & AIDS Project (WRAP) interviews to be published in a digital archive – my day to day hasn’t changed much at all in light of Covid-19 (apart from the added anxiety of worrying about the health and safety of everyone in the entire world, the plummeting economy and the precarity of academia and its dwindling job market, a.k.a my future prospects).

My role in the Reanimating Data team has been to format, anonymise and catalogue all of the meta-data for the interviews and their accompanying field notes. Much of this has been fairly mundane, ritualistic and monotonous – open document, select all, change font to Arial, size 11, justify margins, adjust margins, insert, page number (top of page, Plain Number 3), find all double spaces, replace all double spaces with single spaces. The trickier, more interesting bits have been treating the data with a feminist ethics of care. Simple, I thought. I just need to give each interviewee a pseudonym and redact any identifying information. We’ve been taught a fairly generic, blanket rule of adhering to ethics like this, but I’d never really had the chance to properly engage with and reflect on these principles.

I had tried to be completely objective when changing the names of research participants. I had a list of ‘Popular baby names 1970s UK’ open on Google, and once I’d exhausted those would stare longingly at my book shelf, mentally scouring stories for names that might fit the stories being told in each interview. Naming each interviewee couldn’t really be objective – there’s too much tied up in a name, and assumptions around social class, ethnicity, place, time, gender. Which names would the original interviewees have picked for themselves? What a certain name might mean to me most likely has totally different connotations for someone else. AMD18, for example, I had named Tonya. On reflection, I think I had been struggling to think of any new names, but had recently seen ‘I, Tonya’ with Margot Robbie, where Tonya Harding (whose life the film is based on) was born in 1970 – just the right age for a WRAP participant. Rachel, PI for the project, said that she wouldn’t have chosen Tonya for that particular interviewee. There’s lots of literature on the importance of naming in social science research (Moore, 2012 – The politics and ethics of naming, for example), and I’m sure I could quite happily write an entire dissertation on the topic.

Archiving the dataset for public use raises a whole host of other ethical issues, too – does informed consent obtained in the late 1980s still count today? Could any of the original research participants ever have imagined that the archive would be taken out of a box in someone’s garage and uploaded onto the internet for anyone to stumble across? Would they want the intimate thoughts of their younger selves out there for themselves to find? – I still can’t decide whether I would like the goings on of my teenage years and the way that I would have framed them then to be known by anyone now, other than maybe my therapist.

Of course, any potential identifying information from each interview has now been anonymised or redacted. Originally, this had meant working with some key principles for anonymising the data, but in practice it didn’t feel right to be as consistent as that. Names of places, parent’s job roles, bands and musicians that young women listened to became [NAME OF TOWN], [CARING PROFESSION], [COUNTRY SINGER]. How much of these young women’s stories are woven into the industrial town in northern England that them and their parents, and probably their parent’s parents grew up in, and the opportunities that had been afforded to them there, and how much context do we lose when we omit these? These are all different parts of these young women’s lives that have, to some extent at least, shaped the way that they think about their sexual identities or how they position themselves against the opposite sex.

Anonymisation principles used by the Reanimating Data team.

  • Change all people’s names (in CAPS). This includes the name of the interviewee, any partners, friends, family members etc. This is the only ‘fake’ information we will include.
  • Where a participant gives their name, change this and create a pseudonym. Where they don’t do this, don’t create a pseudonym.
  • Put all new text / changes in caps.
  • Remove names of schools, workplaces and colleges. Do not replace these with made up names. Instead put [school] [company] etc.
  • Leave all place names and details of neighbourhoods where possible. Where a participant has moved around or where the combination of different places feels identifiable change these. Do not make up alternative places but replace with [town] [city] etc. Or make specifics more generic. E.g. ‘Singapore and Hong Kong’ can be replaced with ASIA.
  • Change details of locations and jobs for third parties (parents, partners) unless seems highly relevant.

One particularly tricky interview to work with was Amanda’s, or LSFS23 as she had been known for some time. Amanda had unfortunately been through an incredibly traumatic sexual experience in her home country before moving to London in the late 1980s. She went into quite a lot of detail about this in her interview and it felt significant for her story to be told, and to be heard. After much consultation with Rachel, Niamh and Sue Sharpe, who had originally interviewed ‘Amanda’, we decided that it was important to keep more than the bare bones of her story alive. Redacting some of the finer details of what had happened, and the names of countries, of work colleagues, of slang terms particular to where Amanda had grown up, felt like enough. We know through the wake of the #MeToo movement that there is power in storytelling, but of course we don’t know whether ‘Amanda’ would still want her story to be heard, and how she would want it to be told. Omitting the entire experience felt like silencing her, which happens far too much to women in society more generally, so I hope that we have done enough here to retain the eloquence and bravery with which Amanda’s story had originally been told.

The Women, Risk and AIDS Project had been conducted in both London and Manchester yet publications that came out of the study blurred any geographical differences, identifying young people only by their age, gender and social class, rather than by their location. Reading the data now I notice striking differences between the data from each of these places. There are different norms and youth sub-cultures operating in the different cities, and in different places within each city. Growing up in Brighton, my own adolescence was spent outside Borders on a Saturday afternoon during my ‘emo’ phase, in Saltdean park with a bottle of Glenn’s vodka on a Friday night, shortly graduating to club nights like Shameless at Audio on a Thursday or Pound Dance at Digital on a Wednesday. These are some of the places that helped shape my identity as a teenager.

The concept of place had been an important way of framing the Reanimating Data project. We have chosen to work with just the Manchester data and to work with community groups in the city to ‘rematriate’ some of the original data – to give it back to the city and communities it was abstracted from. In anonymising the data, we have clearly labelled each interview and fieldnote as ‘Manchester’ or ‘London’, offering the opportunity for future researchers to explore the significance of place to these young women’s lives, relationships and sexualities. Through anonymising the data and employing a feminist ethic of care we’ve had to remove many of the details of the streets, neighbourhoods, clubs and areas of Greater Manchester that the young women evoke. It feels almost sad that some of the defining parts of the WRAP young women’s stories (their Friday nights spent in Saltdean park making some questionable choices concerning boys in the year above) are lost through the anonymisation process, and that a new generation of researchers are unable to explore how, why and in what ways place might have mattered to them in forming their own sexual identities.

What we have made possible is that current and future researchers can now read these stories which have been handled carefully – twice. Once in 1989 and the early 1990s by the team of WRAP researchers and again in 2019 and 2020 by a new team of feminist researchers trying to balance our desire to both protect these young women from unwanted exposure or harm, whilst ensuring that the stories they chose to tell are still being heard.

INTERVIEWEE: What do you hope to gain from this research?
INTERVIEWER: I think we hope to be able to give young women themselves some kind of voice in terms of the sorts of things they find difficult or important or are concerned about. And to have some kind of input and feedback into health education. Some people design the leaflets and posters all the time, but they don’t necessarily talk to the people who are supposed to be reading them about what is important to them.
INTERVIEWEE: Will it go back into like sex education in schools do you think?
INTERVIEWER: We hope so yes. Because sex education in schools is pretty rudimentary and most people we have talked to say what you said. Something like it’s about babies and it wasn’t very helpful. It’s trying to say something about that. And more general things about how women feel about relationships and sexual relationships in particular, and what’s important. Generally nobody asks you, so we would like to do that so there will be a number of different things that we hope to relate to that and of course it should have been explained to you already …………..so you don’t need to worry about that, but we would like to think that the women we interview might actually read some of the things we write.
INTERVIEWEE: You actually write in magazines?
INTERVIEWER: Well we might do, we haven’t, but we think that might be quite a good idea to do that……….. not about…..so you might see something about it in Nineteen ……….

Interview with Lucy (MAG18) for the Women, Risk and AIDS Project.

To download and read the anonymised material visit the archive here. And if you don’t know where to start? Here’s Tonya’s interview.