Interviews
A repository of conversations we’ve had with researchers and artists across the life of the project. These interview transcripts are produced using the Reduct transcription platform. The transcripts contain artefacts and politics of machinic listenings, such as the occasional Miss Herd word, and ah vocalised hesitations and filler sounds.
Bernard Mont-Reynaud
Bernard Mont-Reynaud is a computer scientist who worked on and around machine listening for more than 50 years before his retirement in 2023. After a PhD at Stanford in the 1970s, he moved to Berkeley, before returning to Stanford’s CCRMA to pursue work on source separation and music. After time at Sony, Audience and several other companies, Mont-Reynaud became Principal Scientist at SoundHound in 2010. We speak to him about his long career in machine listening, including his involvement with the histories of computer music and music intelligence, source separation, auditory scene analysis, and voice assistants.
Interview conducted in December 2022
Santiago Rentiera
We speak to Santiago about his amazing work on birdsong analysis, acoustic ecology and artificial intelligence, and especially his critical and artistic engagements with Australian magpie archives. Along the way, we discuss biosemiotics, Solomon’s Seal, and how, by starting with magpie calls, you can get to a strong critique of AGI.
Max Ritts
We talk with Max about the politics of ‘smart oceans governance’ and the growing application of machine listening to ‘conservation acoustics’. In the process, we also cover his amazing work on ‘military cetology’ and the social construction of whale song, among other things. All this is going to feature in Max’s forthcoming book A Resonant Ecology (Duke UP). If you like this conversation, you might also like Smart Forest Radio, a podcast Max has been involved with, which includes many episodes on digital bioacoustics.
Interview conducted on 7 July 2023
Beth Semel
We talk with Beth about her ethnographic work on computational psychiatry, vocal biomarkers, and the automation of care. This means following the AI hype machine into privatised healthcare and the politics of voice. Along the way, we talk about ‘heteromation’, ‘listening like a computer’, and the history of the DSM. We also get stuck into some amazing stories from beth’s fieldwork: including ‘MRI theatre’, the ‘humans in the machine’, and the strange continuities between the ‘grandfather passages’ used by linguists and dadaist poetry.
Interview conducted on 10 October 2022
Audrey Amsellem
Audrey talks to us about her work on sound, surveillance and ‘the making of the neoliberal ear’, Audrey’s term for the auditory dimensions of surveillance capitalism. We discuss the histories and politics of streaming platforms like Spotify, music analytics pioneer Echo Nest, Amazon Echo, and the Smart City communication hub LinkNYC.
Interview conducted on 6 October 2022
Guillaume Heuguet
Guillaume tells us about his new book *How Music Changed YouTube* (Bloomsbury 2024), and in particular the crucial role played by ‘audio fingerprinting’ in the operation of Content ID, YouTube’s system of automated copyright management. We talk about the ‘mathematics of originality’ and the automation of judgment, along with the political economy of music streaming. Audio fingerprinting, it turns out, is a technology about which the public knows very little, but with major significance for the culture industries.
Interview conducted on 8 September 2022, with [Anabelle Lacroix](https://www.newschool.edu/parsons-paris/faculty/anabelle-lacroix/)
Dan McQuillan
Dan talks to us about his work on the automation of (mental) health care from voice analysis. Along the way, we also discuss algorithmic thoughtlessness, data luddism, the automation of care, people’s councils for machine learning, and Dan’s latest book on anti-fascist approaches to artificial intelligence.
Interview conducted on 22 August, 2022
Jonathan Sterne and Elena Razlogova
We talk with Jonathan and Elena about their collaborative work on AI and the automation of music mastering, with a particular focus on Montreal-based platform LANDR. But the conversation inevitably drifts towards much bigger questions on the politics of automation and machine listening.
Interview conducted on 12 August, 2022
Sara Ramshaw and Paul Stapleton
Sara and Paul talk to us about their work on improvisation as it connects with questions of justice, composition and performance. We range from George Lewis’ early experiments with interactive music systems on Rainbow Family to family law in Northern Ireland, and what it might mean to ‘humanise’ algorithmic listening.
Interview conducted on 27 May, 2021
Liz Pelly
Liz talks to us about the cultural politics and political economy of Spotify. We talk through some of the ideas in her *amazing column* for The Baffler, along with some of the listening experiments she’s conducted on Spotify’s algorithms (and herself), before turning to her argument for ‘socialised streaming’.
Interview conducted on 23 March, 2021
Mara Mills, Xiaochang Li, Jessica Feldman, Michelle Pfeifer
Mara, Xiaochang, Jessica and Michelle talk us through the history and politics of machine listening, from ‘affect recognition’ and the ‘statistical turn’ in ASR to automated accent detection at the German border, voiceprints and the ‘assistive pretext’. This is an expansive conversation with an amazing group of scholars, who share a common connection to the Media, Culture, and Communications department at NYU, founded by Neil Postman in 1971 at the urging of Marshall McLuhan.
Interview conducted on 19 February, 2020
Lauren Lee McCarthy
Lauren talks us through some of her many works concerned with smart speakers, machine listening and social relationships in the midst of surveillance, automation, and algorithmic living. We discuss: LAUREN, for which she attempted to become a human version of Alexa, SOMEONE, which won her the Prix Ars Electronica 2020 / Interactive Art +, and a range of related works and political questions.
Interview conducted on 22 September, 2020
Alex Ahmed
Alex talks to us about Project Spectra, an online, community-based, free and open source software application for transgender voice training. We discuss speech pathology and the politics of pitch, along with the importance of grass-roots led tech projects and community-centred design.
Interview conducted on 21 September, 2020
Stefan Maier
Stefan’s 2018 dossier on machine listening for Technosphere puts the work of artists like George Lewis, Jennifer Walshe, Florian Hecker, and Maryanne Amacher into conversation with Google’s wavenet. We talk about these and other works along with Stefan’s own compositions which treat machine listening as a prepared instrument, ready to be detourned.
Interview conducted on 11 September, 2020
Yolande Strengers and Jenny Kennedy
Yolande and Jenny provide a “reboot” manifesta in their book The Smart Wife: Why Siri, Alexa, and Other Smart Home Devices Need a Feminist Reboot, which lays out their proposals for improving the design and social effects of digital voice assistants, social robots, sex robots, and other AI arriving in the home.
Interview conducted on 11 October, 2020
Jùnchéng Billy Lì
Billy tells us about his research on ‘adversarial music’, and in particular an attempt to produce a ‘Real World Audio Adversary Against Wake-word Detection Systems’ for Amazon Alexa.
Interview conducted on 11 September, 2020
André Dao
André talks to us about UN Global Pulse, the UN’s big data initiative, and in particular one program which ‘uses machine Learning to analyse radio content in Uganda’. We discuss the increasing entanglements of big tech, the UN and human rights discourse more broadly, as well as an emergent right to be counted.
Interview conducted on 4 September, 2020
Angie Abdilla
Angie talks to us about Old Ways, New, the Indigenous owned and led social enterprise she founded, based on Gadigal land in Redfern, Sydney. We discuss Decolonising the Digital, Country Centered Design, a methodology which applies Indigenous design principles to the development of technologies for places, spaces and experiences, and how this contrasts with the ‘placelessness’ on which so many machine learning/listening systems are based.
Interview conducted on 1 September, 2020
James Parker (w Jasmine Guffond)
This is the first of three radio shows as part of Jasmine’s guest residency at Noods Radio. It features an interview with James about his research on machine listening, this curriculum, the project with Unsound, and a selection of electronic music.
Interview conducted on 1 September, 2020
Vladan Joler
Vladan walks us through Anatomy of an AI System, his 2018 work with Kate Crawford, which diagrams the Amazon Echo as an anatomical map of human labor, data and planetary resources. We talk about the politics of visibility and method as well as Vladan’s work with Share Lab, ‘where indie data punk meets media theory pop to investigate digital rights blues’.
Interview conducted on 1 September, 2020
Halcyon Lawrence
Halcyon talks us through some of her work on the politics of voice user interfaces: in particular accent bias, ‘Siri discipline’ and the ways in which smart speakers reproduce and hardwire longstanding forms of linguistic imperialism.
Interview conducted on 31 August, 2020
Thomas Stachura
Thomas is CEO of Paranoid Inc, which makes devices that block smart speakers from listening. The company’s mandate ‘earn lots of money by increasing privacy, not eroding it’ imagines an emerging privacy industry, as data mining and surveillance continues to become the dominant business model in silicon valley and elsewhere.
Interview conducted on 28 August, 2020
Mark Andrejevic
Mark’s recent book Automated Media considers the politics of automation through the ‘cascading logics’ of pre-emption, operationalism, and ‘framelessness’. We talk through some of these ideas, along with the limits of ‘surveillance capitalism’ as an analytic frame, ‘touchlessness’ in the time of Covid, ‘operational listening’, what automation is doing to subjectivity… and how all this relates to reality TV.
Interview conducted on 21 August, 2020
Shannon Mattern
Leading off from Shannon’s essay “Urban Auscultation; or, Perceiving the Action of the Heart”, which addresses machine listening in the pandemic, we talk about the stethoscope, the decibel and other histories of machine listening, along with its epistemic and political dimensions and artistic deployments.
Interview conducted on 18 August, 2020
Audrey Amsellem transcript
James Parker - Thanks so much for joining us, Audrey.
Audrey Amsellem - Thank you so much for having me.
James Parker - Would you like to introduce yourself however you see fit?
Audrey Amsellem - Sure. So I’m a lecturer at Columbia University, where I teach in the music department. And I’m an ethnomusicologist working on sound and property. Okay, that’s the headlines.
James Parker – And how is it that you go from, you know, being a musicologist to working on sound property surveillance? I mean, I’d love to know a little bit more about your backstory and how you arrive at the problems and concerns that you have. Is that a story you feel like telling? Audrey Amsellem - Sure, absolutely. Yeah, it was definitely a progressive route. I’ve been interested probably since I was a teenager in this concept of the possibility and impossibility at the same time of owning sound. This idea that there’s this fundamental tension with sound. something that exists in the air and that is somehow subjected to property laws are actually quite rigid. And it, you know, music to me, particularly as a teenager, and I think this is the case for most people, right, that it feels like a space of freedom. And so any kind of constraint that was that is put on that was sort of an affront to my freedom, right? It became very political to me. We were chatting a little bit about this earlier, but I’m not a musician myself. I never really wanted to be a musician myself, but music was so important to me since forever that I always felt a sort of deep sense of injustice maybe towards the ways that various musicians are treated. And because I felt indebted to musicians, right? I felt grateful for growing up at a time also where music was so widely available.
So I was a teenager when piracy was all the rage. And so this was a very great moment to be a teenager because all of a sudden we had access to this enormous amount of music and it felt like it was endless and it felt so democratic in that particular moment of my life. And so you have these various things I suppose are converging. And that was really sort of starting point of my reflection around music and around music and power. So the power that music had over me, but also how power is exerted through various politics of access to music and circulation of music. And so then I was thinking, how do I make this my life? Which at first was more, maybe I’ll go into the music business and, you know, try to help out musicians or maybe I’ll become a copyright lawyer. And I moved to the US, so I’m from France originally, and I moved to the US and went to community college. And so I studied music at community college But I didn’t really know what I was going to do with that and I wasn’t even in a four-year college yet. And it was actually meeting, well, I didn’t meet him, but it was going to a talk by George Lewis, who’s an amazing composer who teaches at Columbia where I am now. That really, actually, he introduced me to musicology. I didn’t know what it was. And I thought he was just so incredibly smart and inspiring. I thought, you could do this kind of music philosophy. And I just saw, you know, how interesting is it to think about music this deeply and this intensely. And so that’s how academia came to be, you know, little by little, I would say. And, you know, first, I started working on piracy when I was more of an undergrad and sort of various politics of circulation of music. And then that translated into no thinking more strictly about copyright law and then thinking about various ways in which sound is constrained by by laws and then so surveillance came sort of progressively through that route.
James Parker - That makes a lot of sense and listening to you speak I’m suddenly thinking I wonder if you/we were the first generation for whom the mode of circulation of music was sort of really really highly politicized and present and continuous with our encounter with music per se. I just don’t remember my dad or like even myself as a child, talking about records or tapes or CDs as a political object, as part of their very encounter with music. But the way you’re talking makes me think that, from the beginning you learned to love music through the medium of a conversation about piracy, freedom. Of course, those issues are always there. Of course, the CD is a politicized object, an object of enclosure, and you can go back much further… But I don’t feel like this was as much a part of the public conversation. I mean, I’m making this up, but the way you described it, that really jumped out at me. How could you not think of music and property together, growing up when you did, and encountering music through file sharing services. Of course, there’s a whole generation for whom that’s true. I’m riffing…
Joel Stern - But even in the 80s, sort of in the period before the internet, I remember, you know, the messaging that home taping is killing music is destroying the music industry. And there was all the anxiety around people recording things off the radio and circulating it between themselves through mix tapes, even though that was just a sort of grassroots, anti-commercial kind of culture. Yeah, the growing up in that sort of era of Napster and you know, Soulseek and kind of platforms like that really was a sort of inadvertent political education.
Audrey Amsellem- Yeah, but I do think you could probably retrace various important moments in history in which music and as you said the way that music is distributed actually impacts its sort of political appeal in some sense. I’m thinking also in the 50s with the radio and and rock and this idea that the first time because you have the growing middle class, the first time teenagers are able to buy their own music also. And that is sort of a move, rock becomes this anti-conformist and rejecting the traditional values of the parents through buying Elvis, like the white middle class Americans buying Elvis records as actually a way to form their own identity that is very radically opposite of that. parents. I think there are these moments in history that you could probably pinpoint to and that are deeply tied with technology, obviously.
James Parker - Yeah. I’m now thinking my instinct there was a bit naive. I had naturalized radio, like the whole the relationship between pirate radio or just youth radio per se and rock music, obviously, was absolutely vital. And we’re only talking about the West here and you can think about the role of radio in so many other cultures. But it is still interesting in any case that for you music and property relations sort of come together. You know, and obviously you can see how that bears out in your thesis.
So yes, we’ve read your brilliant thesis at Columbia. Was it in the musicology department, ethnomusicology? Yes, ethnomusicology. So it’s a music department, but within the music department, ethnomusicology is for a subfield. Right. And it’s sort of all about, on the one hand, the politics of ownership, the history of recording and so on, but very much kind of situated in the present around, I guess, I mean, how would you summarize it? It’s your thesis. I was going to say the relationship between sound and surveillance or something like that without getting too much into the details, which we’ll obviously do.
Audrey Amsellem - Yes, so the dissertation is a series of case studies. It’s three case studies. First one is Spotify, then it’s Alexa, so voice assistants in general. And then the third one is an object called Link NYC, which is the sort of communication hub that is part of the Smart City Initiative in New York. And picking these three objects, the idea was to try to show the relationship between various forms of sonic properties that are historically informed and legislative, and how these various forms of sonic properties have actually informed surveillance capture. and how these three devices are actually sort of products, right, of surveillance capitalism in the way that they are listening to people and collecting massive amounts of information about people. And I was trying to show in this project the relationship between particular specific notions of Western property that have been enacted through the history of the West and surveillance capitalism and show how there is this relationship of continuum, and how sound and music is relevant, let’s say, a way of looking at the history. It’s not the only way. But I think it’s a particularly poignant one, it reveals particular things about about the way the West actually conceives of property in a more general way.
James Parker - And so you end up saying that, or diagnosing the present or, I don’t know if that’s quite the right word. In terms of the ‘neoliberal ear’. So you say that these three different case studies are all kind of exemplary or productive of this thing you call the neoliberal ear. So maybe as a starting point, just to kind of inform our conversation about the case studies as we go on, could you say a little bit about the neoliberal ear? It’s very suggestive. It’s quite close to, you know, in some ways, the topics that you’re describing are quite close to what we’ve been investigating with machine listening, but obviously, like the neoliberal bit is quite, it just has a different inflection. So, I just love to know what you’re thinking of when you talk about the neoliberal ear and the consequences it has for your analysis or maybe the opposite, like how you arrive at the neoliberal ear through your analysis.
Audrey Amsellem - Sure. Yeah. So, the neoliberal ear is this concept that I developed, which is defined by is defined by being the set of listening practices within surveillance capitalism. So the various ways in which surveillance capitalism listens, the various ways in which it can be a listening entity. And that specific way of listening to the world, it doesn’t arise out of nowhere, right? And it actually predates neoliberalism. So it’s called neoliberal era, but this modality of listening is a consequence of colonial ideologies, of collection and dispossession and extraction. So it’s the way that the West has been listening through this tripartite process and rooted in colonialism and in capitalism and rooted in the idea that Western powers can sort of come into, when we’re talking about colonialism, we’re talking about a place, we’re coming into a place that isn’t theirs and sort of observe it and seize and dispossess, right? And this is where the framework that Dylan Robinson, who’s an ethnomusicologist, And so this is a framework that he develops in his book called “Hungry Listening.” And that framework is crucial because he deconstructs the constant state of starvation of the colonizer, which I use his framework also to compare that to the constant state of starvation of tech companies for data. And so that’s more of the historical part of it, but then you have neoliberalism that emerges and this idea of the free market profit seeking over public good, like valuing profit over public good, individualism, the emphasis on that’s more actually something Robin James talks about the emphasis on the quantifiable and statistical and the sort of creation of a normalized subject.
So the ideology of neoliberalism, but also its mode of governance that is coupled with the recording and tracking just capacities of the modern technology and together that forms a specific way of listening. And so we see emerging not just the ability or I guess the perceived ability to listen for and to collect subjectivity, but actually we see the deep rooted belief that this collection of subjectivity has inherent value and that it can be marketed and quantified and sold for profit. And so this particular ideology is actually guided by politics of extractions that are characteristics of markets and late capitalism that are characteristic of neoliberalism. And so if there were regulations, there wouldn’t be or stronger regulations, there wouldn’t be private companies that reach this nation state power through controlling the means of communication. Right. So I thought it was interesting also what you said about, you know, in a way, what is the difference between neoliberal ear and machine listening? And I don’t know, that’s a different thing, really. I think it’s It’s more maybe about where the focus is. So when we say machine listening, it’s implying technology. And the term ‘neoliberal’, I guess, is not necessarily implying technology, and then to say technical, it’s more…
James Parker – Yes, I think it is a matter of starting point or something like that. When I think about machine listening, I mostly begin with machine listening, this is in my head, I don’t know if I speak for you, Joel, because I think that machine listening is a very… It’s a term that sort of echoes machine learning and is kind of part of a vernacular. I feel like people get it. So this partly just this kind of, you know, it’s an easy term to convey what you’re talking about. It doesn’t sound too technical. But as a matter of fact, it’s also literally the term that computer scientists and musician computer musicians use to describe the field. And the thing is that the immediate move that you make next or you need to move in my thinking is to say, well, when we say machine, we don’t just mean that there’s some kind of technical thing, you immediately need to situate it in its cultural, political, economic, social context, so that the technical is always and the machinic is always embedded and mutually constitutive of the social. So that machine listening becomes pretty close in some ways to what you’re describing. Once you hash it out on the page or in conversation a little bit. But the orientation is maybe a little bit different. I have to do the work to explain, you know, well, actually, I mean, something that emerges under but also helps to produce surveillance capitalism or one of its synonyms is embedded in neoliberalism but you’re sort of beginning at the other end and that’s that’s fine and interesting I think.
Joel Stern - We get into tricky territory sometimes sort of having to distinguish between human and non-human listening and sort of the the implication that listening by machines is automatically sort of extractive whereas listening by humans might be sort of automatically reparative, which are, you know, which are sort of tropes within sort of sound studies. James Parker - It’s not just that, I mean, machine is like a very flattening word itself, even amongst machinic techniques. I mean, so, you know, machine learning is basically a hype term for in industry. And so the moment you actually dig down into the practices and the history of machine listening, it’s like, well, what kind of techniques are we actually talking about? Are we talking about hidden Markov models? Are we talking about statistical led AI practices or rule based ones? Or are we even why are we focusing on AI at all? Like is recording part of the history that recording technology as such? So, I don’t know, that they all you go down all of these rabbit holes.
Joel Stern - Both of our projects owe a debt to George Lewis. That’s also something we have.
James Parker - Yeah, that’s right. That’s right. Okay, so with that kind of framework in mind, so I guess it’s a study of, yeah, I mean, the proliferating technologies of the present and their, for listening and their embeddedness in, you know, some of the most extractive and exploitative political and economic systems of the time. So what, I mean, that sounds like a good segue into Spotify. Would you like to say a little bit about, you know, Spotify, how it sort of exemplifies what you’re calling the neoliberal ear? I mean, however you want to begin, just in general, or with an anecdote, or whatever you like.
Audrey Amsellem - Yeah, you know, I think Spotify is a great example of the neoliberal ear in my mind, because it actually allows us to see the history of it. So some of it, I was talking that I was talking about a bit earlier, but this idea that if you sort of trace the history of it, you could see the way that music has been subjected to various forms of control in the West, right? And so this is where someone like Jacques Adagli or Catherine Bergeron and those people are important because there are these historians who are looking at the various ways in which, you know, the Holy Roman Empire in the ninth century or the Venetian state in the 15th century, they’re strategically controlling the kind of music that gets circulated and who circulates it. And this keeps happening, right? And so then you have copyright law in the 18th century, and that will serve its own kind of form of control, especially the way that And I think piracy was a form of rebellion against that, right? That’s how I interpreted. Obviously, there’s various people interpreted differently. But the fact that Spotify is born out of piracy, and that is pretty clear, there’s the same people involved. It’s not a coincidence, right? So what I try to
James Parker - I didn’t know that Daniel Ek was literally ran a torrent company before he sort of, you know, went clean and started, you know…
Joel Stern - He didn’t go clean. He went much dirtier.
Audrey Amsellem - Yeah, I know. It actually kind of makes sense, right? When you start to see that connection, it actually makes total sense because there’s a lot that Spotify is borrowing from piracy and actually a lot is borrowing from piracy surveillance, which was the response, the sort of, yeah, the response against piracy. So I guess I tried to highlight that the history of music is I think I mentioned this at some point in the disc that the history of music is sort of burst of creativity and burst of controlling response to that creativity, right? So piracy I consider to be a burst of creativity and piracy surveillance, which is a concept that was theorized by legal scholars, Sonia Katiyar, and that would be a burst of control. So it’s this idea that because piracy surveillance, the basic idea is that because you have piracy and that’s that’s illegal and that’s wrong according to the RIAA and according to some people, then it gives, it normalizes and gives a form of legitimacy to actually track people’s behavior online and do it in a way that’s unprecedented and do it with somehow a legal structure that supports it, even though some people would argue some, I think pretty serious legal scholars might argue that is absolutely going against the Fourth Amendment, right?
James Parker - So can I just can I just interrupt because I find this idea of piracy surveillance really interesting. I hadn’t come across it before. And I just I just wondered if we could slow down and hash it out a little bit because so it seems like you’re making the argument in the dissertation that piracy surveillance sort of sets the groundwork basically for forms I mean, this is sort of well before we’ve got, you know, Google, sort of, you know, in the Shoshana Zuboff story, we’re sort of pre surveillance capitalism proper, right? You know, but, you know, early days of the sort of publicly available internet, enormous explosion in music piracy, and that these companies the music the big the big record labels start to surveil. Download basically and so that’s that’s piracy so it’s perhaps you can give a more nuances definition but then that technique of surveilling music consumption habits originally in the form of piracy becomes normalized i think is the argument and then becomes foundational not just to the music industry that goes on to be. No spotify but this technique. sort of, you know, goes, goes wild. It’s not that, you know, cookies and other things are part of a similar story. But, you know, I think in the strongest version that you put it in the argument, you say that it’s like some foundational to surveillance capitalism. I haven’t heard that argument before. I thought it was really interesting. So I just wondered if you could hash it out a little bit.
Audrey Amsellem - Yeah, I think and, you know, I do believe it’s foundational to surveillance capitalism. I do believe it’s this particular moment. But, you know, again, you could go back in history and trace similar moments, but just focusing on piracy surveillance for a minute. Yeah, absolutely. I think this idea that to just be able to justify something that I don’t think without piracy would have been justifiable. It might have had to do also with the fact that people were not, most people were not too tech savvy at the time and didn’t really also know maybe that they were being surveilled, but there was a whole discourse that you may remember that was saying, you know, how wrong piracy is and how it’s ruining the life of musicians and how musicians can’t make a living anymore. Not just musicians, but it was also about movies and various other forms of I think pre Cambridge Analytica, I would often hear people when we would talk to them about surveillance and privacy in general. And they would say, you know, oh, I don’t really mind surveillance. I don’t really mind that the government knows what I’m doing on Facebook because I’m not doing anything wrong. And that whole way of thinking about surveillance, which I think is very actually quite harmful, but that was pervasive for a long time. And I think it’s changing now, but it was pervasive for a long time. It was this idea that there was somehow the justification, things that were justified, that it was justified for the government to enter into your private computer and your private network and see what you were doing on your computer. And part of that has to do, I think, with the way they legitimized that discourse. And part of it had to do with the fact that people weren’t really, maybe really thinking about the implications of that. But to me, the implications are very, it’s very direct to surveillance capitalism, because all of a sudden, it’s just already, it’s often how it works actually with technology and with the maybe with the law too, is once something has happened once, then it becomes okay, right. And so this idea that this was already a system that was in place, then it became just acceptable and accepted for government entities and then private corporations as well, to just get your information in that way that just became the norm in a way. James Parker - So there’s a whole series of like major lawsuits. At first like individuals, you know, sent all of these letters and they, you know, there was that phenomenon of the record industry just going after, you know, some kid. And then that sort of that a lot of that energy got directed. I mean, I don’t know how simultaneous they were. I can’t I don’t recall, but to, you know, the big platforms like Napster and LimeWire and so on. And then sort of subsequently after Napster shut down and things we get, you know, it’s pretty quick that we get the data, the data tracking industry kind of pairing up with the emergent streaming platforms. I mean, they sort of emerge simultaneously. It feels like in your story, Echo Nest, which is a company we have a little bit of familiarity with, sort of does a lot of the work for you because so Spotify, you know, sets up this streaming platform. the promises that you get all of the access, but none of the illegality, lots and lots and lots of venture capital, more, more money, money, money, money. And at the same time, or very, very quickly, they start to become a effectively a data analytics company more than a music company. And echo nest is the sort of, I don’t know if it’s I don’t I can’t I again, I forget the timeline. It’s not my thesis, but echo nest is is a company they buy up in order to help help drive that turn. Is that right?
Audrey Amsellem - So, yeah, so the Echonest is a music analysis tool, right? So what it’s going to do is it’s a tool that analyzes the song for its musical events, but also for its cultural cache. It also goes and looks at the way that people talk about the songs in a blog or something like that. So that technology is used to make playlists. various playlists and recommended targeted playlist as well as more general playlists. Some of the plays have Spotify or human made and some are made through the algorithms. And so it’s very much, you know, part of this like quantifiable statistical part of neoliberalism, obviously. That in itself has been theorized, the equinox has been theorized by several, many musicologists actually I’ve talked about it. Eric Drott has a fantastical article about it. Robert Prei is a great, I think his dissertation actually was specifically on this. Nick Seaver also talks about actually algorithmic recommendation as traps, as entrapments, which I think is a very interesting way to thinking about those as well.
So the idea of being able to take a song and sort of decompose it in some sense, analyze its various events and then attribute a sort of quantifiable element, quantifiable sort of ranking to it, right, in order to then associate it with a person who recommended to a person is a particular form of musical analysis, which is this idea that you can sort of infer the preferences of somebody based on what they’ve listened to in the past. And but that information then gets sold to advertisers. So this is where the surveillance bit of the aconess comes in is that that information gets sold to various kinds of advertisers who are then using this information to infer things about us. And so there was some journalists who were actually looking into this and talking about how musical taste has been used to infer political allegiance. So like some apparently if you listen to Pink Floyd, you’re more likely to be Republican or something like that, right? It’s this idea that somehow you could infer people’s politics based on what they’re listening to. So that seems, you know, on the first degree, kind of not very harmful.
But then when you look at this, you know, post Cambridge Analytica and you look at the Cambridge Analytica relationship to this, it does seem pretty concerning and also, I mean, idiotic as well. It’s obviously nonsense. But it’s this idea that somehow music has this particular power. that it can somehow our musical taste and our musical listening practice can tell they can tell us things about ourselves that we don’t even know right that is somehow a portal into the unconscious and it’s a very seductive idea that idea it’s made to be true it may not be true but it’s very seductive and so that’s very easy to sell to an advertiser.
James Parker – I mean, it’s a staple of youth culture like since… I mean it’s about which band T-shirt you’re wearing. I mean, music as a tool of like self-identification, like it’s not just foisted on us by, you know, surveillance capitalists. It’s a sort of a detourning or capture of something that people sort of feel quite strongly.
Joel Stern - Yeah, I mean, even going right back to the sort of teenagers buying records in the fifties and it’s both an assertion of an authentic youth culture, but also the construction of a new consumer sort of demographic of the teenager with certain consuming habits and sort of providing that data. But just hearing you sort of talk about this, you know, mobilisation of Spotify, which the majority of people in the world would think of as a music delivery platform, but which in fact has its own sort of listening and practices of extraction… But then the listening preferences of people on Spotify are so shaped by the platform itself, you know, increasingly so, that the idea that it kind of provides access to some sort of unconscious sort of internal kind of information about that person, it’s just, you know, that that sort of listening subject is already fully entrapped in the kind of limited range of preferences available to the platform and its recommendation algorithm.
James Parker - So it’s like a paradox that they can’t, they have to leave it unresolved because they want to sell to the listener perfect recommendation according to their authentic tastes. And meanwhile, they’re selling to the advertising industry, the ability to sculpt and entirely determine those tastes. And they’re both essential to the marketing of the product.
Can I just – you probably have some responses to that - but one of the things that I find interesting about Echo Nest is that they do audio analysis. So, Echo Nest comes out of, partly comes out of music information retrieval and it’s one of the first sort of like companies to really commercialize… I mean, music information retrieval in general like explodes right after the Napster take down and music copyright protection becomes, you know, we see audio fingerprinting explode as a subfield of music information retrieval at exactly the same time. So music information retrieval is one of the few examples where there does seem to be a sort of listening, a sort of true, I’m using air quotes like ‘listening’ going on. So there’s an attempt to extract something from the audio sort of quote unquote itself. And then that sort of is put alongside or together with, you know, analysis of tags and all of these things that are sort of getting at the fact metadata. Yeah, various sort of user produced metadata.
So I’ve always been intrigued by the fact that like, you never really know, like, I’ve found it very difficult to disentangle how powerful the quote unquote music analysis that sort of audio analysis actually is for companies like this? I mean, I don’t know if you know the answer, but I found it interesting because I kind of weirdly - and it seems like you don’t have this hang up - like searching for examples and it’s partly this framework of machine listening. I’m like, well, that’s an example of ‘listening’. I want to sort of follow that example. And for you, it’s sort of, you know, what’s the difference really? They’re both extractive processes.
But I’m fascinated by the sort of the rhetorical appeal to be able to analyze the music itself, quote unquote. And the fact that we actually have no idea. It seems like often what’s driving algorithmic sort of playlist production is simply what other people have listened to. It’s not even like tagging of data. It’s like, well, you know, another 30 something year old male, you know, listen to this kind of song. And so here’s another song they listen to. It’s sort of a bit dumber. A lot of the kind of actually used techniques are dumber than then they seem or dumber than the sales pitch.
Audrey Amsellem - Yeah, it’s interesting to think of when is a machine actually listening and when it is actually just transcribing, when is it actually just transcribing into text and then speaking. There always seems to be that medium of text that has to be in between. But the data from the Echo Nest can actually be used even beyond what is maybe particularly interesting for composers, I think, but maybe it’s also scary, I’m not sure. But it’s how actually this data is sold to record labels. So data from the econ is sold to record labels in which they will give, you know, their their econ as musical, like music analysis of a hit song, let’s say, to the record label. And then they can compile this information and say, well, in 2022, the top 10 songs all had a BPM of 120 and all had minor chords and whatever it is that they come up with. And then that actually allows potentially a record label to say, well, now we can plug that information to actually make music that will fit the taste. So it’s sort of dystopic in every way you look at it. It keeps feeding itself.
And now it’s not just the curatorial platform that actually creates a musical taste, as Joel was just mentioning, was saying that the musical the habits of people are actually shaped by the way the platform is actually designed and how the user is sort of forced to interact with them because of the way the platform is designed. There’s actually the music itself, right? Because the more you’re going to listen to a certain song and that becomes a hit song. And then, you know, then we, this formula is supposedly going to sort of emerge. So yes, it’s a never ending process, I suppose.
Joel Stern - Audrey, one of the artists that we’ve been working with quite closely on machine listening composer, Tom Smith, who’s very interested in sort of cultures of automation. One of the works he made for our program is called Top 10. And what he’s done is he’s taken the top 10 tracks on Spotify from each country and created a single track that averages out all of the qualities of those top 10. So pitch, tempo, rhythm, frequencies to create a kind of generic track for each country. And then he performs these DJ sets where he will sort of play these generic pieces and they sort of simultaneously, you know, quite revealing of what happens when you average out musical material. You know, it becomes increasingly both generic and sort of unlistenable, but also I suppose the absurdity of the kind of logical endpoint of some of these practices. But yeah, might be.
I wonder James, did you put the link in the chat or if not, I’ll do that. But that, yeah, that’s the work I was immediately thinking of. Yeah, and it’s sort of, it’s just a really lovely sort of counterpoint to the discussion we’ve having. So anybody listening, I’d encourage them to go and look it up just because, yeah, of the show, not tell, you know, we’ve been talking a lot about the techniques and just as an artist who’s working with the exact sort of similar concerns, but sort of in the medium of music. It just is. Yeah, I’m sure there’s many others.
James Parker - I wonder if this is a good point to segue on to smart speakers and voice assistants. Although I was also struck by there’s another little anecdote in your it’s not really an anecdote, but yeah, an anecdote in your thesis about music for mums. Because you talk a little bit about the different ways in which we’re sort of bracketed and, you know, so techies, like it’s just interesting. You give the example of techies and moms, like what are the categories that are used to segment us? Like how does that you know, it’s not just men or like people who listen to Pink Floyd, but the idea that there’s something called a techie and then obviously a mom. But in the thesis, you give the example of the this category of the mom who’s constantly recommended Disney songs. And despite all of this rhetoric of being sort of incredibly refined and smart, that somehow Spotify and these other streaming companies can’t distinguish between a parent or a mother’s own listening preferences and the ones that they are doing as I do on the constant demands of their children. Can I listen to, you know, I don’t know, whatever, some Disney song. And it’s just, I don’t know, just yeah, I guess both of my point about the echo nest and this anecdote just point towards the dumbness and the extent of the hype. And yeah, the way that like hype and misdirection are sort of really deeply bound up with a logic of neoliberalism. Um, you know, that’s something that’s not at all brought out if you’re thinking in terms of machine listening, but once you start to think in terms of the neoliberal ear, it becomes much easier to think about just the real extent of marketing bullshit, uh, as like a fundamental feature of the listening practices and the systems of power and control that are going on.
Audrey Amsellem - Yeah, yeah, that was also sort of funny to me when I saw that. And it’s not like I’m working with some raw data and I was able to extract this. This is part of their marketing discourse in which there’s like, well, moms are more likely to listen to Disney songs. Like, are they? Are you sure? Not only have you come up with that, but then you also advertise that as a thing to sell to advertisers as this very exceptionally smart and sophisticated data that is obviously none of that, right?
James Parker - And you requested your own data from Spotify. What was that experience like?
Audrey Amsellem - It was very silly. So basically I requested my data when the GDPR came into effect because Spotify is a Swedish company, right? So they are bounded by the GDPR. And the GDPR part of it is that there’s a duty of transparency. So I went to that contact form on the website I asked for my data, I’m talking to a bot probably, right? And they sent me back this tiny file which has very basic information. Basically the information I already knew they had because it’s information that’s visible, like my playlists, my billing information, what is already obvious that they have. So I asked the bot again and I said, no, you need to send me the real stuff this time. And then I get a much larger file that was essentially incomprehensible unless you know how to, you’re a data scientist basically. So to me that’s not, it’s not just not actually abiding to the GDPR because it’s not actual transparency, right? It’s the idea that the GDPR was to empower people with knowledge, right, about their digital lives and their habits. And this is not only not doing that, it’s actually doing the opposite. It’s actually precisely disempowerment because when I looked at this, you know, hundreds of pages of numbers and letters in random order to me, you know, you look at it and you feel profoundly disempowered by both sort of sheer amount of information, but also your complete lack of ability to understand it without a computer scientists or an engineer department and transcribe this data in a way that makes sense to me. But yeah, hopefully I’ll be able to do that at some point. But on my own, I was not able to find that out.
James Parker - All right, so now let’s segue into smart speakers then. And you know, maybe as a starting point, how we should think about the relationship between streaming platforms like Spotify and assistance. I mean, you know, I have my own like slightly, you know, underdeveloped ideas about this, like, you know, in the same way that I think you quote Lessig at one point saying that illegal downloading is the kind of crack cocaine of the internet’s growth. I sort of feel like without, without music streaming, there’s no, there’s no smart speaker market. Like I just, I mean, whatever you can ask for the weather. But really what a smart speaker is, I mean, at least at first, maybe they’ll become more deeply integrated into whatever, but at least at first, it’s a slightly crappy kitchen radio for just listening to some music, not not in high fidelity or anything, just, you know, what while you’re doing the dishes or something like that, you know, and if you can’t stream music on it, I just don’t see how you’re going to sell any.
So they seem like really closely related maybe. But yeah, I mean, that’s just one possible way of thinking them together. I mean, how do you understand the relationship between an entity like Spotify and its listening practices and something like Alexa? Yeah, maybe we can go from there to ways in which they diverge as well.
Audrey Amsellem - Yeah, no, I think that’s exactly right, because it has become a pattern that this tech companies are basically using music as a way to entice people into buying various device. And you can see even the history of the iPhone like this, you know, before the iPhone, there was the iPod. And, you know, this idea, the iPod was so revolutionary. I mean, this idea you could just like carry in your pocket, 10,000 songs that you know, probably pirated. And then you have the iPhone. And actually remember when the iPhone came out, it was maybe I was in high school, so it was maybe I don’t know, 2008 or something at least when it came out in France. And I remember seeing the commercial for it on television and I thought, it’s an iPod, that’s a phone, that’s so dumb. Who would want to buy this? This will never work. And so obviously, you know, I’m not a techy for this reason. To be honest, that’s all futurists make, you know, catastrophic mistakes. So I say no different from Elon Musk. Sure. So yeah, so this idea of these companies are using first music to entice people. And so once, you know, the iPod was in everyone’s pocket, it was much easier to sell the iPhone. I think this is absolutely what you said also with Spotify and with streaming services in general, actually. And a smart speaker, it says that people were buying them to listen first as a speaker, and then as a speaker that you could command with your voice. So you could say, Alexa, play this song. And so Spotify is obviously sort of embedded within Alexa. in that context. So it is this pattern.
And as you mentioned, the Lessig quote, I really like that quote, this idea of music has being so, being always perceived as harmless in a way, always says this thing that can only bring positive things, right? So in that sense, what could possibly be the harm of having a speaker? Because that this main point of it is to have music in your home and what is more beautiful than that, right? And actually, because sort of taking advantage in a way of the pleasure of music by embedding it with invasive technology, it seems to be a common pattern. And taking that further.
James Parker - Sorry to interrupt, but just this sort of association of listening with care is a kind of further extrapolation of that or even somehow sort of of the same order. Do you mean that the the device is caring for you by sort of attending to your atmospheric needs or what have you?
Audrey Amsellem - Yeah, absolutely. I mean, you know, in the way that comes through, for instance, in the in the smart wife or in that kind of the continuity with sort of domestic labor and servitude. But yeah. Yeah, no, absolutely. I think it’s definitely part of this idea of listening with care and then in our point of view, a perversion of that care, right? There’s a perversion of what music is in some sense. There is a perversion of what the female voice is supposed to be. It’s not supposed to be this sort of submissive servant. So yeah, it’s always sort of playing with these poles and it’s quite dystopic actually in some way.
James Parker – So… to be really crude… streaming music streaming is a kind of Trojan horse for getting a listening device in your home in a certain kind of way. And then once it’s in the home, the key difference, although Spotify also is partly involved in this game, I think you mentioned at one point, is that suddenly the voice becomes… I mean, yeah, when I think about my own encounter with speech recognition technologies, it’s basically like I remember kind of the end of the ’90s. I think I had to — there was some call routing that was done with speech recognition. It was pretty crap and frustrating. And then a few friends, like, who had computers, like, had, you know, maybe their dad or something had bought an automatic description thing, but it was all pretty rubbish. And, you know, I mean, the history of speech recognition is a very long history. It’s as long as there’s been computing, there’s been people trying to do speech recognition. But it’s sort of undeniable that like the voice assistant is the moment of it sort of mainstreaming. And although Siri is, you know, often it’s always talked about as the first one, uh, Really, it’s the smart speaker, I think, that kind of takes it, you know, takes it really, truly big. And so it’s the, you know, music is the route into the home. But when what happens when the smart speaker enters the home is that the voice and speech becomes an interface really for the first time. And then you have a lot of analysis of the ways in which the voice and and speech is exploited and sort of, I mean, yeah, you also talk about a lot of other dimensions of the smart speaker, its gender dimensions, the labor practices, the diverse sort of forms of labor that are kind of hidden in the production of this tiny little object. Yeah, so, but yeah, where would you, what should we talk about? Should we talk about the labor stuff or the, you know, capture and instruction?
Joel Stern - Let’s put that voice as an interface and as a commodity that is sort of, you know, becomes viable via the smart speaker.
Audrey Amsellem - Yeah, so there’s the voice of Alexa or the smart speaker, whatever, you know, that voice assistant is. In my case, it’s Alexa, but we could talk about other ones. And the voice, obviously, of the user. And I think Voice recognition is among the scariest things that are happening today. We’ve, in the past few years, there has been a really fascinating amount of scholarship on facial recognitions and the danger of facial recognitions. And I definitely see voice recognition as a continuation of that. I am quite concerned with the technology that’s being developed both in, I spent quite a bit of time talking about patents in the dissertation. There is also sort of already existing technology, obviously I’m talking about Alexa and Amazon Halo is another example that I give and this idea that voice it’s completely sort of normalized for these devices to record your voice and then create voice profiles and also determine your identity based on the sound of your voice and so gender is an is an obvious issue here and not just the gender of the voice of Alexa, right, but actually gender recognition or misrecognition often through the voice of this idea that somehow, generally speaking, you know, male voices are going to be lower than female voices, which is completely going against all the ways that we’re actually trying to change these things, right, that we’re actually trying to have people define their subjectivity. And I think anyone in sort of voice studies or sound studies would be actually quite critical of this. And I cite some some people like Nina Hindsine who did like incredible work on this.
James Parker – This is similar to the point I was making about echo nest in some ways, but one of the things about a smart speaker is that it’s a portal to a form of listening that is always changing and which you never really know. You never know what kind of listening is doing when you sign up to the, terms and conditions, you know, it’s not like you get an update. You don’t get an email that says, now we listen to, your emotion now we and in these ways and we do it exactly like this. Now we’re listening to whether you sound sick now we’re listening. You know you never really know and that’s part of the whole point i think you say at one point in the thesis that like you know these are lost leaders like they don’t make any money off smart speakers the only reason you’re gonna sell a smart speaker is because the initial whatever like 20 alexa skills are now 100 000. The forms of analysis the forms of products are offered on this platform or this eco this sort of emerging ecosystem is sort of massively more than just being a, you know, a music delivery device in your home. So one of the things I find, you know, I always find hard with these things, is that when I read a patent that says Amazon’s gonna offer, you know, that they’re imagining being able to sell you chicken soup because you sound sort of sick. When how much of this stuff is actually happening? I haven’t done that analysis myself and I know and again I don’t really understand. I mean I presume that it is or very soon will be. But do you have a sense of like what are the actual concrete forms of attention to the voice that are in fact being used right now and which are more speculative and future oriented?
Audrey Amsellem - No, and actually that is part of the, that’s actually part of the framework I develop around the neoliberal era is that it’s always opaque and always ambiguous and that you never really quite know. And so legally that would be complicated for them to implement this. That’s actually, I mean, it’s obviously something that should be regulated, but it’s not quite the same as doing facial recognition. in the US at least, I’m talking about the legal context that I am familiar with. But it’s legal in the US to have a CCTV. It is not legal in the US, and on the street, it is not legal to have a microphone on the street. It’s completely illegal. So in that sense, there is a legal sort of gray area here, which is a good thing. But we don’t actually exactly know what the technology, how the data that how their data is collected is being used, we don’t really know because there is no system of sort of checks and balances. You have a privacy policy, obviously, and it’s one you have to agree to. And if you read it attentively, you’re probably going to find out stuff that you didn’t know were happening when you were using the technology. But the reality is that there is nothing that’s obligating these companies to be fully transparent in their privacy policy. And also, these are policies are written by very smart lawyers who have a particular way of wording things that may, you know, sometimes I talk about this idea of saying collecting information such as and this idea of such as well, but you should actually cite everything because such as is not satisfactory here, right? So we don’t really know what’s going on. It is, I know even if you requested it, you presumably just get sent a giant log that you need to hire a computer scientist to untangle. Yeah.
Joel Stern - Audrey, I really loved this chapter. It was amazing to see the analysis of the advertising and the particular kind of images of domestic normativity that they echo. I wanted to ask you about the section on speech emotion recognition. because I loved this, I mean, I was horrified, but also loved the fact that the Echo has a frustration detection tool. And then you sort of go on to talk about the halo as a form of tone policing. Could you say something about this sort of relationship between these devices and the sort sort of extraction of information about people’s emotional states and then how that information is then sort of exploited and used.
James Parker - Yeah, I think it’s also part of this normalisation of ubiquitous listening, this normalisation of surveillance essentially, is to sell a frustration detection tool or even Halo as a sort of independent device whose goal is mainly to monitor you and then somehow and I don’t know how successful that device is. I haven’t heard too much, you know, I’ve never seen people, anybody with this. I don’t know how successful it actually is. But this idea that somehow this would be desirable, right? And so they’re selling this to you as something that is desirable. You should know the tone of your voice, you should, you should have Alexa be able to detect when you’re frustrated and to have this information. And that is sold to you as something that you need, right? But it is, yeah, absolutely. you said, you know, it is scary. It is scary because also, as most of this technology that actually we’ve discussed throughout this, most of it is doesn’t work right. And most of it is incompetent actually. Is the device scary? Or is it the fact that people desire it? Even scary? Yeah, it’s both. Yeah.
Do you know Chris Gilliard’s work? He’s written about this idea of luxury surveillance and he’s trying to get at exact. So the way that you situate the history of all of this especially with recording in the history of, you know, US but not just US race relations and the relate, you know, obviously, like, the recording industry. But just more generally, like surveillance is very closely tied up with capitalism. And, you know, obviously, Ruha Benjamin, but various people have sort of made a theorize list. Anyway, Chris Gilliard has this idea of luxury surveillance as a sort of white, but not just, white desire.
And so, you know he shows how that sort of surveillance as domination becomes a kind of way for testing out various different things or rolling out various forms of surveillance that are at the same time or just subsequently become objects of consumer fetishization. And then the racial dynamics of those play out really differently. Yeah, so i think that’s the desire of surveillance has a luxury good
Joel Stern – Do you mean in the context of the high end smart home?
James Parker - Right exactly. But the halo device is surely the perfect example because it is literally a prison-like home-surveillance bracelet. An ankle monitor effectively. But now it’s being sold by Amazon as a as a luxury product you know. So yeah.
Audrey Amsellem - No I think it’s very interesting that you’re tying this to whiteness. There’s something I mean there’s so much fascinating work that’s been done that’s being done also I’m thinking of Thao Phan’s work on and on Alexa as well and the way she’s sort of taking the sort of paradigm of the way we’ve been talking about gender and Alexa and sort of reversing it on its own head in a way and her racial analysis of the politics of Alexa and the process of listening is fascinating.
But yeah, I think there’s a small part in the dissertation where I talk about how the field of civilian studies in my mind in the way I’m seeing it, the way I interpret it, it emerges as a field because surveillance, it’s not because surveillance is a new phenomenon, right? It emerges as a field because all of a sudden it’s concerning everyone… including people who are not traditionally subjected to surveillance. Like white people. And people with means, and people in power. It becomes this serious field of inquiry only then. But the reality is it’s been the quotidian for black people obviously in America, also in the colonies… Jewish people. It’s been the quotidian for lots of populations. And so there is also that question. And I don’t know how much there is actually thinking through what is the race of people who buy these devices. And I don’t know if this is even information we necessarily want, but I would be interested to see if this is not a concern because white people didn’t have historically to worry much about surveillance. And so they’re buying these halos and these devices. and other people might be more skeptical of them and justifiably so.
James Parker - I think I’ve read that about Amazon ring that it’s brought up mostly in white neighborhoods and mobilizing a politics of like racialized fear around security and home security and so on. And especially also because part of the point of Amazon ring was to stop theft of Amazon delivered packages and you know, and then there’s a racialized dimension to the delivery people for the packages. And so, you know, often you just got a thing targeted at a brown person delivering a box, you know, let alone the neighborhood. So, I mean, maybe this is a way of segueing out from the domestic sphere, you know, into public space, you know, because the ring is kind of on the threshold of the house. Now, I think Amazon Ring is or can be hooked up to Amazon’s guard, I think it might be called, or maybe that’s the Google one. But yeah, there’s like a kind of a voice activated and microphone assisted kind of dimension even to home security cameras and so on. So you can see the logic sort of spilling out into the home. I haven’t encountered a voice user interface in public space space yet. I’m sure it’s coming. I’m sure that that’s what people are imagining when you know, the Toronto, the Google Toronto, you know, smart city and whatever. But you talk about LinkNYC. And that wasn’t an example that I knew anything about until I read your thesis. Yeah, so could you just sort of introduce it? I’m sure it’s less familiar to people than then Spotify and Alexa and you know, what it is. and how you think about it.
Audrey Amsellem - So LinkNYC is a communication hub. It’s basically these massive kiosks that are scattered around New York City, and they provide an array of services. Free Wi-Fi is a big appeal, but also access to city services. And originally, when it was launched, in 2016, it also had and it still actually has a tablet, but it also had access to a browser on the tablet such that people could go and browse the internet. There’s also a microphone because you can make phone calls and there are three cameras that are on top of the screen. So looking on the street and a one camera that’s on top of the tablet. So the cameras that are facing the street, they were originally, I think, equipped with sensors. This was removed unclear exactly what happened there. But the cameras are set to only film according to employees of Citybridge, which is a larger company that LinkNYC is part of, to record if there is somebody trying to vandalize the kiosks.
Obviously, it could, even if this is a case now, that could you know, change over time and this could be important data for companies to have. Where LinkNYC was, it became complicated and became an interesting object for me and for a lot of New Yorkers is that it was financed by a Google company called Sidewalk Lab. And you mentioned Toronto and that’s part of that same, that’s also Sidewalk Labs in Toronto. And that was sort of the first sort So big question that was happening saying this is all funded by Alphabet, which is Google. And the second thing that arose with LinkNYC is the removal of that web browsing option because a lot of homeless people were actually using the web browsing to just browse the web in general, but also to play videos on YouTube, to play music on YouTube. And that was loud. And so a lot of— local residents had complained apparently about how basically homeless people gathering around the kiosks and that the kiosks were attracting homeless people. And so some residents were unhappy about that. And as a result, they actually removed the web browsing. So that was sort of the departure for the larger exploration of LinkNYC.
James Parker - Can i ask a couple of follow up questions before we get into the listening specific dimensions. What this is a gonna sound silly is very basic. What did they think. Yeah this was for because obviously i mean twenty sixteen is. It’s only six years ago but sort of quite a long time ago in the history of smart technologies and smartphones and so on it seems. completely obvious to me that I would never have any reason to use a LinkNYC kiosk. I mean, I’m going to always have my smartphone with me probably. I can’t think why I would ever need free Wi-Fi because I’ve got 4G and lots of people now have 5G and everything. So, when I think of LinkNYC, I think, well, they’re embedding cameras. Everywhere. And microphones. And the only people who are going to use this are people who are sort of on the other side of a tech, you know, dividing line. And so part of me is like, well, is the point of this to increase accessibility to smart, you know, to the web and stuff is kind of like a library service or something, you know, and part of me is like, well, that’s surely the purpose because anybody who’s rich and never going to use it. So obviously, homeless people and people who don’t have endless free data and so on, poor people basically, and that’s going to be racialized and so on and so on, they’re the users for this. What else were they imagining? Surely that’s part of the sales pitch. So how did that— what was— but that’s like, was Google running a kind of— we’re empowering the poor with digital technologies line or what was the what was the story being used to sell Link NYC and why were they surprised when when homeless people started to use it? It seems obvious.
Audrey Amsellem - Yes, absolutely. And so this was the other that was completely part of discourse, this idea that this was a device that was going to bridge the digital divide. And they kept using that term over and over again. So it’s definitely part of the marketing discourse. And then when they cut off the web browsing, they sort of released that they had a series of tweets and then they had a a press release as well, and they said, you know, it’s to make it more accessible because if people are just taking over the tablet, then it’s less accessible. So this is sort of weird irony here. But this is where also I think Shannon Mattern’s work is important because she talks about how these tech companies don’t look at decades of good practice in public good spaces, right? Like how libraries have dealt with this problem. They’ve been dealing with this problem for a long time. But what tech companies, they don’t look at what libraries do, they look at other tech companies because public good may not actually be the main goal here. So you put hundreds of giant casts with speaker and a browser in the city, and it doesn’t really occur to you that this will significantly alter the daily life. And so all of a sudden people are just like misusing the technology, right? But were they really over there simply using it?
So their response becomes very punitive actually, becomes very exclusionary, becomes quite harsh and sad. I mean, that point about like, where’s the knowledge in relation to the digital divide in the public use of technology? And it’s like, it’s in libraries, duh, is like, it’s just such a great point. It’s just, it sort of seems so obvious now you mentioned it. And, and I’m not remotely surprised that tech companies didn’t tap libraries on the shoulder. But it’s like, yeah, it’s just such a perfect example of the kind of myopia of so many tech companies who are claiming to solve every imaginable problem. And, you know, you saw the same thing with like, you know, COVID and whatever. It’s like, we literally know how to do this. We’ve got public health. I mean, you know, there’s so many public health measures that are known to work. Meanwhile, tech companies are selling completely untested bullshit because because it’s a tech hack.
James Parker So yeah, so interesting. How do you think about LinkNYC? What is the relationship between the technology and its politics?
Audrey Amsellem - One of the aspects of it is that it presents itself as a public good. As I just mentioned, although it may have some positive application, it is not the primary purpose for these devices. because the primary purpose is data gathering. So throughout that chapter, I actually used frameworks of noise and listening to sort of structure my metaphorical understanding of these devices. And I talk about, for example, listening as silencing. So this idea of sort of exploring the tension exploring this tension between public good and private entities by basically arguing that the Neil Beale ear has a starvation for data, which actually in the case of LinkNYC, we say at least actually isolates and excludes people. But I think LinkNYC, we see and Spotify and Alexa, they’re very much tied in because of that sort of constant starvation, that constant thirst for data. And this idea that, that to normalizing the ubiquity of listening and recording and the various nuance between all of these. What kind of data does LinkNYC want specifically? What are they, what’s the untapped data resource are they harvesting? So it depends who you ask, right? So the first more obvious thing is that LinkNYC is equipped and I actually did mention this, but is equipped with two huge advertising screens on each side. So when you’re walking the street of New York, if you haven’t been in a few years or haven’t been, if you’re walking the street in New York, you’re going to see these like huge light sources with these huge advertising screens. So you have advertising here. One of the potential uses is to be able to target advertising So the free Wi-Fi component, you can still, even though you don’t have the web browsing, you still connect with your phone, right? So there is this idea of potential, there’s potential of this data to be connected. So if you know, if you’re using Google Maps, let’s say, and you’re on the LinkNYC network, then potentially you could have targeted advertising, sort of the image appears in front of you as you’re walking, right? So that could be—
James Parker - So people who watch Minority Report and didn’t realize that it was a dystopia. basically.
Audrey Amsellem - Yeah, and so this as far as I know is not happening now, but it was, so there was, we know that they developed that technology. So this was actually this undergrad student who discovered this by going through GitHub and found some of their code. As far as we know, this is not technology that’s implemented on the chaos right now, but it very easily could be. So yeah, this idea of, when we talked about Alexa earlier and said you’re gathering information on domestic space, Linkin.Wisee is just gathering information in a different space, right? It’s a Google project. Google has been gathering information on our online behavior. You also have Google Home that gathers information on your domestic space and they can now also gather information that’s in the public space. So I think it’s part, it’s tied in that it’s just part of this larger thirst and starvation for having data in pretty much every facet of our lives.
James Parker - Would I be right in saying that when you turn to LinkNYC, you’re beginning to move towards thinking about listening as data collection. So a kind of slightly more metaphorical, I mean, I think Robin James also does this in her work on acousmatic listening. So is that what’s going on? We’re no longer so much concerned with microphones but of the kind of signal detection through a field of noise as sort of the dominant way of thinking about data.
Joel Stern - She calls it acousmatic dataveillance. I remember being really struck by that essay she wrote, I think it was on the Sounding Out. Yeah, it must have been 2014 or 2015 and sort of trying to connect it with Pierre Schaeffer’s idea of acousmatic sound. It was always a question of what kind of listening are we describing here? Is the metaphor of an ear still relevant when we are thinking mostly in terms of data rather than audio or sound itself?
Audrey Amsellem - It is not a physical ear. It is not an ear that is actually able to process sound the way that a human can process In that sense, it’s listening in the broad sense of the word. And so on the one hand, in some of my work, listening is very much literal and very much sort of an action. And in other ways, it’s very much metaphorical. But it’s just trying to highlight a general system in which sound can function as a modality of power, essentially. But yeah, I mean, I remember that blog actually, I think I cite that a couple of times. Acoustmatic dataveillance. And it’s yeah, it’s super interesting. And I I’m very curious to see what more comes out because well, she’s she’s the I don’t know if she would say she’s a sound studies person, I guess probably. I mean, she’s a philosopher who writes on popular music mostly. But I’m very curious to have more people that have that interest in that background also. So think about the various ways that listening can be applied to civil rights capitalism. In her case, I think she was working…talking about the NSA so it’s slightly different but just think about surveillance and listening as a framework and I’m sort of looking forward to having more of this work out there to enrich our understanding.
James Parker - It’s funny that you mentioned the NSA because I was wondering like, this is not in any way a criticism because it’s an enormous project already, but where, yeah, where this is not a story for you about government surveillance. And that’s partly because, you know, you’re trying to tell a story about neoliberalism, although, of course, you could probably make the argument that neoliberalism and the NSA go together pretty well and so on. But yeah, just as a, I mean, I want to talk a little bit about where you end in the thesis and this idea of trustworthy listening. But you must have you must have confronted the question, how am I going to deal with the state surveillance? And I obviously, it’s not a centerpiece of this work. And I’m also confronting that myself. So I’m basically asking you to help with my own work! But like, I mean, yeah, how did you think that through?
Audrey Amsellem - So yeah, I actually started this project with LinkNYC first and that initial my initial master’s thesis actually had a lot more about um Patriot Act and about movement and particularly about the history of New York City and movement and control of movement within the city and control of behavior and it had actually a bit more of that angle. When the dissertation had to be formed I had to you know I had to impose myself limits which I didn’t impose too many but that was the one I had to impose just because that literature is so vast and I was not going to be able to really tap into that and I’m already very interdisciplinary. But yes, I don’t have a solution at all for you, but I absolutely agree with what you said that it is very much could very well fit into this because it very much ties with neoliberalism, especially when we consider that a lot of the PRISM program relied on data from private corporations. It relied on data from Google, on data from Facebook. So that there is no sort of intervention to be made. This is a fact, you know, it’s a fact that there is an interplay between the two.
James Parker - Well, you could go back further. I mean, you know, all of the speech recognition stuff and a lot of the music, early computer music stuff was directly funded by DARPA or it’s kind of, I mean, a lot was happening in big, you know, very government backed kind companies like Bell and IBM, but a lot of it’s happening straightforwardly through military funding, including the work on music, like the first work on music transcription, I think is, you know, DARPA funded. And so, you know, there is no feel, I mean, it’s just such an obvious and bordering on the point of dumb point to say, you know, AI is like funded by the military, like, but it’s sort of it’s just so clearly true. I mean, just look at everybody’s PhD thesis when you read like, it’s like, DARPA funded, DARPA funded, DARPA funded. And so, you know, there’s a sense in which the commercial applications all come out of government/military funding. The whole paradigm. So, I don’t know where that gets you really. Maybe this is where we can turn to the possibility of hope or reform… You do that thing in the thesis where you say Oh god, that’s a bit grim. What now? And then you suggest this term trustworthy listening as a kind of antidote…
Joel Stern - But even before you get there, there is a sort of section on the detournement of Link NYC by hackers and artists who kind of find a way to creatively sort of misuse technology and then, you know, in doing so open up some kind of other possible ways of thinking about these infrastructures in more emancipatory ways. But could you just, you know, say a little bit about that work? I think it was Mark Thomas, maybe, who, you know, made these kiosks play, you know, Mr. Softy ice cream music.
Audrey Amsellem - Yeah. Yeah, yeah. Yeah. So I wish it was similar from the NSA question is I wish I had a bit more space to talk more about this. So with LinkNYC, he just seemed very important that I do. But there’s obviously other examples with Alexa, and I’m sure you’re familiar with with many of them. But yes. So there’s been all kinds of like artistic vandalism and various forms of hacking and hacktivism in a sense. Although this particular person you’re talking about, he didn’t think of himself as a hacktivist. So what he basically did is that he recorded a distorted version of the Mr. Softy, which is this ice cream truck song, and a sort of distorted, uncanny version of it that was particularly creepy. He created a delay and when he would launch it on the Kiosk. And so, and he called this minute between the moment he programs it and the moment and simply the magic minute, right? Because if it’s playing as he’s there, it’s not as uncanny, right? Because it’s not the device playing on its own. And so it starts, you know, having this, so obviously people were starts resonating the sound. So obviously people were around, you know, on their smartphones are like taking videos, posting it on Twitter. And so it becomes a viral in that way. And I think it was perfect.
It was really perfect conclusion for the chapter for me, because it was this moment in which actually it becomes a sonic object again, because when it was used by people, including homeless people, but not just, it was this like sonic musical moment. This, this captain was this moment of maybe a different identity for the city. You had all these sort of questions that would pop up. And then, you know, that was completely silenced. And then he’s bringing that music back again. But the music he’s actually bringing back at this very eerie, very weird, kind of sound and throughout my work actually with the Alexa I talk about the uncanny as a potential actually as an actually potential to reveal things that are don’t quite work right that are not quite right and and this is what we’re talking about when we talk about surveillance capitalism right this there is a moment in which we’re confronted with these devices and there is something eerie about them and there’s something that’s not quite right and I think we should trusting that instinct in a way. And so, can the neoliberal ear be detourned?
Audrey Amsellem - You know, it’s an open question, definitely. But I think where my work has taken me is more that I think we can regulate more. Really, it’s where and maybe, you know, in a few years, I’ll have a different answer. But I think at this stage, in my research, this is where I see the potential sort of of answers. There’s a lot of stuff being done, you know, just yesterday the White House, it’s a bit tangential, but the White House released a blueprint for an AI Bill of Rights. And so you have a bunch of organizations, lawyers, activists, artists, a lot of them you work with, right, who have hope and who build alternatives. Yeah, but for me, I think regulation is probably, you know, art is always going to be the most important for us to be able to think through and live through the situation. But as far as, you know, really kind of deturning in a practical way, I think regulation is really the main way to do it.
James Parker - Yeah. I mean, it’s tempting to end there. I as a legal academic, I’m lacking in faith in regulation. Maybe I’m meant to say the opposite. Maybe I meant to say that I have a lot of faith in regulation. But I mean, yeah, we finished quite a few conversations on this question, not the what is to be done question as in like a program, but like how to even situate or handle the sort of the hope story.
Joel Stern - Like where you don’t really want to see sort of the legal sort of infrastructure as an answer.
James Parker - You know, because you know, one obviously one thing to say would be well, you begin with copyright and property as the way that we arrive at Spotify and Alexa and whatever and you know, the establishment of private property and its sort of defense, or infrastructuralization by the legal system in concert with capital. That’s the originary moment of global capitalism in a certain kind of crude way of thinking. And so it’s like law isn’t the answer, right? I’m not saying that you think that law is the answer, but it’s just like it’s hard to… you know, it’s interesting because Shoshana Zuboff absolutely makes the argument that regulation is the answer… and a lot of people have been critical of her on those grounds, but that’s partly because she’s defending capitalism. She just wants a less surveillant capitalism, whereas I don’t, I definitely don’t read you as defending capitalism or neoliberalism. So, so yeah, but it’s just, yeah, I don’t know, I haven’t really got anything to say. It’s completely unresolved for me. We certainly need more regulation, but I don’t know, I just don’t know where else to point my energies.
Audrey Amsellem - No, yeah, I think, you know, I mean, I know what you mean, you know, copyright law was also, you know, invented in a very specific kind of government, right? I mean, this was, it was invented for censorship, right? We’re not quite like that anymore, thankfully, but it’s true. It’s true that there are limits to what the law can do. There are limits to what regulation can do. And I think at the end, it’s individuals who are doing the best they can to think through these questions and trying to promote at the individual level what they think our lives should be like. And we do that through community. So I actually think initiatives like machine listening and various kinds of initiatives that I cite also in my work, this is the way to do it.
James Parker - Do you want to name check some of them?
Audrey Amsellem - So for example, Rethink Link is an organization that I am citing, specifically talking about LinkNYC. So any kind of community building initiatives that is going to inform people and is also going to empower people. Because as we talked throughout, it’s just various forms of disempowerment. And I think any moment in which you can empower people with knowledge or with community or with art or whatever is, you know, little by little, I think that does produce significant change. And so that’s, regulation is the more sort of practical way to do it. And I, but I, you know, I agree with you that there are other things that we have to look at. And there are other avenues that may not be sort of a direct, abrupt solution, but it’s actually at least a way that we can cope with this, right?
Joel Stern - Yeah, consciousness raising, community organizing, cultural production.
James Parker - We also need to seize music as a Trojan horse for the revolution.
Audrey Amsellem – Absolutely.
James Parker - On that note, it’s been fantastic to have to speak with you Audrey. I mean, yeah, you too. People can find your dissertation online. Is that right? Yeah, on ProQuest. Or show me at a DDU or something like that. We’ll put it in the machine listening library.
Audrey Amsellem - Amazing. Thank you.
Bernard Mont Reynaud transcript
NB time codes are a rough guide only. they are retained from the original audio, which has since been edited for clarity and continuity.
James Parker: OK. Fantastic. Well, if you’re OK to begin, shall we just begin? Maybe you could start off by introducing yourself, however feels right to begin with, and we can go on from there. OK. Go for it.
[00:01:19] Bernard Mont-Reynaud: Let’s see. Well, I was born in France, in Marseilles, you know, up in the South. I studied there until I went to Paris at the Ecole Polytechnique, and I hesitated between multiple professions. Was I going to be an architect or was I going into electronic music or computers? And computers won, I would say, very much so, based on this idea that you could capture elements of thought, then encapsulate them and pull them together.
[00:02:00] Bernard Mont-Reynaud: It was sort of an early notion of AI, if you want, the pull of AI, but the pull of just modularism thought and assembling thoughts together. And that was the winning thing. And I went to work on, actually, computer-assisted instruction for a number of years. You may not believe, but as early as 1968, which was the time I’m talking about, there were a large system doing CAI. Professor Soupez at Stanford had some large computer-assisted systems at the time.
[00:02:41] Bernard Mont-Reynaud: What we called CAI in those days was computer-assisted instruction, whereas now in my work at SoundHound, this is conversational AI for the same initials. In any case, so I went to do this at a place called IRIA, which became INRIA, the National Institute for Information and Automation in France. I worked there for four years doing AI, CAI. And then, but the appeal of Stanford was enormous for me. And so I went to Stanford to do a PhD.
[00:03:26] James Parker: Was the appeal because, you know, Stanford is a famous university in America, or is it specifically because that was the place to do that specific work in computing or a bit of everything?
[00:03:44] Bernard Mont-Reynaud: It’s sort of in the middle. Stanford was, remember back in those days, computer science wasn’t what it is today. It was pretty much a pioneering field. And Stanford was one of the pioneering institutions. I also applied to MIT and UCLA and what was the fourth one and was admitted everywhere. But Stanford was where my heart went. And I had been on the West Coast a few years earlier, visited, you know, San Francisco and the area. I was just fascinated. So I just, yeah, Stanford was my choice.
[00:04:24] Bernard Mont-Reynaud: It was because they had the best computer science programs along with MIT. And you quickly or maybe not so quickly got into working on computer music at Stanford, or was that the, because you said before briefly that that was your sort of alternate love music.
[00:04:48] James Parker: And so how did those converge or what was the relationship between those two in those early days at Stanford?
[00:04:58] Bernard Mont-Reynaud: Yeah, good question. Back in the days of thinking about, it was about electronic music and old computer music, which I did not know about at the time, and I’m speaking 1967, 68. It was just out of an interest in music. But, you know, one thing I always regretted in my life is to not have a formal education in music. And this was something I wanted to do for my own enjoyment, for my own passion even. But I was a computer scientist. I had a very rigorous training in math and science, and was becoming a computer scientist.
[00:05:40] Bernard Mont-Reynaud: The idea of music was not there yet, and it wasn’t for a while. What happened was, while I was doing my PhD at Stanford, I took my time and I sampled every topic under the sun in a computer science curriculum. I was involved also in natural language understanding and other things, and very solid training in optimization and so on, like hardcore topics of computer science as opposed to applications. And let’s see, I’m losing the thread here. Yes, I wasn’t… I wasn’t doing anything without music.
[00:06:28] Bernard Mont-Reynaud: I went to be an assistant professor in Berkeley. And I wasn’t, how can I say, I had lost my passion for the core CS, you know, ultimate optimization kind of quality, and I started to be attracted to more perceptual fields like visual and auditory, and that was my interest. And this is when, after four years in Berkeley, I left behind a tenure track position to go do research on computer and music. Not in producing computer music, but there’s a story on this of how come I went to CCRMA (‘Carma’), and I could get into that story separate stage.
[00:07:21] Bernard Mont-Reynaud: But basically, I went PhD plus four, or almost four, when I left the sort of core computer science track to go into the audio and music area.
[00:07:39] James Parker: I’d love to hear more about how that played out. I mean, yeah, what work were you even doing on audio and music in that period? What work was being done in the field? And then how did that take you to, well, I’ve been calling it the CCRMA at Stanford, but Carma, as you put it?
[00:08:03] Bernard Mont-Reynaud: Yes. So, yeah, I want to take a deep breath on that one, because now I’m restarting about CCRMA and all the people and then how come I joined in. But basically, if you go back to the founding of CCRMA, that was around 1977. And there were four founders, John Chowning, who had become the director, and two other people, and the fourth was Andy Moore, a student who at the time was completing his PhD in 1977, at the same time I was, and we had known each other at the Stanford AI Lab.
[00:08:53] Bernard Mont-Reynaud: But he wasn’t involved with music at the time, at least that I know, and our relationship was based on music, it was based on AI, but we happened to have esteem for each other. And he independently had applied for a grant to the NSF, if you know about the National Science Foundation in the US, a grant to do musical intelligence, which at that point, you know, and when you’re a pioneer, you have no limits, you have no boundaries, everything is open field.
[00:09:37] Bernard Mont-Reynaud: And so in that grant, he was going to address everything that could possibly have to do with computers and music or computer analysis of music, and that involved the signal processing, the polyphony, the, you know, musical intelligence in its wider form, and so on and so forth.
[00:09:56] James Parker: So, I know, I know, is that James Moore, is that the same person that I know as James Moore?
[00:10:06] Bernard Mont-Reynaud: Yeah, James Andy Moore. Did you speak with him?
[00:10:09] James Parker: No, I haven’t spoken with him, but I read some of his work on music transcription. It seems like in the mid 1970s, that was, he was really the first person publishing on music, automatic music transcription.
[00:10:26] Bernard Mont-Reynaud: Exactly, exactly. And this is where the connection happens, because he completed his PhD in 1977, same year I did mine, I completed mine. And his work was on, involved the separation of two notes on the guitar, guitar player, and doing the signal processing just to separate two notes, you know, in a succession of notes. That was state of the art at that time. And yes, so, so indeed, that’s the same person. That’s the research he was doing for his PhD. And he became a co-founder of CCRMA.
[00:11:15] Bernard Mont-Reynaud: And he had applied for a grant at, with NSF, where, based on extrapolating his research on separating two notes on the guitar, if you want to put it that way in a humorous way, he was then going to solve just about, I mean, the grant, the project was going to solve every problem under the sun for, for, you know, music analysis by computer. And this is, I guess what you, what you do in grants, you oversell a bit and, but especially when this is a completely pioneering field and there’s nothing to limit your ambitions.
[00:12:01] Bernard Mont-Reynaud: So then, okay, so this is what happened. And, but then he went to Paris to a place called IRCAM, which you may have heard of or not. And, and he was, he’s also, I mean, Andy is a, I don’t know whether to use the word genius or, or to say he’s a, he’s an incredible engine that could, but he’s a, you know, he could write not only incredible software, but he could design incredible hardware and so on and so forth.
[00:12:40] Bernard Mont-Reynaud: So he was at, at CCRMA, he was involved in all sorts of software, but later on he was hired by Lucasfilm and he built hardware for George Lucas that would help with the, the Star Wars series, the Star Wars sequence. Yes.
[00:13:06] James Parker: Trilogy.
[00:13:09] Bernard Mont-Reynaud: Trilogy. I was going to say trilogy and say, but wait, there were more than three. This is my hesitation. It wasn’t a trilogy, was it in the end or, you know, but I didn’t want to misrepresent the number, but at the time it was probably a trilogy. And, but I, I get slightly ahead of myself. So he applied for this grant and then he went to IRCAM in Paris to install, you know, a certain software situation similar to CCRMA. They were duplicating the CCRMA environment in Paris for Pierre Boulez.
[00:13:46] Bernard Mont-Reynaud: And then he got hired by Lucasfilm to go build the software for Star Wars. And by the time this all happened, the grants got granted and he wasn’t around to do this. So he was not going to deliver on any of these promises.
[00:14:09] Bernard Mont-Reynaud: And somehow I heard about this and I kind of managed my way into leaving my faculty, my position at Berkeley, my assistant professor position at Berkeley, where I was getting a bit frustrated and kind of a little tight at the elbows because I wanted to get in something more perceptual, you know, possibly sound or music or visual stuff, as opposed to, you know, core technique, core map, you know. And so the timing was good. I don’t know. I don’t know exactly on what basis I tried to convince them that I was the person for the job, but it worked.
[00:14:58] Bernard Mont-Reynaud: They hired me to do this to actually head that project. And it was all new to me. And because it was all pioneering work, you know, I didn’t feel I had enormous constraints about that.
[00:20:31] Bernard Mont-Reynaud: I forgot that I was supposed to give a biography and I just started unfolding the whole story as opposed to like a quick summary. So I could give you, try to give you a quick summary, but chances are I get lost along the way. I mean, quick summary, what would be, I spent 10, 11 years at CCRMA. During that time, I did various consulting. I could not, let’s see, I went on to industry for the rest of life, been in a lot of places like Xerox PARC and Sony and others.
[00:21:14] Bernard Mont-Reynaud: I went into SoundHound over 12 years ago, I’ve been 13 years at SoundHound, it’ll be 13 years. And, but basically, it’s all around the valley, but some of my best times have been at CCRMA and SoundHound and Sony.
[00:21:36] Speaker 3: So, yeah, so this is a quick collapse.
[00:21:40] James Parker: That’s very helpful.
[00:21:41] Bernard Mont-Reynaud: I’ve spent 50 years, 54 years.
[00:21:45] James Parker: Oh, no, that’s incredibly helpful. I was just wondering if it would be possible to do a slightly, a similar overview of your time at CCRMA, like, you know, in those 10 years, you know, where did you begin? Because you said that it was kind of quite blue sky thinking, you know, the industry was involved, but not sort of driving any specific commercial outcomes. So do you sort of remember, like the broad arc of your time at CCRMA, like what sort of things you were working on the beginning and where you ended up?
[00:22:21] Bernard Mont-Reynaud: Yeah, sure. Of course, I remember all that stuff, you know, near to my heart, I remember way more than we have time to go over. Initially, you know, it was all pioneering and I was discovering that too. It was like, okay, where do I begin? What did I do? I found at CCRMA, one of the other co-founders, Lauren Rush, who helped me a little get started and collect some examples that seem of manageable complexity. And the focus at that time was music transcription.
[00:23:02] Bernard Mont-Reynaud: How do we, and I discovered among other things that finding the pitches, hearing the pitches and the sound was not the hardest thing, or at least I had somebody else as a consultant doing it. Basically, you say what frequency, what time, what amplitude, and you’ve got those few numbers coming out. But I discovered that transcribing is a problem of itself. Once you have the events, you have to figure if there’s a tempo, if this isn’t, you know, you have before, is there a quarter note somewhere or are there triplets and so on and so forth?
[00:23:52] Bernard Mont-Reynaud: That’s the key. And there are a number of issues that have to do with musical intelligence, as opposed to just extracting events. So I discovered for one thing, the structure of this, started developing some representations for notes, little representations starting with events, and then gradually getting enriched with properties that were going towards transcription until eventually we’re able to put note names, note headers, note durations.
[00:24:30] Bernard Mont-Reynaud: I found out about, there was at CCRMA, early pioneering music printing, so I could create a notation from Professor Leland Smith had a system for musical typography, and you could feed this data to his program to get a terrific looking musical score. So anyway, at the beginning as putting those pieces together, figuring out what the components of problems were, and that’s pretty much, that’s what I achieved within that first time, as well as applying for a follow up grant.
[00:25:13] Bernard Mont-Reynaud: So, there was round one, on round two, and after I returned from Paris, I got deeper into the capabilities of the system, in particular, the musical intelligence, but I realized that I had to think about the following grant, you know, that I had to think early on about planning the sequel, and I realized I had to choose one of two paths from that point. At that point, the project had been about everything, from detecting events, to separating the notes, to creating musical score, and so on and so forth.
[00:26:00] Bernard Mont-Reynaud: I thought I had to either go deeper into the music intelligence, or deeper into the event separation intelligence. Call it hearing, or call it musical listening, or musical understanding, you know, as you go less of a pioneering effort, and the effort, you know, deepens, it diversifies also. And so, I thought about this for a while, and I decided to go into hearing.
[00:26:37] Bernard Mont-Reynaud: Part of my reason was, it goes back to the fact that I did not have a, I had some musical training, but I did not have the entire depth of kind of having been to music school and having all of that, to totally do musical intelligence justice, and so I felt that however interested I might be, and also the field of sound separation, source separation, also looked very much a pioneering area, like wide open, as opposed to, musical intelligence was quite open, but I wasn’t as prepared for it, and I felt my limitations.
[00:27:30] Bernard Mont-Reynaud: I didn’t know what my limitations would be in source separation. Anyway, that’s where I aimed my next round, for the third round of funding.
[00:27:47] Bernard Mont-Reynaud: In the meantime, I should probably mention, there’s a couple of students, you know, when you have a research grant, you hire some students, and I had a couple of PhDs happening, on the project, one that involved the transcription of Afro-Cuban drumming, which I used some of my software that did event detection and quantization of durations and so on, but he supplied his own event detection, and then we put it all together, and he’s been a music professor now for quite a while.
[00:28:36] James Parker: Who is that?
[00:28:36] Bernard Mont-Reynaud: That’s Andy Schloss, his name is Andy Schloss, S-C-H-L-O-S-S, and I can give you a contact. Victoria in Canada. Yeah, I will provide follow up on that. Yeah, the second PhD was later on, that’s David Menninger, and that happened, you know, on the next round. So, I’m getting ahead of myself again, but basically, I said there was the second round, where I was still pursuing this double goal, you know, but I felt I had to move towards either the music intelligence or the auditory intelligence, and I went into the second, for the third round.
[00:29:41] Bernard Mont-Reynaud: Do you remember what year that was? So, this is where the story, I would say, 86. Okay. 86. Yep.
[00:30:02] Bernard Mont-Reynaud: Right. So, and in 87, I also had all the interest, I had developed a, an approach to event detection, my own approach, and it was very visually oriented. And I would create spectrograms. I know if it can get a little technical, but you know, there’s such a thing as a spectrum and a spectrogram, and it’s normally put in the frequency time representation. Now, imagine that the frequency dimension, instead of being the linear frequency in hertz, is going to become the logarithm of frequency.
[00:30:50] Bernard Mont-Reynaud: So you, that means if you’re in linear frequency, you have, say, zero, 400 hertz, 800 hertz. By the time you compressed it, the factor of two from 400 to 800 occupies the same amount of space as 200 to 400. It’s like an octave. And by the time these are like octaves, this is like semitones. This is like a keyboard. Okay, so this was more appropriate for, for music, but not only that, but you can actually, I discovered you could do pitch detection on that representation, because the harmonic series is a fixed pattern, is a fixed vertical pattern.
[00:31:39] Bernard Mont-Reynaud: And I could do, I could render audio into this representation, which involved a spectrogram and the log f, a semitone spectrogram, I’m going to call it. And I could do pattern recognition as in image convolution on this image representation and have a, and obtain event detection that way. And so I wrote some of this work for a while, and then this was, this started at, you know, 85, 86, 87, where those things were happening, which was sort of the core, which was where I had taken control of the event detection with my own approach.
[00:32:33] Bernard Mont-Reynaud: And I decided to go into source separation, like if there were multiple pitches at the same time or other, okay. So now we had mentioned Al Bregman, and the fact that he came, what was wonderful is that Professor Al Bregman came, spent one year at CCRMA, while he was writing a book on auditory scene analysis, a 900 and some page compendium of his research in psychoacoustics, the psychoacoustics of how do we detect streaming?
[00:33:16] Bernard Mont-Reynaud: How do we know there’s a source here and a source there versus any psychologically that we have separate sources versus, you know, other events that don’t group as sources. And this was fascinating. And it was great for me to hear what he had to say, it was great for him to have somebody really interested to hear about his thoughts. So we had, we had this little club, he and I, before the book came out, where I got to receive all his ideas on this, and this fit exactly in the right place.
[00:33:58] Bernard Mont-Reynaud: Because I wanted to know his principles for bringing sources together. I don’t know if you’re familiar with the book, but they basically are grouping principles that allow sources to combine as separate auditory streams.
[00:34:20] James Parker: Yeah, the coining of that term, the sort of the auditory stream seems to be one of the sort of key things like, from the book, to start thinking in terms of auditory streams. He has a whole section. And it’s really interesting to, I’ve read the book, or I’ve read, tried to read the book, I’m not a scientist, but everything you say really correlates with my understanding.
[00:34:49] James Parker: It does seem that he was doing this work in psychoacoustics sort of before he came to CCRMA, but that, you know, he’s very explicit in his acknowledgements about the influence of his time at CCRMA, you know, your influence, John Chowning’s and others.
[00:35:14] James Parker: And it’s really interesting to me, it’s always been interesting to me to think about the relationship between this work that’s on hearing, more generally, and music specifically, because it seems like around the time that you’re describing, there’s basically two major streams in, to use that word stream again, in, you know, auditory oriented AI, there’s basically speech recognition. And then there’s, or speech understanding or whatever. And then there’s work on music. And then around the time that you’re describing, suddenly, or not so suddenly.
[00:35:56] James Parker: it starts to broaden out, it starts to become much more about hearing, as you’re saying, or much more about, you know, the separation of different auditory events, you get to things like auditory scene analysis, auditory event detection start to come out. And then it seems from the outside that what were relatively distinct fields, I mean, I should ask you what the if you had any relationship with people doing work on speech recognition, and so on at the time, but but it seems like they were relatively separate.
[00:36:33] James Parker: And then they start to come together under a larger umbrella of people working in, you know, because suddenly all the speech people start to need source separation. And they, they need to do scene analysis in order to, you know, separate out speech from office sounds, and, you know, environmental noise, and so on and so on.
[00:36:59] James Parker: So, so that’s the story that I seem to be sort of, I seem to be finding digging around in all of these reports and PhD theses and stuff, but it sounds like what you’re describing, but I, I don’t know if I’m misrepresenting it at all.
[00:37:16] Bernard Mont-Reynaud: Well, well, yes, I know. First, first to clarify one thing, I, I was not involved in speech at the time. My interest was in source separation, which is a sort of, which is a broad interest, but the examples were musical examples, right?
[00:37:38] Bernard Mont-Reynaud: They there’s no doubt in the speech community, people have been very interested in separating the voice itself from the background noise, which is also a kind of separation, but it has a different focus in the sense that it’s not on per se streaming, for example, noise is not considered a stream, you basically try to find a dominant stream, the dominant source, and everything else is kind of, is what you want to remove.
[00:38:14] Bernard Mont-Reynaud: So, so I think it’s taken a long time for, and in many ways hasn’t completely happened, speech to be treating the voice as just one of multiple sources in the environment, and also recognizing other sources in the environment. They, as far as I know, that particular phenomenon hasn’t happened to treat the voice as just one equal, you know, one of many, of many events. If you use a visual analogy, you know, you, you can recognize chairs and tables and, you know, and people in the, in the image.
[00:38:59] Bernard Mont-Reynaud: And although maybe you’re particularly interested in people or people’s faces, you might still recognize a chair or a table in, in, in the vision field. This object formation is the same thing, you know, at one level, whether you’re in vision or in audition, you’re doing object formation, right, you’re doing object separation, object formation, and you recognize sources, give them properties, okay, but, but in speech, the equivalent, you’re still only interested in maybe the face, you know, it’s not in all these other things.
[00:39:39] Bernard Mont-Reynaud: So, so I, I, I don’t fully see the same type of parallelism that, that you see about speech in music, also that there has been a tendency, I would say for people to, for a field of source separation, auditory source separation to, to exist on its own, and it would pick examples in sound or, or image as the case might be. But not be completely attached to either speech or music. If you want, there are the people who look at this psychologically and the people look at this from the point of view of applications. They’re not entirely the same. I mean, in a research lab, they can be, but the source separation has been driven by research interests and only to a small degree by applications in the real world. It’s still expensive.
[00:41:00] Bernard Mont-Reynaud: Now, this has started to change, but and also we’re beginning to see maybe neural networks that have some capability in that area, but it’s still not very widely spread out.
[00:41:22] James Parker: Okay.
[00:41:23] Bernard Mont-Reynaud: Maybe I could return because we’re on Bregman here and on auditory scene analysis. And that was a very important book. And yes, he had been interested in this topic for a while. Then he decided to write a book on it. And that was a very important book. He put it all together. And I was familiar with a lot of this work. Before the book was written, he was just dumping this stuff on us at CCRMA. And I was one of the most attentive listeners of his ideas, because that’s the field I was really getting into. And I must say, he did a lot of psychoacoustics, which that wouldn’t be within my capability to do all of this.
[00:42:18] Bernard Mont-Reynaud: He’s that kind of experimental psychologist. I am not. And not only that, but it’s not truly my interest to carry out. I don’t have the patience to do this, but to hear him talk about it, that was wonderful because he had done all these experiments. And I came to summarize a 900 page book out of a couple of sentences in my head, which is to say, if you take any criteria, such as suppose you have two sources, they’re based on, you have a source that goes beep, beep, beep, beep. And the other goes, bop, bop, bop, bop. So there’s one high, one low, and you do beep, beep, bop, bop, bop, bop.
[00:43:06] Bernard Mont-Reynaud: Yeah, I cannot do the productive process at the same time. I mean, I can do beep, boop, beep, boop, beep, boop, beep. And if I do it too slow, they start to, but the kind of experiment he would do is to vary the distance between those frequencies, the high and the low, and to vary the timings. And he would show that if the timing makes them very close, and one goes like, beep, beep, beep, beep, the other, boop, boop, boop, boop, they separate as two streams, a high and a low stream. Versus if you have one that goes, beep, boop, boop, beep, boop, then when there’s a long time, they become a single stream.
[00:43:57] Bernard Mont-Reynaud: They become a single note going up and down. And his experiments, he would do this and show them, and show when does it stream together, when does it not. And he basically showed in every case, any number of dimensions of these kinds of things, that there are trade-offs, that the same mechanism that makes it stream can make it not stream as a not great.
[00:44:30] Bernard Mont-Reynaud: So, okay. There’s not one kind of example, say you’re going to stream on frequency, and then it will be on time difference. No, they always trade-offs, trade-off between the two, if you will. There’s multiple factors, each can, by varying them against one another, can cause streaming or not. And so that’s first sentence, if you want, of the summary. And the second is, this implies, seems to me, that there is a central mechanism for separation.
[00:45:10] Bernard Mont-Reynaud: Because everything that goes from the signal processing or from the raw data, goes and can have the outcome of separation or not. So, to me, that speaks for a uniform mechanism that decides whether sources belong together or not. And that’s my summary of a 900-page book, and is essential, because it had implications for the architecture of building that kind of system.
[00:45:39] James Parker: Right. So, do you begin to then operationalize that in the systems that you’re building?
[00:45:50] Bernard Mont-Reynaud: Absolutely. Absolutely, yes.
[00:45:52] James Parker: And does it suddenly lead to significant improvements in…
[00:45:55] Bernard Mont-Reynaud: Well, okay, so you’re assuming that the system was already at full capability, but it wasn’t. It guided how we think about it, how we start building the pieces, but not all the pieces were there. But yes, indeed, it did lead us in that. And this is where that second PhD thesis I was talking about came along. His name is David Menninger, and he did his thesis where sort of showing trade-offs. And there’s another parenthesis on this. There’s a phenomenon that you may have heard or not about.
[00:46:50] Bernard Mont-Reynaud: It goes back to some work of John Chowning from way back, and you know about the frequency modulation. And you may have heard about this phenomenon called frequency co-modulation, where you have a collection of partials. I’m showing things in the spectrogram, and each of my fingers is a partial, and it goes like… And it has this sort of artificial sound, and it’s not well separated from other things if you have multiple. But suddenly, you’re going to put frequency modulation on it. It goes… And the moment you do this, the sounds belong together at one.
[00:47:42] Bernard Mont-Reynaud: It sounds natural. It sounds like a voice, and it separates from anything else. So one of the many dimensions that Bregman talks about is the co-modulation, the fact that those various partials go up and down at the same time. They modulate at the same time. It’s called co-modulation. But it goes way back to the synthesis technique in FM that by putting this frequency co-modulation, it created this voice effect that was overwhelming to the auditory system. This effect is overwhelming because it immediately causes source formation.
[00:48:32] Bernard Mont-Reynaud: Okay, so this was a parenthesis from David Menninger’s thesis into, you know, and Bregman and all this with a parenthesis back to John Chowning. And now we’re back to frequency co-modulation was one of the features that David Menninger has focused on, as well as some others. And yes, we use these ideas in his thesis, which were based on Bregman and on the architecture, which is a central mechanism for grouping features. But the whole system was not built. It’s like it was, again, pieces of there’s so much to be done.
[00:49:25] Bernard Mont-Reynaud: The auditory system is extremely complex and capable. And we only built some pieces of it. It’s the same as a vision system. Vision has so many things in it. And you build pieces of it. You might build a piece that had to do with occlusion. Another piece has to do with texture and color. And there’s perspective. And I could go on and on. Well, the same is true in the auditory domain. In Bregman, you’ll see that he uses visual metaphors all the time for what’s happening in the auditory.
[00:50:02] Bernard Mont-Reynaud: And the reason is like we see the stuff and we understand occlusion. But auditory masking is much harder to represent. But it’s essentially the same thing. It’s essentially the same as occlusion. But we hear it, but if we can’t see it, we can’t easily talk about it, point to it. Because sound only exists in motion. Sound does not exist at a frozen moment. An image can exist at a frozen moment. You can point to this piece and that piece and talk about it. You cannot do that in sound very easily. It’s only moving. So I go parenthesis within parenthesis.
[00:50:49] Bernard Mont-Reynaud: But to answer your question, yes, we put what we could into the architecture of the system. And that plans to continue. But let’s see.
[00:51:04] James Parker: So, this is a fair way into your time at CCRMA. And soon in the story, you must leave this work, sort of, you know, unfinished in a certain way and then move on. Is that right? Like, are we getting to that point in the tale? So you sort of leave, you left this sort of, I don’t know how you would describe it, but sort of basic research, maybe, into machine hearing and moved more into industry? Is that the, is that how you think of it?
[00:51:42] Bernard Mont-Reynaud: Yeah. Let me talk some more about the transition and how it happened. Because by then, by then I was ready, between all the ideas of Bregman and all what I had built over a succession of systems, I was ready for a major onslaught onto auditory machine analysis, right? And I envisioned a larger grant than I had had before. And I went to DARPA, I wrote a grant proposal, I went to DARPA, and there was some interest, you know, there’s a question of how does this connect to industry or applications?
[00:52:35] Bernard Mont-Reynaud: And you know, at DARPA, when you go defend, I went to Washington to defend my proposal, and there are people representing NSF and ONR and the NSA and the different, you know, different agencies that might be interested in supporting the grants. And I felt it was a strange feeling. I felt they were interested, and their mind wasn’t quite there. They both were present and absent. It’s kind of strange. And then I found out later on, it turns out a week later, was the war on Iraq.
[00:53:17] Bernard Mont-Reynaud: So, you know, at the Pentagon, they’re, this is what, you know, they’re also involved, they’re involved in research and they’re involved in the defense department very much. So that explains part of that. And yet I had had interest in my research. And there was, the question was, would you also be interested in applications of this research on degraded monophonic signals? Now, what’s a degraded monophonic signal? It, this is a telephone tapped line, right? You have, it’s mono and it can be arbitrary and anonymous.
[00:54:10] Bernard Mont-Reynaud: So I had specific interest for like secret work. And I knew once you put your foot into doing secret work, you’re kind of, you go underground and this is the end of the research. You’re now working for the spooks. So to, don’t quote me on that, but. Okay, so I found that this huge effort I had put into having this very wide open research and that I was asking for $5 million at the time, which was a fair amount, but I felt it was building stage upon stage where I said I wanted to build a large system to do this.
[00:55:00] Bernard Mont-Reynaud: I couldn’t get the funding and I tried to survive a bit. But this is because I made a mistake as a, as a professor, which I had become by then an associate professor, research. I should know better than to go for large grants. I should also have also, so a little bit grant, so to get, to continue funding while hunting for a big grant. I shouldn’t have small ones to survive. I didn’t do that. My mechanism for survival was sort of on a personal basis.
[00:55:42] Bernard Mont-Reynaud: I would do consulting outside the university, but I hadn’t, so anyway, I made the strategic mistake of, of not having small grants to stay in the game while waiting for a long grant. And this is where I had to leave, just to put it in perspective.
[00:56:04] Bernard Mont-Reynaud: And so it was a while until I was able to work on source separation again. It wasn’t until I was at this company called Audience, where the ambition was to put a chip into telephones to do foreground background separation, to separate the voice of interest from all of the noise around it. So Audience would be another story a number of years down the line. In the meantime, I’ve been at many different companies.
[00:56:40] Bernard Mont-Reynaud: And then again, quite a bit later, I went to SoundHound, where I wasn’t doing sound separation at all, but I’ve been involved with sound and music and then speech, sorry, and natural language. At Audience, I did work on source separation. I finally got to build a new system based on these principles that I got from Al Brickman. And I pulled that together.
[00:57:16] James Parker: What year are we talking about now? At Audience?
[00:57:22] Bernard Mont-Reynaud: Audience, that’s going to be maybe 2000.
[00:57:25] James Parker: Okay.
[00:57:27] Bernard Mont-Reynaud: Yeah, I could go, maybe I should send you a resume.
[00:57:36] James Parker: Well, I’ve read bits and bobs.
[00:57:38] Bernard Mont-Reynaud: I would say it’s about 2000. Yes, I would say 2000 if I picked it like that. So when you were… And again…
[00:57:51] James Parker: No, go ahead.
[00:57:56] Bernard Mont-Reynaud: Even at Audience, hold on. Even at Audience, we had the tension between the broad research angle on this, which is source separation. And the product focused just separate that voice right here from ambient stuff by something cheap, something that works most of the time, but it does not have to do source separation. And Audience initially was addressing broad goals. And there was a lot of interest in this source separation business. At some point, the investors came down and said, hey, what’s your product focus? You don’t have a product yet.
[00:58:43] Bernard Mont-Reynaud: And they just let go of 50% of the company and said, you now focus on something that goes out to market quickly. And at that point, they let go of me as well as several others. Which is the kind of… Even then, of course, this was 2000. And I think to a large degree, even now, we still have this distinction between a system that represents the psychology of understanding multiple sources versus one that achieves a specific engineering goal of just delivering on a very specific task.
[00:59:30] James Parker: What was the specific task that Audience was trying to do this, the voice separation for? Was it like what I didn’t… Did you say telephony or what? Maybe I didn’t catch that. What was the specific industrial context? What did you say?
[00:59:50] Bernard Mont-Reynaud: Smartphones. Well, phones.
[00:59:52] James Parker: Okay.
[00:59:53] Bernard Mont-Reynaud: Phone, phone, smartphone. The first… They eventually did a chip and the first chip went on to the iPhone.
[01:00:02] James Parker: Okay.
[01:00:03] Bernard Mont-Reynaud: And after that contract with Apple, and you know, Apple does like many companies that once they got the… Oh, okay. We’ll do our own, you know, and…
[01:00:12] James Parker: Right.
[01:00:13] Bernard Mont-Reynaud: And then they went on the Samsung with a chip. Yeah. So, but basically, the idea was to separate. So, you have a front microphone and a back microphone, or a primary microphone and secondary one. You can assume the primary microphone captures more of the source of interest than the other, than the secondary microphone. And what they did eventually, instead of a general source separation, was to sort of emphasize the primary with respect to the secondary by what’s called spectral subtraction, which is not an exact arithmetic subtraction.
[01:00:55] Bernard Mont-Reynaud: But basically, you see what stands up. So, yeah, the purpose was to improve the quality of voice in noisy environments, which is obviously an application of great economic interest.
[01:01:13] James Parker: And with quite a long history in.. you know, obviously Bell Telephone was quite invested in similar kinds of techniques for a long time. I mean, it’s a lot of, anyway, you know the history of audio technology and its relationship with telephones better than I do, but it’s an important one.
[01:01:38] Bernard Mont-Reynaud: Yeah, and denoising has been a big interest all along, and to do this once you have two microphones is much easier than with one microphone, especially if you know that one microphone captures more of the signal of interest than the other one, which is more kind of a bit of everything.
[01:02:07] James Parker: Yes. Look, I’m conscious that we’ve already been going for quite a long time, so I’m very grateful. So, I’m wondering, because obviously my interest personally is more on the sound end of things. So, part of me thinks, well, is this an appropriate moment to jump ahead to Soundhound? But then I’m wondering if I insist on that, am I missing a hugely important part of the story? You know, that’s a big jump in time. Is there something really crucial during your period at Lucasfilm or Xerox or whatever that is sort of crucial to understanding?
[01:02:54] James Parker: Obviously, it’s important to your biography, but understanding the evolution of machine hearing or machine listening more generally, or where you end up? So, I don’t want to foreclose those stories if you think that they’re important.
[01:03:16] Bernard Mont-Reynaud: Yes, I think by and large it’s time to jump over to Soundhound. I was taking a look at my notes to see if there’s something I still wanted to cover. I did want to mention, going quite way back, that I skipped over the whole story of Imperius. And yeah, let me skip it together, because there were some pretty interesting things, which were more in the political.
[01:03:47] Bernard Mont-Reynaud: Something I achieved, which was a technological success with speech recognition, but it was a human failure in the sense that we were acting as technologists and not understanding the whole application context. It involved a demo given to President Senghor of Senegal, who is also a poet, by the way. And I thought this might have appealed to your audience and from a political angle. So, maybe I should tell a bit about that.
[01:04:25] James Parker: Please do, please do.
[01:04:27] Bernard Mont-Reynaud: Quickly. Well, you remember, going back to, flashback to the first time I got funding from NSF, and I didn’t have the grant in time to continue the research. So, it turns out I went away for a year and I went to, I almost went to IRCAM, it didn’t happen. I’ll skip the reasons, but there was improper timing, put it that way. So, I ended up being at this place called Centre Mondial Informatique, which was headed by Nicolas Negroponte and Seymour Papert and this French politician, Jean-Jacques Servan-Schreiber. And, let’s see.
[01:05:16] Bernard Mont-Reynaud: Anyway, they were waiting to hire me. And then one day they hired me. Finally, I said, look, if you don’t hire me by Tuesday, I’m gone. Okay. And then Monday night, they said, okay, you’re hired. But Friday, we’re having a demo to President Senghor. He was no longer president by then. He was the first president of Senegal. But then he had retired from that. And you’re going to show him how you can use the voice to command things on the computer. And I came up with this idea of having shapes and colors and numbers. And you could say three red squares.
[01:06:07] Bernard Mont-Reynaud: And on the screen would show three red squares. And you say, make them blue.
[01:06:15] Bernard Mont-Reynaud: in blue, you say, you know, triangles, and they would turn into triangles, or you’d say seven balls, and now you, et cetera, you get the idea. And normally this was not done in English, or it was done in English, but you had a version done in the Senegalese language. Suddenly I don’t remember what the name of the language was, but because this machine was trained to individual speakers, you can get the person to train their own vocabulary in whatever language it was, and then have it work in that language.
[01:07:02] Bernard Mont-Reynaud: So anyway, by Friday and by staying up all night a couple of times or whatever, by Friday I had the demo done. This was four days, it was just amazing.
[01:07:13] James Parker: And even though you didn’t work on speech at the time?
[01:07:18] Bernard Mont-Reynaud: Okay, I should explain, I thought I should explain this. For the speech itself, we had this big box, it doesn’t fit in our screen here, maybe a box this wide and the same depth, and about this high. This was the signal processing you could put into this box, the sound of your voice speaking certain words. You would train it to your own, your vocabulary and your own voice.
[01:07:50] Bernard Mont-Reynaud: And this box was then capable of doing the analysis, both of the words and of the, it was a two-stage dynamic programming algorithm, which as the sound came in, it would give you back the transcription. So the signal in that machine was called the NEC CSP-200. That was the machine doing the signal processing. I had a Symbolics Lisp machine that communicated via a serial line, RS-232, with that external signal processing box. And that’s how I did it, by just route.
[01:08:37] Bernard Mont-Reynaud: The audio was going to that machine, it gave me a transcription, and then I processed the transcription, the demo. So that, it was just at arm’s length, if you want. This is the technology of the time. There was no integration, and there was no way I could have done this in four days, you know, if it weren’t, you know, by having a self-contained piece of equipment to do the speech recognition. And what was interesting about this is that the demo succeeded. It was working.
[01:09:17] Bernard Mont-Reynaud: I mean, there was an incredible technological feat in some ways that, oh, all was together. But President Senghor, didn’t see the point. What does it mean that you can do that? You know, how does it contribute to humanity, to the problem of my country, to anything that you can come in a computer to do something like this? He couldn’t see the point. And if you look at it, it was just a technology demo, was not connected to any real need of anybody, right? And this was 1982, remember?
[01:10:08] Bernard Mont-Reynaud: And Senghor was a poet I mentioned, and a philosopher, and he’s one of three people who had founded this movement called Negritude. He talks about what it means to be black, you know, in the world, and a movement that also wanted unity of the African countries and so on. Very interesting character.
[01:10:32] Bernard Mont-Reynaud: But the Centre Mondial was trying to have all sorts of connections with Senegal in particular, was a main area and other third world countries, but with mixed success, because this was a raw and pushed rapidly forward application of technology to countries that weren’t necessarily receptive to it.
[01:11:01] Bernard Mont-Reynaud: I thought this would connect with the other angle.
[01:11:06] James Parker: Oh, 100%.
[01:11:06] Bernard Mont-Reynaud: Yeah, yeah.
[01:11:07] James Parker: It’s fascinating.
[01:11:08] Bernard Mont-Reynaud: So I have to mention that story.
[01:11:13] James Parker: I mean, I didn’t even, I also didn’t know that Negroponte was, you know, I’ve looked at his, you know, book, The Architecture Machine and some of his influence at MIT. And I don’t know, there’s just, it just sounds quite Negroponte-ish.
[01:11:40] Bernard Mont-Reynaud: Yes, indeed. Indeed. Well, let me give you more on that that you may not know. So Negroponte had The Architecture Machine Project at MIT. He was the director of that and so on. At some point, he wanted to grow that into a university department and have the whole complete independence. And MIT wasn’t quite doing what he wanted. There was resistance. So he’s gone to, he said, oh, it is so. Then I’m taking my team away. And he went like this to MIT and went to France, got money from 12 different departments of the government, 12.
[01:12:25] Bernard Mont-Reynaud: He was going to solve all the problems in the world, the lack of an alphabetism, how do you call it, the literacy. He was going to solve literacy, he was going to solve this and that and the rest, no, no. And got lots of money, got a wonderful center in Paris. There was this French politician involved, Jean-Jacques Scherber, who was very well, very powerful in the government. And they had this thing going on for two years, more or less, during which, you know, they organized conferences and this and that.
[01:13:06] Bernard Mont-Reynaud: And Negroponte was telling MIT, you see, if you want to have me back, you meet my conditions, OK? But that didn’t happen for a while. This is where I end up, you know, going. But ultimately, the game, they didn’t really care about this. They just took all this French money and they played Negroponte’s games.
[01:13:29] Bernard Mont-Reynaud: But eventually, he got what he wanted from MIT, and this is what became the MIT Media Lab, OK? So this made the transition between the Architecture Machine Project and the Media Lab, which was bigger, a whole new building, a whole new department, everything. So we were at a pawn in Negroponte’s game, and a pawn paid for by the French.
[01:13:57] James Parker: And when you were at CCRMA in the 80s, what was your relationship with the Media Lab? Because I know, obviously, the Machine Listening Group that sort of came, eventually appeared there, sort of had some familiar names to your project and stuff. But was there, I don’t know, any tensions at all?
[01:14:21] Bernard Mont-Reynaud: Not at all. I mean, CCRMA and them were in completely different spheres. However, there were some contacts between researchers, for example. I mean, a lot of people met at the Computer Music Conference. And this is, for example, when I met Barry Vercoe, who was at the Media Lab. And he probably, I don’t know if he had been at the Architecture Machine Project or not, but he definitely became part of the Media Lab. And I and Roger Dannenberg at CMU, we were doing research in related areas and sort of developed personal contact.
[01:15:09] Bernard Mont-Reynaud: But there was nothing institutional between CCRMA and the MIT Media Lab. Not even feelings one way or the other. It was just as individuals, we related to each other. Since I mentioned Roger Dannenberg and this idea of score following, you may be familiar with the fact that Barry Vercoe was involved in that. But later on, I was one time doing a paper with Roger Dannenberg of CMU, where we did the first time, instead of following a score, we were following an improvisation in real time. And so I did part of the system.
[01:15:57] Bernard Mont-Reynaud: And Roger Dannenberg was also a trumpet player, was playing his trumpet. And the system had basically blues grid. It would figure out where it is in the blues progression and then start accompanying the blues based on the trumpet solo, you know. And so that was work I did with Dannenberg. So this shows the kind of interactions that were happening across labs, you know.
[01:16:31] Bernard Mont-Reynaud: Okay, so maybe we should turn those parenthesis back, but I thought the Centre Mondial parenthesis was really interesting. And just to close it, they weren’t really caring about my research. The name of the game was, there were two names, was one for Negroponte to get what he wanted and Pappard to get what they wanted from the Media Lab, and number two for the Jean-Jacques Avant-Sherbet, the French politician, was also the director or, I’m not sure, owner of the socialist journal, L’Observateur, and what they wanted was news.
[01:17:11] Bernard Mont-Reynaud: So this was a case of fishbowl research, where you’re doing some research and there’s cameras on it and news articles about, oh, they’ve done that. And then it started to happen, you could see, they would make news before the research is done. So a couple of times, I worked really hard because they had announced we do X, and then I rushed to do X. And then after a while, I realized that they don’t care if it happens at all.
[01:17:43] Bernard Mont-Reynaud: I don’t have to rush to do X because they’ve announced they’re doing X. I was in charge of the audio and speech and I created, the name came back, the language Wolof, the language from Senegal, I created sentences from text for Wolof, because they had announced they were doing it, and in two months, I did that. In two months, I built this by going to Stockholm with a linguist coming from Dakar that was to a place in Stockholm where there was this thick of snow and ice on the ground, and he was frozen out of his wits, you know.
[01:18:23] Bernard Mont-Reynaud: Anyway, we did it, and he was there for a week. I put together that system in two months, which is an incredible feat. You think they cared? Not at all. The payoff had been two months ago when they announced they were doing it. Okay, anyway, once this happened, and they did all the projects, and at some point, my grant got funded. I went, bye-bye. I went back to continue my research at Cormac. Okay, so that was one parenthesis. One thing I had wanted to mention in the second round of NSF, I did pattern recognition in rhythmic material.
[01:19:07] Bernard Mont-Reynaud: You have to follow temporal variation. Music can have what’s called rubato, right? You slow down and accelerate, but there’s also a lot of fluctuation of events themselves. Maybe it’s due to signal processing errors, or it’s due to the fact that quarter notes aren’t all equal in performance, and you have to decide when is it more like temporal variation, and when is it kind of micro variations, and so on. There was a bunch of work aimed at doing that, and there were also in the software layers that would look at, well, what happens to those patterns?
[01:19:47] Bernard Mont-Reynaud: If I, what makes a good quarter note, or what are triplets, triplet eighth notes, and you would create those clusters, and so on, and be able to fix errors by looking at those statistics, and this is kind of what, one thing I was involved in, but okay, I mean, those were some of the things. I think we can go forward to SoundHound.
[01:20:17] James Parker: Yeah, let’s do that. I mean, my understanding is that SoundHound began as a music recognition or humming, melody recognition company. And now it’s like a big, huge speech recognition, voice assistant company, but I was guessing that maybe it was the music connection that somehow brought you to be involved with them, but I don’t really know. So, that was my best guess. How did you come to be involved with SoundHound?
[01:20:53] Bernard Mont-Reynaud: Yeah, you’re guessing quite correctly, that at one point, I found out about Soundhound, and I picked up the phone. And I had a wonderful conversation with, and next thing I was talking to CEO, and the next thing I was hired, they had a hiring freeze. They had had a hiring freeze, but the part that wasn’t so easy is that they were not hiring because they had been on a stretch. You know how companies get into this thing called, how would they call it, the desert of funding. They launch a company, and then there’s no product, and they are dry.
[01:21:40] Bernard Mont-Reynaud: They don’t have any income, and the investors are tired of giving them funding. So this is no man’s land until something happens that makes them, anyway. So they were at that stage, they let go of a few people and had a hiring freeze. They begged the board to hire me, and I went in and so on. And yes, they were entirely focused on music pattern recognition, and the, what’s that song, name that song issue.
[01:22:17] Bernard Mont-Reynaud: It’s good to know for the story that from way back when, they were interested in actually doing the speech recognition and the natural language understanding. It had been on their mind, but there was no traction for that, there was no product. So it was an interest of theirs, but it got put behind because they had some traction on the song recognition, and some in the humming, right. And they had just a dialer as far as speech recognition, just dialing application.
[01:22:59] Bernard Mont-Reynaud: So yeah, it’s quite right, it’s my work in music and music pattern recognition that made me a fit.
[01:23:09] James Parker: So you came on relatively early, it sounds like around 2010 or something like that.
[01:23:15] Bernard Mont-Reynaud: Yeah, 2010, February, yeah, it will be 13 years in just a month.
[01:23:20] James Parker: And then they, I think they were originally called Midomi or something?
[01:23:28] Bernard Mont-Reynaud: Correct.
[01:23:28] James Parker: Or the app was Midomi maybe? But then…
[01:23:33] Bernard Mont-Reynaud: The app was, maybe the app was Midomi, the company definitely was Midomi. I suppose that is Midomi, you know what I mean?
[01:23:45] James Parker: Oh, I did wonder, okay. That makes sense.
[01:23:49] Bernard Mont-Reynaud: Yes. And yeah, I’m not sure if the product was called Midomi at the time or not, but it became called Soundhound. Before the company was called Soundhound, I believe the product was called Soundhound.
[01:24:14] James Parker: Oh, okay.
[01:24:15] Bernard Mont-Reynaud: Or maybe it was at the same time.
[01:24:17] James Parker: My understanding is that, I mean, I don’t know this history very well, but it seems like they had a lot of traction in the sort of early days of smartphones, because sort of name that tune was quite a sort of cool thing to be able to do with a smartphone at the time. It was like in the very early days of the App Store, you know, and they were one of the biggest apps on the iPhone for a while there, and then in direct competition with Shazam and so on. And then they sort of, yeah, it seems like they just moved on.
[01:25:02] James Parker: And now they’re a totally different company. So this is what it seems like from the outside.
[01:25:10] Bernard Mont-Reynaud: Yeah, most of what you say is correct, except for one thing, that the Soundhound application, the music recognition application continues to this day, it hasn’t gone away. It’s just that the music market was kind of this big and has remained kind of this big, you know, the speech and natural language is much bigger. So if you wanted growth, and also it was their original love. It’s not that it was a complete reconstruction of the company, it was not a complete pivoting.
[01:25:50] Bernard Mont-Reynaud: It was, they had been interested in this all along for like 10 years, or something, right. So they actually wanted to do that. But so at this time, the music application Soundhound still exists, you can still download it and use it and so on, it’s gotten better over time. But the conversational AI is really the dominant one.
[01:26:26] James Parker: What work were you doing? I mean, it’s a long time to be with a company. Did you also move from music towards conversational AI? Is that sort of your main field recently?
[01:26:42] Bernard Mont-Reynaud: Yes. Well, I’ve done both. I did start with the music. I see my power is low. I started with the music. I started optimizing it, even using hardware to optimize it and so on, or your software. One of the things that happened, at some point I started doing patents. I started having an idea for one thing, an idea for another. I wrote a couple of patents. And then they had a need for a patent person, you know? And I became Mr. Patents, Patent Guru, they called me.
[01:27:19] Bernard Mont-Reynaud: And so I ended up, you know, doing a lot of inventions myself, or helping other people doing inventions, and doing the sort of responding to the patent office. You know, you have to usually, it’s very rare when they say, oh, good, you have a patent here. You have to defend it. It’s called prosecution. Sometimes you have to adjust things, and so on and so forth. Anyway, I’ve done a lot of patents, but also I’ve been involved later on with the natural language intelligence, natural language understanding. And so I’ve been very busy in that as well.
[01:28:01] Bernard Mont-Reynaud: And so I ended up having multiple hats over time. I had a patent in natural language understanding, and some mentorship in other areas, and speech.
[01:28:19] James Parker: Can I ask, because this is a, I don’t know if this is really your field of expertise, but you’ve been close to it in a way that it’s much closer than me. You know, I read recently that Amazon had just burned up a hell of a lot of money in one year. I could be wrong, but it was a huge amount of money in any case. And I can’t, it’s hard to understand where the market like is going or where the investment, the smart investment is. It seems like SoundHound as a company, I’m not really asking you to do PR for SoundHound.
[01:29:23] Bernard Mont-Reynaud: I can explain, I see where you’re headed. Let me try and answer your question. First of all, there are many companies, they have virtual assistants for their own purposes. You know, Amazon is maybe hoping that Alexa is going to make them money with something, let’s see. Now Apple has Siri and Siri helps them sell equipment. So they, I don’t know how they evaluate their budget, but they have a reason. Now we are not selling anything. We have this application called Hound, which is a virtual assistant. You can try it, you can have it for free.
[01:30:08] Bernard Mont-Reynaud: This is not how we make money. One thing we do with it is we collect voices and those voices. And by the way, we, speaking of a privacy issue, we mask them, we store them in a way you cannot recognize the original people, the original voices. But this is helping train our systems. You need a lot of voice data to train the voice recognition. So we use it for that and we use it to give demos and we use it to, but where the money is, is none of that. The money is in applications.
[01:30:47] Bernard Mont-Reynaud: And in particular right now, SoundHound focuses on the automotive market, you know, having those kinds of system in cars. And on the restaurant voice kits, the ordering, you know, you order from your car and so just generally restaurant ordering. So these are the two. SoundHound has had an evolution from being this tech company, which of course still remains behind the scenes, but to have this market-driven areas. And now it has specialized on those two largest areas. This is where we are, company focus-wise. So maybe that answers your question or not.
[01:31:36] Bernard Mont-Reynaud: But for a while we were carried.
[01:31:43] James Parker: Because it sounds like you’re saying that, you know, Amazon wants to become infrastructure, basically. It wants Alexa to be…how you access everything, but there’s not that much money in that, maybe, whereas…
[01:32:04] Bernard Mont-Reynaud: I don’t think they’re succeeding in that. They don’t have the internal capability. Alexa is very compartmentalized. Each of the capabilities, they don’t form a network like we do. You can’t have one of these… I’m trying to remember their competencies, their applications. They don’t talk to each other. They don’t have a full background of intelligence. We have what we call… I can’t remember the branding word.
[01:32:37] Bernard Mont-Reynaud: Integrated AI. No, it’s collective AI. That means that we have all those things talking to each other and sharing parts. They don’t have this in Alexa, so they are failing on the infrastructure as far as we’re concerned by not providing this crosstalk of all the different apps, which skills… The name came back. They call them skills. Well, those skills don’t talk to each other. They remain isolated skills. If they don’t solve that problem, they can’t have any conversational AI of any power at all. It’s kind of silly to have invested so much.
[01:33:27] Bernard Mont-Reynaud: I think the smart speaker is very good. We wish we had one. No, but the Alexa capability does not compare in terms of its intelligence.
[01:33:43] James Parker: That’s really interesting. I mean, I’ve never knowingly used a Soundhound voice assistant. I probably have because I think the idea is that you’re sort of the engine behind many, many, many different voice assistant systems, aren’t you? So, I probably have used them without knowing.
[01:34:10] Bernard Mont-Reynaud: That could be.
[01:34:12] James Parker: Anyway, I guess I’m wondering if there’s a way of drawing the two conversations together so you come from music through into SoundHound, you end up working more on this sort of very product-driven voice assistants or sort of application-driven voice assistants. Is there a through line between that most recent work you’ve been doing? Is it that auditory scene analysis is somehow crucial?
[01:34:48] Bernard Mont-Reynaud: So far, you scored 90% on your intelligence and understanding. This one, I need to make some corrections. First of all, let’s see. First of all, I am not deeply involved in what SoundHound has taken that pivoting, I could say, or this new focus towards a very market-driven organization. I understand they need this for growth. We’ve gone public six months ago or so, and they need to do that. I haven’t been part of that.
[01:35:28] James Parker: Oh, you were doing the patents and the mentoring.
[01:35:33] Bernard Mont-Reynaud: I’ve been doing the patents. I’ve been building pieces of the deep architecture of the natural language understanding. I’m still deep in technology. I have not been at all involved personally with any of that marketing effort. So, that makes sense to me. So, that makes sense to have done that. Well, I’m now retiring. A lot of people have to focus more their work onto the marketing areas of interest. I haven’t had to do that, which works fine for me because that’s not my temperament to go and work on a large market.
[01:36:08] Bernard Mont-Reynaud: It’s my temperament to build technology or to have ideas or to be a mentor to people. The other thing is, yes, I wanted to tie it back in another way. I’m doing this work on natural language understanding, but if you go back to I was talking about my days at Stanford and doing my PhD.
[01:36:33] Bernard Mont-Reynaud: Among other topics, I had gotten kind of very interested in natural language understanding. At the time there was Professor Terry Vinograd coming fresh from MIT with his PhD thesis on the system called Schwergelu, Schwergelu, you know, go find a name like that. But he was speaking a mile a minute about NLU and that was just fascinating. I had a big interest in that. So that was the, the natural language AI was already on my mind. That’s in 1974, 75.
[01:37:15] James Parker: But now you must be doing, you know, data driven and statistical methods in a way that weren’t so prominent back then.
[01:37:32] Bernard Mont-Reynaud: Yes and no. And my battery could be dropped out at any moment. I don’t want to start moving out to a place where I could plug it in. So just so, or would that be? Let me see if that, if that wire happens to be just the right thing and plugged in. Doesn’t seem to be plugged in. So with a warning that my computer could die. Yes and no. The speech recognition part of the system has gone over into the statistical and neural network.
[01:38:14] Bernard Mont-Reynaud: Basically neural network architectures, one type or another, language models, partly statistical, and now they’re vastly neural network built also. So we’re going, we’re with the rest of the industry in that type of approach on that part. But on the part, excuse me, on the part that’s natural language understanding, there are front models and the system that we use at SoundHound as a language that has a grammatical component and it’s semantic grammar.
[01:38:59] Bernard Mont-Reynaud: So you have a syntax structure that is being extracted and then the semantics are hanging off the syntax. This is more in a way the old AI approach as opposed to putting everything into a neural network. There’s a big battle in the field or a big, it’s not a battle so much. People choose one camp or the other, but they don’t battle, they just invest in whatever they do. But it’s a big difference of approach. There are ways to get the two talking to each other and we do some of those ways and I believe all those like Google.
[01:39:36] Bernard Mont-Reynaud: Google, they’re primarily neural network, but they have a lot of linguists providing information. For us, we are using a programmable grammar if you want, but we have linguists and we have neural networks helping some of this, but we have a different mix from other people. But the core, the backbone if you want, of the natural language understanding is grammar-based and semantic grammars that is.
[01:40:04] James Parker: That’s incredibly interesting.
[01:40:05] Bernard Mont-Reynaud: As opposed to, you know, yeah, yeah. And I forget.
[01:40:13] James Parker: I mean, you’re gonna run out of batteries. You’re gonna run out of batteries and we’ve talked for nearly two hours. So, you know, do you wanna just wrap it up here or do you have any concluding thoughts or I mean, you could talk a little bit about where the field of machine hearing is or machine listening or sort of, I don’t know if you have any general observations, you know, I don’t know, blue sky thinking or if you’d prefer to just draw a line and say, say thanks.
[01:40:53] Bernard Mont-Reynaud: Yeah, well, I mean, I can say thanks anyway because I feel like, you know, I’m now retiring and I’ve been lucky to have a wonderful career, you know, that has brought me a lot of interest, curiosity, joy. I’ve been involved in many different fields and applications. I haven’t mentioned many of those things, but certainly, certainly audio and music and natural language and the auditory system have been like a big part of what I’ve done and that’s been wonderful. So the gratitude is here.
[01:41:40] Bernard Mont-Reynaud: In terms of kind of the future, you see, I had this idea back when of this quote, programmable Bregman to build a system that does auditory scene analysis in a general way. And it’s been one of my dreams as I’ve seen people go to build new, to, as opposed to the kind of systems I build, which were the old AI, right? The feature-based, you just build the pieces yourself as opposed to have it be learned out of the statistical, you know, the large corpuses of data. Well, people have started to develop a lot of cleverness about putting together architectures of neural networks, of sub-networks and so on and so forth. And I personally kind of missed that turn. I was deeply involved in one thing and another.
[01:42:41] Bernard Mont-Reynaud: I did not take a jump into that type of research. And I think the kind of principles that I was working on after my work with Bregman, were asking for a neural network architecture, which would do this, which would take the sound apart and give it components, which basically would do grouping based on the same type of reasoning, the same type of psychological principles that Bregman uses. But it would have been a dream of mine to actually build that type of system.
[01:43:19] Bernard Mont-Reynaud: I never got to it because I had too many different responsibilities and maybe I wasn’t quite deeply versed enough in the neural networks. I don’t know, but anyway, I think there’s somewhere in the future, some systems that will address and solve or better solve source separation, you know, with neural networks. But I see beginnings of that, but I think there’s still a bunch missing. You know, it takes a lot of work to reproduce what the auditory system is capable of.
[01:43:58] Bernard Mont-Reynaud: But now I’ve become convinced, which I wasn’t for a good many years, I’ve become convinced that the time will come when neural networks can do that. Neural network combinations of neural networks in a more complex architecture. And of course we have a lot of knowledge about how that happens in the brain, you know, and it’s a very complex architecture that has all these pieces. But there’s evidence, I think, that this will be possible in time. And so that’s…
[01:44:34] James Parker: That seems like a really good note to end on. If you hear of anybody doing it, let me know.
[01:44:45] Bernard Mont-Reynaud: Yes, well, I’ve seen a few papers. If you ask, you see, I’m no longer in this because I’ve had to be a little focused and stuff. But yeah, let’s see. If I did some digging or if you ask around, there are bits and pieces, or even, you know, there were searches on neural networks and auditory separation or something like that, or sometimes you have to refine your search, but you may see a few things coming up. It’s not completely there, but there’s definitely a beginning of it. So, and again, I don’t exclude the…
[01:45:34] Bernard Mont-Reynaud: Once I have more free time and after I take some traveling, you know, I may go back into this and say, okay, so what’s going on here?
[01:45:41] James Parker: You never fully retire.
[01:45:44] Bernard Mont-Reynaud: Yes, but I may not, you know, I don’t think I’m going to contribute to this. I’m not one of these people who are like never going to retire, who are going to continue as like emeritus professor, never stop. No, I’m ready to turn the page. You know, I’ve had a 54-year career and that’s good. I want to travel and do art.
[01:46:15] James Parker: You should. You absolutely should. Thank you so much for your time. It’s 10 o’clock at night here, and I feel like you’re in Spain and on holiday and you should go and enjoy yourself. This, I’m going to click stop on the recording, if that’s okay with you.
[01:46:39] Bernard Mont-Reynaud: Yeah.
[01:46:40] James Parker: Thank you.
Beth Semel transcript
James Parker (00:00:46) - Thanks so much Beth. I was wondering, maybe would you like to start off just by introducing yourself however seems right to you?
Beth Semel (00:00:55) - Yeah sure. So my name is Beth Semel. I’m an assistant professor in the anthropology department at Princeton. And I like to say that I’m an anthropologist of science, technology, and language. So my training is in science and technology studies, linguistic anthropology, also dabbling in medical anthropology. But the science and technologies that I study as the first part of that label is communication, engineering, speech signal processing, this thing that we could call machine listening, and then also the science and doing of psychiatry, mental health care, and the kind of various technologies that are involved in that practice. And the language bit is me kind of thinking about, you know, not just modes of speech, not just the production of language, but also the that modes of interpreting, receiving language, which again is where I see the kind of, when I think about machine listening, I think about it always in that kind of relational interplay between both speaking and listening. So, I guess that’s me.
James Parker (00:02:13) - Amazing. I mean, yeah, it sounds awesome. Do you want to say anything about how you arrive at that set of concerns? I don’t know if that’s, I don’t know if you’ve got a kind of a story, some people that we’ve spoken to, you know, their sort of biographical information that really sheds light on their kind of series of concerns. I don’t know if that’s the case for you.
Beth Semel (00:02:36) - Yeah, it’s, you know all anthropologists kind of have their practice origin story, but I’ll maybe veer off of that a little bit. Thinking about the personal side of things, you know, my mom is a retired speech therapist. And so, yeah, so, you know, growing up, I was kind of aware of this idea that there is a need for, you know, bringing people up to a kind of like standard threshold of intelligibility. And that’s not, you know, that standard is something that can be kind of fluid, right? It can be set between the client and the therapist, but they’re also kind of like broader systems that set that standard in place, right? Standard kind of acceptable quote unquote professional or whatever ways of being intelligible, sounding intelligible. And just the idea that that’s, you know, that has to be put into practice and made, I think is something that I realize more and more really does like play a big role in the approach that I the critical rather critical approach I take at thinking about this weird, strange field of focal biomarker research, which I sometimes call machine listening and mental health care.
But I mean, the more standard origin story is that I was at MIT because I was studying, pursuing a PhD in science and technology studies and anthropology. Because I was really interested in mental health care as a thing that’s both technical but also interactional, right? The primary tools are talk and interpretation and interaction. And at the same time, there’s this whole idea of like evidence-based therapies, evidence-based treatments. So there’s this quantifying matrix that practitioners are trying to push patients’ talk through. And it just so happens that I entered into the PhD and was kind of chugging along at this moment of change. American mental health care and the kind of primary funding bodies of American mental health care kind of rejecting the standard technology of mental health, which is the diagnostic and statistical manual of mental disorders. And saying instead well we should be hanging and collaborating with engineers and using data driven methods instead of ones that are more based on clinical wisdom, clinical know-how. Being at MIT, there were people there doing this data driven research about ‘ok, can we use functional magnetic resonance imaging to predict which patients will respond well to cognitive behavioral therapy.’ And so this is like a type of research that falls under the umbrella of digital psychiatry or precision medicine. And so in kind of like hanging out with those people, getting to know them, talking with them, they mentioned, oh yeah, you know, it sounds like you might be interested in this one lab that does vocal biomarkers. research. And I said, excuse me? What is that? What is a vocal biomarker? That doesn’t really make any sense. How can the voice be biological? And this person said, well, you know, your brain controls everything, including your voice. So if you analyze the sound that the voices make at a kind of, you know, the level of the physical waveform of speech, then you can find out about the source that made it, the brain. And so through there, I just went on this kind of wild goose chase to find people who were doing this work because it kind of miraculously allowed me to unite an interest I had in thinking about interpretation, mental health care, but also technology. And again, that impetus to push talk and interaction through, in this case, not just a quantified matrix, but a specifically computational one, which is like almost like hyper quantified.
James Parker - Right. And when did you sort of when does that origin story like take place? Like what year?
Beth Semel - Yeah. Around 2015. Yeah, so that so that was the time that the director of the National Institute of Mental Health, Thomas Insel, he, he was a director at the time, and he put out this really, you know, this funding call that really kind of upset a lot of people that said we’re not we’re not going to be supporting any research that uses the DSM which was kind of like earth-shattering because people use the diagnostic categories and the diagnostic criteria in the DSM not just like in a clinical context but in a research context so they use it let’s say you your lab wants to study bipolar disorder you need patients who who have been diagnosed or presumably like exist within that diagnostic category. So in the past, people would typically use the DSM as a way to create a research cohort. And in so basically said, no, we won’t fund that research anymore.
James Parker (00:08:09) - But this was like early days, right? So I mean, so it sounds like you’re saying that this is like a big driver or something that was kind of like a sort of a sleeping giant or something for a while. Because I mean, I didn’t hear about the kind of vocal biomarkers research and sort of computational vocal diagnostics until much, much more recently, like really only the last couple of years. And then there’s been some kind of huge like US national announcement about like sort of pouring money into this project. So it’s or related projects, it seems like sort of you were you were studying studying a field? I don’t know at its birth, or it’s sort of I don’t know. I don’t I don’t know. How would you describe it? And like what’s changed? I suppose? I mean, we should get into the details. But just in terms of laying out the landscape, like, it seems like it would have been, you would have had to explain yourself to absolutely everybody when you said what research you were doing when you started. And now maybe you don’t need to so much.
Beth Semel (00:09:21) - Yeah, it’s really been remarkable. Indeed, when I first started this research and I would tell people that this exists or that people have a desire for this or think it’s a good idea, they would say this is absurd. How could anyone believe in this as a concept or want it? And as you say, this giant, I think it’s, I don’t know, three, four million dollar NIH grant, that’s, it’s for voice, it’s called voice as a biomarker for health and they’re kind of different arenas of health that they want to study and one of which is mental health care, the other are more like I mean, I think, you know, thinking about what’s that the question of what has changed or what has led to the coalescing of these two kind of paradigms or two kind of orientations towards like language and specifically like language and mental health care. That’s been a question that I’m trying to think through now through archival research. So the archival material that I found, I found papers that are doing what we would today call vocal biomarker research, are trying to say like, can we, you know, ring something of the neurobiological, psychopathological, biologically speaking, from the voice using, you know, computational methods. I found papers that are doing that, like in the 1930s, trying to do that, proposing that as a concept. A bunch of dissertations in Germany, which is slightly concerning. Why were people interested in this in 1930s in Germany, PhD students in particular?
But even in the US as well, I found papers kind of parroting this concept, engineering papers parroting this idea that you could get something meaningfully biological about the psyche through the voice in the 1960s. I think, you know, maybe to to conjecture a little bit and kind of pull the, you know, the camera lens outward, I think there’s a kind of, in many ways, a broader acceptance or really a kind of, like, capitalist, like, market hunger for computational things. There’s a kind of inertia that I see happening that, you know, I think voice stuff is just one way that people are trying to do precision psychiatry, but I think because there is this imaginary about the voice as being easeful, right, as being immaterial, as being kind of a public object, freely available, freely floating. I’ve heard, you know, startup people who talk about the voice being a very cheap signal, right? The face is so multifaceted, the visual world is so multi-dimensional, the voice is flat, it’s a singular dimensional thing, it’s just one waveform that makes it not just easy to analyze, but inexpensive. So I think, again, there’s
James Parker (00:13:05) - microphones are cheap, like still, right there, you know, that it’s it’s sort of cheap across every dimension, isn’t it? I mean, just to give one sort of very stark example of the kind of the drivers or the kind of capitalist orientation or the market for this kind of stuff. I mean, an Australian based company, we’re based in Australia, just got bought by Pfizer for $100 million on the basis of a claim that it can do COVID cough diagnostics. So it’s not exactly what you’re talking about in terms of computational sort of psychiatry specifically, but the kind of the sort of voice body nexus and the sort of the idea about the the the sort of the way in which the body speaks through the voice, even when it’s not speaking, if you know what I mean, you know, it’s sort of very clear in the cough is kind of the perfect example of something on the fringe of the voice, the voice that is sort of without it’s kind of not laden by speech. So I mean, yeah, I mean, it seems like there’s and they’re not the only company to have done that. I mean, it’s interesting that Pfizer did that, you know, in the wake of COVID. but like there are a number of other companies, Sonder Health in India and like a few others that have been pushing this. It feels like the pandemic context has, you know, and the sort of the riskiness of the voice and the relationship between breath and contagion and stuff has really kind of been a bit of a kind of, yeah, kind of a trigger for an explosion of interest in voice diagnostics. I could just it could be coincidence, but it does certainly feel that way that that, yeah, the pandemic is a kind of a driver into this field.
Beth Semel (00:15:03) - Yeah. And I think, you know, even before the pandemic, a lot of these, the, you know, emerging vocal biomarker companies are people who work constantly, you know, in like the initial starting, there’s quite a few actually vocal biomarker companies that were cropping up in this like 2015, 2016 time that ultimately ended up pivoting to a different offering like, you know, something like, oh, we’ll give you personalized recommendations for a therapist, or we’re a therapy app now. I know a few companies have done that. But I think this, this impetus or this desire to like be tethered to the patient, even in the absence of any kind of physical connection, and this imaginary of easefulness and easefulness particularly of capture and of knowing the patient. I think that’s, to bring it back to your question about why did this all happen at the time that it did, that’s where I think we can start to see the kind of connective tissue with really a very like, biologizing impulse in psychiatry and mental health care that’s driven by capitalism but Also really kind of about like capture, right? Like pinning down something essential about not just the person but trying to funnel down mental illness, psychiatric suffering, mental suffering as like a definitive object that can be held onto and known and done something to which, you know, from like a disability studies perspective, that’s, you know, sounds, it sort of rhymes a whole lot with eugenics and other kinds of modes of social control that are pretty oppressive.
So I think there is, you know, because the aim or the impetus or the goal here is like health, right? Health benefits. Of course, we want to do whatever we can to mitigate the spread of COVID. Of course, you know, asterisk in the absence of like actual state infrastructure to help mitigate COVID, we need, you know, something of a band-aid that might do the best that it can to help. But I think lots of really smart people in the critical code studies, race critical code studies like Ruha Benjamin and Safia Noble have done a lot of great work to show how that kind of beneficent intention isn’t enough and sometimes can ultimately draw attention away from looking at the kind of unintentionally harmful side of things. Yeah.
James Parker (00:17:48) - I’d love to get into some of the nitty gritty of your, I mean, you’ve written a lot about this topic or these topics. And you have a wonderful thesis, which has all of these amazing ethnographic case studies. And I’d love to get into the nitty gritty of at least some of them. Talk through your own experiences, sort of navigating this strange new and emergent world. But I just kind of feel like it’s worth pointing out as a segue into that, that those examples are all of kind of research labs, sort of really in the kind of prototyping phase, like none of the, it seems like, am I right, that none of the projects that you investigated sort of gone to market? And I just wondered if it’s worth setting up a little bit, you know, what the contemporary landscape is other than kind of hype and flooding investment and so on, you know, are there any extant companies that are really already doing vocal biomarker stuff? Are they are they all pivoting out because it doesn’t really work, you know, into being other kinds of apps? How do you how do you understand the sort of the lay of the land as far as this field goes in the contemporary moment before we dive back into those case studies.
Beth Semel (00:19:16) - Yeah. Yeah, it’s a, it’s a complicated question. You know, I think, on the one hand, the field of like real time sentiment analysis, which is like a branch of affective computing, that stuff exists, and we could call that quote unquote, vocal biomarker research or kind of machine listening for sentiment analysis research, right? I’m not sure that people would connect the two, but to me, I think they are connected. And those technologies are, I mean, I don’t want to, I want to be a little careful naming companies just because I don’t want to be slapped with a libel lawsuit. One of the other benefits of working with the university is you don’t have to deal with that scary corporate boogeyman. But there are companies who use this real-time sentiment analysis stuff for call center workers, right? To help not only help the call center worker kind of manage the affect of the person that they’re on the phone with, but also as a way for you know, managers to surveil the workers, right? And it creates a kind of quota making system. And it does work under a similar premise, which is that there’s like emotion, affect, sentiment exists as a kind of physical feature of the voice that you can track and, you know, hold on to in a sense.
But in terms of like specifically vocal biomarker stuff, you know, there’s a lot of companies. There are some research labs that are pivoting into or kind of spinning off into companies. I’m assuming too there are people who are sharing or selling AI models or data sets with kind of companies. But as to whether or not it works is, that’s a question that I used to get a lot, especially in the earlier kind of times when people were less, you know, this wasn’t like a regular headline in the US news. It doesn’t work.
But is it real? I would get that a lot and can I hear it right can you have an example for me that you can play for me. You know, I used to say like no, of course, it doesn’t work. It doesn’t work according to you know, all linguistic anthropological framework, which says you you can try to disentangle language from social context, from history, but doing so requires ignoring a lot of really important, essential things about how people make meaning or have meaning imposed onto them through language and talk. But does it work in the terms that it’s trying to work, which is, again, to capture something essential about the mentally ill speaking subject? It sort of doesn’t really matter because, like the paradigm there is one of capture, right? One of essentializing. So, you know, those technologies are being developed. They’re being developed by being tested on people in the same capacity that I studied in my field work. And, you know, while my understanding is that the labs that I studied, you know, kind of close up shop or shelve their companies, I can either confirm or deny that they didn’t hand over their data sets, that they didn’t integrate their data sets with other people. And it’s really hard to prevent people from building models from voice data sets that do something slightly different, a little bit different, or very different from what the data set was originally gathered for.
So it’s hard to say what the field is doing right now, there has been this cycle of kind of hype and then no one really producing great statistically sound, you know, results, but somehow it feels like this NIH grant is somewhat of a turning point. And I don’t know several other things. I don’t know if it’s too much detail to go into happening in the US, like legal cases and revelations that like McDonald’s old is collecting voice biometric data through the like the drive-through kiosk, you know, stuff like that just makes me, you know, my like initial kind of naive anthropologist like, no, of course it doesn’t work. It’s like, well, I can’t deny that it’s working for some end, right? Right. And that it exists. People are investing lots of time and money and, you know, investing voices, right, their own voices in it. So it’s enrolling people into systems of surveillance whether or not whether or not it quote unquote work. So it’s doing a kind of work in the world and it warrants study as a result. I mean, yeah. Yeah, I think if I could just like add on to that a little bit. Shoshana magnets concept of biometric failure, I think is really helpful because she essentially says, you know, these biometric technologies, the failure is not like, okay, that’s it, it’s done. But things are constantly failing by virtue of working. So they worked by making this kind of misalignment or doing that kind of decontextualizing work. That’s how they work. That’s what they’re designed to do is to decontextualize. So they are working, but from another perspective, they’re failing.
James Parker (00:25:01) - Right. Should we do some of the case studies?
Beth Semel (00:25:09) - Sure.
James Parker (00:25:10) - I can’t help but begin with your your chapter on depression. And there’s this.
Beth Semel (00:25:19) - What an opener.
James Parker (00:25:20) - Right. Well, it’s not because of the depression bit. It’s because of the scene that you describe. I use the word scene deliberately. I mean, you know, you should introduce it yourself of but of patients or in one case you I think being inside an MRI and then being asked to recite or perform, you know, a script and I couldn’t help get have in my mind this idea of a kind of an MRI theater and then the researchers are kind of the idea is to study your the speakers brain through the medium of their voice via the MRI so there’s this kind of an incredible kind of. Performance dynamic you describe of the the best sort of conducting all the direction of the speaker inside the MRI. As a kind of in order to solicit speech that would yield insight about their brain and i just it’s just such an amazing sort of scene and scene and I just wondered if yeah if there’s a way of telling us a little bit about, you know, vocal biomarkers and through the medium of this case study that you solve. It’s just so sort of amazing.
Beth Semel (00:26:43) - Thank you. Yeah, I mean, I’ll try to do it justice, but so, you know, biomarker, right? Some people, some in startup people say that’s an incomplete metaphor. It’s a metaphor. It’s not necessarily something there isn’t something concretely biological there. And a lot of focal biomarker people, they’re not actually looking at the brain. They’re kind of doing the engineering cheat sheet thing where they’re like, “Well, it doesn’t matter. We don’t really need to see the thing that’s causing it. We can just, it can be made knowable through our techniques without us having to, you know, go in there.” But in this particular lab, you know, they really are doing, or we’re trying to do like basic science work to say, “Okay, well, what is happening at the brain level when people are producing speech and when they are under the diagnostic category of depression or not.
And so in order to do that, you can’t just have people talking, speaking, right? You have to have them speak in a particular way, right? A particularly kind of regimented way. And so there’s this really wild parallel story about these standardized vocal tasks that are used in speech therapy, but also used in more looking like Parkinson’s brain studies, ALS brain studies, where the tasks are supposedly, and they’re designed explicitly for English speakers, right? So the task wouldn’t necessarily, it can’t travel globally, right? And the tasks are designed to make you use as many of the articulators as possible to kind of maximize your articulatory action in one go so that you can get as much brain data as possible. But they’re very bizarre sounding, like Dadaist poems, I call them. Like, “Pah-tah-kah” is one of them.
There’s this one passage called like grandfather passage, that’s, you know, it sort of lights up all of the articulators you use, the full range of phonetic features in English in order to like say it out loud. But you know, the catch is that you, it’s not even enough to just say these like highly stylized poems, right? You have to say them in a particular way, like a correct way. So there’s just this kind of like narrowing of precision that’s really like a horizon point because people are not standardized. Their vocal apparatus is not standardized. Even the way that the researchers kind of wrote directions about how to say these tasks were, it was a constant source of frustration of, okay, this person isn’t saying it right. And the setup of the way this was done in the fMRI machine is there’d be like, the research subject in the and the machine in one room and then the researchers in another room and they would every now and then because the subject’s mic’d up, they would intercom them through the control room as it’s called and would say okay this guy’s not saying it loud enough or like you’re supposed to say these vowel sounds at different pitches, but when she says he’s supposed to be sitting and he’s like a manly man.
So he’s only saying it like because he doesn’t you know want to compromise his masculinity by making this like girly quote-unquote girly high-pitched sound. So again, like things like gendered expectations about vocal performance kind of like bleed into what’s supposed to be this like highly controlled setup and like constantly destabilize it, right? And so I think it’s another example too of the tension between, again, the precision that this whole thing is supposed to produce, right? Okay, once we have the data that we need, we’ll just be able to capture these signals from your voice without even touching you with our special machine learning magic, but it’s actually quite haphazard and really kind of full of weird noises and noises both in the sense of non-language sound, but also like error glitches, fuzziness, things that can’t be captured because they don’t fit quite neatly into the boxes that researchers are that they’re requiring.
James Parker (00:31:33) - Could you say a little bit more about like what the specific sort of line of critique? I mean, I know you’re just doing description at the moment, but because there’s a couple of different things going on in your writing about this. One is to draw attention to the obviously and overtly non-machinic in the production of the data set. So you talk a lot about all of the care work and often feminized care work that is sort of co-opted into these sort of scientific systems and then immediately kind of excised out in the name of kind of objectivity. So it’s sort of re on one hand it’s like re-inscribing the human and the and the careful and the feminized and so on into the system.
So that’s one kind of and that’s a care that involves a certain kind of listening always right because it’s not just the machines are listening, it’s the quite sort of highly tuned, careful listening on the behalf of the researchers. But then also, like there’s a line of critique that sort of, well, what is the status of the data that’s being produced since in order to that the idea is that you can capture an authentic depressed voice, but the performance of depression is so highly stage managed, then it’s sort of hard to understand like that it’s really a performance of depression. And then I mean, on that point, there’s this kind of amazing moment in the thesis where you describe somebody or a system whereby for ethic, on ethical grounds or financial and ethical grounds combined, the researchers routinely turn away people who are too depressed because they don’t actually have the facility and the risk is too high.
You know, they can’t manage somebody who, you know, they’re not doctors, they’re not clinicians, right? So, the subject that is producing the data is a sort of this kind of weird kind of subject that’s sort of depressed enough to have met a DSM threshold, not so depressed that they’re actually in crisis, sort of like Goldilocks depressed and who is also heavily sort of directed in their vocal performance. And so now you don’t, I think you don’t get to the point of like saying this means that the data set is a complete nonsense, but you could. And so I was just wondering if you could tell us a little bit about how you think through those different dimensions of the other kind of the critic i mean i probably missed some other dimensions of the critique that you sort of draw out like what do we make of this strange theater that is. You know you producing all of this data that sort of then goes out into the world and has all of these sort of strange and potentially harmful after lives.
Beth Semel (00:34:48) - Yeah i mean that’s a good question. I think it is important in and of itself to really emphasize that the kind of concreteness of depression as a thing is not just fabricated, but also fabricated in a somewhat arbitrary way that does any kind of claim to that the bio part of vocal biomarker should always, again, you know, I think we should be inherently skeptical that there isn’t necessarily a kind of, you know, it’s not a representational relation, it is a performative relation, right? It’s being made to be in connection to each other rather than actually like, okay, this is what it is, this is depression. Then, you know, I think that is, again, it’s really important to emphasize that that’s sort of how mental health care works in general. I mean, even outside of, you know, we can do this highly elaborate, put someone in an expensive machine, ask them to stay very still, ask them to speak in a precise way, get, you know, fancy, beautiful brain pictures and do, you know, fancy high-powered stats on them. But at the end of the day, just like, you know, the experience of moving through the mental health care system, it involves, right, a kind of reductiveness and a kind of, you know, treating the depressed person as if there is something inherently, stably, definitively pathological about them, even though that object is always kind of slippery.
And, you know, at the end of the day, too, it’s kind of, it’s slipperiness doesn’t, how much does it really matter, given the stakes of the situation? I mean, one thing that I don’t think it really got into, it was something that was sort of haunting both the field work that I did with these labs and also the dissertation and that I think something that I’ve really been wrestling with, which is just the whiteness of all of this, not just that the data sets were primarily comprised of white people. But also just even the kind of imagined user, the person who would benefit from this type of technology or benefit from this mode of intervention, or even would be, you know, there to receive it, is, I think, a kind of normatively white subject, right?
So thinking specifically about the context of the US, and maybe this is taking your question in a different direction, but, you know, who gets a kind of nice mental health care interaction who begrudgingly seeks care or who has care kind of imposed on them versus who doesn’t even have a choice as to what kind of intervention they receive, right? So thinking about not just things like non-consensual, like police intervention, like when someone is in a crisis, they have to like, you know, they call a crisis line or somebody calls, you know, a crisis unit to check in on them, but also the way that, you know, I mean, mental illness is a very, or, you know, trauma and anger, psychosis is a kind of very reasonable response to like an inherently anti-black world, right? It’s not, it makes sense in many ways. So if the, if the mental health care system is, you know, really built in a way that kind of is always not kind of naming race or, you know, acknowledging race or trying to, in the case of vocal biomarker research, really kind of trying to push race out of the picture and say, you know, this is really like the hidden asterisk that people use all the time. They’ll say, oh, vocal biomarkers, they’re language agnostic, right? It doesn’t matter.
James Parker - That is bananas, isn’t it?
Beth Semel - That is bananas, but it aligns with a very kind of, you know, liberal democratic way of like doing like, like this is a, like this is a non discriminatory form of medication, right? And it’s if we look at like something like the pulse oximeter, right? Very, very kind of banal healthcare technology, right? That a lot of kind of in an emergency medical context, a lot of medical decisions depend on the pulse oximeter. Lo and behold, it turns out the pulse oximeter is calibrated towards skin with like less melanin, So it’s calibrated to work best with lighter skin tone than darker skin tone. So it’s the kind of white supremacy, if I can be so bold, of not just the mental health care system, but the medical system in the US is just so much in the background, that I don’t know. I mean, in many ways, In many ways, it’s like Anthony Ryan Hatch, a sociologist at Wesleyan, he has this distinction that he talks about between liberal science and liberatory science. Liberal science says, okay, this is a band-aid. We need a quick fix, we need a patch, but doesn’t really disrupt anything or change radically alter the conditions. of things, but liberatory science does do that altering. And I think like, yeah, in many ways, like focal biomarker research is kind of like doing the same kind of normative work of the mental health care system. I mean, very, that was a very like, you know, runaway response to your question, but I think, you know, again, getting back to like the, is it real? Does it work? Like, again, like how does it work? What are the stakes of it working in that particular way for whom is it working right.
James Parker (00:40:59) - I mean what one thing that occurs to me is. So there’s like a double universalization going on on the one hand the voice is universal by the biological voices universal across languages across i mean we haven’t talked about disability or sort of. Unstable voices and how. Like obviously there are people trying to measure, you know, Parkinson’s in the voice or whatever. But what do you do if you have Parkinson’s and you’re and depression? And like, how do how do these things like intersect with each other? That’s like a huge problem. And then there’s the universalism of like the the idea that there’s something called depression that’s located in the body that you situate that you obviously mentioned.
And one of the things you do in the thesis, and maybe this is a good time to just sort of bring it out, is to sort of decenter in a way that the sort of big data contemporary context, something a little bit that we’ve done in some of our work with machine listening to is like a lot of these imaginaries are in place before we get like massive, you know, sort of the kind of brute force machine learning that arrives and, you know, precisely around the time that you’re talking about, like a lot of a lot of this, you know, one of the things you do in the thesis is say, well, look, computational psychiatry kind of arrives before computational psychiatry with a certain kind of empiricism of the body and the, you know, the placing of the mental health diagnosis in the body.
I mean, I know that you’ve sort of gestured it before, but it’s like quite an important move, right? Because just to be able to say, yes, I’m doing work on this flashy new thing but the flashy new thing is really largely this old thing and it’s you know white supremacy and it’s um, biologists um and sort of hubristic science or you know whatever however you want to put it I mean I feel like science and technology studies sort of really good for that kind of move um But anyway, it came out strongly from your work. I don’t know if you have any reflections on on that aspect of it.
Beth Semel (00:43:18) - Yeah, I’m the kind of like, is this just like, what is it? Old lion and new bottles or something? Yeah, I, I mean, I do think, you know, I think there is a particularity here to the fact that the object of analysis is both mental illness and the voice. And, you know, I think while on the one hand, like in my, in the work of mine that you’ve read, especially, you know, as I was writing the dissertation, like fresh out of fieldwork, like it is just kind of like a big giant extraction machine. a big kind of like, let’s continue to do this like hegemonic way of doing mental health care machine. But on the other hand, I do think there’s something interesting going on with the interactions that researchers are having with research subjects, which is, I think, a kind of scale of observation that’s really lost in a lot of the kind of top down accounts of local biomarker research. So actually looking into like, okay, what does this work involve? What are the kinds of relations that people are put into in doing this work? What kind of relations do they make with each other in kind of producing these voice data sets? There is the example you talked about the FMRI theater. There is also these, I think, really powerful instances across the three field sites I was working at. One was the neuroscience lab. Another one was in the top floor of a psychiatric hospital. They were trying to find vocal biomarkers of bipolar disorder.
Another one was looking at PTSD and depression also in a kind of, you know, away from research subjects altogether, away from like any kind of hospital. But you know, research subjects, right, so they have to, in order to like make the data happen, they have to talk, right? So researchers have to figure out ways to like cajole them into speaking. And sometimes in those interactions and in talking with each other, you know, people are having these really kind of cathartic moments, like really actually transformative moments in their interaction, interactions with each other, where, you know, it’s not, it’s not like capital H healing or like capital C care that would happen in like an actual, like official mental health care context, but there’s something, the person leaves that interaction, that encounter, transformed in some way. That doesn’t, I think, fit kind of neatly into again like a hegemonic model of like what cure looks or sounds like, right? And a lot of the times I would hear research subjects say like, it actually just feels good to be listened to, which, you know, might not be much and it’s, you know, maybe speaks to like how like sad it is that people don’t feel like they have someone available in their lives to do, to just do that kind of being with. Um, them, right? But I don’t think that that can be discounted as something, you know, something is happening there. And yeah, it’s sort of like, like, there are these moments of access that are produced within these, the, you know, big extraction, like vocal biomarkers, sausage making machine that I’m currently like very, very interested in.
James Parker (00:47:21) - But there’s a real irony there, isn’t there? Because It sounds like the examples that you’re giving are the examples where the subject has been listened to by a person. Whereas the way that they’re being listened to is precisely in order to, for reasons of efficiency and insurance companies and la la la to prevent them from having that kind of listening as care, you know, in the future or people like them. That comes through really strongly in the chapter on bipolar. And because that the sort of the end game there is kind of, sort of, in contrast to the MRI one, is much more about us in explicitly about producing a surveillance architecture. So like, it’s about your cell phone, your cell phone will be listening to you constantly for you know the possibility of a flare up or whatever I don’t know the language in your bipolar and that will trigger some kind of system which will get you the care you need but in other words like wouldn’t it be amazing to have to be listened to non machinically almost all of the time as a kind of a substitute for occasional human listening.
And then in that chapter, you also talk about this idea of listening like a computer and the way in which the, um, the annotation and marking up of the data set that you have, um, or, um, requires inattention, a certain kind of inattention, what people mean when they say listening as a computer is sort of not really listening. And so you’ve got this kind of weird, um, dynamic where. Your saying there’s something real and therapeutic in a certain kind of way coming out of this process but the whole aim of this process is to leverage. A non listening as a form of listening or an inattention and inattentive listening at scale within a surveillance architecture that by the way is heavily corporatized via like API is. from Apple and Google and yada yada yada.
I found that chapter extremely rich. I’d love to talk about the annotation process. It sounds like some of the examples you’re giving are come from this, you know, listening to these people, the people providing the data and you’ve been you’ve been told not to not to listen to their stories. But you talk a lot about how well you you can’t help but listen to the stories. And there’s a kind of an irony there. And could you could you tell some of that story?
Beth Semel (00:50:19) - Yeah. Yeah, it’s you know, that was a I will say I don’t I think it’s in the chapter a little bit and also kind of vaguely in the in the research article that came from that that chapter but um that was a very hard field work for me to do it was pretty I mean ironically I became very depressed while doing that work because you know most of what I did during the day was listen to these you know the the team had been doing this longitudinal study of bipolar disorder and as part of they kind of like hooked onto that study, this voice data gathering stuff. So research subjects would agree for a six to 12 month period to have this souped up phone that the study provided, like a nice smartphone, which many of the subjects, otherwise, they didn’t have a smartphone, or they just had one person in their family had a smartphone that they shared. So they got their own smartphone. But the smartphone had this app that was recording all of their phone conversations, including this conversation they would have once a week with a social worker on this research team. And so I was floating back and forth between sitting right behind the social worker while they’re doing these phone calls and not listening to what the person on the call was saying, just listening to how the social worker kind of coax these answers out of them. And then I would literally walk down the hall and go sit down and begin annotating voice data that had been gathered before I got there of these same phone calls. So listening to these chopped up segments of the calls, but listening but also not listening, as you say, right? So trying to listen and assign a label to the sound of the person’s voice, but try to not have that label correspond at all with the content of what they were saying.
James Parker (00:52:25) - Could you say a little bit more about that annotation process and method? Because I mean, I’m just, there’s that phrase from Foucault, like the micro physics of power and you know, like, it just seems like that labeling process, like, what are the labels? Where do they come from? What’s the process of you sitting there trying to assign a label? What happens to the labeled data afterwards? It just seems like, there’s like a whole world of political decision making going into this practice of labeling of voice audio. I’ve never spoken to anybody before. I know that lots of Amazon Turk workers and so on have to do this, but I’ve never spoken to anybody before who’s sort of gone through that process. So I’m especially not an ethnographer, especially not one who works on STS. critical voice studies and so on. So I just would love to know like the incredibly finely grained detail of that process and what you were able to learn about it and understand about it because it just seems fascinating to me.
Beth Semel (00:53:38) - It’s so funny that you asked that because as part of my field work since like as the anthropologist you are you know pressured into doing all the stuff that no one wants to do So what I did was they were like, okay, we need to make a training video to train other people to do this annotation task. We’re all going to collaboratively write a script, but Beth, we’re going to have you read the script and be the narrative voice of the video. So there’s this video of me actually walking through step by step, as if I was telling another annotator how to do this. It made so much sense and was very clear to me then. In listening to it now, I was like, this is so bizarre and I can’t believe that this is like at the time I was like okay yeah I see how this makes sense within the kind of what what the study is trying to produce it makes sense that they would you know chop it up in this way that there would be these so there’s these like strict not strict but fabricated criteria of when to exclude a segment from annotation and to say that okay this is fundamentally and unannotatable you couldn’t understand the person.
That would be when in the span of the segment, so the segments would be five to ten seconds long, the person was just saying “um” or “yes” and dead air or “no”. Or if they were laughing, or if they were coughing or sneezing, or if they were talking to someone else, or if they were talking on speakerphone and that kind of distorted the audio quality, or as was the case with one research subject, they had many pet birds and the the birds made it hard to hear the person’s voice, right? Or if they had their child sitting in their lap and you could hear the child’s voice on the audio. So you would have to say this is not, this has sensitive information. We have to remove it from the corpus. So, you know, already there, that’s, you know, these again, like the MRI theater, there is this kind of in notation theater of, okay, like let’s, you know, Let’s dig into this thing. Let’s all agree upon these rules, which otherwise seem, you know, quite arbitrary. And the rules about laughing, the rules about crying are ones that we kind of made up, but what would be the tipping point of them being unannotatable and not?
But, you know, I have to, you know, be a bit of a downer and say that the I explained the kind of setup of the experiment in a great detail in the article on science, technology, and human values. I feel like to really explain it all would be very, very boring. Like sort of why they’re doing it the way that they’re doing it has to do with these theories about the relationship between mood and emotion and kind of studies of bipolar disorder and why they chose the, you know, the scales for the annotating tasks like activation and valence, right? So you’re supposed to rate the segment based on how the segment feels. Does it feel really negative? Does it feel really positive? Or is the speech really, really energized? Or is it lacking in energy? Is it dull? Those come from the dimensional model of emotion. So there are all these theories jam-packed into this one task. And despite that apparatus, it was a very— Again, emotionally intense work. And also very, you know, when I first, I remember when I first started doing it, I was just stressed out. Like, how am I going to do this correctly? Are these other annotators who are literally undergraduate students, how do they seem to have more confidence than I do in doing this work of just kind of ignoring the content and just only like kind of trying to feel the feelings of this voice, right? And, you know, I got better at it as I, the longer that I did it, because we all just kind of developed this like tacit knowledge of like what, you know, for instance, like on the scale of one to nine, there’s always the five, like five activation, five valence, which is like the neutral speech. So we would just kind of develop this like collective sense of how neutral speech feels that I don’t know I could replicate or like I couldn’t tell you because we would, you know, we’d be listening and that the annotators would, we all annotated the same subjects like multiple times over. And, you know, we were supposed to annotate on a, like person to person basis.
So instead of like invoking some kind of like generalized notion of what neutral speech sounds like, we have to figure out, okay, what is neutral for this one, you know, person who has a lot of pet birds? What does their neutral speech sound like? How does their five-five speech feel? And, you know, that involves, like, listening to their segments over and over and over again before annotating them, really kind of like getting to know how this person sounds in which, yeah, you just so happen to without kind of meaning to, like, learn their life story, right? This is where they work. This is, you know, they’re having a fight with their sister and this is why, you know what I mean? It’s a year’s worth of phone calls with a social worker, right? But you know, over time we really started to develop these like very kind of, you know, assured feelings of like, yeah, this person has a lot of neutral segments. This person has a lot of segments quote unquote good for depression, right? So they have a lot of segments that are low activation low valence meaning like negatively charged speech kind of very like very like this like I don’t know I just don’t know how I feel that would be that’s like okay that person has lots of depressive segments great we’re really excited about those very clear data points right um but then you know I don’t think if I don’t think I could go back and do the same work that I did then right and
I think that’s because I have these strong feelings about what this whole apparatus is supposed to be producing but also too because it was something that was collaboratively made with these specific engineering, junior and senior computer science, right? I’m working alongside the grown-up squeeze into a little desk with them.
James Parker - That is absolutely wild. I’ve got so many things I’d like to ask and sort of unpack a little bit, but just I hadn’t come across the phrase hetero-amation before as a kind of antidote to or alternative to automation. And it just seems like that’s the perfect example of, you know, just as far as I understand it, the term hetero-amation is trying to get at the sort of the non-automated in the production of automated systems and like everything you’ve described there is like so profoundly social. So like, like, like, just like, the kind of negotiations that you’re doing with your colleagues over coffee on how to understand a thing like it’s just that’s where so much of this is really happening. And, you know, in the designation of the people as undergrads who have a particular economic and sort of attitudinal like context they bring to anyway, it’s just amazing.
Beth Semel - But they would always say Like the PIs would always say, Beth, we’re so glad that you’re here because we bet you’re really good at this because you’re an anthropologist. We’re engineers. What do we know about other people’s feelings? But you talk to people for your job. So you’re probably really good at noticing how they’re feeling and stuff like that. So you’re actually— hence, they’re like, you should be in charge here. You should— we want to know what you think, which I oftentimes, you know, very honestly, would be like, this is like, what? Are you guys serious?
James Parker - Right, because we don’t know anything about this thing that we are doing and rolling out, you know, potentially at scale.
Beth Semel (01:02:15) - I mean, that site, just one more thing, that site in particular is really interesting because, you know, I really don’t want to like reveal the identities of these people, but someone who is working very closely in that context, you know, kind of said to me, like, look, I’m doing this because I don’t think it can be done. I don’t necessarily want to challenge my PI and say that it can’t be done because of my own life experiences as someone, you know, who grew up, this person grew up outside of the US.
They had come to the US fairly recently, you know, they didn’t grow up speaking English. And they said, I can’t tell when I oftentimes think my supervisor is mad at me. I know that they’re, I later learned that they’re not, but I always think that they’re mad at me. So what does that say about, you know, this kind of like innate ability? So yeah, they said like, yeah, I’ll do it kind of to show that it can’t be done in a somewhat passive aggressive way that protected their position. So the idea that it can be done is sort of rooted on some level in this idea of the dimensional model of emotion.
James Parker - I’m sure that that’s got an incredible, you could just do like, write a whole book probably about that history. But could you give two lines like what, what, where does this come from? It sounds like that’s something that comes maybe out of a certain branch of psychology that’s not necessarily computational. Like, is it people accept that? Is it some kind of crackpot fringe theory that just happens to have been taken up by this particular, like, what is I mean, could you just say a couple of lines about that because it’s doing so much work in this study.
Beth Semel - Yeah, I mean, oddly enough, my understanding is that in the kind of psych world, it actually is considered a more capacious model than some earlier models of emotion. So, like, rather than like being there being like a set kind of list of emotions that people could possibly experience, like anger or sadness or fear, that, you know, like big subcategories that smaller categories of emotions fit into. The dimensional model of motion says like, okay, let’s make this four quadrant graph and plot valence, like high valence on one axis, axis and then low valence on the other, like on the, let’s say the vertical, vertical axis and then on the horizontal we have, you know, low activation, high activation and so you can rather than kind of add these labels, which are, you know, so non universal, what if we instead just figure out how to we could, what if we just plot these emotions in this kind of much more capacious four dimensional space? Got it. And so is, I think, seen as something that is more capacious and allows for more granularity and more kind of difference and it’s got numbers!
James Parker - You know numbers are very.. But I mean… I’m serious because we were looking at for example this Toronto emotional speech set tests recently where they they have all of these They have actors perform the phrase. Yeah, say the word sheep. Say the word fish. But like and then they got these seven sort of core emotions, obviously like not only just like call for who like whatever but also they’re literally being performed by actors and they’re like but they’ve got musical training it says in the in the paper like oh okay well then that’s fine then so that’s a totally different model so that’s that’s a data set that’s being used to train systems based on yeah like a sort of linguistically grounded or something idea of emotion or a sort of socio cultural idea of emotion, whereas this is, yeah, it’s sort of, it sounds like, I mean, I’m sure it doesn’t achieve anything particularly different, but it’s a it’s trying to somehow circumvent that problem of like, what do you mean by fear was like, Oh, it’s not fear. It’s like a four and a two in the like, in the kind of make spatial mapping of like, you know, la la la. So it’s sort of, yeah.
Beth Semel - I think it’s another way to kind of coround things into this very kind of like biological, like thermodynamic model, right. So activation and valence are like, you know, energetic states, right, or like charges, like valence is like this very kind of, you know, what valence is like, I don’t know, like electrical engineering or something Like it’s some kind of, it’s like more like sciency and like, oh, okay, yeah, no, we’re not talking about feelings. We’re talking about energies. We’re talking about like, you know, physical things, right, that can be measured with math. So I think on the one hand, yeah, that is doing this more, this work of like, okay, we’re outside of the kind of hand of language and culture, bringing it back down to the body. But again, you see that same kind of gesture to the body as being someplace that’s, like outside of history, like outside of white supremacy, right? Outside of all these other, you know, kind of differences, somehow it’s like a protected space.
James Parker (01:07:43) - Can I ask, like for example, you’ve got all these university undergrads doing this annotation, like is there any attention to the like biographies of those listeners? So on one level, this person you’re mentioning, like wants to sort of prove it wrong. But if you get enough for the sake of argument, 20 year old white boys who have a particular socio economic background, probably are going to get a fairly kind of normative listening kind of emerge that will look sort of stable. So it’s obviously like the kind of the histories of the annotators are like crucial. So it’s like, is there any attention to that at all? Or are we just, or is this just an example where you’re going to reproduce? If you produce anything, you’ll reproduce and re-inscribe and then automate and roll out a kind of whatever, like whatever the demographic of the listeners is that that are enrolled to do the annotation.
Beth Semel - Yeah, I mean, I think in many ways that’s what a lot of them said. Where they were like i guess we’re kinda just making this like what one of them said like an american culture machine like this is for making american culture machine that’s amazing. Yeah i mean those are you know the words like an engineer not mine i just happen to be in the room you know. Yeah after you leave we would have loads of conversation about this thing and it was why was so weird and. But, you know, I think that the point that you’re making is one that I believe really bears repeating, kind of in response to this real kind of, I don’t know, fortification of vocal biomarker stuff, which is that actually, and this is something that Nina Sun-Eidsheim talks about in “The Race of Sound,” right? That actually it’s not that, like the vocal biomarker stuff up says the voice in and of itself can be meaningful. Whereas, Aidshem is saying like, no, that voices are made meaningful through techniques of listening, right? And I, I actually think that that like consideration, like you’re saying of the kind of who are these annotators and biographies of them.
James Parker (01:10:14) - I think there is some importance in saying, hey, actually, look, that is doing something here that is not just, you know, identifying, calling out features that are inherent to this speech signal. Because that is what the vocal biomarker is promising, right? Precisely by kind of papering over or downplaying or, I don’t know, misplacing the role of, like, you know, actual listeners and not, you know, like mathematical modeling waveforms, which is what machine listening is in this sense, quote unquote machine listen, listening, right? The listening of the machine is like fundamentally different from the actual, like, again, human listeners who are trying to mine this idea of how the detached way that, you know, listening is supposed to take place. But again, it’s sort of like, I don’t know, I don’t know the best way to say it, but like the this thing it produces its own critiques, right? So we see even from within the context of making this stuff that machine listening as an object is like one that depends on human listening and that even the those categories are not kind of neutral things that like have any kind of like inherent thing to them that make them human or machine like, right? In the same way that like as you were pointing out, about how like, well, isn’t it ironic that the way that people actually feel something kind of reparative or transformative in these encounters is through, you know, like a person listening to them talk, like even though the point, the reason why they’re there is to basically, you know, discourage people from listening to other people, like yes, precisely, like the critique actually, it’s like, you know, like the blueprints of its own undoing is actually like, like That’s how to build this.
Beth Semel - Exactly. Yeah. Exactly.
James Parker - So, you know, that’s the part that, you know, maybe it’s like a, to make that critique is maybe like a little bit annoying to people who are making these things and, you know, perhaps too much like you played yourself. But it’s all tantalizingly there, right? I’m conscious that we’ve had a lot of your time already. I’m wondering if it’s worth talking about like, yeah, where you end up sort of with this research, you know, I don’t mean to say, can you give us a list of five empirical kind of conclusions or regulatory, you know, reforms or whatever, but sort of where you end up or maybe it’s easier to talk about what you’re doing next and what seems like sort of fertile territory to be thinking with now? Yeah, I’m just sort of wondering where you exit or where you’ve begun to move towards after doing this research.
Beth Semel - Yeah, I mean, you know, I have a second project that is very exciting to me, but I am trying to, you know, it’s on the shelf for now as I focus on really, you know, turning the dissertation into a manuscript. that I could talk about, but yeah, I’m trying really hard to, I guess just to give a preview, in kind of doing this vocal biomarker research, I found this really less than savory branch in the history of this field that involves voice-based lie detection, voice stress analysis, it’s really treated oftentimes as occupying, again, a very separate branch in the family history of this type of technology. But the second project is looking about how they’re actually quite intertwined. Looking at the history of this one particular technology. The field of voice stress analysis, likewise, like vocal biomarker research, So many people have talked about how spurious it is and yet it keeps churning on. And so I think this question of like, well, if this stuff is just the same old kind of extractive, harmful apparatus, then why is it still going? I think in the vocal biomarker case, it has a lot to do with the kind of liberal versus liberational science, right? People are genuinely trying to help. I don’t want to discount that people really, you know, want to do something to help people. It’s just, you know, the grammar of that care is one that, you know, really rhymes with control, right? And I think with the, this kind of other what I’m calling like a carceral prehistory of algorithmic clinical listening, I think is kind of the the double whammy there to say, you know, this is there’s something about this illusion of care and control that begs, I think, closer attention and greater scrutiny to the kind of field of the mental health care system in the US and it’s kind of close relationship with Carcerality and policing really which is yeah, so I mean that’s you know, I Don’t want to say that’s where the first project Leaves off.
I think with the the first project, you know what I am hoping to really you know get people who are working in these systems to think critically about them. And, you know, to perhaps question whether or not this is like the best use of funding, right? That’s the horizon. I mean, but also, too, you know, I think that a lot of the again, like with the whole band aid approach that the reason why it’s just a band aid is because there are fundamental issues with the way that mental health care is done and like arranged in this country. And so I think in many ways the fact that Amazon is like, there’s some hope that Amazon is promising that it can be your therapist or your nurse. Yes, Amazon is gross and evil, but also what can people do? People who are experiencing mental illness, people who also don’t want to go to the doctor and don’t trust um, you know, the mental health care system, like again, for good reasons, right? So, um, I think, you know, asking people to sit with that tension and to say like it comes in part not just from the kind of technical apparatus, but also from the, you know, the broader system that it’s being built to like be hooked into is, um, you know, what I’m hoping is the kind of takeaway. I mean, we, I think that The crucial difference too with the vocal biomarker stuff, which you know, it’s like maybe a whole other conversation is that it’s really not, the technologies are really not diagnostic technologies. They’re always being put out as screening technologies.
James Parker (01:18:06) - Same with COVID stuff.
Beth Semel (01:18:09) - Right. In the US, it’s because if you are building a clinical decision support tool, actually you, you can bypass certain federal regulations. And so I think thinking about the paramedical and the parabiological, this space that’s at the margins of the mental health care system, this place where people are interfacing with the mental health care system and how that encounter hooks them into all other kinds of systems in the US, like child protective services, the welfare state, the carceral system.
James Parker (01:18:43) - Insurance.
Beth Semel - Insurance, yeah, exactly. I think those are the legal, criminal legal system, like, you know, the fact that you, you know, you can, these things can be used against you in the court of law like years down the line, right? So they’re in many ways like evidence producing technology. So I think kind of asking people to see how, to look beyond again, the actual technology itself, itself and see all the kinds of systems that the technology kind of draws together and will draw people through is I think the kind of at least for now the kind of you know bigger takeaway of the of the book. I mean for what it’s worth um I thought that stuff came through really strongly in the finely grained sort of description like Giertzian sort of thick description in the thesis like it’s just impossible to read your account of these things and not and feel like the main event here is computation. I mean, I’m not trying to say that there’s nothing going on, but like, yeah, it’s just like the way it’s described just draws the reader draws the reader into the embroilment of these systems with everything that’s ahead of the game on this stuff.
James Parker - A lot of the people we’ve been speaking to are sort of just finished a PhD or kind of, you know, in that sort of stage. And then there’s a lot of people sort of turning their mind, you know, more established figures turning their minds to this as well. But it seems like there’s always more people out there that we haven’t spoken to or don’t know about. And so if you’ve got any ideas who people should read or listen to, that would be great.
Beth Semel - Yeah, I mean, I, so Edward Kang is a PhD student at USC in the Annenberg School of Communication. And he has a really great article about voice and bodies and race and specifically like contemporary voice print technologies, where he’s like critically reading through patents. I mean, you know, Nina Sun Eidsheim is a sound studies person, but I think a lot of the stuff that she has to say about the racialization of timber, I think, you know, it’s not necessarily, well, I think there are some dimensions in her book that, where there is a kind of machine object in question, right? But, you know, I mean, really the people that I like to read and kind of bring into this conversation are like, increasingly like, sound studies people, like Dylan Robinson who does work on indigenous sound studies, thinking about alternative paradigms of relationality that happen through listening, trying to think of others. Yeah, so really, again, the computational is kind of like a receding object in the way that I think. But and like who I’m kind of wanting to be in conversation with. You might have already talked to Xiaochang Li. Yeah. But she’s also, yeah, I mean, she’s, her work is like super helpful. Likewise, Mara knows, I’m sure you’ve also talked to, right? Those are like my, you know, my people and linguistic anthropo, I mean, there’s linguistic anthropologists too. I don’t know if that’s like the jam of this interview series, but
James Parker - Well, it’s not not the jam. And you know, that’s already so helpful. And I kind of put you on the spot as well. Look, that was an absolute pleasure talking with you. Thanks so much.
Beth Semel - Yeah, thank you. This is it’s been it’s been great. And I appreciate, you know, you having done such a close reading of this, of this material that I, you know, I need to go back and revisit now. So it’s been a good it’s been a good call/inspiration to revisit. So thank you.
Guillaume Heuguet transcript
James Parker (00:00:00) - Well, thanks for joining us Guillaume. Perhaps you could start by introducing yourself. You know, however feels right to you.
Guillaume Heuguet (00:00:10) - Yeah, thank you for having me. So, at the moment I’m teaching art school in Clermont-Ferrand. I have a PhD in Media Studies from La Sorbonne. And I’m running a publishing house named Audimaté-Elycian. And I have been doing this journal of critical essays about musical experience and different genre and scenes for like 10 years. And I just recently started another journal about critical thinking on technology. And that’s about it.
James Parker (00:00:46) - So, a PhD, several journals and that’s it. That’s a pretty comprehensive package. What exactly is the journal? Can you say a little bit more about them? Guillaume Heuguet (00:01:10) - Yeah, so Audimaté, the main one, the one I started 10 years ago, I wanted to feel what I felt was a lack in the French landscape of writing about music. I was reading a lot of essays coming from the UK or America, which I felt were reinforcing the subjective experience of music. And mixing it with theory or analysis. And in France, I was feeling like you had to choose between a very academic approach or journalistic approach that was a bit more superficial. And essay criticism was not so much of a thing when it came to popular music.
Guillaume Heuguet (00:02:06) - We have a long tradition of that with Cahiers du Célème and movies, of course. But I felt that I was missing something. So, we did a lot of translations of English authors like Simon Reynolds, Kodwo Eshun,… in sound studies as well. And we tried to bring people who wouldn’t write in this fashion to a kind of writing style that’s more critical essays. So, people who will be in journalism or in academia will take the opportunity of writing for Audimaté to share their passion and their love and their understanding of a very niche scene. Guillaume Heuguet (00:03:01) - Or interesting listening experience, their relationship to nostalgia in music. There are a lot of different topics that we cover. But the idea is really to try to show the diversity of writing that is possible about music. Because… are writing styles. And we want to bring all that diversity of relationships to music. So, that’s Audimat. And the technology journal is very much younger. And we are still trying to find our own style, I guess. But so far what we want to bring is probably a more materialistic approach to technology.
Guillaume Heuguet (00:03:58) - Both on the technical level and also on the economic and political level. Because we feel like in France what dominates the discourse about technology and the criticism of technology is more coming from philosophy and from civilizational perspective. From Heidegger and stuff like that. So, it’s very much like apocalyptic idea of how technology comes from society. And I share the political concerns of those people. But I want to anchor the criticism in a deeper practical knowledge of the apparatus that we are discussing. And so, we are trying to…
Guillaume Heuguet (00:04:49) - We just launched this journal as a platform to gather people who want to focus on that.
James Parker (00:04:58) - They both sound amazing. It’s funny the people you mentioned with Audimats, that sort of all my favorite music writers. I mean many peoples, but it does seem like there was a kind of a sort of English language heyday of work that came out of that melody maker period in the 1990s. It’s been incredibly influential and a lot of people I know kind of aspire to.
Guillaume Heuguet (00:05:28) - Yeah one thing that I did not mention is that I think in France too we had like in the early 2000s we had a very big scene of music blogs and music messaging boards and the conversation about music was very lively and actually when the platform moment came out all of this a bit dried up a little bit so the journal was a way to keep up the conversation and not to lose those styles of writing.
James Parker (00:06:03) - No I think that’s absolutely right because you know we’ve seen as you know the blogger sphere has kind of faded away so many of the interesting forums for writing online have disappeared with it like some of my favorite websites you know are defunct now or sort of only barely live.
James Parker (00:06:25) - But anyway this is all fascinating but I want to get to your amazing book and I think we’ve also got Anabelle Lacroix here who’s a nearby neighbor of yours and she’s a curator and thinker and writer in her own right but is also going to help out if we encounter any translational difficulties. Would you like to introduce yourself Annabel?
Anabelle Lacroix (00:06:56) - Hi everyone, I’m a curator and researcher. I used to work with you both at Liquid Architecture for many years. It’s nice to join the conversation and I’m looking forward to it.
James Parker (00:07:18) - Thanks so much Anabelle. Now Guillaume your book is brand new right? Like it was out last year and it’s called YouTube et les metamaphores de la musique in my best French accent. Is that right?
Guillaume Heuguet (00:07:38) - like yeah the the French title is YouTube and the Metamorphosis of Music kind of literary style and the English title will be way more catchy will be How Music Changed YouTube.
James Parker (00:08:01) - How Music Changed YouTube okay well that’s a really provocative turn of phrase because obviously everybody talks as if it’s the other way around right? That it’s YouTube that changed music and it’s platforms that changed music and you know what would be the you know Napster and iTunes and so on. So I mean I take that as a kind of that must that’s a sort of a central provocation of the book presumably to to switch that order around. I mean I should say I haven’t read the book because my French isn’t good enough.
James Parker (00:08:38) - So could you tell us a little bit about the book about that sort of that flip and anything else like what what kind of what’s the book about what does it cover?
Guillaume Heuguet (00:08:49) - Yeah so the book is kind of a multi-dimensional platform biography of YouTube through the lens of music.
Guillaume Heuguet (00:08:59) - So I cover a lot of dimensions and topic from the early style of music videos and sound based videos on the platform up until the infrastructure for the counting of views and the tracking of views and the control of copyright to content ID. But I also analyze the way YouTube address musicians on the platform to tell them how to promote their music and I’m really interested in the whole history of YouTube and also the I’ll say the every corner of their strategy from the very technical technological aspects of
Guillaume Heuguet (00:10:03) - To the discourse of changing our music or videos or culture should be made. So it’s very multidimensional. And yeah, it’s a book I wanted to do because I was in the music journalism, I don’t want to say industry, I’ll say culture for a while. And at some point everybody was starting their reviews with a line saying, know that YouTube or technology or platform change everything. And then they didn’t feel like they couldn’t review the music the same way.
Guillaume Heuguet (00:10:48) - And it was very, very strange to me how we were ourselves as journalists or music, not fans, but music listeners. We were thinking that technology should have this bigger fun impact on our own experiences. So I wanted to study that and have it not just as this kind of vague idea that everybody was repeating all the time, but I wanted to really tackle that and take one platform and see. So how are we able to say that something like YouTube has an effect on music?
Guillaume Heuguet (00:11:34) - And of course, it’s really not easy to draw causes and consequences when we study one platform or stuff like that. But a big part of my work is to try and unpack all these layers of intervention of YouTube and see to what extent YouTube is really the main force of the changes that we see in music culture. So that’s kind of the broad picture of the book.
Anabelle Lacroix (00:12:19) - Do you listen to music on YouTube?
James Parker (00:12:27) - You do?
Guillaume Heuguet (00:12:28) - Yeah, I don’t.
Anabelle Lacroix (00:12:29) – Me neither, that’s why I was a bit nervous about being part of the conversation.
Guillaume Heuguet (00:12:35) - So when I started my PhD, I wasn’t. So I knew that a lot of friends around me were, but I was going to check one song from time to time. But it has never been my main listening device or tool.
James Parker (00:12:54) - So Guillaume, what made you choose YouTube as the platform to focus on amongst the many different platforms through which people access music?
Guillaume Heuguet (00:13:09) - Well, I have been doing my master thesis on boiler boom. So I was interested in the interaction between image or video online and music, which was one thing. And also, I was noticing that most of the industry reports in France, at least, were interesting in the streaming industry and the streaming music market. We weren’t including YouTube at the time because it was supposed to be a video platform.
Guillaume Heuguet (00:13:49) - And as a researcher, I was very interested in this way that the academic research, industry research defines the border of a market, of a given market, and how technology makes the borders of markets and culture evolve with time. So I wanted to, I chose YouTube to be able to also have a part of a critical approach to the shaping of economical research. And then I knew that with YouTube being that massive in the music space at the time, I would have a lot of things to unpack.
Guillaume Heuguet (00:14:43) - And this was very interesting as well to me to be able to talk about music videos, to be able to talk about audience measurements, to be able to talk about all those things.
James Parker (00:14:56) - I mean, what is the sort of rough market share that YouTube has in streaming? Because I think it’s really big. Like, I mean, whether or not, I mean, I did actually subscribe to YouTube music for a while because I once upon a time had a Google Play music account because it allowed you to upload your own MP3s. And they eventually tried to force me onto YouTube music. And for various reasons, I moved off it. But I remember around that time reading that YouTube is a very significant player in the market, like not as big as Spotify.
James Parker (00:15:30) - But from what I understand, it is pretty big in its own right. And then one of the interesting things about YouTube as a music streaming platform is that it’s not just that. And so, you know, I mean, this is something that we’ll talk about in due course, but from the perspective of machine listening, it’s a really interesting platform because the kinds of automated listening practice that are rolled out on YouTube move between music and other forms of listening. And so it’s sort of interesting because of that breadth.
James Parker (00:16:06) - But I do think it is a large player, isn’t it?
Guillaume Heuguet (00:16:09) - Yes. So the last numbers that we have is that YouTube, not YouTube music specifically, but YouTube is like 55% of consumers watching music videos. Or like want to consume music go on YouTube, whereas it’s 24% on Spotify, for example. So larger than Spotify. So it’s larger than Spotify as a place to go to consume music. Like the market share of Spotify might be bigger with a subscription, but as a destination, YouTube is still bigger. And in France, it has been the leader as a music destination as well for 10 years now.
James Parker (00:17:02) - So it’s a big deal in other words. And, you know, you beautifully sketched out some of the ways in which your, the things that interest you about YouTube. And I have some ideas about what interests me from the perspective of machine listening as I’ve been getting at but where do you end up? I mean, is there a sort of, you know, a conclusion or a destination that you arrive at or is it, or are there a few that you’d like to talk through maybe a couple of examples, or one example where maybe something unexpected sort of comes out of the work?
Guillaume Heuguet (00:17:45) - Yeah, one thing which is interesting to me is that when people want to talk about the book and about the streaming economy and its impact on music with me. It’s mostly because they feel like because I have a critical approach, I’ll be able to tell how much YouTube is responsible for this or that. And one of the main result of my research is that actually the music industry was very complicit of the growth of YouTube from the very beginning, because they were licensed from the very beginning actually but we didn’t know it at the time.
Guillaume Heuguet (00:18:36) - But it took a bit of time for this thing to be public. And also, I discovered that the disposition between the tech industry and the music industry from a deeper historical viewpoint is kind of irrelevant, because the music industry has always been a tech industry as well. Edison and all the earlier labels were like, yeah, RCA, Sony, all were tech products sellers.
Guillaume Heuguet (00:19:20) - So I tend to focus on this, maybe a larger approach where what’s interesting to me is how music is made into a commodity and it’s not so much about is it tech against music industry or the music industry against tech, but it’s more about our basic things of our musical culture, like the commodity form is being enforced again and again.
Guillaume Heuguet (00:20:00) - And so there is this layer of the shape of the two industries. This is one aspect of my research. The other aspect is that I was interested in the fact that, to me, YouTube was trying all the time to give shape to the listening experience, to a kind of music market. And the people on the platform always redefine the border of what listening is, of which music should be made into a commodity or not.
Guillaume Heuguet (00:20:42) - And I was quite interested in the aesthetical dimension of this research and the idea that when you publish videos that are supposed to be consumed as music, actually you realize that the idea of what music is or should be is not so easy to define.
James Parker (00:21:11) - Could you give an example?
Guillaume Heuguet (00:21:13) - Yeah, for example, I was… So that’s not in the book for now. Maybe it will be in the English version, but it’s in the PhD. I spent a lot of time studying ASMR. And do you copyright ASMR? Is ASMR a musical experience or not? You know, stuff like that. But also, is or end? Also,
James Parker (00:21:57) - the foreground-background distinction, I suppose. One of the things I was thinking of before is that YouTube is a platform in which music is sometimes the foreground and sometimes the background in a way that just simply isn’t on other music streaming platforms. And so, yeah, that’s sort of another really interesting layer, I suppose.
Guillaume Heuguet (00:22:20) - Yeah, a great example of that is channels like Cecile Roberts. I don’t know if you know those channels. Like Cecile Roberts will publish Africa from Toto. The song Africa from Toto, but recorded with the effect of an American mole with a lot of reverb. So, it’s the song by Toto, but it’s also a big splash of reverb with Toto playing in the back. And this kind of edit doesn’t count as a remix or as a new work of music. But a lot of our Niko remix can be just a song sped up.
Guillaume Heuguet (00:23:07) - And all those edits, to me, are really interesting because it’s kind of a new vernacular of sound, maybe music. And it’s not playing to the rules of what singularity of originality in music is or should be. And I find that there is a continuity with the world of music online. Just before YouTube appeared, we had a lot of different remix edits and unofficial versions of songs. And we never knew if we had the right version of a song.
James Parker (00:23:51) - You mean like on Napster and LimeWire and things where you would download something that said it was, I don’t know, The Eagles, and then it would just be some prank?
Guillaume Heuguet (00:24:02) - Yeah, and it was not so much about the idea that the right song was not so much relevant all the time. Music was not in that state of a thing or a commodity. And that’s really something that’s interesting to me, how the form of the work of art applied to music or the commodity form is not something like the industry has to fight all the time to get back music to this really specific shape of a commodity or form.
James Parker (00:24:40) - Guillaume, do you see the continuity going back further to sampling and plundaphonics and music concrete and the other avant-garde approaches?
Guillaume Heuguet (00:24:52) - Yeah, of course. And when I’m talking about this, of course I will like…
Guillaume Heuguet (00:25:00) - I will find support in like going back to John Cage or to any, or to Fluxus or to Max Neuhaus and every avant-garde artist who pushed the limit of the sound or the music experience. But I find it really interesting that it’s kind of a sonic popular culture. And I want to insist on the popular, like it’s not aimed to deconstruct the limits of the roles of the academic or the institutional art world. And it’s not to push the definition of something that has been constrained by the medium of the museum or the medium of visual arts.
Guillaume Heuguet (00:25:46) - But it’s like I think there is a lot more spontaneity to it as well as reflexivity. But there is not a big agenda to define it as a nice technical gesture. And I think to me… aesthetics are not so well defined. And the difference with what the avant-garde in the visual arts or sonic arts makes of this is that there is no need for negativity or for the gesture of pushing the boundaries or something that’s very robust and rigid because it’s way more fluid than that.
James Parker (00:26:41) - It’s so funny that you say that because I did some writing on vaporwave, whatever, 10 years ago when vaporwave was a thing. It’s still a thing, but whatever. Anyway, my point is that so much of the writing that I was drawing on for thinking about vaporwave was from art history, basically, critical art history. Because so many of the moves that those artists were making were really recognizable from the kind of conceptual art and gallery arts of the 20th century. But nobody within vaporwave was interested in that at all, basically.
James Parker (00:27:23) - And so I think that’s a really, I mean that might just be because my writing is crap, but I do think that there’s something about the popular, the fact that it wasn’t an intervention within contemporary art, was a really important feature of vaporwave as a movement or a practice that was doing something that was very recognizable within the language of conceptual or contemporary art. I think that really chimes with me.
James Parker (00:27:53) - Now, I kind of want to jump cut now to Google Content ID, if you don’t mind, because I think, I mean the book sounds amazing and I can’t wait for it to be published in English. But the way in which we kind of came across your work for the first time was through your writing about audio fingerprinting and content ID, which now that I understand a little bit more about the book, I can totally see why this would be sort of in there.
James Parker (00:28:28) - But from the perspective of machine listening, you know, audio fingerprinting is a kind of, is a technique that comes out of computer science, basically, we should talk about it, that gets kind of scaled up to be one of the main methods for sort of tracking copyrighted works on YouTube and the internet more broadly. And so when I came across your work on this, I was really excited. And I’d love to have a little bit of a conversation about that part of the project, if you don’t mind.
James Parker (00:29:06) - Would it be worth maybe beginning by explaining very roughly what Google Content ID is and why people should care about it, like in broad terms, why artists, musicians who are on YouTube do care about it and spend a lot of their time engaging with it, which I think they do?
Guillaume Heuguet (00:29:27) - Yeah, I think one way to define a Content ID is that it’s a management system for the use of copyrighted material on YouTube. So that’s the main aspect of it. And a part of it is that there are devices, semi-automated devices to detect this copyrighted material in videos and especially music, but not only music. So that will be the end of the presentation.
Joel Stern (00:30:00) - Yeah, the outline of what Content ID is.
James Parker (00:30:03) - Yeah, and could you say a little bit about what the experience is of Encounter? I mean, I have my own ideas. From what I understand, if you’re an artist or a musician putting work up on YouTube, Content ID is like a sort of Kafka-esque sort of nightmare that not only will kind of often lead to a video you put up being taken, but because of the automation and because the way that Content ID works sort of changes so often, it’s just completely opaque. So you’ll find that the video will not make it up.
James Parker (00:30:47) - It will be found to have been in breach of some copyright rule and you won’t know why. And it seems like there’s a subculture of YouTube videos of people explaining how to sort of hack or work around Content ID, that it’s generated this kind of enormous discourse on Content ID from people trying to negotiate it. So I’d love to know what Content ID is doing to music culture in that sort of very obvious way.
James Parker (00:31:22) - It seems like if you’re putting music or videos onto YouTube, you just have to learn how to engage with and deal with Content ID. What’s your experience of that? Guillaume Heuguet (00:31:35) - The main thing is that when there are two aspects of it, there is the aspect of you being a kind of general casual user publishing videos on YouTube, and music is the main focus, but you want to use copyrighted material in your video.
Joel Stern (00:31:58) - So there is this… being a musician, an artist or someone who represents a musician or an artist, and using Content ID to track the use of your music in other videos. So that’s the two separate main experiences.
Guillaume Heuguet (00:32:18) - If you are a casual user or even what they call a creator, a video maker on the platform, you’ll want to put out a video with copyrighted music and you’ll learn in the process if the people who have the right on the song want to block your video, monetize it or just track the statistics of your video.
Joel Stern (00:32:48) - These are the three options that you have on the other side of the apparatus.
Guillaume Heuguet (00:32:55) - And so the complicated thing is that sometimes people want to just monetize the video that you make on the basis that it uses this song, this copyrighted song or record.
Joel Stern (00:33:13) - So your experience will be that you’ll have to share a good portion up to 60% of your revenue with the owner of the song. But also maybe your video will be blocked, just blocked, and then you… why it was blocked and discuss through the messaging device on YouTube with the people who blocked your video.
James Parker (00:33:45) - And am I right that sometimes the blocking can be retrospective, so you’ll have a video up, it’ll be fine for years, and then suddenly it’ll be blocked because of some new sort of…
Guillaume Heuguet (00:33:56) - Exactly. The reason for that is that on the other side, you have artists working with online music distribution companies who in most cases will handle the tracking of your music on YouTube.
Guillaume Heuguet (00:34:19) - Sometimes you can do it through your record company, but mostly it’s music distribution companies.
Guillaume Heuguet (00:34:27) - And the policy you have can evolve with time. So you yourself can tell this music distribution company that you want all user-generated content using your music allowed because you think it’s good promotion, for example.
James Parker (00:34:47) - Or
Guillaume Heuguet (00:34:48) - I don’t know, you want to sell the rights to your catalogue to a big publishing company, and you think that it’s giving you more leverage to get everything off of YouTube.
Guillaume Heuguet (00:35:08) - So, and the way this works is that you can set up very finely the parameters of the detection. There is a whole technical aspect of it on the side of the music distribution company. So, for example, you can define the length of the extract that are authorized on that, and each music distribution company can set up different parameters for different song or catalogue.
Guillaume Heuguet (00:35:38) - So there is a whole complexity of it on the side of the music distribution company and artists as well.
James Parker (00:35:45) - And what about the other example where artists are sampling or using copyrighted… I mean, I sort of… I actually do want to get to that, but even this idea that there’s a thing called copyrighted work and that’s knowable and it’s obvious what… It sort of puts the cart before the horse in some ways because I’ve read of so many examples of people having white noise in the background of something and then their video being taken down because that’s found to have been copyrighted.
James Parker (00:36:20) - You know, like, in some ways part of the whole question is when is audio copyrighted? So to say that an artist is working with copyrighted material or whatever sort of already presumes that they know. But part of the point is that you sometimes find out that according to Google Content ID, you’re working with copyrighted material, but you didn’t know that in advance or it’s simply wrong. And there’s nothing you can do about it.
Guillaume Heuguet (00:36:52) - I think a big move on the part of YouTube on this is that you’re right, it’s not copyrighted material in the general sense. It’s official artist…
Guillaume Heuguet (00:37:11) - Official artist is very much a phrase that YouTube uses.
Guillaume Heuguet (00:37:15) - Music that has been registered on the platform. So if you want to be… YouTube, you have to have like… I think there has been an update on this last year. I think now it’s like 3000 subscribers or something like that.
Guillaume Heuguet (00:37:36) - Or you have to go through one of those big distribution companies or one of the registered labels.
Guillaume Heuguet (00:37:42) - So one interesting thing is that it’s not about having registered your music with a copyright society outside of YouTube.
Guillaume Heuguet (00:37:55) - It’s being able to fulfill the parameters that make you official as an artist on YouTube. And part of this, of the… the fact that you’re being registered with a big label or music distribution company.
Guillaume Heuguet (00:38:15) - But there is also the fact that your music registers… Sorry, your music doesn’t go against the policy on the platform and the appeal to advertisers. So when you want to register your music with YouTube, you have to make sure yourself that your music is not offensive, your music is good to place ads on and stuff like that. So all of this is what can be defined as a copyrighted material on the platform.
Guillaume Heuguet (00:38:50) - And if you’re doing music that can be copyrighted outside of YouTube, but you don’t fulfill those criteria, someone can possibly register something that sounds a bit like your music before you.
Guillaume Heuguet (00:39:09) - And you can’t oppose your own copyrighted rights to the one of the people who register another recording close to yours. It’s a bit complicated to explain.
James Parker (00:39:24) - I wish our collaborator Sean was here because he has an old artwork that you might be interested in where he uploaded a… He automatically generated and uploaded a series of videos to YouTube with a plain color, nothing but a plain color and was it like a sine wave at different pitches? And he just automatically generated and automatically uploaded them and then the work is basically them being progressively taken down over many years.
James Parker (00:40:00) - Because they would sort of pick up all… So the work is sort of meant to be a kind of like an… I mean, I shouldn’t speak on behalf of him, but it’s kind of like an index of the strange encounters between these non-works and a copyright regime and the kind of artifacts of the… what they document is a changing copyright regime as opposed to… because it would be
Guillaume Heuguet (00:40:31) - the proof to my thinking. So that’s exactly…
James Parker (00:40:35) - Well, we’ll have to put you in touch properly because I think you can find it. There are residues of it still on YouTube. And yeah, I mean, I don’t know if you know off the top of your head, Joel, how… where to find that. But anyway, we… Guillaume Heuguet (00:40:50) - I’ll find it.
Joel Stern (00:40:51) - I mean, all I remember is that the username that he used for uploading the videos was Alexander Rodchenko. And that there were many thousands… there were tens of thousands of videos which were just monochrome images and sine waves. And the speed of generating and uploading was sort of roughly consistent with the speed of blocking and removing. So there was a constant recycling of them. But one interesting aspect of the work was… that was not expected…
Joel Stern (00:41:28) - was that beyond the automatic uploading and removal, there were actually human viewers who were watching the videos and speculating in the comments as to what was the possible meaning of these monochrome sine wave videos and, you know, what kind of reason they might have for existing and they were getting into some quite complicated theories and speculations as to what it could be.
Guillaume Heuguet (00:41:59) - In the visual field, there has been a debate about… I think we even have a law in France and maybe in Europe about landscape exception, which means that you can have, for example, a very well known French monument in photography and nobody… no photographer or no representative of a photographer can take down your image because there is copyrighted material in the background. And I don’t think there is such a thing for music or sound.
James Parker (00:42:35) - Well, that, I mean, we might have time to talk about things like parody or various different sort of ways of… that copyright has sort of dealt or some regimes of copyright have dealt with musical quotation or sort of genericness and so on. But before we do that, maybe we could talk about audio fingerprinting because you said before that, you know, with tech that your journal and in the project in general, you’re quite interested in sticking to the sort of the technical and material dimensions.
James Parker (00:43:11) - I mean, one of the things that’s come out of these processes, one of the things that’s really struck me in the way that you’re talking about content ideas, how… I mean, so little of what you said at the moment is technical. It simply describes a bureaucracy and in a huge economy. I mean, I’m just… every time you speak, I’m just imagining all of the lawyers and administrators and, you know, the kind of…
James Parker (00:43:42) - there must be huge businesses whose exclusive role is to interface with the parameters on the back end of Google Content ID. I mean, it’s a huge sort of political, economic, social material process, but that’s all built as far as I understand it or largely built as far as I understand it on top of this technique or set of techniques called audio fingerprinting. And that’s how I first came across your work. So could you say a little bit about audio fingerprinting, you know, what it’s got to do with Content ID, maybe why it matters?
Guillaume Heuguet (00:44:20) - Yeah, just one thing before that, you’re right. Like I’ve met those people. And I think one of the main things about this way of handling copyright on platforms like YouTube is that now you have teams of people at music distribution companies, which are doing YouTube’s work, like who are trying to work around their catalogue and the roles and the back office of YouTube to make it so that this dream of controlling every copyright on the platform works. And so I think it’s very important to stress that the…
Guillaume Heuguet (00:44:59) - The economic model of YouTube works because it externalizes pine activity to music distribution people in music distribution companies. So your question was how I came into studying audio fingerprinting?
James Parker (00:45:23) - Well, sure, but also what it even is and how it’s related to Google Content ID. So audio fingerprinting, again, I think is the industry name for something that I will say is an attempt to make a singular record of each song in a given database and be able to match it to any other file uploaded on a given platform.
Guillaume Heuguet (00:45:54) - So I’ll say the most visible form it takes for us if we approach the technical layers of those platforms is a hash, which is a kind of a serial key for our software. It’s just a series of numbers, which is supposed to be singular to one file, one record, one… James Parker (00:46:21) - And not just one file even, is it? It’s one segment of a file. So the copyrightable work doesn’t just have a number or a title or something. It’s got many hundreds or thousands of micro identifiers or something.
Guillaume Heuguet (00:46:39) - Yeah, so that’s why I was saying supposed to be.
James Parker (00:46:44) - Ah, supposed to be.
Guillaume Heuguet (00:46:46) - Because, yeah, obviously to me the interesting thing is the fact that we call that audio fingerprints when it’s more of a mathematical model with a series of numbers or a small data file as a result. So I think audio fingerprinting is focusing on the semiotic, on the effect, on the use of something which when you go behind it is a machine listening and a feature extraction process, which is trying to make a model of a song or a portion of a song with several different techniques.
James Parker (00:47:42) - Can I just paraphrase? Is it what you’re saying is that there’s a sort of rhetorical or branding move effectively being made to make this thing familiar? Make these very diverse and complex and heavily mathematized and quite constructed processes seem… just make sense to people, right? So we’re familiar with the metaphor of the fingerprint and it sounds like a kind of… it sounds plausible that a song could have a fingerprint. But it doesn’t. And so is that what you’re getting at?
James Parker (00:48:25) - There’s a kind of a critique of the kind of the rhetoric of the fingerprint as a way of describing what’s effectively… I mean, at one point in the piece that I’ve read, you talk about a mathematics of originality, which I think is a really provocative phrase and very different from rhetorically to the idea of an audio fingerprint.
Guillaume Heuguet (00:48:46) - Yeah, I’m really amazed by this idea that we could call audio fingerprinting something which is trying to make an approximative statistical model of a song, which is… it’s two totally different things to me. If you listen to audio fingerprint as a phrase, as a couple of words, you will think like something of an index, something like… is a part of a bigger thing and you can’t separate it from it, you know, it’s like your… it’s the tip of your finger.
Guillaume Heuguet (00:49:29) - So you have the whole body and you have a singular person and it’s like all those things are impossible to differentiate. And it’s a very different task from the perspective of… work for engineers and data scientists to try to make a good enough mathematical model of a portion of a song, which will be different enough.
Guillaume Heuguet (00:49:58) - From any other mathematical model of any other song. It’s like two totally different tasks. And one thing that’s truly fascinating to me with the idea of coding that audio fingerprinting as well is that there are a lot of work in STS, in sociology of science and technologies to stress that fingerprints are not even a good method to make sure of the identity of one person. There are a lot of work about the failure of fingerprints as an imprint and a record of someone’s identity.
Guillaume Heuguet (00:50:43) - There have been several criminal cases with failures of people going to prison because we thought that we got their fingerprints and stuff like that. So I won’t delve into that. But even fingerprints, it’s not a good idea to a good way to define the singularity of a given person.
Anabelle Lacroix (00:51:06) - I agree with what you just said, it just made me think that the way we think about identification has totally changed. So audio fingerprinting tells us something about how we think about identification today.
Guillaume Heuguet (00:51:21) - Yeah, I agree. Really, I want to stress that I think we have to acknowledge that the idea from the perspective of data science and stuff like that, and even the policy of those platforms is to have good enough probabilistic results. And it’s very different to the idea that we all catch the essence of a given song. And I think it’s interesting because I don’t want to have this critical approach of saying there is the real definition of a song and not the human definition of a song. And now we only have the mathematical model.
Guillaume Heuguet (00:52:07) - So it’s not true anymore because it’s mathematical. You know, it’s not about that. It’s that I think we are the very, I don’t know, like the definition of singularity, originality of what a musical work is, of the borders of a work of art are very soft in a way and very relational. So any idea to make that solid is already kind of a challenge, even in a human context.
Guillaume Heuguet (00:52:47) - And when it translates to data science, I think we have decided that we can do a probabilistic account of the definition that we as listeners, artists, and any other kind of experience out of these songs. And it just has to be good enough for the system to work. And I think one interesting thing is that we have to unpack this idea of good enough for who, what means good enough of a probabilistic definition.
Guillaume Heuguet (00:53:30) - And when you look into it, we don’t have the specific data, but if we want to see what it means to make a statistical model of a song, you have to choose between different parameters and you can either make it wide enough to catch every analogy with any other recording that might be close to that specific model, or you can make it tighter with fewer and more singular, more granular parameters. And then you might miss some songs or some videos or some other files that might resemble this one, but might be some other thing.
Guillaume Heuguet (00:54:19) - And I think in the case of YouTube, it’s very clear that the agenda, the political agenda, the economic agenda is to please the copyright holders. And so I think that by design it’s set up so that it catches too much analogies between a model and another recording. Hence the case that you told us about some white noise being similar to any other white noise in any other video. And so there is a kind of ontological limit to those attempts because
James Parker (00:55:12) - there is I mean, it might be worth specifying the whiteness and the westernness of that political and economic agenda, because obviously, you know, the classic critique of the of copyright as a regime is, especially in relation to music, is that it comes along to work very well for the Rolling Stones and much less well for Muddy Waters and every blues musician who’s ever existed and very well for, who’s the guy who did Graceland? Paul Simon.
James Parker (00:55:50) - And not so well for all of the traditional African musics on which he’s heavily drawing and producing this mega hit. So, you know, obviously there’s a very long legacy of the copyright regimes drawing capital to the West and to white musicians. And I mean, it just seems like a pretty obvious fact that this is kind of bedding that down and making it more opaque and harder to critique. Because at least if somebody says, well, look, only a melody is copyrightable and a sort of a groove.
James Parker (00:56:40) - Well, no, that’s not or like a sort of chord structure as generic as the blues or something that’s not copyrightable. But you can say, get fucked. But when you have a mathematical model that nobody understands and a Google Content ID finding that’s possible to appeal, but very difficult and time consuming, and you’re probably never going to find out exactly what the back end looks like anyway. I mean, you. Yeah.
James Parker (00:57:14) - I mean, it just seems sort of obvious that it makes it harder to the mechanisation of it and the automation of it isn’t the isn’t like the problem compared with some more authentic humanist. But it makes it more difficult to launch a sort of a strong critique. Although on some level, it’s just easy because, well, of course, YouTube is drawing money to the existing powers that be. But yeah, I don’t really know. That’s not a question. That’s a comment, not a question.
Guillaume Heuguet (00:57:47) - There are so many things that you just said that resonate with my research. One thing is that I think it is interesting to see two parts of the world where copyright is not working the same way. To draw the same critique and to look in another way to our own cultures and realise that even in Europe or in the US, we can make it weird that our definition of music is so linked to the idea of the individual artist or the originality of songs and records. And so I would say that copyright is weird in any part of the world.
Guillaume Heuguet (00:58:49) - It stresses out the definition of music as a common, maybe latent in every musical practice and listening experience. And maybe the symbolic visibility of the industry and the categories of originality, authorship, commodity. But we have to ask ourselves if that’s really at the core of why we like music and why it’s so valuable to us as a practice, I think.
James Parker (00:59:32) - It’s interesting that you use the language of commons because it strikes me that it’s not just the music that’s being privatised here, but the system for… There’s a privatisation of the copyright regime itself. Of course in the 20th century, who has the means financially to enforce copyright or whatever? It’s only rich people.
James Parker (01:00:00) - or in record companies anyway. But effectively YouTube has entirely privatized a system of governance, and it’s not just YouTube, but that platform governance as a privatization of the very regime for even encountering, disputing, regulating music. That seems like a really big move. It’s a bit of a problem. It strikes me as a big problem.
Guillaume Heuguet (01:00:32) - It was a kind of a speculative hypothesis in my work, but it is interesting to me to discuss this with you as a law scholar. I was like, one of the unexpected result of focusing on these topics was to realize how much of the copyright debates about music were taking advantage of the jurisprudential dimension of law, at least in a European context. The idea that it’s not just about arbitration and like arbitrage, It’s not just about threats of censorship and removal of a given song. So you give back money to the real owners and the case is settled.
Guillaume Heuguet (01:01:25) - It’s about public discussion, taking advantage of the publicity of law to define where lies the value in music and where lies the value in musical practice. And I think even before YouTube and the platforms we have seen a kind of transformation of this philosophical debates happening within law towards the idea that it’s a competition between companies who own music, kind of like they will own technological patents, threatening other parties, getting money from their capacity to use law to threaten other parties.
Guillaume Heuguet (01:02:18) - And law is not the space where we collectively define what is allowed and what is not allowed but more another way of making money. And I think what you call the privatization of governance is another step in this process where law is not a space for debates and the shaping of culture as a whole, but just a tool to further economical, what they will call innovation and from another point of view, economical profits.
James Parker (01:03:01) - I feel like I’m supposed to provide some jurisprudential wisdom. I don’t know if I have any. I mean, it was very well put. I’m conscious that we’ve taken up a lot of your time already Guillaume and it’s getting relatively late here. I’m trying to think of what a nice sort of closing question or two might be. I don’t know if either Anabelle or Joel, you’ve got one. I mean, I suppose I’m interested a little bit in where you think things are headed.
James Parker (01:03:36) - We could talk a little bit about resistance or hope or whether you think vibrant music cultures can be sustained in the margins and the kind of the inabilities of something like audio fingerprinting or content ID to catch experimental forms or what have you. I don’t know. I’m just interested when you look at all of this, what do you see? Do you see a story of like limitation of creative opportunities online or do you see something different and where do you think things are going?
Guillaume Heuguet (01:04:22) - So I think there are several trends which are interesting at the moment. I think one interesting thing is that music is a kind of driver of economic innovation and economic power. So for example, TikTok is big today because you can put a few seconds of music behind.
Guillaume Heuguet (01:04:51) - Any video easily and stuff like that. So I think there is a strong incentive from the part of the platforms to make the most of music for their own agenda. And that’s one trend. And what’s interesting to me that on the other side, on the side of people who listen to music, of artists, of musical culture. I think there is a. So there are real shifts in the definition of music and what makes sense as a song and music today and all those practice.
Guillaume Heuguet (01:05:31) - Of I don’t know, being it can be the importance of ambient music at the moment of all those edit and remixes of musical experience that are not defined by copyright. I think that is a whole shift in the music. What was the musical culture which is being expanded at the moment, even in popular culture. So that’s kind of two contrasting trends at the on the one part, the platforms are making music as the ultimate community to sell technological products.
Guillaume Heuguet (01:06:07) - And on the other side, people aesthetically and socially are relating to music and sound in the way that has less and less to do with the. All the industry community forms, I think so. That’s one thing. Another thing is that. I think the. The machine learning, the rise of machine learning. Is allowing to create a lot of forms on the. On the scale that was never possible before, like if I want to set up a few parameters to tomorrow and create 10 or 100 or 1000 different version of a song and put them to YouTube, I can do that.
Guillaume Heuguet (01:06:56) - And so I think there will be a clash between the generative aspect of machine learning and the use of machine learning to control and detect. Copyright. And I’m curious to see how these two things will interact in the years to come. And the last time that I think is interesting in regards to what we discussed. Is the NFTs because NFTs, one way to look at it. Is a way to extend musical copyright to a lot to a variety of forms way beyond any idea for a singular song. And potentially you can make NFTs about everything and nothing.
Guillaume Heuguet (01:07:49) - And so I think there are a lot of trends at the moment which are already pushing the limits of copyright as a way to create a social. But mostly economical value. And I think I want to finish with that is the fact that even from an industry perspective, even if you’re not critical and you want to and you think that copyright is just a market device which can help make money for a lot of people. I think copyright might have done its time as an economical device. So maybe from the side of academic artists and the general people.
Guillaume Heuguet (01:08:46) - The question we might ask ourselves is if we are escaping out of copyright, do we want a market escape out of copyright? Well, people who like find new ways to make money without copyright, but maybe with new forms of control and new form of markets and stuff like that. Or do we want to seize this opportunity to push new definition of music and of social practice outside of markets?
James Parker (01:09:17) - I choose the second one.
Anabelle Lacroix (01:09:19) - Me too.
James Parker (01:09:26) - And Joel, did you want to ask anything or comment?
Joel Stern - Not for me. I mean, I thought that was a really elegant conclusion.
James Parker (01:09:44) - Just a few minutes ago. Anabelle?
Anabelle Lacroix (01:09:45) - I was thinking about something more around listening, it’s a question for the three of you in relation to the prompt of machine listening and the question of ‘What music is doing to YouTube?’ would be to continue the question with: ‘What does music do to YouTube, that in turn affects our listening? And how?
Ahead of the conversation, I was thinking about the work of Bernhard Stiegler and this idea of the technique, which links back to the philosophical weight of the approach to this topic you mentioned at the start of the conversation Guillaume. We could think that YouTube has changed music if we think about it in the same way as the record, as Stiegler has shown that the record has changed how we listen and turn, it has changed how we made music. And so it’s the same thing with YouTube. How does YouTube has changed our listening?
Guillaume Heuguet (01:11:06) - Yeah, I, I’m really interested in this question but I’m a bit frustrated that I didn’t discuss it more in the book so thank you for the question. I can talk a bit about it. So is a part of the of… view the YouTube view as a audience measurement. And I found it interesting that, you know, all those discussions about. Do you need to stream a video for 30 seconds for for the thing to be considered as as one view of the video or the song. And so, I think there are several.
Guillaume Heuguet (01:11:52) - There are several lines of the way which in which YouTube shapes, the consumption of music and the listening of music. I think one thing is the consumption, not even the listening, the listening. The consumption is like the main thing on the main YouTube YouTube music but YouTube as a video site. Is that the channel is the way to publish to publish video music music videos. So music is news. Music is always like it’s something of a daily conception with notification, you have a new music video out, and you check it, and you have checked it.
Guillaume Heuguet (01:12:33) - And I think streaming really pushed this way of seeing musical culture, as something of a social news.
Anabelle Lacroix (01:12:47) - It’s also like an event.
Guillaume Heuguet (01:12:49) - Something you subscribe to, and you, you get to know, and then you pass to the other thing. So that’s one layer. I think it’s also with the view is decided that we have this share of leisure time you know, and so, so you have this relationship to music so how much music music fits my share of leisure time. And so that’s why you have all those playlist, even on YouTube music with like music for your Friday night. So, that’s what extent this music is a good use of my leisure time. I think it’s a way to look at it.
Guillaume Heuguet (01:13:31) - And so, I think, but at the same time you have all those videos by, by, I’d say regular users even if it doesn’t mean anything, like push the definition of what music is and and and edits music in a way, and if I do those ideas music music as a as a user value and stuff like that.
James Parker (01:13:58) - Can I make an alternate suggestion which is that when I think of the difference between YouTube and every other streaming platform is that YouTube always stands out to me as archival in a way that like not archival in the sense that like it’s got everything, but much more historically oriented than every other music platform so of course like things that are news. But if I want to find like a rare, like,
James Parker (01:14:32) - Jungle like pressing from 1992 from time to time that I do, or like a specific performance, you know, in like Japan or something. You get that on YouTube, right?
Guillaume Heuguet (01:14:49) - YouTube.
James Parker (01:14:50) - And so, you know, some of the people that you were right, you were talking about publishing in Audimat, like Simon Remnalds, for example, and his book Retromania, you know, are really, really specifically thinking about something like YouTube when they make that argument. Not so much like Spotify, I think. It’s the fact that those songs all come ready-made with the music videos or the hairstyles or the padded shoulders, you know, the suits that they were wearing on stage and whatever.
James Parker (01:15:30) - There’s a sort of an archival or an historical orientation to YouTube, which seems very different from other music platforms to me. So I totally accept what you’re saying about the new, but I also think that YouTube is oriented towards the old and the historical in a way that has really shaped my music consumption. You know, like I’m on the academic sort of orientation. I do think there’s a kind of academic like tendency in YouTube music consumption.
Joel Stern (01:16:00) - Well, there’s also a lot of explicit remediation on YouTube where you’re watching something taped off a television show from the 80s, or you’re watching someone drop a needle on a record or put a cassette into a tape deck and press play where the reproduction of the media is actually foregrounded in the video in a certain way. So there’s a sort of media archaeological dimension to it that is much more explicit than in other platforms.
Guillaume Heuguet (01:16:37) - Yeah, I agree with both of you. There is a chapter in my book when I try to order the other relationship and mediation of music on YouTube and archival part is totally one of them. The media aspect is in it too. And there is also the more social, you know, artist fan relationship aspect. relationship to how TV is remediated in YouTube as well.
Guillaume Heuguet (01:17:20) - So I agree with the fact that there is all those things, but I think it’s interesting to see how it shifts over time in the dominant discourse on YouTube or on the main page, for example, what is being shown first and how the other dimension are more like a vernacular culture of YouTube that’s been there for a while, but not the top priority of the company itself.
James Parker (01:17:54) - I’m sure there’ll be a lot of interest in the book when it comes out in English. When is that happening?
Guillaume Heuguet (01:18:00) - Probably September.
James Parker (01:18:00) - And how was the interest in France? Or French speaking countries when it came out? Because what I’m wondering is, like, I feel like, academically, it’s an extremely interesting topic, but I do feel like it sounds to me like the kind of book that has a sort of a potential broader readership.
James Parker (01:18:23) - I think a lot of people are interested in YouTube, and I’m interested to know, just to tie us back to the very start, like how you see this in relation to the kinds of critical writing, you know, that you’re sort of platforming in Audimat, which are sort of para-academic or like maybe for a more popular audience or, you know, in some cases. Well, I mean, yeah, I don’t know.
James Parker (01:18:56) - I don’t want to do too strong a distinction between academic and non-academic, but for some of those writers, it was very important that they weren’t writing from within a university context. And so, yeah, I’m just interested to know, like, what the reception was and how you thought about that. But we should wrap it up. But I mean, yeah, it seems like there’d be a big readership anyway.
Interviews
A repository of conversations we’ve had with researchers and artists across the life of the project. These interview transcripts are produced using the Reduct transcription platform. The transcripts contain artefacts and politics of machinic listenings, such as the occasional Miss Herd word, and ah vocalised hesitations and filler sounds.
Kathy Reid
Kathy talks to us about her work with Mycroft, Mozilla Voice and now 3AI on open source voice assistants and the technics and politics of automatic speech recognition, along with a couple of utopian possibilities.
Interview conducted on 11 August, 2020
Max Ritts transcript
minor edits by James Parker. time stamps reflect unedited recording, so are a little off throughout.
[00:00:20] James Parker: All right. Thanks so much for joining us, Max. Yeah, would you like to maybe kick off by just introducing yourself, however feels right to you? Sure.
[00:00:32] Max Ritts: Well, I am a geographer. I’m a first year professor. Just finished my first year at Clark University in Worcester, Massachusetts. But I guess I really consider myself a person who works and thinks along the North Coast of British Columbia most of all. So that’s the sort of context in which I approach issues of sound and sound studies and also the other things that shape the work I do, like indigenous politics and political ecology.
[00:01:02] James Parker: What’s your disciplinary background? Oh, yeah? OK. What’s your disciplinary background? Because what you’ve just described there is kind of pretty sound studies. Is it or isn’t it a discipline? But perhaps you have a training or something that’s relevant.
[00:01:33] Max Ritts: Sure. My background is geography. I consider myself an environmental geographer, most of all, I guess, although I definitely have other interests and passions and areas of curiosity.
[00:01:47] Max Ritts: I’d also say that I was lucky to have some really wonderful advisors in the world of sound studies, including Jonathan Stern, who was a committee member, and also a guy named Jeff Mann, who is known more as a political economist but wrote a really wonderful piece on country music back in 2008, which really introduced me to some of the ideas that I would then work into my own kind of project of geography around ideology and sound and power and even capitalism and the way that those issues manifest as acoustic issues. But yeah, geography is my field.
[00:02:19] James Parker: The first piece of yours I came across was called Military Cetology, a piece you wrote with John Shiga. And you’ve got work on the social construction of Wales song and then obviously more recent work that’s to do with digital bioacoustics and various other topics. So it’s a pretty broad palette. Could you describe some of the topics you’ve worked on? I mean, I’ve mentioned a couple. You said British Columbia, but what’s the sort of umbrella view, the broad overview of what you’ve worked on previously? Because I know you’ve got a book coming out.
[00:02:59] James Parker: Are these all going to feed into a book? How does the story kind of hold together, I guess I’m wondering, beyond geography, beyond sound studies?
[00:03:10] Max Ritts: Yeah, it’s a good question. And I guess that’s why I also introduced myself with respect to this part of the world that has really shaped my training and my interests. Because the North Coast has been a space that I’ve worked through in relation to a number of different kind of sonic problems, I guess you could call them. Wales song being one of them, ocean noise being another, but also music, also indigenous popular music and acoustic ecologies and its legacy, and smart technologies and smart oceans, which I’m sure we’re going to talk about.
[00:03:47] Max Ritts: Each of them in different ways beyond that region. But the region has kind of maintained this kind of cohering effect in the sense that these different topics that I’ve studied in the world of sound studies and geography all kind of related to my initial time in this part of the world, where I struggled, and I’m still struggling in some ways, to kind of see the connections.
[00:04:07] Max Ritts: But that’s really what the book that I’m working on is about too, is sort of seeing in this diversity of different kinds of sonic cultures some underlying questions and underlying problems that give those things meaning in that part of the world. I’ve written about bioacoustics, and I’ve written about calving glaciers and aesthetic forays into listening to ice.
[00:04:28] Max Ritts: And I have a piece coming out around urban sustainability politics, which initially had a whole bit on surveillance actually, eavesdropping, which we kind of had to unfortunately remove from the final draft.
[00:04:38] Max Ritts: But all these topics kind of came out, again, of this very different part of the world, even though they’ve taken place in places like Denmark and the Beaufort Sea, which is a bit of an idiosyncratic way to do geography, but I guess it maintains that idea of situated knowledges and place-based research and just ways of thinking problems through the articulation of different spaces in question.
[00:05:04] James Parker: Do you have a title for the book?
[00:05:08] Max Ritts: I do. Yeah. It’s called A Resonant Ecology.
[00:05:13] James Parker: OK. So it’s an ecology is kind of going to be the, you know, the thing that holds it together.
[00:05:21] Max Ritts: Yeah, well, I can get into it right now if you want. I mean, but the the reason, the initial reason I called the book a resonant ecology is because I was thinking through some problems around the term resonance, which comes up a lot in in sound studies, and also in kind of eco theoretical discussion involving sound.
[00:05:40] Max Ritts: And I was unnerved by the way in which resonance was almost uniformly being described as a kind of positive trajectory or valence like it was a good thing to achieve resonance, when at the same time, we noticed that, you know, big data and IBM acoustic program and Microsoft and Google are also very interested in, you know, the relations of sound and space for very different purposes than, you know, democratic futures.
[00:06:04] Max Ritts: And the way in which discourses of resonance appear in Silicon Valley and in big tech in general, and you can look at auto charmers work on resonance, for example, as an example of this, this kind of moving into that space, made me realize that the term is much more better conceived as a kind of uncertain site of social mediation of sound rather than as this uniformly positive thing.
[00:06:28] Max Ritts: And in the end, the book doesn’t really take on resonance directly too much, but it does maintain that interest in kind of critiquing this sort of uniformly positive turn to sound that I see pervasive in the tour even write these like very prominent thinkers who kind of posed ideas of resonance and listening as almost inherently virtuous. And of course, there are many beneficial aspects to thinking relations of nature through sound. And I want to celebrate those in the book as well.
[00:06:57] Max Ritts: But I want to kind of critique or push back against this notion that it’s a positive thing, through and through. And so that’s one of the sort of through lines running through the book is kind of looking at these different case studies that take on different aspects of that problem. And another one is sounds digital is digitalization. So the period in which I was on the North Coast, roughly 2012 to 2019, kind of coincides with this ascendant moment of digitalization of so many things in culture, but, you know, sound among them.
[00:07:25] Max Ritts: And that’s the period that we see this explosive growth in things like eco and bioacoustics. And the leveraging of those sciences through again, big tech, through things like smart governance technologies through surveillance, you know, things that we now see in the form of like racializing urban surveillance, right. So that wonderful book by Brian Jordan Jefferson digitize and punish, all this stuff begins to kind of gather steam in and around this period.
[00:07:47] Max Ritts: And I was kind of witnessing it play out in the North Coast, which is, of course, also unseated coastal First Nations lands and you know, a space where resident relations are very much, you know, things that are being valued and celebrated alongside these other kinds of terms. So the book is trying to make sense of this kind of competing and intersecting discourse of sound at a time of global unrest, but also at a time of like political economic innovation in sound, and the way that capital is moving into sound in new ways through digital technologies.
[00:08:18] James Parker: Sounds absolutely amazing. I know that this wasn’t primarily conceived as an ad for your book, but when’s it coming?
[00:08:26] Max Ritts: Fall 2024.
[00:08:28] James Parker: But for 2024, do you have a you have got the publisher all lined up and everything?
[00:08:32] Max Ritts: Oh, yeah. Yeah. So it’s the Duke University Press. Signal Storage Transmission Series, which Jonathan and I believe also forgetting the other, the other author is a part of the couple of really cool sound studies people who edit that.
[00:08:51] James Parker: Yeah. Yeah. Amazing.
[00:08:53] Max Ritts: Lisa Gitelman.
[00:08:53] James Parker: Lisa Gitelman.
[00:08:55] Joel Stern: Okay. Yeah. Okay.
[00:08:57] James Parker: So not messing around then. Look, that sounds absolutely amazing. Before we continue, we did have a couple of like slight drops there. I think it’ll be okay. But should we maybe continue with a video off just in case it makes a difference?
[00:09:14] Max Ritts: Yeah.
[00:09:14] James Parker: Did you also lose him there for a second, Joel?
[00:09:17] Joel Stern: Yeah, just cutouts for like a second at a time. It sort of seems like every every four or five minutes, there’s just a drop for a second or two. But yeah, let’s see. Let’s see.
[00:09:42] James Parker: The book project sounds amazing. And it seems like it has a lot to do with this machine listening project. You know, everything you’ve been saying about digital bioacoustics seems absolutely spot on for us. You know, could we maybe dig in a little bit to that research? You published some of it. I imagine there’s more of it that makes it into the book. But for the sake of argument, perhaps we could start off with this piece you co-authored with Karen Bakker on conservation acoustics.
[00:10:18] James Parker: Yeah, I mean, it seems like it’s what the work that you’ve done that is closest to machine listening, but it’s not just about machine listening. So, yeah, perhaps I don’t know what the best way into that project is. Perhaps you could start by explaining what conservation acoustics is. Or perhaps it’s better to explain how you arrived at this as a problem or a field, you know, through this kind of located sort of methodology that you have. Like, yeah, how do you end up working on conservation acoustics basically and what is it?
[00:10:51] Max Ritts: Yeah, so I’ll define the term as sort of we saw it first and then I’ll kind of backtrack a little bit. It’s really meant to kind of combine two innovative kind of techno sciences of sound, which are both concerned with ecological conservation issues. And those are bioacoustics, which has been around for decades and has really kind of, I wouldn’t say matured as a field, but has really kind of advanced in the last decade and a half for reasons I’ll get into. But also ecoacoustics, which is far newer, which kind of got consecrated in around 2011.
[00:11:23] Max Ritts: I think there was a paper by Brian Pijanowski and colleagues that called it Soundscape Ecology. But then other people like Alma Farina and Stuart Gage using the term, you know, ecoacoustics. In both cases, it’s kind of like a landscape scale evolution or other study of sound. You know, the idea of the soundscape as a kind of measure of the ecosystem’s health, whereas bioacoustics is more focused on individual species interactions.
[00:11:49] Max Ritts: And so what both of those fields are doing is kind of maturing at the same time through this digitalization story I just mentioned. Mitigating emergent threats in conservation areas, detecting poachers, learning more about the subterranean or hard to reach nature of certain animal interactions and all sorts of exciting work happening through the study of these complicated ecologies that are changing as a consequence of climate change. Climate change and deforestation and so many other things. And so there’s a lot of hype behind these fields.
[00:12:23] Max Ritts: And they’re being supported by, you know, not just scientists and universities, but also by Huawei and IBM and Microsoft and Google. And you can look at Google’s Under the Canopy project for an example of a kind of virtuous, quote unquote, merger of, you know, the work of rainforest conservation with indigenous peoples in Brazil, in the Amazon. And the idea of kind of saving the rainforest and saving indigeneity with all the problematic consequences of that idea through the white male savior who can, you know, wire the territory with listening devices.
[00:12:55] Max Ritts: And so all of that, I think, is this moment that we’re trying to make sense of. And it was especially kind of strange for me because in 2016, I published my first ever piece with the GitGat Nation, with these scientists doing something very similar in their territory, which was this ecoacoustics baseline.
[00:13:12] Max Ritts: That basically put up these sensors along the Douglas Channel to listen for the ambient kind of profile, ambient sound profile of the region to generate an inventory that the nation could use in their engagement for a process with these industrial proponents with Enbridge. Basically to show that the territory has all these different rich sonic relationships that would be affected by all this noise from the tankers. So we did that work in 2012, 2013, and then it was published in 2016.
[00:13:40] Max Ritts: And between that period and the period of the work that I did with Karen, there was this huge transformation in what we’re calling conservation acoustics. Because of all this stuff happening with these big tech companies taking a new interest in the topic, because of the support from the governments, because of the, you know, widening crisis of climate change.
[00:13:58] Max Ritts: And so it was remarkable to me to kind of consider how much that project had changed, how much the kind of DIY citizen science thing had become a kind of institutional thing that was written about in Washington Post and blogged about on Forbes. And so that’s also what’s happening in this paper.
[00:14:18] Max Ritts: We’re trying to make sense of that transition, which is why there’s these ideas from Jason Moore and other kind of critical political economy people, kind of helping us make sense of the way in which this is a story of innovation through techno science as much as it is a story of responding to the crisis of nature.
[00:14:35] James Parker: That’s a great overview. There’s quite a lot that I’d like to dig into, but I wondered if, you know, I’ve read a lot of, you know, in your work and read a lot of the scientific papers. And I think it’s still quite unfamiliar to people, like the idea that certain sections of rainforest or tropical reefs and various other kind of ecosystems have like fairly densely networked with microphones. You know.
[00:15:07] James Parker: Perhaps it would just be worth explaining just a couple of kind of really famous, or famous or doesn’t it, they’re not famous, but that’s the point, but kind of a couple of flagship examples that you think demonstrate, you know, what’s going on with conservation acoustics and a little more detail. And then perhaps we can dig from there into questions of political economy, the state and big tech backing of these kinds of, you know, this sort of latest push in terms of the science.
[00:15:38] Max Ritts: Yeah, yeah, for sure. I mean, one example is from Australia. It’s the, you know, the Australian Acoustic Observatory, which is, I think, based or funded through the University of Queensland, if I’m correct, and it has a whole bunch of support from the government and it’s, you know, 360 sensor locations across the subcontinent.
[00:15:56] Max Ritts: Like it’s a huge, ambitious project to map different kinds of, you know, faltering ecosystems and to measure, you know, changes that could then be routed back to a kind of centralized, you know, repository for analysis and for response, even in certain cases. Most of these projects that were much smaller in scale, you know, it’ll be like a network of 12 to 15 microphones, you know, oftentimes with a university project that then takes up the gear, you know, afterward and, you know, is done with it.
[00:16:29] Max Ritts: But there’s also been, you know, certain projects that have lasted for longer than that, like the work that is done at the Cornell Laboratory for Ornithology, which I think just got renamed for a donor recently, so don’t quote me on the title, but that, you know, that work there, the Elephant Listening Project has been doing this kind of work for decades, specifically in Central Africa, where, again, these national parks become wired up with, you know, sensors, you know, Kruger National Park, Korup National Park, for extended periods.
[00:16:58] Max Ritts: And there’s an ability to monitor how species move and how they respond to emergent threats or, you know, again, poaching being a big one. And more and more, we’re finding it’s celebrated in accounts of how, you know, conservation can save the world by better understanding the soundscape, which, again, is a kind of really strange, you know, proposition, right? Because I mean, the problems continue, the deforestation continues, and there isn’t a lot of attention.
[00:17:25] Max Ritts: I think part of the concern is that this stuff is channeling energy away from the real political processes that are problematic, even though I totally support, you know, innovative, interesting science. But yeah, to get back to some more examples of it, I mean, underwater, you see a lot of examples of this kind of work now, especially in heavily shipped areas.
[00:17:44] Max Ritts: On the west coast of Vancouver, just off Vancouver, you have, you know, this thing called the Echo Project, which is listening for not only whales, but also shipping noise, and trying to figure out, you know, basically how to co-locate whales, would depend on, you know, sound for navigation and communication and noisy ships, which produce noises of an effective moving through the ocean, how to co-locate those things in a kind of heavily trafficked area, such that you can allow both whales and ships to passage through.
[00:18:12] Max Ritts: And so in that case, the network, the acoustical network becomes a kind of guardrail for industrial development, basically. And you can adjust the shipping lane through what’s called lateral displacement, you know, this or that degree to the left or the right of the proposed original lane to reduce… So that’s another example of one that I think really highlights the way in which this stuff is linked to political economies as much as it is to conservation ambitions.
[00:18:42] Max Ritts: And then I guess just finally, like, you know, there are lots of examples of this work in this specific case of new kinds of firms that have emerged, like new kinds of conservation outfits, like Rainforest Connection, which has worked in 20 different countries. It’s based, I think, in Texas now, but was founded in the Silicon Valley area by a guy named Topher White, or Pantera, which does tiger conservation and leopard conservation in Latin America.
[00:19:06] Max Ritts: And again, like you just go to these websites, and you’ll see they have, you know, very kind of heavily promoted reports about the work they’re doing with these networks in national parks or, you know, in places like northern Canada, where there are also really more hopeful versions of these kinds of things happening in collaboration with First Nations, you know, around management of caribou herds and just the biodiversity of the avian populations that are trying to migrate through the tar sands and things like that.
[00:19:33] Max Ritts: So there’s a bunch of different examples of it. But I think you’re right that, you know, even though these articles can be found on Washington Post and The Economist, which is still kind of surprising to me, it is not a huge thing. It is not a huge story in the world of conservation. But I think it’s growing. I think it’s getting more important. I think it got supercharged by the pandemic, because, of course, people were more online and more unable to go to certain places.
[00:19:56] Max Ritts: And because, of course, of mediated technologies, everyone’s much more connected. So you can listen to these rainforests on your phone now, right? You can get the Rainforest Connection app and listen into their networks on your smartphone. And as another example of how these things are becoming more normalized.
[00:20:11] James Parker: Yeah, great. I mean, the most recent example that I saw a bunch of online was this one, this Google one calling in our corals with this reef scientist, Steve Simpson. And that’s explicitly a citizen science project. The framing is that, you know, there’s this enormous, there’s all of these recordings of reef sounds, but there’s too much of it, right? Because now that it’s so easy and cheap to put, you know, hydrophones underwater and in the reefs, we’re collecting all this data, but we can’t analyze it all. And so what we need is people to annotate it.
[00:20:50] James Parker: And so it’s an attempt to sort of enroll citizen scientists via social media, basically, and word of mouth to listen to these recordings of the reef. Click when you hear a fish sound, you get trained in how to hear a fish sound, and you click, and then Google, in collaboration with Steve Simpson and his collaborators, will take this annotated audio and use it to train machine learning systems to identify the sounds of fish automatically.
[00:21:30] James Parker: And they say, ultimately, to produce what they call mega mixes to play back into the reef that will be, you know, specifically tuned and tailored healthy reef sound with the right amount of fish, you know, perfectly calibrated to the specific needs of that particular reef at that particular time. And so in that process to kind of retune the reef.
[00:21:53] James Parker: And so there’s an interesting connection of like, you know, big tech, state-sponsored science, the enrollment of citizens, possibly via a kind of greenwashing process, and all in response to this kind of argument that as, you know, microphones are more and more embedded in nature, the only possibility we have is to automate the process of listening and ultimately to automate the response and the form of intervention, you know?
[00:22:30] James Parker: And so I was thinking like, that the acoustic observatory project that you mentioned, the Australian Acoustic Observatory Project, they’re very clear about that because they’re not, you know, backed by big tech. As you said, they’re a sort of vicariously a state-funded project because the researchers have got a big Australia Research Council grant to do it.
[00:22:54] James Parker: But they’re just, they say, well, the problem is once you put in, you know, 360 audio sensors in the environment, and they’re recording 24-7, and they’re solar-powered so they can just keep taking on the data, if you wanna do monitoring at this scale, you don’t have any choice but to automate it. And so how do you do the automation?
[00:23:15] James Parker: Then suddenly big tech becomes the infrastructure, you know, provides the infrastructure via whatever it is, Amazon Web Services, but maybe it’s also, you know, a convolutional neural network trained on audio set, or whatever the case might be. So you quickly get these entanglements, and it’s sort of every, all roads lead towards automation, it seems like at the moment because of the scaling.
[00:23:42] Max Ritts: Yeah, it reminds me so much of Mark Andreevich’s constraint, right? He has the automation loop critique that this is all just automation begets more automation. And I guess to return to the sound moment, I mean, it does also kind of depend on this virtuous promise of sound, like I was saying before, right?
[00:24:01] Max Ritts: This kind of inherent human curiosity that is what enrolls us in these projects of citizen science that, you know, allows us to work as humans in the loop technologies effectively, to commit to free labor, is this kind of interest and curiosity in apparent beneficence of like sonic nature, that we wanna be a part of it.
[00:24:20] Max Ritts: And so we’re invited in, and we maybe are doing important things by listening and contributing, but we’re also doing a lot of free work that is ultimately training, you know, the algorithms better to do the work without us, or so we’re told, although I wonder about the energy costs of all this kind of thing down the line, like, you know, this all depends on power and electricity and all that kind of thing.
[00:24:39] Max Ritts: But I do think that there’s an interesting kind of combination of like cold economic calculations going on here, and then this kind of soft, you know, suggestiveness, this kind of interpolation of sound that enrolls people into these projects.
[00:24:54] Joel Stern: Yeah, because it’s hard to understand in that story that you’ve just told about the calling out our coral, James, what the actual role of the citizen science is in the operation of that technology and complex. I kind of did the exercise and I’m clicking to identify the fish, but it’s so abstract and so imprecise that it was a real struggle for me to understand how that data would be implemented.
[00:25:36] Joel Stern: It did feel, as Max sort of just put it, more like a kind of way of familiarizing people with the… kind of tech no solutionist kind of approach and this kind of innovation rather than doing real scientific work for the project.
[00:26:08] James Parker: But yeah. That reminds me, Joel, of a great line that I wanted to highlight in your piece on conservation acoustics, Max, where you say, for big tech, conservation acoustics provides opportunities to experiment on sonic data and with minimal legal oversight. So, to me, on one level, the value for a project like Calling in Our Corals or any of the others that you’ve mentioned that have big tech involvement would be, well, it’s sort of greenwashing. It’s like good advertorial.
[00:26:43] James Parker: Every time one of these projects comes out, they send the press release to everybody and there’s hundreds of these sort of micro pieces everywhere. And maybe it gets picked up by The New York Times or maybe it doesn’t. But you’re not going to get people to freely annotate your Alexa home, whatever, surveillance recordings. But you can do it if it’s fish, right?
[00:27:12] James Parker: And then you have these big data sets that people don’t feel the same kind of privacy concerns, that maybe don’t feel the set, the need for as much legal oversight, rightly or wrongly, probably wrongly in certain ways. And yeah, this is a this is just a great place if you’re a big tech company to start experimenting with audio analysis that where you’re not going to run the risk of regulation since that’s hotting up at the moment. So I thought that I’d never thought of that before. And I just I think it’s a great point.
[00:27:45] James Parker: And it draws directly what you are saying, Joel.
[00:27:50] Max Ritts: Yeah, I totally agree. And I think another person I know that you guys have covered in the series is Mara Mills, and she has that great phrase, assistive pretext, you know, but the resourcing of disability for techno science. And she’s dealing with disabled people in the history of the 20th century in America.
[00:28:06] Max Ritts: But you can see a similar logic with acoustically injured animals, right, with deafened whales or animals that might lose their acoustical niche, who we are listening for ostensibly, but actually helping to train the data sets for the companies that are using that information ultimately for their privacy and products.
[00:28:27] James Parker: Should we talk a little bit more about the sort of political economic critique? I mean, we sort of covered it a little, but, you know, I was struck. So, yeah, I’ve been doing some research myself on this topic of conservation acoustics, although that wasn’t the term I was using. And I came across your piece and it just absolutely stood out. First of all, there’s not many people writing at all about this that aren’t the scientists themselves.
[00:28:54] James Parker: So basically, if I was like trying to describe the state of the literature on bioacoustics, ecoacoustics and whatever, there’s a bunch of work being done by scientists, quite a lot of it. And as you say, it seems to have really sort of exploded in the last 10 years or so, partly because of, you know, cheapening equipment, smaller, cheaper sensors, but also because of the ability to automate the analysis where, you know, after the kind of image net and the kind of explosion of neural nets and stuff, we start to see them being used much more widely.
[00:29:26] James Parker: So there’s this explosion and it does get picked up in the papers. But there isn’t any, there’s not a critical, there’s not a body of critical writing about it. So you’ve worked with like Jennifer Gabris, for example, and there’s a couple of people writing about smart environments. But this is the only piece I’ve come across, I could be wrong, that is writing critically about digital sonic ecology or digital bioacoustics or what have you. And so I just wanted to dwell a little bit on that. Like, what do you see the critical intervention here as being?
[00:30:05] James Parker: What were you trying to do with the piece? I know we’ve touched on a few points and sort of where do you see the field going and what do you, like how urgent do you feel the need for intervention is? Because to me anyway, it really stood out that there’s not many people worrying about this and that maybe you’re one of them.
[00:30:26] Max Ritts: Yeah, I mean, I’m sure there I know that there are more folks asking these kinds of questions. I know that in Australia there’s.
[00:30:39] James Parker: Max, we lost you. We lost you there. So you said in Australia, there’s somebody. I was like, with bated breath waiting to hear who it was.
[00:30:49] Max Ritts: You probably know AM Kanngieser
[00:30:51] James Parker: Yeah, yeah.
[00:30:52] Max Ritts: And then there’s some folks at Queen’s University in Canada working on extinction ecologies. I should get the name and send it to you because it’s important that I reference this person’s work. Just blanking on it right now. But Jennifer, myself and Trishant Simlay are also working on a piece that kind of takes up some of these topics as well. I do think it’s coming in more and more with the smart environments question. And I think you’re seeing more work that looks at the specific question of sound in that broader story.
[00:31:21] Max Ritts: And it’s, you know, in geography, for example, there are people writing about digitalization and sound more and more. Like there’s a wonderful piece about policing sounds that was just published by Lally, Nick Lally, which is looking at not the ecology so much as the weaponization of the environment, of the acoustical environment, which is, you know, using, again, similar technologies, similar kinds of algorithms to do the work of annexing sound to political economies of surveillance.
[00:31:47] Max Ritts: So there is definitely, like, I think a lot of work in and around this topic now. But yeah, I mean, like, I think anybody who does, you know, a project is always surprised that it isn’t the most interesting thing in the world for everybody else. And I’m definitely like guilty of that. And I’m glad to know that you’re working on this stuff, too, because it’s fun to have conversations about it. And I do think it’s important. I do think, you know, a lot of kind of the politics of nature goes unnoticed in the world of sound.
[00:32:12] Max Ritts: And I think there’s a lot of reasons for that, including, you know, the apolitical nature of earlier discourses of sound, and just the ideas that we have around, you know, the ways in which politics work through nature, and how these smaller, more subtle sites aren’t often taken as seriously as they ought to be.
[00:32:31] James Parker: When you look at conservation acoustics, or the, and especially, you know, the front, the current frontier, which seems to be machine learning’s kind of pairing with digital bioacoustics or conservation acoustics. What, what is it that worries you most? Like, is it like, big tech is gonna, you know, just be monitoring environmental systems forever, and they’re going to be our point of access to the problem of ecology, you know, a kind of monopoly, scientific, technoscientific kind of monopoly kind of problem?
[00:33:15] James Parker: Is it a sort of an extractive exploitation, like, problem? Is it a distraction problem or a greenwashing problem? Probably the answer is all of the above. But like, yeah, when you when you look at conservation acoustics and its trajectory, what is it that you find most concerning or worrying or you think that we really should be applying our attention to?
[00:33:40] Max Ritts: Yeah. Before I answer that, I just want to do a quick, I hope you don’t mind me doing this. A quick kind of review of some of the folks I wanted to just briefly mention in this area. My last kind of response. I mean, I think it’s important that you know, Hannah Hunter’s work at Queen’s, Sandra Jasper at Humboldt University, John Pryor and Michael Gallagher are also doing work in this area, not necessarily in the political economy, but kind of asking critical questions around listening and conservation.
[00:34:06] Max Ritts: And I would just want to like reference those folks as well. And there’s many others. I’m just thinking of those two right now, or those four. So yeah, your question. I mean, I think definitely the automation question in all of this is concerning to me. The depoliticization of political processes, because we are more and more reliant on systems we don’t fully understand. And by we, I mean, you know, society in general. The automation of surveillance, right? Which Mike, excuse me, Mark Andreevich writes about.
[00:34:41] Max Ritts: Or the work that, you know, is being done around cloud ethics by Louise M. Moore and like, you know, facial recognition technologies. Just the way in which more and more of our natural world is coming under logics, which we have to trust implicitly because we can’t understand it or unpack them. I think there’s always been a level of like black boxing with science and conservation science and just trusting.
[00:35:05] Max Ritts: …differences in degree, and we just seem to be inching toward an extremely strange situation where entire ecologies are being managed, you know, through remote systems based on, you know, what are ultimately, you know, computationally intensive data sets or rather statistical prediction systems rooted in, you know, you know, evaluation data sets. Like, that’s what machine listening really is, right? It’s kind of a very, you know, ideological term. It’s not really as quote-unquote smart as we often attribute it to be.
[00:35:38] Max Ritts: It’s about, you know, pattern recognition based on evaluation data sets that may be very biased because, of course, we don’t know what the initial assumptions were with the people who collected the initial training data is necessarily. So there’s a lot we don’t know about what we’re seeding here in terms of the authority and the power to manage these kinds of processes. And, you know, I think that that’s one of the major concerns. Of course, that ties into the second thing, which is that they’re not going to be very effective.
[00:36:03] Max Ritts: That because we don’t know how they work, it’s hard to know how they’re going to do the job necessarily of predicting emergent threats and risks, or whether we agree with the definitions of what those things are, risks and threats, whether a poacher who may be a racialized, you know, farmer who’s looking to recover some sort of area that he had worked before is in fact the poacher that’s being declaimed as such because the gunshot that he’s using has been identified by the acoustical system as a threat to an animal with the idea that, you know, gunshots are indexical of poaching.
[00:36:35] Max Ritts: So there are just assumptions like that that are worked into these systems that I think are deeply unsettling. And we don’t know how effective they are for actually achieving conservation goals at this point.
[00:36:46] Max Ritts: There may be very well, you know, good, you know, future work on this area, I’m not denoting that, but I just don’t think that there’s a lot of evidence yet that like, the investment in these systems versus the investment in, you know, obviously, like dealing with the structural problems, or even dealing with the local problems is the way we should be going about things.
[00:37:05] James Parker: Right. I mean, to put it very bluntly, like, how much do we really need to know about how a reef sounds to know that the only real way we’re going to save the reefs is if we put this breaks on carbon production, really bloody fast. And, you know, on some level, it’s like, knowing precisely how screwed the planet is, is, or knowing computationally in real time, you know, is on some level missing the point in terms of the application of energies.
[00:37:42] James Parker: Is it you know, is that is that one of the that’s what I’m hearing from you, Max, on some level, I put it Yeah, no, totally, totally.
[00:37:51] Max Ritts: And I kind of go back and forth between just kind of being like, you know, this is all a distraction to know there’s actually important things going on here. And I think it’s a bit of both. Like, I think, you know, we do know enough. This is like, abundantly clear enough to like, you know, intervene more directly in the way that conservation happens today, just to take that one case example. So there is, in some ways, there’s a waste of resources going on here. But in other ways, like, you know, there are also new questions that are being asked.
[00:38:17] Max Ritts: And I don’t want to denounce the science, because the science can be really cool and interesting and important. And so that’s the tension.
[00:38:25] James Parker: Just so I think this is a good opportunity to segue, but I just want to get something on the table in case it’s not completely obvious. But I just want to clarify when you said, you know, you’re worried about the surveillance dimensions. I think that the kind of New York Times can style liberal concern with this stuff is that if you’ve got microphones in nature, then if you go into nature, they might hear what you’re going to say, right? And that the problem of surveillance is a problem of human surveillance.
[00:38:54] James Parker: I just want to be, I think that we’re, I think that you would, that’s not what you mean. And I certainly know what I would be concerned about here. But I just want to, when we talk about surveillance, we’re talking about knowledge production in relation to or via sound, you know, rather than, which sometimes involves, you know, spying on people or whatever. But that’s not really the concern here that this is like a Trojan horse for, you know, listening to people’s conversations.
[00:39:29] Max Ritts: Yeah, I think so. I mean, I do think that like the human dimension of surveillance actually is interesting here. I mean, a colleague has a really interesting story about how in, you know, these camera trapped forests in Northern India, you know, women no longer sing as they would have to alert predators of their presence, not because of the camera tracks hearing them, but because they confuse the camera traps sensing tools with ones that could hear them. So there is a kind of strange, you know, self-regulation going on here.
[00:39:55] James Parker: That is so interesting.
[00:39:57] Max Ritts: Yeah, in surveillance spaces. And I think, you know, that’s also something that happens in urban spaces now too, right?
[00:40:06] James Parker: Right, so it’s not that you specifically, oh, sorry, I think we got cut off there. Continue the thought.
[00:40:13] Max Ritts: Well, just that, you know, I think that the human dimensions of surveillance, I think, are concerning for that reason. So are the animal dimensions of surveillance for more kind of broadly, not necessarily philosophical, but like for animal rights issues, I think that there are real issues around how exploitation can be made possible through surveillance of animals and how bad ecological decisions can be made on the basis of exploiting relationships around animals. I mean, there’s this whole thing about, you know, data poaching right now.
[00:40:41] Max Ritts: An ecologist in Australia named David Lindenmeier wrote a piece about this a few years ago now, but I think it’s still a relevant issue where poachers can, like, you know, basically eavesdrop on the sensing tool to know where, you know, animals that would be worthwhile for poaching are located.
[00:40:59] James Parker: Right, because there’s a sort of a political thing playing out in the, who owns the data in relation to this stuff. So you see like a big push, like the Australian Acoustic Observatory, the whole thing is like open source, open source sharing. You know, there’s the ocean. I can’t remember what the, is it called, GLUBS, the global, the ocean, the library of ocean sounds that they want to create that is all kind of meant to be open source. Of course, you know, the problem with that is that anybody can access it.
[00:41:33] James Parker: So the political response is to close it down and to host it on a trustworthy, you know, gate kept server or something. But even the open source ones are all on Amazon Web Services and what have you. So big tech wins either way. But you could certainly see how you might end up seeing a kind of an enclosure, basically, of the sort of the commons of acoustic knowledge, like precisely on the kind of basis that you’re suggesting there, Max.
[00:42:08] Max Ritts: Yeah, totally.
[00:42:11] James Parker: I don’t know if this is, you know, I hope this is an appropriate question, but yeah, I’m still there. Are you here?
[00:42:20] Max Ritts: Yep.
[00:42:21] James Parker: So, you know, we’ve been talking about the sort of the critique of this, you know, field and maybe the kind of ambivalent politics of it and feeling, you know, wanting to honour the scientists who are doing interesting and innovative work at the same time as acknowledging that maybe the scientists aren’t always the best at thinking through the politics of what they’re doing and that there might be, you know, reasons for concern here.
[00:42:54] James Parker: And I just wanted, since she was your co-author, to ask you about Karen Bakker’s new book, which was sole author, The Sounds of Life, which I think is the most high profile book on acoustic ecology that I’ve, certainly digital acoustic ecology that I’ve come across. And it’s very recent. And yeah, it just seems like that is a kind of a work of advocacy and from sort of promoting a new scientific field that doesn’t so much align with the critical story that you’ve been telling, Max. And I just wondered if you, you know, you’ve worked with Karen.
[00:43:35] James Parker: I’m not really asking you to, like, I don’t know, take a position on Karen’s book, but just to, yeah. Or did you, what do you, how do you, what do you think about that book or that angle? Or, it’s a bit of a sensitive question, I realise.
[00:43:52] Max Ritts: No, I know. I mean, I’m glad you kind of brought it up, actually. I mean, I haven’t read the book, but we’ve had back and forth about this. And I like having back and forth with collaborators. And, you know, we disagree on the politics or the political valence of these things. And I think, you know, it’s cool to kind of know the other ways in which you can interpret these developments and not necessarily agree with them, but understand the logic. And I think for Karen, there’s a lot of optimism in the ideas of innovation and emergence.
[00:44:28] Max Ritts: I think that’s also why I hesitate from kind of being too denunciatory or too kind of classically Marxist about it and just seeing it all as a story of like appropriation and, you know, enclosure. I mean, I think that it’s just a mistake to read it that way. I mean, I think that we both agree that Donna Haraway is a kind of wonderful exemplar of that kind of tense position, right? I mean, where she is both pro-science and anti-capitalist. And I think it’s a tricky position to occupy.
[00:44:53] Max Ritts: And since I haven’t read Karen’s book, I can’t speak too much to her detailing of that issue. But I do think that, like, you know, there are people who are going to like have a different opinion on these things, and it’s really important that they be part of this conversation. And that’s certainly been, you know, the experience that we had in writing that conservation acoustics piece that, you know, it was a debate and certain arguments kind of won out. But they’re by no means like the kind of definitive arguments.
[00:45:16] Max Ritts: And I could be persuaded otherwise, too. You know, I’m kind of I’d like to be more hopeful about it in a way, too. I mean, I think, you know, it would be a shame if all this innovation and creativity in science was just being completely annexed by big tech. Yeah, I think I don’t think it is. But I think that’s, you know, that’s one of the problems with studying these things is you form a kind of a kind of hunch. And, you know, it’s important to kind of revisit it at times, too.
[00:45:47] James Parker: I wonder, is it worth saying, having a quick conversation about the ocean, specifically, Max, because most of your work, and certainly the way I first encountered your work was via your work on Ocean Sound, and you briefly mentioned the Echo Project before, and some of the stuff on whales, but yeah, I mean I guess I have two thoughts on this. First of all, it’s just like most of the world is ocean. And historically in machine research and machine listening, there have been always been people working on oceans there, like right from the beginning.
[00:46:26] James Parker: And I just wondered if you could, if you wanted to talk a little bit about your work on smart oceans, and how some of these political questions or techno-scientific and political questions play out in relation to the ocean specifically, or whether there even is a kind of ocean specificity here, or whether we should just be talking about like, you know, acoustic ecology and the forest, the ocean, you know, doesn’t matter too much.
[00:46:55] Max Ritts: Yeah, no, I mean, I think, you know, Melody Ju’s really wonderful book, Wild Blue Media, and I think that’s… analytical categories from like land to sea and not expect changes and transformations. And so I definitely think the ocean poses different analytical questions around sound and listening and surveillance and governance than does the rainforest or the urban space for that matter.
[00:47:23] Max Ritts: And the reason I kind of engage with marine spaces a lot is because again, a lot of the work that I’ve done is kind of been oriented around my time on the coast and the North Coast and working with communities that really hold marine spaces as central to their ways of life.
[00:47:37] Max Ritts: And the story that I tell in the last chapter of my book is really about the Smart Oceans Governance Project, which is, you know, a kind of a utopian idea of like, you know, real time regulation at a period of intense environmental risks and increased shipping and the faith and the ability of data systems to like coordinate responses to emergent risks is really the story here, you know, it’s part of a bigger story of sustainable marine development.
[00:48:03] Max Ritts: And this ONC based project or Ocean Networks Canada based project of Smart Oceans, you know, has been thoroughly supported by the state of Canada, by the government, and by all these tech entrepreneurs and by people who are interested in sort of seeing data and big data, firmly a part of marine governance.
[00:48:23] Max Ritts: And it’s also important that I think a lot of coastal communities have, you know, hope in this kind of a project, because they see, you know, not a lot of government support for other kinds of ways of doing governance, and also because they’re concerned about their lands and territories.
[00:48:37] Max Ritts: So writing a paper with Mike Simpson about the topic, kind of brought us to the realization that it is again, kind of too easy and analytical argument to kind of come down on this is pretty a story of enclosure, when communities are finding reasons to kind of take these new approaches seriously in the way they manage their territories. And sort of it sort of moves the site of politics to a different space, right, with like indigenous communities, and indigenous data sovereignty movements, more at the center of how a smart ocean should evolve.
[00:49:06] Max Ritts: And wanting to kind of create space for that in our arguments and in our analysis is something we were sort of struggling to do in that paper, but we hope hopefully did do. Because again, I think, you know, these systems pose a lot of risks, but they also are ones that people are having to use because they don’t have other options, but also maybe using because they find new ways of, you know, enacting their stewardship obligations through those systems in ways that the analysts like myself wouldn’t have, you know, anticipated.
[00:49:31] Max Ritts: And so kind of keeping open that door for like some surprises, I think really important as well.
[00:49:37] James Parker: So just to be clear, what is what you’re saying that indigenous communities, marine or sort of custodians who groups who understand themselves as custodians of marine environments are quite pro automation in certain forms? But.
[00:50:09] Max Ritts: Yeah, I think it’s tricky territory, especially as like a white settler to speak, of course, on behalf of other communities. But I will say that there’s been a long history of innovation and experimentation and creativity in the part of the world I’m talking about among First Nations and in taking technologies and, you know, enacting counter modernities or what one historian calls more additional economies that combine traditional and modern.
[00:50:36] Max Ritts: And I think that’s a story of indigeneity under modernity that, you know, to use new systems for the benefit of the nation is something that these communities will embark upon if they decide that they will be rewarded for it in the sense of better governance. So I wouldn’t say it’s automatic. Those can include smart technologies as well.
[00:51:02] James Parker: I guess the reason I’m interested is because of the scale problem, because you mentioned like data sovereignty. And I suppose what I see with digital bioacoustics in oceans and otherwise is like a real drive towards scale. It’s the Mark Andreevich point you mentioned before, right? That like automation begets automation and that the only players capable of feeding and really sort of governing that degree of automation and scaling are really, you know, the big tech companies and states and really mainly the big tech companies.
[00:51:48] James Parker: And so I wonder what the front lines are in terms of holding or pushing for data sovereignty in the face of that kind of a problem.
[00:52:07] Max Ritts: Yeah, these are really tricky questions. And I’ve spoken with critical indigenous studies scholars, really brilliant people about this topic. I think for some people, refusal and turning away entirely is the only way to go because there are so many risks and so many potential spaces of error effectively. But there are also communities that don’t have that choice anymore because of political economic circumstances, because of the desire of the developmentalist state to ship things down their territories.
[00:52:44] Max Ritts: And those communities, in many ways, have to engage with these systems because they’re kind of caught in these contradictions. And I think part of our job as kind of researcher allies is to look for spaces of potential opening in those conflicts while we’re mindful of, like you said, those huge, you know, prevailing logics that are really buffeting against any kind of hopefulness. But nevertheless, I think we have to look for them and not invest too much in those systems and other systems can be cultivated as well.
[00:53:13] Max Ritts: But also not abandoning those systems entirely because they are being utilized in those communities and because there is a capacity interest in using technology for local governance. So I guess it’s a very tricky position to think through. And I like to say it’s like, you know, it’s one that I think invites very situated kinds of findings because there are so many different particularities for different contexts. But. That makes sense.
[00:53:44] James Parker: We lost you just at the end there, Max, but, you know, obviously, you’ve you’ve you said situated quite a few times, which obviously, you know, a really important part of your work as being geographically situated. You know, obviously you’re interested in Parraway and the. The grounding of that situatedness in a feminist practice of science studies, you know, against the kind of machismo view from nowhere kind of logic that we often see in this kind of research. And I thought perhaps that’s an interesting note to end on because.
[00:54:26] James Parker: You know, part of the logic of automation itself is sort of on the one hand, like intense situatedness because it always claimed that the logic is always to be responsive to the particularities of the context, but on another level, like radical desituatedness, that the same neural network can be applied literally anywhere to anything in the world. To anything all of the time.
[00:54:56] James Parker: It’s a completely they couldn’t you’d you’d struggle to find a system that’s claiming to that is more comfortable with the kind of God’s eye view than a convolutional neural network that is can be applied to birdsong, but it can also be applied to whales and the sound of typing and gunshots and, you know, everything right there. A what, you know.
[00:55:25] James Parker: a domain agnostic, it can just do what it can be applied anywhere. And that’s the like a radical desituatedness. So it seems like on some level, the. Yeah, like, I don’t know, I just I guess I’m just hearing that from what you’re saying, that like the political intervention here is ultimately about a kind of situatedness. And insofar as there’s a critique of science, maybe that technoscience, it’s really along those kinds of lines.
[00:55:56] Max Ritts: Yeah, I’m going to just try to riff on this a little bit. It’s late here, but I think this is, you know, really interesting stuff to think about. I don’t know if it’s situatedness for me. I think that you’re right. You’re totally right. I’ve been using that word a lot. I’ve been using the North Coast as a kind of figure a lot in this conversation. But I guess when I think back to the book project, there’s two other words that kind of come to mind that I think are also about what I’m thinking about here is that kind of politics of this stuff.
[00:56:22] Max Ritts: One of them is is limits. And, you know, a lot of work has been written about ethnographic limits. Audra Simpson’s term in the context of, you know, indigenous refusals and so forth. And I think, you know, that word can be too easily used to kind of designate things that are off limits when really what what Simpson’s talking about is the community’s right to self-representation.
[00:56:45] Max Ritts: And I think, you know, that as a kind of principle is something that I think we need to insist upon in the context of machine learning and big data, that what those systems do is try to represent people. It’s a kind of one of the oldest stories of colonialism, right? It’s like, you know, the way in which people can be counted and assembled and to insist upon limits is to insist upon, you know, community’s right to self-representation.
[00:57:08] Max Ritts: And the other thing that I think about in the context of the North Coast is this idea of possession, because a lot of, you know, the history of anthropology and salvage anthropology begins in this part of the world. Right. Like Franz Boas wrote that piece on alternating sounds in 1898, I believe, which is, in my opinion, maybe the first ever kind of modern piece of sound studies.
[00:57:28] Max Ritts: And it’s all about him as an anthropologist trying to make sense of these Simshian words that he had to kind of possess within his linguistic framework to kind of understand them. And you can kind of trace a line between salvage anthropology and the kind of possession of like artifacts and sounds and music and songs and later on and other things as well.
[00:57:46] Max Ritts: And, you know, modern, you know, kind of reconciliation discourse that happens in both Australia and Canada as a way of kind of possessing things for other people, because you’re better at taking care of them. And that’s, of course, the story of museums and is the story of our practice as much as it is a story of state policy. And so I think one of the other political kind of questions is how to kind of enact a discourse of counter possession, of not needing to possess things to appreciate.
[00:58:14] Max Ritts: On bioacoustics, you know, and not needing to like possess massive datasets of animal sounds to like have some claim over the ability to manage them or live next to them even better than manage them. So I guess the question of possession is also one that I’m interested in and that I think about through the specific histories of the North Coast, because, again, that salvage logic is such a big part of the story of colonialism. And also, I think, sound in that in that region.
[00:58:40] James Parker: Amazing. That’s a that’s a great note to end on, I think, unless you want to raise anything else or table anything else, Max.
[00:58:52] Max Ritts: No, I mean, I, you know, like I said, it would be great to, you know, have more conversations. I really appreciated the questions. I hope I I hope I was giving you giving you guys some coherent answers. But otherwise, good.
Santiago Rentiera transcript
James Parker: So thanks so much for joining us, Santiago. Could you maybe begin by introducing yourself just briefly, however feels right to you?
Santiago Rentiera: Well, yeah, I’m from Mexico, Mexico City, and I’m living in one of the most isolated capital cities in the world. And I also think it’s a great place to be in Australia, despite of the time differences and all the logistics that imply flying in and out. But I got a scholarship at the University of Western Australia. This is an ARC scholarship on a project that is researching the cultures of automation. My research is mainly concerned with how automation is impacting the way that we manipulate and standardize representations of non-human sound.
Santiago Rentiera: So mainly I am using as a case study the Australian magpie sounds because I think it’s a super cool bird that can do so much things that I thought like birds couldn’t do before. So they are very social. They are also complex vocalizers. They can mimic. And so this inspiration of biology is a guide in my intellectual journey. Besides that, I have also been researching the intellectual history of listening and how this can be mapped to the scientific methods in bioacoustics.
James Parker: So, I mean, that’s all sounds amazing. I kind of wanted to, I’m tempted to just say, let’s talk about magpies for a bit, but we’ll get to that. We’ll get to that. I mean, what’s your, like, can we, can you tell us a little bit about how you end up in that place, you know, working on this? Like what’s your background intellectually, institutionally? I mean, are you, you know, an artist first and foremost, or, you know, a computer scientist?
James Parker: What’s the sort of the training or the institutional formation that like ends you up here rather than somewhere else?
Santiago Rentiera: Yeah, I think it’s a little bit funny because I’m like this kind of multiple hat type person or maybe a no hat. So I think, well, my starting point was music because my bachelor’s degree is in music. And also due to the influence of my family, because I come from a family of musicians, formerly trained musicians. So I think I have like this musical enrichment in my childhood. So that influenced my artistic perspective.
Santiago Rentiera: But I think curiously as my parents wanted me to be like kind of this concert musician, I ended up in the sciences and now doing something that is more in between. So I did music and production engineering and then moved to the computer science field to a master’s in computational science. in me of listening to birds. One of them is from the UCLA and the other one now rests in peace. But I am very grateful to his drive and impulse in the study of birds.
Santiago Rentiera: So I think that’s one of the main reasons I ended up doing birds because before my master’s, I was thinking about doing something more like kind of music interface design or this creative computing field. And with the master’s I discovered that there was this field of the use of algorithmic techniques and sonic methods to study the communication of birds.
Santiago Rentiera: And yeah, that’s where I came up with this model with the artificial neural network model, which is a few short model capable of dealing with various small data sets, which is one of the common problems with some of the species that are not, don’t have labels or are very few recordings that are actually labeled, have labels. So yeah, so that’s how I ended up doing the cross-disciplinary bioacoustics and computer science.
James Parker: Could I just rewind back a little bit to that, your supervisors or that context? So was it that you arrived in a computer science department with a musical background and you just so happened to find that in this computer science department, there were already people who are working effectively on machine listening, whether or not they were calling it that. Is it, is it, there was a, yeah, that was sort of something that was sort of sitting there ready and waiting and you were kind of inducted into? Is that what I’ve understood? Is that right?
Santiago Rentiera: Well, I think I simplified it a bit because, yeah, actually to get into that master’s, it was kind of a bit of a journey in kind of talking to people because I was coming from music. So when I kind of enter into the computer science, I was taking us or auditing classes during my bachelor’s in different topics of computer science. So then I kind of had to gain the trust of the people, you know, I know what I’m doing. So that’s how I met the computer science researchers Tecnologico de Monterrey. This is the university. I did my master’s in Mexico.
Santiago Rentiera: And one of them invited me to apply to this master’s program funded by the ARC of Mexico or the equivalent. And there, two of my supervisors had this long ongoing project connected to the UCLA research. So the UCLA researcher is Charles Taylor. So he started this group of Charles lab, the study of birdsong and using sonic methods to develop techniques to map and understand sequences.
Santiago Rentiera: And then Edgar Vallejo, who was my supervisory in the master’s was the only one I think at that time in the faculty that had this project crossing biology, acoustics and machine learning.
Santiago Rentiera: So, yeah, it was pretty unique. I wasn’t even expecting that, because you apply to the program but you don’t have, you don’t choose your supervisor until you kind of submit your proposal, kind of like six months after. So, yeah, and I didn’t know, I didn’t know Edgar before.
James Parker: So what sort of year was that?
Santiago Rentiera: I think I finished master’s around 2019..
James Parker: Okay, so this is, this is all post sort of deep, deep learning, right? You know, this is like quite an important period in the history of machine listening where suddenly everybody’s transitioning across to machine learning techniques sort of at en masse. That’s precisely the moment you’re doing it, is that right?
Santiago Rentiera: Yeah, I think the well, I wasn’t very aware of the field itself because everything was machine vision and even people in the department treated classification of of bird sounds as a problem of machine vision, because they pretty much turned the sound into spectrograms and that was like, oh well, you just have to just read it as an image. They actually didn’t, didn’t have listening skills, so it wasn’t listening at all. For me it was like just doing this data analytics.
Santiago Rentiera: I think the moment when I began kind of studying that historical progression, when it became listening or when it stopped being listening, was until I kind of wrote my PhD proposal where I had to actually make that argument that there’s a gap in knowledge and there’s this transition from sonic methods and notations to spectrograms which no longer require, you know, certain listening skills.
James Parker: I do want to move on, but can I ask just one more question about that period, if that’s okay, because I’m really, I’m just really interested in, I guess, what it felt like or what the understanding of the institution was around that kind of work did like. Was it that you guys understood yourself as doing something extremely marginal and esoteric at that time and now it’s like less so? Or were you, you know, was this an introduction into an enormous field of bioacoustics? You know, or you know, you mentioned, you know the changing techniques.
James Parker: But I just kind of interested to know, yeah, when a field feels new or part of a longer history or, yeah, how it felt to be in the field at that particular time. Because I’ve just the context for that question is is we’ve been trying to sort of track how machine listening fits in relation to a broader, you know, field of bioacoustics.
James Parker: Basically, like you know, and one thread is that it seems like people in very early days of machine listening, were interested in, for example, whale song in particular, and I just it’s just never been quite clear to me as to like, is you know, was birds been there all along too? Or is this, you know, is this new thing that really only emerges in seriousness, you know, in the last five, six years? Or you know, how did it? How did you understand the work that you were doing at that time as a collective or working in the university together?
James Parker: Does that make sense?
Santiago Rentiera: Yeah, yeah, well, I think the well. The university had three areas of computational science. One was kind of like operations research, management, data science, which was more like kind of using the methods for business. The second one was kind of adaptive systems and kind of a more like bio-inspired techniques. And I think the last one was like just general machine learning, which was mostly dedicated to tasks, tasks of vision and anomaly detection.
Santiago Rentiera: So the one in which my group fit was the bio-inspired ones and there were a bunch of researchers there working or collaborating with the Medical institute in problems like kind of medical machine learning by detection of cancer or other kind of diagnosis techniques that relied on big data. And there was also a very interesting group doing bioinformatics, which was, kind of thing, one of the most solid groups that had.
Santiago Rentiera: This group well, division had and they were pretty much doing sequence analysis, alignment, a bunch of different techniques that I think are kind of generalizable to the study of other types of sequences, like linguistic sequences or sequences of sounds. So yeah, I was the only one, with Edgar and also his other students, doing this approach to biology using computational methods. But I think machine listening was in the core, the core component, they, they pretty much use listening as a medium, or extracting information that wasn’t available to, to the other sense by the cameras or due to the occlusion of, of some of the birds, or some of the birds are too small, or, you know, all those challenges.
James Parker: Should we turn to magpies? You know, could you tell us a little bit about the magpies project that there’s a lot to cover? Yeah, I mean, I don’t know, I was really interested when you said, at the beginning that magpies are particularly interesting, or I think you might have even said Australian magpies. And I’m aware that you’re Australian magpies, but I don’t really know much about like, the broader magpie world, you know, so I don’t know where you think the right place to begin is in terms of describing the project.
James Parker: Is it like, you know, flesh out that point about magpies in relation to other birds, or, you know, wherever you’d like to begin, I just, yeah, just want to talk about that project, really.
Santiago Rentiera: Yeah, I think, well, my first encounter with magpies was a friend advising me not to feed them when I arrived to Australia, and then followed by this anecdote of a friend of a cousin or whatever that got souped, and kind of ended up in hospital. So it was like, yeah, these birds are evil. And that was kind of my first image. But I have haven’t been souped yet. So, so I don’t fear the magpies. And I think actually played, I made a natural difference when I actually encountered the project, which was pretty much chance because I arrived here to Perth.
Santiago Rentiera: And then I met Amanda Ridley. She has been research and doing behavioral ecology, observation of different animals and birds in the wild. And one of the projects was magpie research project, and they are interested in the comparative behavior of magpies because is one of the one of the birds that they breed cooperatively. So they don’t necessarily split like in pairs or couples and take care of their own kind of keen, they gather in groups that might not be necessarily keen. They take care of each other, and they do the breeding.
Santiago Rentiera: So that’s one of the aspects that motivates his research, in their case. And when I first met Amanda, she was interested in using machine listening to observe and extract meaning from all these recordings that couldn’t be manually screen or listen in order to give support some of the hypothesis that they had, which is that the magpies can combine different types of calls, and then compose meaning that helps to regulate the behavior of groups.
Santiago Rentiera: So this kind of group complexity is in a way the pressure that creates the need or need of a complex communication system, which is a use of call. So for instance, you have an alarm call and a recruitment call. And this combined them, this could be like a mobbing call. So, or you have a call for a ground animal and I call for alarm, and that this could be like someone is running on the ground. So those are, I think, studies that they are interested in.
James Parker: Yeah, can I jump in there? So it sounds like you’re saying that there was already an archive of audio that had been collected by, I guess, were they even bioacousticians? Or they’re just people studying magpies for whom audio was one obvious kind of way of studying them? And if that’s the case, like, I mean, how are the recordings being collected? I’ve done some reading that like shows that it seems like passive acoustic monitoring is quite a kind of common phrase now.
James Parker: And, you know, there’s lots of like pictures of these low powered microphones attached to trees. And, you know, they’re meant to be sort of discrete, and so on and sort of blend into the environment. So I’m just wondering, you know, what was the form of data collection that was going on and into which you entered?
Santiago Rentiera: Yeah, I didn’t do the data collection. I received the recordings and a set of annotations that well like, I cannot well disclose, but because the data set is meant to be for private use, at least for now. But the way that the some of them were recorded wasn’t necessarily passive, because they they have to have it through the Mac Pis and they have to create a community and they have to ring the Mac Pis so they can identify which individual is singing, because it’s relevant to know if the individual, the repertoire of calls- is changing across time.
Santiago Rentiera: So they do the longitudinal studies and also is relevant to see how they interact between each other. So sometimes they respond to the calls and then you can associate the recordings and see if they respond to the same type of call with a another familiar call and so on.
Santiago Rentiera: So, yeah, I wouldn’t say it’s, it’s passive, and previous previous studies of the same group have also developed their own communities and I between them in order to perform other types of tests, like cognitive testing, where the Mac Pi has to kind of decide between two types of food or two types of stimuli, you know, associate the colors and so on.
James Parker: So yeah, and so you arrived with this archive of recordings and you’re told- not being in magpie expert yourself- such and such a recording is an example of a magpie doing or saying, saying, quote, unquote. You know this. Like there’s a- I think you said a ground, did you say a ground predator or something? Yes, sir, so there’s a there, just to be clear. There’s like a whole kind of acoustomology of magpie calls and there are people who… I don’t mean this like cynically, but who claim to know what, what, with quite a large degree of precision, what magpies are trying to communicate through their vocalizations.
Santiago Rentiera: Yeah, I think that I should have made a bit more precise, but there’s, there’s a big discussion on that, on the meaning right there, the meaning on the causality of the calls, and you know the radical interpretation, because we are not magpie, so we are just collecting this lexicon or a bunch of sounds, and also the problem of segmenting them and defining what is the sound atom, because there are no phonemes. So that’s that’s a big issue.
Santiago Rentiera: And, yeah, they don’t claim that they magpies, really mean that they just associate some behavioral observations, do the sounds that are produced and the context, so it could mean a bunch of different things. That’s why they are trying to use the machine learning methods in order to assess if there’s some significance to that kind of correlation that they are making or that the observations cannot be cannot justify the hypothesis of those calls meaning this that from a behavioral point of view.
James Parker: So this is super interesting. So I mean obviously that’s like a big question but like a lot is at stake in it and a really important for thinking about the role of machine listening in relation to this. Are you saying that machine listening is being invited to like, confirm or disprove existing theories about magpie calls?
Santiago Rentiera: It was.
James Parker: Was that the invitation? Like we think XYZ, would you be able to sort of run the data and produce some kind of response to the existing theory? Is that, is that? Was that what you’re saying?
Santiago Rentiera: But the first year, the first invitation, was that first I wanted to experiment with the archive, didn’t necessarily want to do kind of scientific research with the archive. So my interest was: let’s use this archive for something, and that’s where I probably later we’ll talk more about the- an archival use, which for me is an archival use because I I completely disregard the, the scientific part of that. So but they are they.
Santiago Rentiera: The focus of the research is that they study the call combinations and they have been using different computational methods in order to give support or empirical support to the hypothesis of: do magpies combine sounds meaningfully or what the combinations mean, or like sound sonic variations that have kind of behavioral correlations or cues?
James Parker: Yeah, I mean it, it’s. It sounds like it’s a really important segue into your work, because your work, which we’ll get to in a second, is a kind of rejection or critique of that, or at least a bypassing of that as a kind of as the framing question, it seems like. I mean, to give another example, I read a paper recently about, I can’t remember where the sheepdogs were from, but there was a group of scientists trying to train a neural net to recognize sheepdog vocalizations. And so they had the sheepdogs like in a sort of an agitated state and a friendly state. So there was like seven different sheepdog states. And then they were trying to, yeah, like they made all these recordings.
James Parker: And then the idea was that they would be able to have automatic identification of like sheepdog vocalizations. And in the paper that I came across this from, the idea was that this would be used to produce sheepdog vocal synthesis in order to be able to communicate with sheepdogs. And then that was being used to make an argument that we could in principle use animal vocal synthesis to communicate with any animals. And we would do this in the context of the quote unquote fight against climate change.
James Parker: And so there was like a sort of a cascading logic being unfurled where you begin with, well, I think I know that that dog sounds a bit angry, to well, we can probably work out what dogs mean. And if we can work out what they mean, we can work out how to speak to them. And then we can communicate with them at scale through automated and embedded like microphone and network systems throughout the environment. And so, yeah, that was a bit of a wild paper to come across.
James Parker: And it sounds a little bit like you might be, have a similar kind of critical impulse or something towards that kind of literature or that search for meaning or the presumption that there’s a kind of meaning there and the kind of logics that that, or the pathways that that kind of thinking takes you down.
James Parker: So it’s a big question, the question of whether animals mean or communicate, but it sounds like you have a position and that your work is kind of inspired by or motivated by thinking about that in some level, or perhaps I have it completely wrong and there’s a totally different motivation for your work, but maybe let’s segue into talking about your work at the very least.
Santiago Rentiera: But you want me to give more an opinion on that or?
James Parker: I don’t know, like if that’s a factor, it sounds like you have an opinion, I could be wrong.
Santiago Rentiera: Oh yeah, well, I think I sent the last email that I sent was like this question that I keep getting since my master’s thesis, which is like, if we could develop this big data translator of sounds, of animals, and then speak to the animals, and this is kind of like the Solomon’s ring or Solomon’s seal power.
Santiago Rentiera: I was going to write a paper on that and then I didn’t, which is like this kind of myth of speaking to the animals and also thinking that animals do these assemblies and have their politics, the conference of the birds and all these myths, and how these myths are kind of re-enacted in the computational scene, and used as narratives and arguments to convince that we can gather more data and ultimately translate the animals.
Santiago Rentiera: So I think there are nuanced views on the bio-semiotics of animals and how we can talk about to a certain extent of indexicality and perhaps not symbols, but there’s some indexicality there. There are also methodological, I think, limitations on saying that animals have language and the way that language is actually studied by linguists is a different kind of thing from what the animal communication people are doing, even if there’s this field in between called animal linguistics, which is where all these compositionality hypotheses are happening.
Santiago Rentiera: But I think you can see the same in the study of DNA bioinformatics where all these, the letters of the genes were combined to mean something. So I think the narrative is very broad and the extent of meaning and what we mean by meaning is something like very vague that I don’t think I’m engaging directly because I’m not doing philosophy of meaning.
Santiago Rentiera: So for me, I think what I want to do with the work is stage kind of these encounters with the machine that can synthesize sound or that can listen and kind of question this idea that we will be able to ultimately translate something and also open the notion of the archive of these collections that don’t have a clear index because they, or the index is completely arbitrary because you know, the encyclopedic kind of drive of classifying things didn’t succeed.
Santiago Rentiera: And now we are seeing that we have these models that are being fed uncurated data and creating all these kind of outrageous responses from the public, so.
Santiago Rentiera: That’s why I think archives are interesting. I would like to engage more with the archive also as a way of accessing collections, information, and dispel this myth that the AI can understand or mean anything, which I think is a broader goal than what the magpie research is about.
James Parker: Start with magpies, and then you take down AGI. Exactly.
Santiago Rentiera: The magpie is actually a fascinating animal that I think should be respected on its own. Sometimes I have these conversations with the AI people that they are like, we’ll have bird-level intelligence, and it’s like, what the hell do you mean with that? That’s why you need to be in a biology department, and then they know that they won’t be able to classify all the calls because they keep inventing new calls. Some of the birds can mimic other birds, and they can just create repertories endlessly.
Santiago Rentiera: They also sing, and they might be singing for pleasure, so there’s no ultimate evolutionary explanation to know. It’s a big mystery, and I think one needs to be a little bit more humble in having these universal explanations of things.
Joel Stern: Santiago, can you just say a little bit more about inventing and composing new calls and singing for pleasure? Because I’m just thinking, obviously, your training as a musician comes into play when making a claim like that, potentially, but also the resistance to indexing and classifying something that is invented and possibly has an aesthetic dimension rather than can be purely semiotic comes into play there, too.
Santiago Rentiera: Yes, surely. I think I’m perhaps abusing of my aesthetic appreciation of birds. I don’t think I necessarily appreciate them because they intend to be appreciated as bird musicians, but I think it’s an interesting mode of thinking. Why do we create music? There’s this evolutionary argument that music play this social cohesion role, and then one can translate this into the birds that they are bonding with sound, but that’s also the argument that when we sing or when we are playing for ourselves, it’s like this loop and entrainment with our own voice.
Santiago Rentiera: So I think that it doesn’t have to have any meaning, like the music. It’s like music doesn’t have to point to something beyond music. Music can be something like sound in itself, but then there’s also music as communication. So I guess I resist to commit to one idea. And also, you mentioned that this resistance to classify and index things. I think that also applies to music. It’s like when we have all these classifications of styles and genres and what Spotify is doing.
Santiago Rentiera: It’s similar to what to some extent libraries like the Cornell Bird Archive is doing, kind of try to archive and classify all the sounds and then correlate them. So yeah, my concept of an archive is resisting that universal knowledge approach.
James Parker: So let’s talk about the work then. I don’t know how you think of it. Is it an art project or a form of… I don’t know, maybe it doesn’t matter how you think of it, but perhaps you could describe a little bit about the project, however you feel is right. Oh, yeah.
Santiago Rentiera: Well, the project consists of one part on the intellectual history of listening and body acoustics. And I trace a history of listening techniques and also reproduction techniques like whistling. There are a bunch of interesting papers of how imitation whistling was used artistically, but also as a way to record the songs of the birds of the environment. And also used by hunters to call the birds and then attract them and then hunt them. So in a way that was inherited by the ornithologists when they wanted to attract a bird in order to observe it or potentially to capture the bird and then kill the bird in order to study the organs of the bird. Then you had that kind of listening skill but also sound production skills. So that’s one part that kind of line of history and how now there’s this embodiment of listening and also vocal production.
Santiago Rentiera: And the other aspect is how you can use these same techniques to create an experiment with the limitations or the boundaries of archives, of what cannot be recorded in the archive, what escapes recording in terms of meaning and in terms of also expressivity. So one of the ideas is the vocal puppetry metaphor which is about just training this machine learning technique to synthesize sounds of a bird. But instead of using it to reconstruct the sounds, using it to kind of match vocal templates like whistles to the closest sound in the archive.
Santiago Rentiera: So it’s a way of information retrieval that instead of using a verbal index, like say tags, words, like in a search engine, you are using something non-verbal.
James Parker: Just to be clear, I would sort of hum a melody or whistle a tune or sing Bohemian Rhapsody and the algorithm would produce a kind of a form of mimicry or something that is in magpie, quote unquote, but I don’t know exactly what I mean when I say in magpie, like the form of index or something that’s being produced. Cause it’s synthetic. So it’s not a indexing, is it? It’s not like a matching, like here’s a particular piece of audio that is being pulled out and then paired or replaced. It’s a new, never previously existed magpie sound. Is that right?
Santiago Rentiera: Yeah, I think, yeah, well, that’s kind of the idea. I wouldn’t say it’s translating anything, but maybe I could play with that in an artistic way, like absurd way. But yeah, what is happening there is like, they call it timbral transfer sometimes, or they call it reconstruction. Like with this model, the variational autoencoders, for instance, you pretty much approximate a function that can reconstruct the input.
Santiago Rentiera: So it’s pretty much like compression, but instead of reconstructing the input that was compressed, you are using another sound to query a chunk of this compressed representation thing and then retrieve it. And that’s what creates this kind of mismatch between what the device was designed to do and the meaning that we attach to it, that we think that perhaps the easiest way of explaining it is like a translation, but it’s not a translation.
Santiago Rentiera: It’s some kind of decompression process that will give you the closest match in terms of the objective that the machine was trained on, because this was trained on the similarity objective, like the reconstruction has to be as close as possible as the original one. So that’s what the machine will do in the end. We’ll try to give you something that is similar using the bits and pieces of the compressed MacPy sounds.
James Parker: And what’s the sort of critical or artistic purchase for you? I mean, you’ve talked about the Anarchive and you’ve already sort of gestured at it, but what’s the, it’s a very crude to say, like payoff or point or something, but like, yeah, what is it that you think that this process of vocal puppetry is kind of showing or opening up or examining?
Santiago Rentiera: Yeah, I think, well, perhaps when people think about the Archive, they just mean a random collection or maybe the theorists will try to map it to Foucault, you know, this system of transformation of statements. And I think there’s also another word, or similar concept by Zielinski, which is Anarchive. I don’t think I’m using the same concept because this is more like a political undertone of memory, like the memories that were not officially called Archives.
Santiago Rentiera: by the institutions. So I think my archival approach is closer to the techniques, which is more media archeological than, I guess, historical. I think I’m more interested in experimenting with the accidents of the indexes and retrieving sounds using non-verbal inputs and non-logocentric approaches that don’t require this tagging, but also ways that we can also challenge the idea of the voice as something that has to be always human. We have this vast research on speech synthesis that is pretty much about human.
Santiago Rentiera: It’s all about humans, so it’s very narcissistic. And I think exploring what this non-human is of the voice using other archives that are not necessarily human ones can also perhaps add to this discourse on the non-normative voice of what is not original, but is generative in a way that is transcending the archive as something that encompasses everything. Yeah, I think I still have to think about what the artistic import itself, because I don’t want it just to be like this data translation experiment, right?
Santiago Rentiera: So, which is what we risk if we just use it as another filter, like, you know, you could sell this as a Mac by filter to TikTok and people will do silly things with it, but- Karaoke clubs. Yeah, exactly.
James Parker: We’ll take it up en masse.
Santiago Rentiera: So, yeah, so I think maybe also could be used to ways of listening other archives that don’t require us to input things like, or especially knowledge, and then access the birds and enjoy the sonic complexity of the bird using other means that are not words. So, yeah.
James Parker: It’s kind of anti-taxonomical, but like on quite a fundamental level, right? It’s like, it sounds like Joel’s point from before about, you know, the musical orientation or background is like really prominent here. Like, there’s some level that there’s a kind of a critical resistance to the kinds of things that all of the bioacoustics- But yeah, but it’s also returning the data set to something sort of sensory, like it’s like a sensory ethnography approach to this sort of, you know, bioacoustic data.
Joel Stern: And, you know, that’s one of the things we’ve been thinking about a lot too, you know, recent sort of projects in which we’ve been listening to data sets, that just how it is quite radical to actually as a human sit and listen to these sounds at some length when, you know, they have been amassed and accumulated and not necessarily with human listeners in mind.
Joel Stern: And what we kind of can understand and sort of experience through that sensory interface is something different from an indexing kind of interface, even if there’s some overlap and even if as human listeners, we’re sort of categorizing as we listen in certain ways, there are other things going on. But I was just thinking about that process of tone transfer because obviously, you know, Google with like magenta DDSP and have been applying this sort of technique as an effect for transferring from one instrument to another instrument.
Joel Stern: So it’s a sort of, and in your, I was listening to your SoundCloud examples and it was great because there was, you know, the whistle and then the magpie, the whistle and then the magpie, and then quickly it’s followed by the magpie beat box. So, yeah, I’m just wondering if you can say a little bit more about the kind of experimental sort of horizons of these sort of techniques, like what are some of the ways that you imagine applying these sorts of techniques in kind of non-indexical, non-taxonomical ways?
Santiago Rentiera: Yeah, I like how you frame it as it’s very sensory, almost auto-ethnographic approach of your own way of collecting bits and pieces. And yeah, it’s totally anti-taxonomical. I think I like to do more stuff that involves less classification, you know, even if that’s what I’m usually getting, like people that pay me to do stuff are like, I need to classify this.
Santiago Rentiera: Perhaps how I am most useful as a computer scientist, classifying and ordering things. But yeah, I think I find in this art accidents of the index something worth exploring. How could I, you know, take this onto the artistic expression or more like musical expression? I’ve been thinking about what other people have been doing with same technologies. I don’t think technically, I’m doing anything new. This has been done by Holly Herndon, I think, with the Spawn, which is a timbral transfer.
Santiago Rentiera: She did this with her own voice and I think actually did a weird kind of decentralized organization blockchain thing in order to license the reproduction of her own voice in this kind of Timbrel transfer way. But I don’t think I want to do something like that. I think that’s what I’m.
Santiago Rentiera: My contribution here is thinking about this in archival terms and how the notion of these big collections that are going to just get bigger and bigger and bigger, how to approach them in interesting ways that are not necessarily like oh, the computer knows something, or like this big Oracle, big tech Oracle thing, or like the universal index, like kind of the Aleph on Boris or from the library of Babel, that you have these indices of things that are no longer human readable.
Santiago Rentiera: So I think the way of challenging the anxiety of things no longer being readable or humanly indexable is at this point where the art can enter and then the accidental retrieval can be used as a way of creating new situations, the data kind of embodying the data, drawing the attention of people towards some species. That is, perhaps pour it down this big archive and nobody knows about that recording.
Santiago Rentiera: And I didn’t mention extinction, but I think it’s also relevant to think about what’s going to be extinct in the next years and these sounds that we won’t be able to experience without mediation. So, like there’s a very interesting case of the Huya: the only bird recording that we have is an imitation of a Maori elder. So we have this second order mimicry of a machine that mimicked the human, that mimicked the bird and maybe the bird was mimicking other sounds and that’s how the bird learned the song.
Santiago Rentiera: So yeah, the notion of extinction and second order extinction, like when the materials of the extinct bird go extinct or forgotten. I think that’s also another point that I would like to explore poetically, like disappearance and retrieval.
James Parker: That reminds me of another paper we were reading recently for our own project on the sort of bioacoustics meets machine listening, and maybe I could use this example to segue towards a sort of a broader conversation about this field, because you know you were saying before that the archives are going to kind of keep expanding and I’ve been like genuinely shocked by the- I can’t think of another word than like- imperialist impulse of certain bioacousticians or their sort of interface with machine learning researchers, because some of the papers just that I’ve been reading just sort of suggest the kind of the most encompassing listening, you know, data collection that you know the NSA would bulk.
James Parker: You know, almost that’s what I was like Snowden: we want to listen to everything all of the time in the entire biosphere in order to produce a kind of constant and perfect archive of all ecological sounds. That’s sort of what it sort of reads like sometimes. So that’s sort of where I’m heading. But in the context of that I was reading a paper about what was called acoustic enrichment. So this paper was about reef health and the idea was that reefs are extremely unhealthy.
James Parker: You know, the Great Barrier Reef, this was Australian research- but that if you, they were testing whether, if you could play health, the sounds of healthy reefs into reefs, that would improve the reef health and it- and it makes me. It reminds me of things I’ve heard occasionally in like conversations with architects or
James Parker: or, and so on, when they say, well, there’s no urban birds anymore, but we could sort of just play the sounds of urban birds and everyone would sort of feel happier. And they mean it very sincerely.
James Parker: And so it just touches on this sort of point of extinction that you were making before, whereby like in the name of ecology and perfecting, you know, healthy ecological systems, we have this kind of weird simulacrum of healthy soundscapes being constantly reproduced, you know, with extinct species and so on in order to supposedly improve ecology and reduce extinction. And then all of that depends on, you know, this archival impulse or data collection. So it’s just a bit sort of mind bending.
James Parker: Yeah, so I was just wondering if you had any comments on that, because it just seemed like the, what I’ve read in terms of acoustic enrichment sits very closely to what you were saying about, I don’t know, the poetics of extinction or something.
Santiago Rentiera: Yeah, yeah, totally. I think I probably had come across a paper before in the context, I think, of assisted evolution, which is this idea that we can design different infrastructures that support beings that need to evolve but can’t because of our own destruction and impact. So this one, yeah, I read that it was like a soundscape research thing going on there. And also I find very problematic the way that the soundscape health is assessed using these very abstract indices, right?
Santiago Rentiera: Like how do you even determine what is healthy with this very broad and incomplete picture, which will never be complete, right? Of what the animals are doing at the sonic level, right? So is it a matter of it being spectrally diverse or what? Right, so I think regarding the extinction question, yeah, about synthesizing or the simulacrum of the dead, right, which is like, I think I didn’t mention it, but I have this concept of sonic reanimation, which is very close to what they are doing with reefs.
Santiago Rentiera: And I probably maybe should change the name so people don’t think I’m actually trying to do that by reanimating extinct sounds and then just kind of greenwashing destruction. But yeah, I think synthesis of what has passed away is a topic and the copy also, how we use the copy as a way to alleviate decay. And maybe we also should think about how archives decay, right, like how they decay as not human memories, but in their own way, how we are no longer able to retrieve something because it’s lost.
Santiago Rentiera: Like this is kind of the anxiety of the Library of Babel and also the Book of Sand, which is another story by Borges. Like this book cannot, you cannot return to the same page anymore. So, and then you never know what was the original one because the original and the copy maybe just differ in one character and things like that.
Santiago Rentiera: So, thinking about extinction and the visibility of extinction in big data, I think that that’s a critical question for the concept you just mentioned, the planetary listening, how we’ll become perhaps more aware or less aware of extinction with the increasing recording, which they’re recording in the end is not kind of this God’s eye view because the placement of the microphones and the decisions that people have to make in order to kind of put them in certain places that kind of obscures the universality, which in the end is not universal.
Santiago Rentiera: It’s a very particular imperial view of someone that decided to put the microphone set where it’s not the full, it’s never the full picture. So, I think questioning that incompleteness of the archive and the things that are extinct or the awareness of extinction as a way of kind of extinction listening in the archive, I think that that’s relevant from an ecological and poetic point of view.
James Parker: Does it feel like, you know, this is a growth area? Does it feel like people are investing money or asking you to be involved in, you know, bigger projects? You know, the reason I’m asking this is because a number of the papers I’ve read are directly, for example, in response to the UN Sustainable Development Goals. And they have this sort of tone of, like, salvationism, like, and an urgency, which is, of course, appropriate to the climate emergency. But there’s a kind of, like, we’ve got to do this now.
James Parker: And there’s a lot of, it does seem, you know, you cite in your paper, this piece by, this book by Karen Backer. She’s just, like, ecstatically excited about, you know, the amazing and diverse and diffusion of planetary machine listening, basically. Like, all of the, every imaginable ecosystem is now being monitored and modulated and interacted with. And she’s so excited. And is that what it feels like to be working in the field now?
James Parker: You know, that, like, there’s energy being poured in, there’s money or, you know, is, for example, is there industrialization on the cards? Because one of the things I’m interested in is, like, could a system, there’s a lot of talk of, like, the green internet of things, right? But is there a version of this planetized machine listening that isn’t totally and utterly dependent on Amazon Web Services or centralized data power, basically, you know? So I’m wondering, like, what the kind of, yeah, feel is like within the field at the moment.
James Parker: Is it, does it feel like how it reads?
Santiago Rentiera: Well, I think that’s one of the reasons that I think I’ve been trying to see a way of AI as a concept. And I was very annoyed when I saw this letter and brilliant people saying it, like, you know, Joshua Benjus saying this letter that was written with hyperbolic language and trying to obscure other catastrophes that are there are perhaps more evident than the super intelligent machine doom fantasy.
Santiago Rentiera: And I think the way that this is being used to control technology and manipulate opinion, yeah, it’s so problematic, but I think listening is not in kind of under the limelight as the notion of intelligent and language is because perhaps language is more relatable to us humans.
Santiago Rentiera: So even if we can synthesize every possible sound or, you know, map and recognize all the event classes that are necessary for a perfect universal surveillance state even that wouldn’t create the same kind of doom associations as a intelligent machine, because, you know, listening by intelligence kind of experts or the artificial intelligence people is just treated as another modality, right? So it’s just more data.
Santiago Rentiera: So I think they abandoned the drive of the computational auditory scene recognition that was like from the 90s trying to model how the cochlear worked. And they totally abandoned that with the deep learning because it doesn’t work like a brain at all. It’s kind of this massive approximation, hyperparametric devices. So they no longer use that rhetoric.
Santiago Rentiera: They kind of call it sound event detection and sound or scene recognition which obviously have their own taxonomical issues like what I think you mentioned, what is a scene, how a scene is even defined or how do we kind of determine what are the relevant categories and there is no universal index. I think research that probably should be done there to counter the excitement and the wonder of backer could be like linking this to the way that the military industry complex is also funding the projects.
Santiago Rentiera: Like I remember reading from Douglas Kahn beautifully written book on how this infrastructure or international monitoring system that was used during the 60s to guarantee the compliance of the nuclear test ban treaty is now used by well researchers to observe and also to explore other effects of human action on the deep sea. So the deep sea is already being monitored. All these whale research is also an entry point into deeper areas of the sea and resource exploitation and mapping the planet.
Santiago Rentiera: Things that are not only purely scientific nor conservation efforts, but are kind of strategic and have a purpose in the battle space.
Santiago Rentiera: So I have to say that I’m actually preparing a chapter on this notion of planetary listening. And it will include some of these kind of links between conservation and how the mapping of the planet and geography are linked to the strategy. And it’s not only about the animals, right? The animals are kind of like the flagship kind of, you know, the argument is that we are trying to save the animals. But I think there’s more to say about that besides the conservation.
James Parker: I’d love to read it because I’m also writing something right now on the same topic. I don’t know. We’ve had a lot of your time already, Santiago. Is there anything else that, I don’t know, you wanted to talk about? Or I don’t know, questions for us or Sean or Joel, did you want to jump in at all on anything? It doesn’t seem like it. It’s great work, Santiago. Maybe we should wrap it up and say thank you and segue into a more informal conversation about how we might continue to work together or talk at least.
Joel Stern: Yeah, I’d like that. Thank you so much, Santiago. It’s been great to do this interview and to learn about your work in more depth and to, you know, understand more about sort of what is motivating the research and the artistic practice. And I think we should have a conversation about ways to collaborate, you know, sort of creatively and in other ways too, and around how to sort of engage with these collections of sounds and sort of do more experimental work with them.
Joel Stern: Because especially as we’ve said over the next few months, leading up to an exhibition in August, we will be developing what will probably be a multi-channel sound installation with the working title of Planetary Audition. And it will be touching on many of the themes that sort of have come up in the conversation between us amongst other things, but we’re sort of still trying to develop, you know, the material. So, yeah, I’m not sure what the best way to kind of foster a collaboration is, but what do you guys think, James, Sean?
James Parker: Sean is green, so.
Joel Stern: Go on, Sean.
[01:15:14] James Parker: I’m green. Your little box was green, so. And he’s green, yeah. Sorry, I can’t pick up from what you asked specifically, Joel.
Sean Dockray: No, I was thinking, I mean, there were just two things. Sorry, James. There were a couple of things that just stood out to me in the conversation that as being sort of, you know, that you identified Santiago too as the sort of artistically sort of rich, you know, or generative things. And, you know, was the accidents of retrieval and accidents of indexing. And, you know, one of the things I was wondering during the interview itself was a little bit like how those accidents actually come through or come up. I was I, I just couldn’t find a way to like, I didn’t want to interject. But a lot of that, like the instrument, the way that you kind of play with the archive, like to me, it’s like a form of play rather than query, right? Which I love.
Sean Dockray: But at the same time, like the in the SoundCloud, it’s, you know, like, the verbs that you were using in the conversation with James were like, transfer or translation, not necessarily that this is what it was doing. But those are the verbs that come to mind. But the one that came to my mind is like mirror, because you kind of like making a sound, and then it mirrors something back to you, right? And it’s a little imprecise, but it’s still this kind of relationship.
Sean Dockray: And so I was wondering, like, where the accidents, you know, like, if there’s other verbs that make different kinds of accidents, and I feel like the beatbox is a bit of a funny thing. But at the same time, that’s where it sort of like begins to become something else, too. So I was really interested in this question of the accident, not because of my own investment in it, but just because of like, the potential to kind of loosen up the kind of relationship of mirroring or transferring or translating to something else.
Sean Dockray: So yeah, just to riff on what Joel was saying about the accident a little bit. I wonder how, yeah, if we can like figure out ways, like strategies for that. I also love the reference to the Solomon’s Seal, like not something I’d thought of, but like, actually thinking, yeah, like some of the sort of historical theoretical kind of things that came up in the conversation, like would be quite, like, rich to, I think, think through in the work as it goes forward, too.
James Parker: Is that something you’ve written about already, the Solomon’s Seal stuff, Santiago?
Santiago Rentiera: Actually, it’s what I think was a failed paper, because the paper, the, originally, the one that you received, the magpie puppetry one, was going to be about the Solomon’s Seal. And I actually read Conrad Lorenz, a book that has a reference to Solomon’s Seal. So I wanted to go deeper into this kind of mythology of the animal communication. But then I realized that the word count wasn’t going to be enough. And I didn’t want to burn it by doing like this very short paper thing.
Santiago Rentiera: So I ended up doing something more in practice, because it was also kind of, we were prompted to write something about practice, so not so much about history and theory. So I decided to change it, but in the end, it also didn’t get accepted. So I still have that archived. There is perhaps a chapter that I will write at some point.
James Parker: I mean, so I don’t know if you’re noticing what we’re doing. I’m just going to say it. I think we’re, yeah, we are inviting you to collaborate in some way, potentially, and the degree of collaboration is open on this potential project.
Joel Stern: You know, and I was just thinking about the Solomon’s Seal stuff as well.
James Parker: It was just going to be that, like, for example, the moment you said Solomon’s Seal, I was like, this is incredible for the artwork. And we’re not just going to like borrow or steal your ideas, do a whole bit on Solomon’s Seal and just be like, cool, glad we spoke with you. But like instantly, I’m just thinking the Solomon’s Seal thing situates what’s presented always as novel and digital and so on within an imaginary of mastery that’s not unrelated from the imperialism point that we were making before.
James Parker: And I really want, it’s really that kind of, quote unquote, imaginary that, like, I’m most interested in the planetary, like, it’s just wild. I just can’t get over that. Like some of the things that they imagine doing, whether or not they’re, like, going to happen, they’re wild, they’re mad.
Joel Stern: But also, just in terms of, like, the artistic format we established with Afterwards, which was really, you know, a provisional and first attempt at working together in this way. And the works that we produce in the future may reproduce some of that form, but also change.
Joel Stern: But there was a sort of episodic quality to the work and a narrative quality where different sort of, where storytelling is sort of quite important and in which a kind of speculative imaginaries are quite important to those stories, you know, and the blending of fact and fiction and historical material and sort of fabulation. So, I think, you know, there might be a way, even though it’s been difficult to get it up as a scholarly text, there might be a way to rethink how to work with that in a more artistic context that makes a lot of sense.
James Parker: The other thing I was wondering, I don’t know if this is the right time to say it, but have you made a, can you do…
James Parker: Does your thing work in real time? The magpie mimicry?
Santiago Rentiera: Yeah, actually I’m trying to run installation in November for… I don’t know if you heard about the…
Joel Stern: The Cultures of Automation.
Santiago Rentiera: Yeah, the Cultures of Automation. I’m going to present something there and I’m trying to embed this algorithm in a real-time device called Bella. Well, it hasn’t worked because the library… There are some bits and pieces that have to be compiled for the specific architecture. But this model that I’m using, which is called Rape, real audio variational autoencoder, works in real time. It can be used in an interactive fashion and that could open a lot of possibilities of not only using the voice, but using other input signals that are non-audible.
Joel Stern: I’ve noticed that the real-time tone transfer, as a hardware device, that there’s quite a few people trying to…
James Parker: Because one of the things about the RMIT commission context is that it’s a very, very different space to the one that Afterwords was in. So we’ve been thinking a lot about the installation element and what visuals or artifacts or things might be present in the space. And it just strikes me that one thing would be potentially to have something like your vocals. I mean, I’m not trying to say that definitely would work. But I could imagine something that was sort of interactive in that way being present. It sort of depends a lot on what happens.
James Parker: But I don’t know if that’s something you’d be open to developing.
Santiago Rentiera: Yeah, I want to reflect on the concept and also build something that can actually be experienced by people not only in academia. So yeah, that fits perfectly.
Sean Dockray: It could be interactive, not with viewing public, but something else too. There’s interactivity and kind of real-time systems. I’m just kind of interjecting from the point… interactive artworks and seeing a few. Sometimes kind of like public interaction with an artwork is less interesting than other relationships that you can set up. So I’m not saying not to do it, but I’m saying that it could be that over time we develop more conceptually.
Joel Stern: There’ll also be numerous public programs associated with this exhibition. So there’s the possibility of having an installation artwork that has an exhibition form and then performances that activate that artwork in certain ways through interaction or, you know, some form of encounter. Awesome.
James Parker: All right.
Joel Stern: Let’s leave it at that.