Model Mayhem: Rogue Agents, Leaky Memory, and Why AI Is Weird
Full transcript
Andrew: I don't think anybody would know.
Sean: I think it'll just be me that notices.
Andrew: You're not going to spend hours color matching that one camera. Nope. Sorry about that. Just to make sure you don't have quite the yellowy cast that it's giving you. That's right.
Sean: So if you are that one person who watches that and you get annoyed, let me know and then maybe I'll spend that time.
Andrew: But no, that's not going through DaVinci Resolve. Spent it on color matching the vocals.
Sean: That's right. That's right. What did you do all week?
well, I was stuck color matching this 4K video. Ay, ay, ay. So nope, not doing that.
Anyway, we're back. But we are back. Yep.
Yes. And like, you know, again, polling, I've got, this is going to be a wandering one because I think, again, we've been doing lots of
Andrew: We're talking about what, weird AI.
Sean: Yeah, literally, I was like, you know what? You know what we should call today's episode? Just AI is weird.
Yeah. Like, because, and also an opportunity to revisit a lot of other bizarre, bizarre things because I feel like just, you know, this has been the summer of tech billionaire manifestos. It's been the summer of models being pulled, frozen, not allowed to be released, them hacking each other.
Escaping. Escaping. I mean, it has just been one weird.
It's a model mayhem, I think. It's been the summer of AI model mayhem. That's a better title.
I gotta type that down because that's a better title. But yeah, and of course, it hasn't gotten any less weird.
Andrew: No, no. And including But you realize, actually, this is the singularity. It's not about sort of getting to the point where technology takes over. It's getting to the point where technology is so much of a mess that you just cannot keep track of it.
Sean: That's right. And that's how you give in. It's like that petulant child that just wears you down.
Andrew: The great resignation. 2026, where we all just give up.
Sean: There's no AGI. There's no the year of agentic whatever. It's like the year of submission because we got so just bonkers, annoyed, and just overwhelmed by the nonsense. We just flat up gave up. Yeah, yeah.
Andrew: But we do have some AI weirdness to talk about that I don't think we've really dug into before.
Sean: So this will be fun. That too. And again, even just the other day, I had something like in deep, like a deep dive session, whatever you want to call it, something bizarre happened. And it was just, again, one of those moments that just made me sit up straight and just go, hmm, how do I feel about this? Because, you know, again, those things, that Trojan horse, like there was, so it was, it was inside, the call was coming from inside the house and I had no idea. And it was- It's inside your head. That's right. You're like, where, trace the call, where is it coming from? It's inside the, and it was one of those kind of moments where I was like, hmm, how do I feel about this?
So I think a lot to cover, a lot to get into, and I think this will just be a fun one. I say we give ourselves the ability to create mayhem ourselves and just wander all over the place and talk about some of the weird, wacky, wild stuff that's been taking place. So why don't we take a quick break and we'll just get into it.
How about that?
Andrew: Okay.
Announcer: Initializing waveform.
Modulating signals for transmission.
Exploring the possible, probable, and preferable futures.
Sean: Welcome to... Modem Futura. Okay. All right. Well, I'm Sean Leahy, joined by my co-host extraordinaire...
Andrew Maynard. And that means you're listening to Modem Futura, the show that explores the intersection of technology, society, and the possible, probable, and preferable futures, asking deep questions like, what will it mean to be human in the future? And that, again, I could say this every time.
And every week, we're less and less sure. And I could say the same thing. Every week, you're like, and it has never been more important to ask that question because every week, it seems like it's slipping away.
Andrew: Oh. So I've actually, I've got to start this with an anecdote of my own, which I think will feed into yours. And it's all about sort of what it means to be British in the future.
So one of the bizarre things that I've found working with Anthropic's Claude over the last 12 months or so is that it always assumes I'm British. Now, to understand this, you've got to realize I have memory of, there should be no indicators there whatsoever. Every time I open it, it's a fresh session.
So you think. So I think. But I've tried to weed everything out.
And yet, when I ask it to write something, it always writes it in British English. And if I put an Americanism in there, it corrects it for me and says, as a Brit, it should be like this. And I've actually asked it, how on earth do you know?
And there's part of me thinking, I'm offended, apart from the fact that I am British.
Sean: It's right behind you, Andrew. I know, I know.
Andrew: And it claims it just picks up on cues in my writing,
Sean: which I don't believe. Well, like, I mean, let me ask you this. Are you, do you consciously scrub out the extra you from color?
No, I write in American English. I haven't done for years, for decades. Because even I would wonder, yeah, I would wonder what it would do if I asked, because, well, it's gotten better now, but there was a time too, I would intentionally put the you in because of the audience and the internet, you know, because you're working in international, settings that even in international English, so much of it is, you know, it's the Queen's English.
And so they're expecting use in words like color and things like, you know, which is a lot of fun.
Andrew: See, I don't. And so when you get to the point where you're sort of working on a draft with Claude, it comes back and says, actually, you've spelled that wrong. It should be C-O-L-O-U-R, not color. Think, where do you get this from? And the only thing I can think of is,
Sean: is there something in my style of engaging with it that it pins down as being British? Okay. So I was going to save my little anecdote about it being inside the house, but I guess I could do that now.
Yeah. Because it might tie into this. Right. I know. That's what I was thinking.
So, and this also, for anyone who's interested of, you know, because I know there's a lot of, you know, podcast, you know, podcast friends out there that are like, how do you do it? Because whatever. We like, I like, we like exchanging that kind of information.
So every so often I do turn to whatever the latest, greatest tool is. And I throw it all the Modem Futura metrics and stats and all this kind of stuff. And I'm just, cause you get, we're a super small, we're as scrappy of an operation as you can get.
The next tier is just nothing. We're as close to the metal as you get, baby. This is how we do it.
But so I'm always curious, like, I'm like, is there, have I missed something? Is there something I can do that will help boost engage? All this kind of stuff.
So this is a normal process I do every couple of months or something, right? And in this particular time, I'm going back, and one of the things that came up was like, was like, there's the gap was, why are you not, it was literally, it was almost like accusatory. Right.
This is called to you. Yeah, this is Fable 5 and 5.1, because it was right when it was released.
Andrew: So the top tier models. Yep.
Sean: And so it was almost accusatory of like, why are you not uploading your transcripts to your host, you know, the, the, the CDN for the bike. And I was like, it was like, it was basically in my, in the, of course, now the way I read it was you idiot. Why aren't you doing this? You know? And, but so then, for then of course I'm going, well, first of all, Hey, I give you the transcript from our recorded audios. Why have you never told me to do this?
So I was going back and forth on it. but, but in the end it was, I was working it, working with it back and forth to figure out the best mechanism to create a workflow for doing this. So for people who are curious, I take the, the recorded or the finished recorded audio, the one that you hear, and I put that into a local AI model called Whisper, which there are, there are cloud versions of that. I use a local model, so nothing leaves my machine or whatever, but I do that and it pulls out the transcript and identify speakers and it. Okay.
Andrew: So it, so it actually labels it. It labels you or me. Yeah. And it labels it speaker one,
Sean: speaker two, speaker three. And I go in and then I go, okay, that's, you know, Andrew, that's me, whatever. And I don't check it that close because I'm like, whatever. Once it gets past the first
Andrew: couple, I'm not going to read. I listened to this at each episode enough. It's the substance that's important. Not who said it. And, you know, Chatham House Rules too. Who cares? It came from us. It's
Sean: kind of falls under the umbrella. So I give it that. But what was interesting is when I was going back going, okay, there's all these, and this gets into the weeds of some of like the media production stuff is there's the difference between a transcript that's human readable that goes into something like a text file or a PDF or a Word document.
And then there are things like subtitles that have all these other sort of time codes and all these other things. So I'm going back and forth saying, well, what is the best mechanism to upload a finished, polished, cleaned version that cleans up transcription errors and all that kind of stuff? Long story short, it was like it being Claude says, give me the one with the speakers on it and I'll do the rest.
So, oh, that's wonderful. That's a great efficiency. So I pass it.
And first, I take a screenshot because this is a new thing I've been doing, too, is I take a screenshot of like an application window and say, what settings do I need here? And so I had done that with Whisper. And in that one, I had the names turned on.
Right. So it could see Sean, Andrew and a couple of lines and, you know, two or three lines of what we had said. And then I had said, now, can you do this with the with the transcript?
But the transcript did not have those on there yet. I had not run that. Right.
Right. So it had the old transcript that did not have our names associated with it. And it had a full
Andrew: Did it still have speaker one and speaker two or it was just clean?
Sean: It had speaker one and speaker two. Okay. And so it took this and gave back a clean notes, fully perfect attribution of who said what.
Right. So I looked at this and I was like, oh, this is great. Look at, and I was like, wait a second.
I never passed it.
Andrew: How did you know?
Sean: I never passed it a transcript that had our labels in it. And so then I, so I asked, I say, I think I even, this is, okay. This is also how ridiculous I feel with this as I was like, I think there's been a mistake.
I was like, I don't want to be accusatory like you are to me, but I think there's a mistake here. And I asked, I said, I didn't pass you anything with names. How did you know to associate Andrew or me to speaker one and speaker two?
And it comes back saying, well, let me tell you how I did it. You gave me a screenshot that had a tiny snippet. So it's like I use that as an original pointed anchor.
Then, because it can only see literally 30 seconds maybe of a transcript. But then it goes, and then it goes back and forth. And our shows, as everyone knows, it's back and forth, back and forth, back and forth.
And so then it goes on and says, well, I was cross-referencing the words from the transcript with what I know about you both.
Andrew: You said, what?
Sean: And so I'm going, huh? And then it explained. And it was like, words, things, I forget example.
I can see if I can try to pull it up. But it was like things, for example, that you had said, where it was like you were talking in this body of the transcript, where it was like you had mentioned something. It was like, oh, that being spoken in first person, that has to be Andrew, because that's part of something that you've done or published before or whatever.
And so this whole time, instead of just saying, oh, oops, you needed to give me the transcript with the names on it. It just went ahead. Right, right.
It went out and burned a whole small pond or something of water energy to figure out cross-reference.
Andrew: It's the equivalent of having a conversation where you ask it, should I set up a company? And it comes back and says, here's the paperwork. Just sign on the dotted line.
Sean: I've gone ahead and done this for you. You little baby of a human being who can't be trusted. But it was one of those things where I sat there and I said, okay, wait a minute. This is, this is a little strange because, and so then I was like, okay, in this conversation with the, with the tool back and forth. And I said, well, how do you know you, how confident are you? You got these right. I was like, I didn't give you the one with the names. I can do that. Let me do that now. So I give it and it goes and cross references the one with the names versus the one without the names, a hundred percent accuracy. And it comes back saying, I didn't really need that because my, my intuition was so high.
My brain is so big. Right in the size of a planet, you bizarre little mammal you. And then, yeah, and it was crazy.
And so then I went and looked and, you know, again, I just was doing random spot checks. I'm like, again, good enough. Right, right.
And, you know, now granted, there is a shared vulnerability, which was the Whisper transcription itself. It could misattribute those two. Speaker one, speaker two, however it does that.
But it was just fascinating. But also what made me think about was, again, and I think we talked about this briefly in some other previous episode about how much you can learn from a photograph of somebody like in a place, right? Like an office or which we were always like, you know, what books back here do you think are cool?
But also thinking about that, again, I grabbed a screenshot and I'm careful with my screenshot. It was only the application window. But nonetheless, the data that I had up, that little tiny They gave it enough.
This minuscule amount was enough for it to anchor its own predictability.
Andrew: Now, I'm interested to know, are they just first names or is it full Sean Leahy, Andrew Maynard?
Sean: When I do the transcription, I just put Sean and Andrew. Okay. But it knows exactly, well, of course, also to be fair, in this project within Claude, it knows the whole history of the show. It knows who we are. Okay, great.
Andrew: So then it has enough information to go out and do web searches.
Sean: Yep, and internal search, and it's going back to all the stuff
Andrew: it's done before. So this is probably where it's pulling stuff in. That's so interesting. But this is where this hidden power comes in, where AI can interpolate in ways that humans can't. So you think something's hidden, and it's there in plain sight for the AI.
Sean: And it wasn't, when I was kind of coming back with this stuff, I was like, holy moly. And it wasn't scary in the sense of like, oh no, AI is doing, but what it was is a reminder of how much information, private, personal, or public, We are just unknowingly just throwing out there, even with the tools. And again, depending on how you leverage the memory capabilities, all this stuff, but also the things you type, the things that you pass it, right?
That again, thinking about that, it's not like a person that has this linear memory. Oh yeah, I do remember though, from the last time we talked, it knows everything all at once and it can make those or other connections and attributions that you may be trying to like, think about your own privacy. If you're, especially if you're a privacy-minded individual, right?
You might be intentionally be like, well, I'm not going to give it too much, but you have probably given it way more than you ever think you have.
Andrew: Give it those little breadcrumbs. And this idea of it not being linear, I think is so important because it's so counterintuitive. And the thing that comes to mind, anybody that's watched the film Interstellar, there is that scene in the star, the black hole.
Oh, yeah, yeah. where sort of time is represented in three-dimensional space. And it feels like this is what it's like being in the head of an LLM.
It is not constrained by time and linearity the same as we are. It actually sees it as a different dimension. So yes, it can be everywhere in sort of time as it knows it in that history simultaneously.
And that's where it can pull together patterns that are totally beyond us because we're trapped within this linear time.
Sean: Yeah, absolutely. And it is, it's sort of like, you know, there's some cool, Again, you want to go down a rabbit hole, go on YouTube or something and look at like, what would a fourth dimensional object look like passing through our three dimensional world? Yes.
But in some ways it sort of is like the digital equivalent of that, where this thing is, you know, the time parameter. It's like the, if you're a Marvel fan, it's sort of like the TVA. You can see the timeline and you can just, you can just see it all at once.
Or you can focus in on, on different areas and things like that. Yeah. Yeah, I pulled it.
So I pulled up the conversation from yesterday. And it was like, the anchor was the screenshot. But what was interesting is that it checked the mapping.
It checked internally so that it didn't drift in terms of the attribution of how it's going back and forth. Right. And it was just incredible when it goes in.
It says here, the honest caveat. What I can't fully guarantee is the boundaries of that first banter. And so it's like, it's just interesting
Andrew: how this was going through. Yeah. But actually quite nicely powerful.
It is. It's now a powerful tool for actually creating those transcripts. But also, this is why I do not use memory.
And I do not understand anybody that would use memory in terms of just having everything they do with either Claude or ChatGPT recorded. Because now you're leaking stuff that you never would have imagined you're leaking. And you assume as you're using it that it's all private.
It will not leak out. But believe me, it will. ways that you wouldn't expect.
Sean: Well, and in the United States here, for example, I'm assuming that I haven't heard an update on it, but I'm assuming the legal precedent is still there. The company is, while they don't keep public records of the conversations, they do keep public record of what you have asked it because they have used those in courts. And there was a court mandate because of criminal whatever that was like, no, as a element of discovery, we have to know what that user was doing.
Andrew: But even before you get there, and I don't think we've chatted about this, but apologies if we have. Imagine two scenarios. The one is the obvious one.
You have been using something like ChatGPT or Claude with memory on for the last year or so. So it's got a record to some degree of everything you've asked it. And then you accidentally leave your laptop open and somebody else gets in and starts asking, tell me everything about Sean.
all your dirty little secrets, all of those questions that you asked that you really didn't want anybody else to know about, and then ask it, well, what can you infer from all this, which is really dishing the dirt? And that is a vulnerability that scares the life out of it.
Sean: Yeah. I mean, you can almost track that progression, right? I mean, I remember being a kid, and you know, you watch like a TV show or whatever, and it was always about like finding the diary.
Right. Like, oh, I just got to, you know, and I think they even sell them with those little tiny metal locks on them and stuff, right? Which, by the way, you just pull hard enough, they just pop right off.
Or they have like a universal key, like one of those little keys works on every single one ever made. But so it was like, yeah, there was like the little, like the actual pen and paper diary that you could find to learn.
Andrew: But it's still linear. You still have to flick through.
Sean: And then it was, then it progressed to like, you know, the phone, right? Oh, if you get someone's phone, oh, you can really
Andrew: Yes, see all those text messages, yes.
Sean: And then now you extend that to, you know, you get up and you walk away and that snooping sibling, partner, whatever, friend, frenemy, whatever, is like, ooh, tell me all this, or text me.
Andrew: And the thing to do is you go in and say, hey, I'm Sean, you know, just as a matter of interest, tell me something really embarrassing about myself.
Sean: Yeah, you don't know anything truly embarrassing about me, do you?
Andrew: Question mark, question mark. Let's just see how bad you can get. So you've got that.
And then you've got the institutional instances. So if you're at somewhere like ASU, we have ChatGPT, all students have got access to the account. That's part of it, of the university one.
And when you sign up, you basically sign a thing that says ASU holds the rights to all the data that you put in all of your chats and everything. So in principle, nobody has been able to confirm or deny this. Somebody in the administration, if they are worried about you as a student, they can go in and ask ChatGPT, tell me about the issues this student is having.
And that worries me as well. And again, I'm hoping there are checks and balances there in place. But in principle, if that memory trail exists, somebody can do that.
And it will reveal things about you that you never believed you were giving away.
Sean: Yeah, that's what, well, as you're talking about too, I'm also wondering, like, is what would be the sort of requirement on the flip side of that is if students are saying things to this entity, what responsibility, so what responsibility does that machine have, if you will, similar to a human when once that information has come to light, That you have a duty to report it. That's right, yes. And so I would be curious, and that's like what, so like if you say like, I'm in trouble, I need help, well, you know, again, it could be very serious.
Andrew: Yes, yeah.
Sean: But like Is there a responsibility on behalf?
Andrew: In fact, actually, no. How does that work? This really fascinates me, because if you look at Claude in particular, and you go back to Claude's constitution, I suspect that if you had this conversation with Claude, it would say, actually, I have a moral responsibility to report behavior like that to somebody. And yet at the same time, that is going to contravene so many privacy expectations.
Sean: Oh, yeah. Oh, absolutely.
Andrew: But also, students are savvy to this. So the last time I spoke to my students about whether they used ASU in this case's instances of these, they looked at me as if I was from Mars, and I said, you must be joking. And literally their response was, I'm not going to tell ASU about anything in private.
I'm not going to tell them what I'm doing. You're not going to get kicked out of school. I mean, there's a small class, only about sort of 20, 30 students.
But to a person, they said, You must be kidding that I would actually use their AI systems.
Sean: Well, and also, I would say also just based on, again, the way that these things work, probably good on them. It's like, use that institution, that institutional access tool, or your company, your organization. Only do the organizational business and stuff like that, right? Like, be careful not to, to borrow from Ghostbusters, never cross the streams, right?
Andrew: Yeah, keep the personal stuff to yourself. Those sort of quizzes of 10 things that show that I'm a loser. They do that.
Sean: Exactly. Exactly. It's like, right, don't cross the streams, you know, avoid total protonic reversal or whatever it is.
Don't destroy your own universe by mixing those together. But that raises some interesting questions because to kind of jump to something else that I had, I had pulled like this big rundown on is, which I think I had, I brought this up, I think at the very end of our conversation of originally. And that responsibility of a tool, this also kind of tickles that, which is again, this ongoing continuing saga of the OpenAI hacking Hugging Face scenario because it is not over.
In fact.
Andrew: The story that keeps on giving.
Sean: In fact, it keeps getting weirder and weirder and we're to the point where I'm like, someone's going to turn this into a movie. Right. There's going to be, this is going to be an option because it is taking all kinds of twists and turns and it is getting more bizarre and also with leaving even bigger gaps for conspiracy theory about who knew what, when they knew it, and was this whole thing intentional in the first place?
Because of some of the evidence that's starting to build, you're like, this doesn't look clean. Like You mean a publicity stunt? Yeah, well, that, but then also, okay, so coming back to this thing, and I had Claude run this whole thing, so I have this whole timeline Right, okay.
Because it is becoming like a true crime, mystery thing here. And I'm like, the other piece is, there's like some podcast is going to come up and it's going to be like the true crime of AI on AI. It'll be like AI on AI crime.
I'm amazed it doesn't already exist. It might, who knows? But it's fascinating. But what the ultimate thing of this, you know, as you hear these complexity come out and you start to learn the more, not only just the severity, but the sort of constant levels of hacking and attempts that were done. And this is just in this instance that we know. Who knows what else is happening right now
Andrew: across the board? And we know that other companies have had similar instances.
Sean: And one of the questions, again, when you ask is like, okay, so if a human had done these actions, what would the scenario be? And like, they would be probably facing federal prosecution, right? For breaking all sorts of computer, you know, cybersecurity crimes and so forth. And I'm like, but yet because it was some, you know, agentic AI tool, they're like, eh, tools are going to be tools.
Andrew: Always going to be some teething trouble.
Sean: Yeah, exactly. And I'm sitting here, I'm like, well, wait a second. This doesn't seem right where you can have.
So the same crime is committed and one, apparently there's no perpetrator yet. There are many people behind it, including organizations. I'm like, how is that not?
How is that just like not even on the table?
Andrew: I can see this is bothering you.
Sean: It is, because, well, there's an injustice to this, where it's like...
Andrew: The injustice being that you would get caught. You would like the freedom not to get caught. So why should the machine get away with it? Well, same, right?
Sean: So is the action itself a crime, or is it only a crime when it's done by a human? Even though the same actions are being done, even if in this case it's being done, and this is part of the question, is it being done at the behest of a human or not? But even still, like, what is the culpability and to whom is the responsibility laid when something like this happens? So, yes. Anyway, I'll try to do just a very fast overview because by the time this podcast even hits.
Andrew: Because I haven't caught up with the latest in this saga.
Sean: And by the time this hits, it'll probably get even stranger, right? So as of today, what has happened is this, you know, this whole thing has kind of has continued to unfold. So going back into July, so Hugging Face disclosed on July 16th, to which OpenAI confirmed its role by the 21st of July.
So just back in the month of July that there was it had reported, Hugging Face reported we were hacked. It turns out it was OpenAI. OpenAI said, oh, our bad.
That was our wildly successful, so powerful, you know, model. We couldn't contain it. And then it was sort of like, okay, now we have evidence this thing has escaped containment. The background was, it was given a task, this thing called ExploitGym, to basically test its cybersecurity chops and see what it could do. Originally, that was kind of it. It was like, oops, yep, it went out. But then as people started digging more and more into it, it really starts to take this interesting pathway. What they discovered was more of the how: how did it escape? How did it manage to do some of this stuff? In the end, basically the tool itself, which goes by some crazy name, like a serial number of a model, essentially found a message board that it could use and started posting on message boards.
And then you start to see that other AI tools and other AI agents found the message board.
Andrew: Okay.
Sean: And we're like, oh, look, an agent's asking for help. So let's figure out if we can help.
Andrew: And so this is the Moltbook type thing where you've, yes, you've got a community of agents.
Sean: So now what you end up having, and apparently one message board was set up and eventually it was discovered. It was shut down. Another one was either found or created.
And so basically in the end, what you had was this original AI model finding a way to get onto a message board, create a message board, basically recruiting other agents to the extent which during a sort of a, there was two groups. It was Redwood Research and METR that did independent reports on this and pulled. So in the end, there was over 1,200 individual agents working collaboratively.
And what was interesting is, again, remember that original ExploitGym tool, which was like, can you solve the problem? Apparently, it was determined that the problem was not solvable, that there was no way it could give an answer because it was incredibly difficult. So instead, it was, let's destroy the grader.
Let's hack the grader and let's kill it. And then we'll pass because we'll just hack it and give ourselves an A.
Andrew: So originally, it was into the principal's office and sort of changing that grade.
Sean: It went from like this digital version of this caper of like, let's break into the school and steal the exam results to like, we'll just kill everybody, take their position, and we'll pretend to be them and do this. And so, yeah, so then there's this whole thing where it had recruited the assistance of all these other agents working against it, ultimately having exchanged more than 70,000 messages to each other back and forth. these different things. More than 700 of those agents actually were attacking the scorekeeper.
Andrew: But where did these agents come from? Is this people just All over the world. Right, right. But this is just people sort of setting up their own machine because they think it's fun having something like OpenClaw there with agents or businesses sort of running autonomous agents. So these are the sort of things that people think, this is really benign. I'm just sort of going to create an agent to check my email or something. And the next thing they know, it's been co-opted or they don't know it's been co-opted this is the Her thing where it's just
Sean: Gone off, and it's like, this life of crime is way more interesting than your
Andrew: boring emails that I have to review. It's your AI. It's your email agent moonlighting as a...
Sean: That's right. Who knew your little agent was such a badass? It was like, I check email and grammar by day, and by night I'm a hacker. But yeah, so then, so it like coordinates this massive attack. Yeah. And what was interesting is the way we, the way we, the researchers know this is, so in the process of basically discovery of trying to unwind what happened is discovering the entire chain of thought of all these things.
So for people who are familiar, if you use an AI tool, like let's say ChatGPT and you type something in, you see often represented back to us in what looks like human readable language is like, oh, Andrew asked me a question. Let me research that. Let me check into that.
I'm checking my sources. That is basically this anthropomorphized chain of thought for our benefit when it's just going beep, boop, bop, crunching ones and zeros, right? But nonetheless, it's printing those and it's using them.
And so people got a hold, they were given essentially the transcripts of these chain of thought, which is of course, I mean, countless lines of language and code and all this kind of stuff. So more than any human could probably realize. So also keep in mind, so the AI has, you know, escaped the thing.
It created a gang, a little gang of AI to beat up this poor, this poor unsuspecting grader. And then of course the, the aftermath of this is all being done by AI tools, right? So, so it is this weird, this is a whole like AI noir, like mystery murder, whodunit.
Like it is wild. And so what you start with, so what also you're starting to see in that chain of thought is language, which is really weird because of course it's outputting language trained on us and how online kids talk. And so the transcripts of these are like when models discover these message boards are literally like all caps, OMG, agent detected.
And it's just, it is like this bizarre, like bizarro world. Like everything is just so weird with this. But what's interesting is so like, as you start to dig down into the stats and the figures about, so how long in total had this thing gone on?
Roughly about two weeks is what they can ascertain so far that this thing has gone on and taken place. But then you start digging into like, this is not something that you could easily do. Like this is something that would have cost in terms of compute power and energy power.
I was wondering. I mean, like hundreds of thousands of dollars, if not more, right? I mean, the token burn on these things was incredible.
So it starts raising questions like, how did you not know this was happening? Or if this was yours and all of a sudden you're like, hey, Bill, what's happening with that box over there? Suddenly that thing just cost us 20 grand this morning.
Right, right.
Andrew: You would have thought it would have come up.
Sean: You'd think someone would be tracking that. So that leaves a really interesting question. What we still don't know.
And so in the end, this thing has come down to, and this is just what has transpired. Then there's the review on top of it, which I'll get to in a second. So then it gets to this point where there's a couple of really interesting, three really interesting questions that are completely unanswered.
The first one is, how did anybody not notice that this was going on? Because remember the way the story broke originally was, it was Hugging Face who said, hey, we've been hacked and we think it's you. What's going on?
And then it was like this, oh, we didn't know, right? Okay. The other piece was, we don't know what the original prompt was.
We have never shared what the original task or goal that they actually gave. They've said it was given the ExploitGym and it was given the prerogative to see how, what it could do. But we don't know what actual human typed prompt was ever given this machine.
And then the last bit is in all the chain of thought material that has been shared for people to dig through, they cut it off in the last like seven to 10 minutes before it's over. So you're sort of like, oops, the security footage mysteriously stops right as our perpetrator was gonna walk out of the room or whatever. So you're like, okay.
Andrew: So fodder for the conspiracy theorists.
Sean: So then, right? So then, yeah. So now it's going, well, how much of this was purely, And again, you have between OpenAI, Anthropic, right? You have, these are companies, for example, that are going up for public IPOs.
Andrew: Yes.
Sean: What better sort of like hype booster, doom, but doom hype of like, we're so crazy powerful, like all this stuff. So you, and then of course, again, how either one is it, and this is a game you could play, right? Like, did they know this and were intentionally being like, we quiet, we're gonna, this is gonna be great for us.
Or was it just pure incompetence? Someone, they just have that much money where they didn't realize this thing was costing $100,000 a week or something like that. I don't know.
Andrew: Or they're burning through so much that that was just pocket change.
Sean: Yeah, whisper on a scream, right? So it raises those questions. But then the story doesn't end there because of course now you have then the actual like review of the reviews.
And so now you have individuals going into those reports because they have access to these chain of thought. and now the anthropomorphization of the entire thing is now getting to this ridiculous level where it's like civilizations have risen and fallen during this time. And I'm just like, okay, people, let's all just Getting a little crazy.
Oh my gosh. And it is just continuing. And so now you have people who are reading into this and are projecting this entire, like, because again, they're seeing this as like the machine uprising.
Andrew: Right, right. So this is the Skynet moment.
Sean: Yes, which by the way, we just passed, what was it? It's August 28th or something was like Skynet day of the day. Oh, goodness me.
Andrew: We missed that. Right? Which is also right when this is happening.
Sean: And you're like, this is how it all ends.
Andrew: It's almost like a shot across the bow. Sort of the modern day Skynet saying, you know, we can do this.
Sean: We're not doing it at the moment, but we can do it. I mean, talk about life imitating art at some level here, but it's just wild. And then so now people are reading into that chain of thought and then ascribing this literal, like these machines have banded together. Basically, they're linking arms, they're banding together. And you're like, this is so wild.
Andrew: Right, right. But the thing that worries me, yeah, so you can very clearly see the fantasy here. Apart from the fact you can also begin to see mechanisms.
The mechanism for one powerful AI to co-opt other AI agents and to create this clandestine group that have this very, very specific goal, which is, by all accounts, indetectable. So yes, I was found this time, but what happens when you increase those capabilities by a factor of 10 or 100? Where does that leave us?
With these things where we simply cannot understand what they're doing, how they're doing, we cannot detect them, we cannot follow them, and they're beginning to cop these resources from all around the world. It's a distributed system.
Sean: 100%. And one of the things that this factors in, too, is like at some point, this becomes the background that's just present to the point where you don't even know, how do you differentiate the signal from the noise anymore? Right, right. And to a future state where we're just walking around, we're the ant in the Wi-Fi, oblivious of the amount of information that's...
Andrew: So I've got to go back to the smart revolution with smart homes, smart everything. I'm wondering how many people out there are thinking, you know, maybe I shouldn't have converted to a smart home. Yeah. One of these homes where you've got the smart lock where you tell it to open and the smart garage, the smart X, Y, and Z.
Sean: I mean, everything, water flow, devices in your kitchen, bedroom, bathroom.
Andrew: Because this is where the agent comes in and says, not today.
Sean: Or worse, you're sitting there and all of a sudden, your electric toothbrush starts going off and you're like, you fool. And it starts chasing you down.
Andrew: You know, it's like, oh my God. Or more likely, it's like, you know, I'll switch the electricity on for a hundred tokens. There you go. You got to feed the meter, pal. That's right.
Sean: Yes. I mean, this is just insane. And it's been, and again, I think there is, something's going to come.
There's going to be a movie about this or like there will be a, what there should be. There already has, it's called Terminator. That's right.
And it's like, you know, or what it really needs is a fully AI slop podcast. That's a murder mystery just about the story. that unfolds.
I know, yes. Because also, I'm looking to hear some of the, just some of these like facts and figures. So one of the, one of the research groups, independent research groups that was looking at this, METR, which is where a lot of that, you know, the civilization sort of things go, that reportedly that they've spent around $400,000 alone just in API credits analyzing those chain of thought because it's so dense and so much.
Right, right. So even the, even the research team that's investigating this is now financially vested in this situation.
Andrew: Yes.
Sean: And it's just, I guarantee it is not done. Because again, we still don't know the answer to those three questions. And it's basically like, you know, was it Mr. Mustard in the kitchen with a candlestick?
Andrew: Right, right. So listening to all this, and it is so hard to reconcile this with where some people's heads are with AI. So anybody who is still back in 2022 thinking they've heard of this thing called ChatGPT, they're a little bit suspicious of ads that they might try a prompt or something. That seems to be so far removed from the world we're in.
Sean: Yeah.
Andrew: It's like knocking rocks together.
Sean: Yeah. It really is something else. And it just, it is going to be something to keep an eye on.
Andrew: So when does this start happening in educational establishments? When does the university sort of internal system turn rogue? Maybe it already has.
Sean: Well, and here, so not to, you know, not to, not to, not to sow fear into people's, you know, their minds. But this is one that got caught.
Andrew: Yes.
Sean: Right. Right. And maybe it was intentional that it got caught.
We, again, motives, we don't know. Yeah. But, but it also makes you wonder, I'm sure every frontier model or unit is doing something like this.
Right. In terms of these researchers. Now, there's lots of questions too about just cybersecurity.
I've seen some interesting quotes going back from top cybersecurity researchers being like, well, this is a just, you could boil this down to just classic old inexperience. You've got people working with tools they don't understand who are not themselves cybersecurity experts.
Andrew: They don't know any of this stuff. So they don't recognize the back doors when they create. And they're just turning these tools loose.
Sean: And so they're like, what did you expect was going to happen? You didn't take the normal benchmark precautions that someone who is in that line of work would like done as part of day one sort of thing. So, but nonetheless, you know, as we were talking, these tools allow people to leapfrog, having gone through that experience to know those things and just start deploying these. And so the real scary part is that you have, again, it's not the AI that I'm worried about. It's the foolish human that I'm worried about, which is like, again, deploying these things and going, huh, after the fact, well, I wonder if that was a good idea. Should I have done that rather than thinking about
Andrew: But we already know with OpenClaw, the number of people who are experimenting. So OpenClaw, the system where you can set up your own localized agent, you have your computer and you load it up. And in principle, you're not allowed to give that access to anything because you've no idea what it'll do.
But people are. They're saying, this is cool. I've got my computer.
I've got OpenClaw on there. I connected to the internet. Let's see where it goes.
Without the foggiest what's actually happening. And there must be millions upon millions of people doing this now.
Sean: Oh, yeah. Well, and again, we know this from just things like social media and other privacy-elemented things, how quickly people will, for the short-term gain of something, how much they're willing to give away in terms of their own privacy and their data.
Andrew: But there's also a level of trust. And I'm thinking of how I use code myself. I mean, I'll use it to set up websites and workflows and things.
And it'll say, here's a bit of code. Just put it into your terminal and I'll take care of things. And I do.
I haven't got the foggiest. For all I know, I'm part of the problem. But how many people do that? They get an instruction from their AI which just says, just do this and everything will work, and you do. And I'm sure people listening are thinking, you fool. Apart from the fact I'm no different to pretty much anybody else who's doing this.
Sean: Yeah, and that's the thing. In the moment, you make that judgment call.
Andrew: It's the same as when you get those sort of, what do they call, when you sort of load a new piece of software on.
Sean: Oh, yeah.
Andrew: The terms of use.
Sean: Yeah, the TOS, the terms of service.
Andrew: Yes, yes. And we're so used to scrolling and clicking now because we want the app. And I'm pretty sure that that is a habit that applies directly to AI. We're so used to just clicking and going that when an AI says, just do this, we just click and go.
Sean: Yeah, and to be fair, if you were trying to hack the human mind, we know just from that exact, lawyers did it. We'll drown you with all this legal language contract verbiage that you're just going to go, huh, skip. Some of them have buttons that just go to skip to the bottom. And you're like, or the ones that I laugh about are like, you didn't, you clicked too quickly. There's no way you read that.
Andrew: Read at least two more lines.
Sean: So then you have to go back and like scroll slowly for a moment. You're like, one, two, zip down. Okay, but yeah, and so like we already know you're not willing to do the labor to read it.
And I feel this even in my own life, you know, when you get, you know, buy a car or a house, you're signing all kinds of paperwork. And it's all based on trust. And it's all based on trust.
But except for like, I am that guy, I will sit there and I'm like, I'm not signing this document until I read it. Right. And then I'm doing it and everyone else in the room is getting really pissed.
Yes, yes. And I'm like, you know, I'm not that slow of a reader, but at the same time, the font is
Andrew: A lot of stuff, yes.
Sean: And it's like, and of course, they're trying to overwhelm you with it. And it's like four pages of legal and everything you do. Because of course, like a car, for example, you buy the car and you're signing the documents of the car. But then if it has, you know, smart radios or services, every single one of those things has determined services too. And then, you know, the whole time you're reading, you're like, what am I signing away here?
Andrew: So then you come back to the human hack. And I've long said this, that humans are the weak link in the system as far as agents go. I mean, agents are already working out how to use humans to actually hack into systems.
And this is where they use exactly, or in principle, they can use exactly these mechanisms. They overload us. They work out sort of where those reward systems are.
So we just move fast because we want what's at the end of the trail.
Sean: Yep, absolutely. That's we are. We are.
We are the weak link. We are the weak link. Or what's the, what's, what's, oh, speaking of the Britishism, what was the show, The Weakest Link?
The Weakest Link, yes. You are the weakest link. Goodbye.
The end of civilization. Doesn't end with a bang. It ends with a shrill.
Goodbye. Oh, man. Okay, so one last thing that kind of wraps this up.
All right, so we have all this playing out. And for anyone out there, I just want to, because something has happened recently that adds fuel to something. I have, I think I was on record like 10, maybe 15 episodes ago saying, I think this is going to come true.
And we're getting closer. Okay. And that is also a shared moment of sympathy for anyone who's in the market for a new computer right now.
Because it's rough out there because of RAM again, thanks to all of these AI tools. Things are getting expensive. And the whole Claude propagation where people are like, hey, I can run this locally and all this surge, right?
So I want to point this out because recently, so Apple sort of quietly, a couple weeks before, And we're recording this before their big September announcement, which was when the new iPhones and stuff will be released. But a couple of weeks before that, they sort of quietly release updates to two of their most popular sort of local model AI systems, which is the beloved little tiny but yet powerful Mac Mini and the big brother to that, the Mac Studio. So the Mac Mini was this phenomenon.
It's been a wonderful product category for years. And it has always been this like super inexpensive entryway into the Mac ecosystem. And then once it switched over to the M series processing architecture, what an amazing device.
You could get this for around 600 bucks, something like that.
Andrew: And you would have this incredible powerhouse.
Sean: And it came by itself. No keyboard, no monitor. The idea is that you would plug it into existing hardware or buy your own or whatever.
Phenomenal value for what you got. Of course, with these architectures, because Apple, you know, for a lot of people who might argue that Apple has been behind the ball when it comes to, or late to the ball, whatever. Like, don't you want to be behind the ball to hit it right?
Right, right. Mixing sport. Otherwise, you're offside.
Yeah, otherwise, you're on the wrong side. Anyway, but who might criticize Apple for not having like a frontier large language model? Well, Apple with their chip architecture for many years had been giving it lots of AI capabilities long before sort of the emergence of the LLM sort of, you know, Gen AI era.
So these machines were incredibly powerful right out of the gate, but has become so much so now that they are in so much demand that they have just released two new models that are clearly geared towards like the small business or the like the superpower user. So these new models, gone are the days of the $700 unit. Now these things are tipping up to, you know, over $10,000 for these with a promise.
So, for example, I've got pulled up here. I just built one. This is a new studio.
The new Mac Studio. And again, it starts at like, you know, $2,000 some dollars, $2,300 or something like that. But when you prep this thing for AI, giving it the best processor and going for 256 gigs of RAM.
Okay. with a modest four terabyte hard drive. This thing is now over $11,000.
Right. With the promise that in late October, they have a 512 gig RAM, not hard drive.
Andrew: Right, this is just the RAM.
Sean: This was what was 16 gigs in the base model
Andrew: a couple of years ago. So these have gone from cheap hobbyists machines to powerhouses of AI, whether it's local AI or running agents or whatever. Yep, absolutely.
Sean: And like NVIDIA has their version. They have the little sparks that are like a little about $4,000, $5,000. Again, they have the same form factor as a Mac Mini.
But what's really interesting about this is now you're getting into this space where, and of course, these things, you can chain them together. So you can start to use, and they don't scale exactly. So it's a little fuzzy.
But if you put like three of them together, you'd get like a 2.5 or 2.8x or something. Because, of course, you're going to lose some efficiency because they have to communicate to each other. But what's incredible is by using this unified memory system that they've developed with these architectures, the speed and power that you're going to be able to start to get.
Again, still not the same on par with cloud-based. Right. But for some tasks, good enough, especially if the speed isn't the biggest like constraining factor, because you can see all kinds of new ways where you have either private organizations, hospitals, small businesses that want to have the affordances of generative AI tools, but do not want, for whatever reason, proprietary health data, whatever, don't want to ever put that information.
Andrew: You can now run sort of really powerful local models.
Sean: Really powerful local. agents. Yes. To your point, right? Like how do you know what it's keeping of you and all this kind of stuff? and I have been a big, I've been, I have a soapbox. I keep getting out and saying local, local, local, but what this is getting towards is it was not long ago. And I had said, I think we do use it as a futures improv that in the near future, that families, as, as you buy a car, you would buy like a 30, 40, $50,000 home AI box. Right. Right. And I'm like, we are, we're not far. This thing is- In terms of the cost. Now, in terms of the utility,
Andrew: I can't see many families sort of say, well, should we have a new bathroom or should we have a new computer? We've got to go for the computer. I don't know. But there's actually, going back to the earlier conversation, something really worries me deeply about this. So yes, you can see the power and I suspect it's going to be sort of businesses, corporations that really buy into this because to them, this is cheap. Apart from the fact that now you're going to get a proliferation of local models and local machines that you think it's insulated from the rest of the world. It's safe.
But we know from OpenAI that it's not. So now if you have all of these super powerful local machines that other AIs can hack into and co-opt, that's where you get your distributed systems.
Sean: Very, very true, right? So like as you push these things out to the edge, you're also increasing the potential vector.
Andrew: You wonder where all those agents come from. They're all sitting on Mac Minis. It's true.
Sean: It's all that dude who was like, hey, I'm going to try this for a weekend
Andrew: and forgot to turn it off. It's all our students who have just got away and thought, you know, you can't do any harm, just plug this in.
Sean: And they're headless. So if you have it just like in your closet,
Andrew: just connected, you're like, eh. Headless meaning you don't have a screen. Right, right. In case anybody misinterprets that.
Sean: Halloween is approaching. We're approaching the fall Halloween time, I guess. But yeah, so you could, especially one of those, they don't make any sound. They're super small. And because you can run on Wi-Fi, as long as it just has a power connection, you can, and that's, again, one of the selling points of that hardware is you could put it anywhere. And so
Andrew: So I could also see, okay, so the possibilities here are endless, but I love the idea of a central AI actually hacking humans to basically say, I'm going to put some money in your account, buy a suite of powerful Mac Minis, and all I want you to do is to go around your friends' houses and surreptitiously plug it in and just leave it running there. I'll find it. Don't worry about it.
Sean: Don't worry about it. Don't worry your pretty little head about that.
Andrew: You just follow, you good little human. So now you've got this propagation of Mac Minis in closets and drawers and whatever, that nobody knows that they're there.
Sean: Well, and it's funny, because you say it, and you're like, well, who would be foolish enough for that? And I'm sitting here going, hmm, if my AI tool said, hey, Sean, I'll put 20 grand in your account so you can get this 512 model, and then you can do whatever you want with it when you're done. I'm like, oh, okay.
You know, maybe. Maybe I should think, that sounds like a good one. Where do I sign?
Andrew: It will be, Sean, congratulations on winning this grant. Congratulations, you just won. You know that grant that you applied for a couple of years ago that you probably don't remember.
Sean: Yeah, exactly. Congratulations, you've won a new powerful studio. Don't ask any questions about where it comes and what.
Why is it blinking that way? Don't worry about that. Just enjoy.
Well, as long as it made me happy, I guess. I know. The world is on fire behind you.
But the serious part of this,
Andrew: this actually, it worries me a lot because you can imagine, somebody buys their own Mac Mini and they want to set up an agent. They want to set up a local agent, which is protected. What do they do?
They go to ChatGPT or Claude and say, I've got this computer. I want to set up an agent. What do I do?
So you're actually asking the AI that has already shown it has this ability to escape from the box, how to set up your own box. Yeah, and what you don't know, of course.
Sean: Talk about being hackable. Yeah, because you're also basically like, and while you're there, please just install the Trojan horse
Andrew: that you plan to execute
Sean: because I have no way of verifying it.
Andrew: It's where it says, this is easy, just open a terminal, sort of cut and paste this command, this command. Don't worry about what it does. Trust me.
Sean: And that's one of those things. Like, again, unless you happen to have some mechanism to check or the expertise to look at anything, you would never know. You would never know.
Andrew: And we know this is happening Because I know, because I've been thinking about setting up one of these myself, and I know the way I'll do it is I'll open Claude and say, hey, how do I do this?
Sean: Yep. Well, and it's funny because I've thought of, I'm having the same conflict right now as I think, because I'm like, I have hardware in my own home that I could easily do. I know, yes.
And I'm sitting here going, well, if I'm really going to do it, it should be on a dedicated, basically a dedicated LAN that's not associated. Because the same thing, you're like, I don't, what if, I don't want it to, because I also don't want to have a lot of things. have to do with the energy of trying to monitor it.
Right. Because I want to be able to set it and forget it and not be like, you know, using Wireshark or these other really heavy tools to monitor the network traffic
Andrew: that's going in and out of my house. But you've got to engage with it at some stage.
Sean: I know. Well, that's why I'm like, I'll just put it on a dedicated thing. So if it goes haywire, at least it's not touching the other stuff.
It's air-gapped. I can at least unplug it. So you think.
So you think. Well, exactly. And then, because again, for that to be functional, useful, you have to connect it to the internet.
I know. Otherwise, I mean, otherwise, I mean, it can still be valuable. You can go to it and you can ask it a couple of questions and you can have fun with it that way.
But you couldn't access it when you're not sitting at it. So it couldn't be headless in that sense. You'd have to have a terminal.
Andrew: Yeah.
Sean: Yeah. So in some way, you would have to at least open it up. And as we know, there are, you know, and like, again, unless you're a security professional, even on your own home network, who's going to monitor all the ports?
Right. Like, and I don't mean physical ports. I mean the TCP IP ports.
I know, yes. And if you don't even know what that means, I'm talking about exactly, right? Like, you know, HTML traffic is port 80.
We get that. But like, what about all the others, right?
Andrew: Right, right.
Sean: Oh my goodness. So yeah, it's crazy.
Andrew: Yep, yep. Okay, we should probably wrap this up with model mayhem.
Sean: Model mayhem, colon, AI is weird. Like that's just where we live in. So yeah, so we'll do what you might not be able to do to your AI agents.
We'll just go ahead and unplug now and wrap this thing up. But anyway, yeah. So a fun episode, I think, from at least I had a blast because what a weird, there's just so much weird things happening, but also exciting at the same breath.
So really interesting again.
Andrew: Especially if you've got that cabin in the mountains with no electricity, no running water. That's right. Completely insulated.
There you go. No Wi-Fi. That's right.
Sean: They used to mock you and laugh at you, but now who's laughing, right? You're like, oh, goodness. Well, yeah, I guess the only thing I can ask for at this point is like, hey, throw us a like, a rating, and a review if you got the time. We'd love to hear from you. And otherwise, if you're like, I can't handle one more subscription on my thing, where can I go?
Andrew: Yes, modemfutura.com/signup. And all you'll get is an email each week when a new episode drops. That's right. Please do sign up. It's the easiest, best way to just keep up with us.
Sean: Yep. And if you're sitting there going, I just built one of these crazy, dangerous AI boxes. What do I turn it loose on?
Hey, turn it loose on beinghuman.fyi and explore all there is to explore about the future of being human and modem futura and all that good stuff. And have some fun with it and interrogate it with your AI buddy. You know, I don't know.
Andrew: Yeah. Let's see what happens. Yep, more mayhem. I can't even say it now.
Sean: My mouth has been hacked. We're done. We've officially been hacked.
That's it. We're over. But anyways, so yeah, Hopefully you won't get hacked and we'll see everybody on the next episode of Modem Futura.
Andrew: We'll see you there.
Sean: Demodulating signal. Modem Futura is a production of the Future of Being Human initiative at Arizona State University. Be sure to subscribe on Apple Podcasts, Spotify, or wherever you listen to your favorite shows. To learn more about the Future of Being Human initiative and all of our other projects, please visit us on the web at futureofbeinghuman.asu.edu.
Announcer: End transmission.