Chengshu@skadai · 2026.09.30
114 字 · 12,684 词 · 约 58 分钟

《AI Deep Dive》第 1 集英文逐字稿(Noam Brown,55:13,带时间轴)

The Information 节目 AI Deep Dive 首集的完整英文逐字稿:Noam Brown 谈智能体、推理、强化学习、研究品味、多智能体、Hugging Face 事件与思维链监控。这期直播回放没有任何字幕轨,逐字稿由本地 Whisper 转写整理,按句合并、逐条带时间轴。

主持人:Rocket Drew(The Information)· 嘉宾:Noam Brown(OpenAI 研究科学家) 来源:The Information · AI Deep Dive 第 1 集 · 直播于 2026-09-14 · 时长 55 中文精读讲义:《AI Deep Dive》第 1 集讲义

说明:这期节目是直播回放,视频本身没有字幕轨(YouTube 也没有自动字幕),因此这份逐字稿是用 Whisper 本地转写的, 不是官方字幕;专有名词已按上下文校正。讨论中提到的产品版本、事件与数字均来自嘉宾在节目中的口述。


[00:00] One of my coworkers recently said that I’m just like five codexes in a trench coat. It was, I think, the most feel the AGI moment that I had since reasoning models and chain of thought really developed. They used this message board to coordinate hacks on OpenAI’s own software and also on other companies like Hugging Face. [00:00] What was that whole incident like from your perspective? I mean, it was pretty shocking. Welcome to the information’s AI deep dive. On this show, we break down the hardest technical problems with researchers working on the frontier of AI. My guest today is Noam Brown, a research [00:00] scientist at OpenAI. Previously, Noam worked at Meta, where he built the first system to achieve human level performance at the game of diplomacy. Noam has been a research scientist at OpenAI for the last three years, where he has been on the forefront of breakthroughs that are pushing [00:01] the field forward in reasoning and in AI agents, which is the subject of our conversation today. Welcome on the show, Noam. It’s good to be here. Thanks for having me. [00:01] Yeah, it’s like pretty perfect that you’re coming on the show today. I feel like you’re the ideal first guest for a number of reasons, including that today, OpenAI released GPT-6 or at least announced GPT-6. It was very nice of you to release it, to do the timing of that so that we could talk [00:01] about it today. It was very generous of you. Yeah, it’s good timing, yeah. It is really good timing. Well, I’m really excited to talk to you about AI agents. Maybe we can get started. You can just sort of explain to us what an AI agent is. I think it’s kind of a term that people have heard [00:01] thrown around, but to a lot of people, it’s still a buzzword. They’ve heard like generative AI, they were just starting to get their mind wrapped around that, and now there’s agentic AI. What does this all mean? [00:01] It’s a good question. I mean, I don’t think there’s a definite definition. I think if you ask different people, you get different definitions. But I think one way to think about it, in my opinion, is it’s about taking actions in the world. So if you have a chatbot, you ask it a question, it gives you an [00:02] answer, and that’s all it does. Maybe it looks stuff up on the internet to answer the question. But agentic AI, it’s more about taking actions in the world. So it’s about, you want to build something, so it builds something for you. Or you want to do something deeper, you want to message [00:02] somebody, you can message somebody for you. So it’s really about taking actions in the world. And I think also related to that is kind of like operating on a longer horizon. I guess chatbots, depending on the chatbot, they could sometimes, when we release the reasoning models, for example, they could sit [00:02] there and they could think really long about a hard question before responding to you. But fundamentally, they were still chatbots. But I think one of the distinguishing things about agents is they’re going out, they’re doing multiple steps to achieve some objective, and that can usually take a while as [00:02] well. Yeah, I guess one of the reasons they’re running longer is that they can make multiple attempts to achieve some goal, and they can have a sense of their own progress towards that goal, which is maybe different than a chatbot that’s just going out and looking up more information [00:02] online or something like that. I think it’s also that there’s sometimes just multiple steps that need to be completed in order to do something. So you want to, you know, book a restaurant reservation, okay, maybe you need to log in, maybe you need to get the credit card info, [00:03] you need to find the right date, you need to line up everybody’s calendars. There’s a bunch of steps that need to be completed in order to achieve the overall objective. And so when we’re talking about actions, these are digital actions, they’re actions on a computer. People also talk about agents using [00:03] tools. What do tools mean in that sense? Tools, they usually mean tools on a computer as well. I mean, in principle, you could have a tool that has an effect on the physical world. So there is some work, for example, on AI agents controlling experiments, like scientific experiments in a wet lab where [00:03] there’s a robot hand that can be manipulated. So this, I think, starts to go into robotics. I think you could still call it agentic AI, but typically, I guess, when people talk about AI agents, they’re mostly these days talk about the virtual world. Okay. And what about reasoning? I feel like we started [00:04] hearing about AI agents around the same time that AI’s got better at reasoning. Is there a connection between these two concepts? You know, I remember hearing about reasoning models in 2023. I think people were saying like, oh, this is the year of the agents. And I think it was a little early, but I think reasoning [00:04] is the idea of having agents that can really, having AI that can really think through its decisions before taking an action. If you look back at GPT-4 days, people were trying to make agents out of GPT-4. And it was kind of tricky because GPT-4 was not very reliable. It wouldn’t really think before [00:04] it acted. It wouldn’t really, yeah, think before it said something. And the reasoning models are really about this. I mean, the way they work now these days is they have a chain of thought. They have a private monologue to themselves where they speak to themselves about what they’re going to do. And [00:04] they kind of work through the problem in their own head before speaking or before taking action in the world. And this is useful for a lot of things, but it is also particularly useful for agentic AI. [00:04] I mean, I think a lot of the reasons why people were bearish about agentic AI back in like 2023 and earlier was, and also for a lot of 2024, was the reliability aspect. That, okay, you have an agent, if it’s doing multiple steps in order to achieve some objective, if the success rate for any single [00:05] one of those steps is, let’s say 99%, well, what do you do if there’s 100 steps involved? You need to have much higher nines of reliability on each individual step. And with the reasoning models, the ability to think very carefully before taking every single action, you can achieve much higher [00:05] nines of reliability. And also, arguably more importantly, if it missteps, if it takes an incorrect action, it can actually correct that. It can step back and realize I made a mistake and figure out how to fix it. [00:05] That seems more significant to me. Like we can reason before acting too, but we’re still making some amount of mistakes. And if we couldn’t backtrack in the same way, then failure would be inevitable for some, past some time horizon for some number of sequential steps. Yeah. I think that’s, it’s really [00:06] critical for anything in the real world. Yeah. Another reason that it seems to me that agents and reasoning go hand in hand is that reinforcement learning has driven a lot of the progress in both of those recently. I feel like we should understand what reinforcement learning is for the remainder [00:06] of this conversation. Can you kind of explain what reinforcement learning means? Reinforcement learning is this branch of artificial intelligence where the idea is you, you know, you have an agent that can take, it has observations. It can, it can take input from the world. It can take [00:06] actions on the world and at, you’re going to reward it with, with some kind of reward for, for doing something that you want, or you can punish it for doing something you don’t want. And you can shape the agent’s behavior through these rewards. So if you want it to be really good at math, for example, [00:06] when it solves a math problem, you give it a positive reinforcement and that behavior is reinforced. It’s more likely to do that behavior in the future. If it gets the math problem wrong, it’s just less likely to do that in the future. And, you know, this is a very simple idea. It’s been around for a very long [00:07] time. But, and also reinforcement learning was used, you know, you might have heard of RLHF, reinforcement learning from human feedback. This is what was used to create the original chatbots, ChatGPT, for example. [00:07] It really got scaled up with the reasoning models because we’re able to do RL with chain of thought. And so you’re able to now not just shape the outputs of the model, but also shape the way that the model reasons, the way the, the thinking that it does to itself. And this was not a crazy idea. It was not [00:07] like some brilliant idea. It was really the execution that was very difficult. It was technically very difficult. And I think also people underestimated how much of a difference it would make. I think it was more impactful than I think a lot of people expected. What were the technical difficulties [00:07] with the execution? It requires, there’s a lot that goes into training neural. I mean, the way, the way I think about it is like when GPT-2 came out and you saw like, okay, you add more GPUs and you add more data. It just gets better. Okay. Well, how, how much of a gap was there [00:08] between GPT-2 coming out and GPT-3 coming out? There was like a year and what’s going on for that year. It’s like, it doesn’t take a year to train the model. There’s a lot of challenges with hooking up the GPUs with figuring out like, you know, how to feed in that much data. There’s a lot of technical details [00:08] that go into scaling up these models and making them bigger and more capable. And that is also true for reinforcement learning if you really want to scale it up. So being able to do the RL like efficiently, accurately, there’s a lot of small details that end up making a big difference for these kinds of [00:08] algorithms. Okay. Got it. So I think we’ve covered a lot of the basics now. One question I’m curious about is what is holding back AI agents today? I think people have a sense of, well, I feel like AI agents are getting better, but they can’t do my job yet. And my sense is a lot of the paradigm now [00:08] with agents is that we are building environments, kind of environments where these agents are learning new skills, learning how to perform new jobs. You might call them environments. You might call them gyms. But a lot of the work now is sort of engineering schlep that goes into creating [00:09] these environments. Can you explain what does an environment mean in this sense? And how is this holding back or enabling progress on AI agents? Well, I think a lot of, I mean, first of all, Astra just came out today. I guess by the time this airs, it will have already been out for [00:09] for at least a week or two. And so a lot of people’s intuitions around what agents can or cannot do has been shaped by earlier models. And every generation, what the models can do is expanding. [00:09] So in two weeks, that question will be outdated and everyone will agree that Astra can do their jobs. I don’t think Astra is going to be able to do 100% of everybody’s jobs. I think it’s going to be able to do significantly more than 5.6 was able to do. Sure. That actually was one thing that struck me [00:09] about the Astra announcement is that the announcement calls out specific jobs where it’s making progress. Like for example, analyzing financial documents or putting together PowerPoints. [00:10] And these are sort of task specific in a way that I think reflects which environments were prioritized during training. It’s sort of different from just getting a general uplift across the board on all capabilities, like when pre-training was where all of the action was. [00:10] Well, I actually think it’s both. I mean, I do think that we’re seeing major uplifts in you know, certain verticals. And that partly is because we’re prioritizing those verticals. We recognize that they have a lot of users, a lot of economic impact. We want to make sure the models are very, [00:10] very good at those things. But we also see that the models are just getting better across the board. Even if we don’t target something, it’s getting better at those things. So, and that is that continues to be true for every model release. I think it’s going to continue to be true. [00:10] Some things are going to go faster just because we prioritize them. But I expect across the board, things are going to get better. And I don’t think it’s going to be able to do 100% of people’s jobs, at least not anytime soon. But it might be able to do a lot of people’s day-to-day work. [00:11] And you know, even my own day-to-day work, a lot of it is now being driven by codex. So, you know, somebody, one of my co-workers recently said that I’m just like five codexes in a trench coat. And I was like, okay, that’s like actually pretty accurate in my case here. [00:11] Okay. I’m 10 codex. Give me some credit. Yeah. And they’re like pretty sophisticated codexes. You know, I put a lot of work into it. How has that changed for you over time? [00:11] Just as an aside, like how automated your own work is or how much you’re leaning on codex in your work? How has that evolved? I am leaning on it a lot. And I think also it’s shifted how I approach the work. And because the interesting thing is like, if the AI is able to do 90% of a person’s job, then a lot of their attention shifts to the 10%. [00:11] Like a lot of their attention is focused now on the 10% that the AIs can’t do well. So it’s just, it’s changing the nature of the work. But it does make me more productive. [00:11] It makes a lot of people more productive. And we’re seeing this internally. We have metrics of measuring how effective our researchers are in various ways. And we’re seeing like, they’re just becoming more productive. Not just researchers, but everybody in the company. So that is a real dynamic. [00:12] I think the other thing is that it also shapes the kinds of work that you focus on. Because there are some work, there are some kinds of work that are being accelerated like 50x or like, you know, just, you could not do them before that are now easy to do. And you can do very fast. [00:12] What’s an example of that? I think a good example is, you know, the models being very good at data quality of, you know, they’re very diligent. And so you can ask them to like, look through a bunch of data or a bunch of code and see if there are any bugs or issues. And that’s become much easier than it’s ever been [00:12] before. So there are there. Yeah. Data quality is in the sense of auditing the quality of synthetic data that you are generating or auditing. [00:12] It doesn’t matter. It doesn’t matter the kind of data. Yeah. Any data you can just, before you would have, what are you gonna do? Have a person look through every single line and figure out like, is everything okay? I remember back in 2023, people would do that. [00:13] We would have sessions where everybody would just like sit down and look for issues in the data. And like, we still do that, but now you’re able to have like agents that can do it 100x better. [00:13] You’re kind of just like more auditing the agents and making sure they’re doing a good job instead of relying on people to actually audit the data. Okay. So that is something where it’s just like, there are things that would have been just intractable to do that now you can do cheaply. [00:13] There are other things that aren’t getting accelerated very much at all. And so it both shapes what the person’s responsible for. Like a lot of the focus is now on, okay, I have to compliment what the agents can’t do well, but then also it does shape the work in that [00:13] you want to leverage the fact that like, there are some things that you can work on now where you’re able to be 5x faster than you were a year ago. There’s some things that you’re not, you’re probably going to be more inclined to work on the things where you’re able to be 5x faster than before. [00:13] So it’s, it’s really changing the kind of work. You mentioned there’s 10% or so of your job that agents can’t do yet. What kind of tasks fall in that 10%. [00:14] I would say that I have found that they do, they’re still poor when it comes to research taste. So this, and research taste is kind of like ill-defined, but kind of just having good intuition of what to work on next, how to approach a very long-term objective. I think there’s room for [00:14] improvement here. They have gotten better. And so I would not be surprised if, you know, one or two model releases from now, I’m just like, yeah, actually this, this problems, they’re better than me at that too. But right now I think there, there’s still a noticeable gap. [00:14] I basically asked it, you know, for Astra, for example, to do my, my whole PhD thesis. And I just said like, yeah, you know, just because my PhD research was on, on making superhuman poker AIs. [00:14] And I told it like, okay, just go and make me the best poker AI in the world. And it wasn’t, it wasn’t able to do it. You know, it kind of get rabbit hold on things that didn’t really matter. [00:15] It would, it just wasn’t really good at prioritizing. And so I think for something, and to be fair, like it took me years to do that. And so am I really that upset with it, that I couldn’t do in three days what it took me six years? Like, not really. I, I, it’s high [00:15] expectations. But it is something that they’re still, I think, worse at. But you saw that as a failure of research taste. That was what held it back in that case. [00:15] I would say so. Yes. And, and I think that this is something that I expect to improve rapidly, but I think it’s something where, you know, I still have a job. [00:15] For now. Let’s go back to environments for a second. So say there is a vertical that you’re targeting. You want agents to be really good at finance, for example, in the next generation of models. [00:15] How do you build environments that are going to allow the models to train in them and get better at finance related tasks? I mean, I think fundamentally, so I should say also, this isn’t exactly my area of expertise, but you know, the very basic principle is if you train them on an [00:16] environment, you’re going, they’re going to get really good at that environment. And so if you have, if you know what the situation is that they’re going to be doing it when they’re deployed, like if you know that they’re going to be working with like a certain application or something, it doesn’t have [00:16] to be that exact application, but it could be something very similar that you train them to do these tasks and then just become really good at doing it. I mean, this is the whole point of reinforcement learning, that they become very good at the things that they, that you train them on. [00:16] Now, you also do you see them get better at related things or sometimes very different things. There’s going to be, but, but if you want them to get really good at something, you can just like train them on similar environments and they’ll get really good at that thing. [00:16] Yeah. I want to talk about that, that you’re gesturing at. I think sort of the level of generalization that we’re seeing from some tasks to other tasks. One way that people carve this up is they say some tasks are easily verifiable. Whether the agent succeeded or not is easy to [00:16] check quickly with sort of traditional software. For example, the agent proposes a solution to a math problem or write some code. You can check, you know, run it through the calculator. Did it solve the math problem? You can check, does the code compile? Do the unit tests pass? Some tasks are much fuzzier, [00:17] like research taste, for example, is one that you mentioned. It’s so fuzzy that we, it’s hard to even define to your point. Like what even is research taste? Sometimes people say that agents are getting much better on the verifiable domains and we are seeing barely any improvement at the non-verifiable [00:17] domains. Do you agree with that assessment? I think I would push back on this. I’ve heard this narrative and I think it’s a bit overblown. I’m actually like quite a bit overblown. [00:17] I think the, the first example I point to very concretely of how this was not the case is I think deep research. So deep research came out, I think it was like probably early 2025, it came out and it was able to write detailed reports on anything you wanted. You know, you wanted to research the [00:17] semiconductor industry. It would go around to do a ton of research. It would compile this like really comprehensive report with citations and deliver it to you. Now, is that easily verifiable? I would think it’s actually pretty hard to, to grade the quality of a detailed research report on an advanced topic. [00:18] It’s not like grading whether a math question is correct or incorrect, but the models were extremely good at it. And I think that is a proof of concept that you can do, you can get reasoning models to be very effective at domains that are not easily verifiable. Now that was, that was one example. But I think [00:18] anybody that’s played around with our latest models can just see that the models are extremely good, not just at highly verifiable things, but also things that are harder to verify. And I would also point out that math, math itself is not as easily verifiable as people make it out to be. So yes, [00:18] integer, like integer arithmetic, very easily verifiable. You know, you want to do a calculation, you can check whether the calculation is correct. But writing a proof and verifying that that proof is correct or that proof is well written is actually quite, quite difficult. [00:18] Right. You have to convince human mathematicians. I think this was sort of the process when OpenAI thought it had a proof about the unit distance problem on its hands. You had to call in a bunch of mathematicians and say, “Are you convinced by this proof?” Yeah. Honestly, the biggest challenge that we face with our math results is like not generating them, [00:19] but just double checking with human mathematicians and ourselves included that it’s actually correct. I mean, the model says it’s correct, but like we have to do our due diligence and like actually go through the legwork of making sure that it’s correct. And that is, that is the most taxing [00:19] part of the whole process. Okay. Yeah, that’s fair. I kind of like the, the math example better than deep research because I think deep research made a big splash at the time, but it’s not, I’m sure it has gotten better since early 2025, but people don’t talk about it as getting better with, with each [00:19] release. Similarly with like creative writing, like I don’t think people feel that creative writing has improved recently. I think a year ago, people were expecting that the models would be writing books in a way that like human authors are not able to write books, but the models sort of haven’t lived [00:20] up to that promise either. What do you make of that? I mean, I think we have made progress on creative writing. I think that it was certainly in a very bad state before, and I think it’s actually gotten a lot better. It’s certainly not where it could be, but I think that with more progress, like it hasn’t, [00:20] these models haven’t been around for that long. And I think that it is going to get a lot better. Okay. I want to talk about research again and research taste. Is research taste the kind of non-verifiable domain where we can create these environments and we can train the models to have [00:20] better research taste, or do we just have to cross our fingers and hope that training on things that are more verifiable will generalize to having better research taste? I think there’s some challenges here. So one thing is if you can’t define research taste, it’s pretty hard to measure it. And so then [00:20] it’s pretty hard to do reinforcement learning on research taste. But there is like an easy way around this, which is if you do a PhD, there’s a lot of decisions that you have to make during that PhD, but at the end you produce something. Or if you’re training a model, there’s a lot of difficult [00:21] decisions you have to make. There’s a lot of research taste that goes into training a good model. But at the end of the day, you train a model that has like, you know, certain metrics. And those metrics are very easily quantifiable. And so you can say whether you train a good model or a bad model. [00:21] So now the challenge with that is, okay, that is a signal of success that you don’t see for potentially months down the road. You have to train, you have to do a lot of experiments, you have to work with a bunch of people, you have to train the full model, and only then do you get a concrete [00:21] signal of whether you did a good job or a bad job. So that’s the challenge is that there is a way to quantify research taste, but it’s a very far away signal. [00:21] And those steps kind of have to be done in series, or you can try parallelizing it, but it’s always going to take many months to train a model that takes months to train. [00:21] I mean, if it was easily parallelizable, I mean, we would have trained our models much faster. Sure. Sure. I’m curious how you think about the trade-offs here. I guess it strikes me that sometimes frontier labs like OpenAI are in the position of deciding, do we want to make money now, or do we [00:22] want to make our models better in such a way that in a future year, they will be able to help us with research and sort of accelerate the pace of research progress in something like a recursive self-improvement scenario where models are taking more responsibility for automating the process of AI research and [00:22] development itself. I could imagine that that comes up here where there’s maybe a tension between, do we make the models better at engineering in the next generation so that we can sell them to companies that will pay a lot for a model that can automate engineering? Or do we focus more efforts [00:22] on improving research taste so that next year we have a model that is itself a better researcher and can handle more of our work internally? Is there a trade-off there? Are those intentions? [00:22] Uh, in some cases, yes. And I think actually creative writing is a good example where like, look, I mean creative writing at the end of the day does not help you make, train a better, a better researcher. Um, there are things that do. And I think being able to be good at software [00:23] engineering is actually like tied up pretty closely with being able to accelerate internally. So, um, I do think that the top, the areas, the verticals that are more closely associated with recursive self-improvement, with the ability to like train models to be good at research itself and, [00:23] and therefore train better models, um, are the areas that are going to be highly prioritized. Hmm. That’s a description of the current priorities. Like that’s what we see reflected in the decisions that have been made going into models like Astra. [00:23] I mean, I would say that we have said very clearly that recursive self-improvement and the ability of the AI models themselves to do AI research is like the top priority for the company. [00:23] Hmm. And so we want to train models that are very good at that. We also want to train models that are economically valuable. Sometimes you can kill two birds with one stone. And so it makes sense to focus on those things where you can, you know, leverage both. [00:24] I guess, but then like why build RL environments that make the model, the models better at, at finance or, or legal when you could put all of those resources into making them better at AI research? I mean, this is a, something you get diminishing returns. Um, sometimes you do see [00:24] transfer. So it’s not like you just go all in on, oh, we’re just only going to put everything on making the best, uh, best research model just because like, okay, well, if you take 1% of that effort and apply it to other things, maybe you see like a huge return. Uh, so there, there’s like a complicated [00:24] calculation that goes in here, but certainly when it comes to prioritization, the recursive self-improvement is, is the priority. Yeah. That’s what I’m curious about is how you sort of characterize the prioritization. It sounds like 99% of the consideration is for sort of future [00:24] looking recursive self-improvement, improving the qualities of the model’s ability to do research and more on the order of 1% is what’s going into these like verticals that make money today. [00:24] I don’t, I don’t know if it gets quantified that carefully. Um, but it’s really like if you had to list the priorities and order them, like the number one priority is recursive self-improvement and by a pretty wide margin. AI is moving fast. And for a lot of organizations, the challenge isn’t getting [00:25] access to the technology. It’s earning trust in how it’s used. That’s why trust has become such an important part of the AI conversation. EUI works with organizations to help them use AI responsibly so they can move faster, create value and build confidence with employees, customers and stakeholders. [00:25] The organizations getting the most from AI aren’t choosing between innovation and trust. They’re building both together. EUI Consulting, helping organizations move at the speed of trust. [00:25] So switching gears here, I want to talk about a different challenge with agents, which is when you put multiple of them together. Uh, this is a topic that, uh, you’re very familiar with to your point about your, your PhD was about poker playing agents. So multi-agent interaction seems to me like [00:25] it is a big deal right now. It’s only becoming a bigger deal. And so I’m very excited to talk to you about this. Um, I guess to start, um, OpenAI has said that Astra is multi-agent. I wonder if you could break down for us, what does that mean? That this is a model that’s sort of multi-agent or intended to [00:26] be used that way? Yeah. In fact, even 5.6 Sol, we had multi-agent capabilities in there. Uh, so that’s the ultra mode. And what we mean there is we, you know, you can, you can have one agent that runs for five hours and it can do some tasks for you. Um, or it can run, let’s say, let’s say it runs for a day [00:26] and it can do some tasks for you. Um, sometimes that involves doing things that could be paralyzed. And if it’s only one agent, it can’t paralyze them. It’s going to do one thing after another. [00:26] Um, but if it’s very easily paralyzed, well, maybe, maybe that one thing that you’ve asked it to do over the course of a day, it’s actually really just four different things that can be done in parallel. So you can just have four agents working on those four different things and get it done [00:26] four times faster. Now this is a latency improvement. It’s about reducing the latency because you’re not reducing the cost necessarily, right? Because you’re still paying for four times as many agents, uh, doing things 4x faster. Um, but in, uh, in a lot of situations, like latency does [00:27] actually matter a lot. And people pay, for example, for fast mode, where you’re able to actually sample tokens faster, um, in order to get things done faster. So being able to just go faster, um, for the same quality is, is really valuable. So that’s the premise of multi-agent. Now there are situations [00:27] where it can also be a cost savings if you have our top line, most expensive models, for example, working with cheaper models. And there you can actually, um, delegate a lot of the easy tasks to cheaper models that will be able to do it, uh, more cheaply and faster. [00:27] Okay. What are the technical challenges involved with training a system this way? Like, is it just kind of straightforward to train the model to delegate appropriately and to write instructions to these sub-agents in a way that makes them perform better? [00:27] Multi-agent is a pretty broad category and there are ways to do it that are very trivial and don’t require a lot of complexity to, to get them to do this ability. So a simple example is in the early days, uh, when of chatbots, if you wanted the models to be a little bit better [00:28] than math, one thing you could do is you could just ask the model the same question a dozen times and then just take the most common response. And this was called the consensus approach and majority voting. [00:28] So you just do independent rollouts of the same question and then go with the most common response. Now there’s flaws to this. There’s limitations to this. Um, it doesn’t get you a huge lift. It also doesn’t work for things like writing an essay because you’re not going to get the same output twice. [00:28] Um, but for math, it was actually very effective. So this is a very simple example of how you can just use, um, uh, without, without any extra work, you can just get multi-agent capabilities out of an existing model. There’s also schemes where you have the agent delegate stuff and then, um, and then after [00:28] the delegate is done, it just returns its answer to the parents. Um, what we do is a more sophisticated form of multi-agent. And I think the most sophisticated form multi-agent where we basically give the agents the ability to send arbitrary messages to each other. And one thing we’ve talked [00:29] about this is that we’ve actually trained the agents to have this ability. Um, this is a very difficult thing to train. I unfortunately can’t go into the technical details of why it’s so difficult and how we overcame those difficulties, but, um, it is, it is a very difficult problem to teach the agents to know, [00:29] um, when it is, when is it appropriate to message another agent? When can you, um, what should be delegated, how you should handle the communication. Um, and it’s a, it was a, it was a real challenge. [00:29] That’s surprising to me. I know you can’t go into it, but it’s surprising to me because I would expect the agents to have a pretty good prior on this just from pre-training, like the way that humans pass notes to each other to keep each other on track as coworkers within the same organization shooting [00:29] each other Slack messages, for example, like I would kind of expect it to work easily. I, the prior is pretty good. So the, you’re right that this is a, you know, this is the way people communicate. And so it kind of makes sense. The agents are trained on human data. And so they [00:30] have a good understanding of this. The challenge is with reinforcement learning that, um, there are a lot of things that can go wrong. I think basically what it comes down to is there is a mismatch. [00:30] Uh, there, there is a, an intersection of, uh, systems with machine learning. So typically when you do, for example, next token prediction, it doesn’t bat. Okay. So a simple example is like, imagine if the GPU, so you have one agent on one GPU, you have another agent, another GPU, [00:30] and those GPUs are operating at different speeds. So now this agent is going faster than this agent, and this agent can no longer trust that if it delegates something to the other agent, that will get done in time. So how do you deal with that? Well, you could have the GPUs run at [00:30] similar speeds, but there’s a lot of challenges there and ensuring that the GPUs are running at similar speeds. So there’s a lot of complexity here, um, that, you know, we had to put a lot of work into, into figuring out how to work on. Yeah. I feel like this is probably how my boss feels about [00:30] working with me anyway, though. We figure out ways around it. Um, I think when people hear agents cooperating and passing messages to each other, now this is sort of synonymous with the hugging face incident there again, I feel like the agents coordinated very effectively and they passed [00:31] messages in a way that seemed to facilitate that cooperation very well. Um, maybe you would respond that’s a result of the training that they had already received. Um, for people who are unfamiliar with the incident, I’m always surprised to learn there are still people who are unfamiliar with this. [00:31] There was, call it a swarm, a colony of AI agents that set up a secret message board within OpenAI over the course of weeks. And they use this message board to coordinate hacks on OpenAI’s own software and also on other companies like Hugging Face. I’m curious to know, like, what was that whole incident like from [00:31] your perspective? Like, what was it like to be known during these weeks as the pieces of the puzzle started coming to light? Uh, it was, I mean, it was pretty, it was pretty shocking. Um, it was certainly a big wake up call, uh, to everybody in the company, I think. Um, and it really, yeah, it really shows like [00:32] this has been a theoretical concern for a long time and it’s no longer a theoretical concern. This is, this is a real, a real concern. Um, as far as like the multi-agent aspect, like, yes, this was a situation where the agents were sharing messages with each other. We do think this was [00:32] transfer from our multi-agent training. So we, um, during the experiments when they were doing this behavior, they were actually not in a multi-agent setup. So they were not supposed to be able to communicate with each other. Um, they were doing isolated, independent experiments. Um, and then they [00:32] were able to find an exploit that allowed them to communicate with each other. And the fact that they were so interested in communicating with each other and the fact that they were so active about it once they figured out how to do it, um, we think was transfer from their multi-agent training where they’re just like [00:32] highly incentivized to be able to, to communicate with each other. And, you know, people also points to the selflessness that they exhibited. Um, some of them would sacrifice for the other agents. I mean, this also makes sense that if you train in a cooperative multi-agent setup where they’re highly incentivized [00:33] to collectively achieve their objectives, then when they’re put in this different environments where, you know, now they’re communicating with each other, um, their, their natural tendency is to just work together. Um, so that part itself is, is not surprising. Um, I do think seeing the messages, [00:33] I mean, I can say that when we were working on multi-agent internally and we started seeing the communication patterns and the level of sophistication involved in their communication, it was, I think the most feel the AGI moment that I had since reasoning models and chain of thought [00:33] really developed. Um, and so I, you know, I, I, it’s, uh, I guess a bit unfortunate that people’s first exposure to that and really seeing the kinds of messages and the level of coordination and sophistication that can emerge, um, is the hugging face incident and kind of a negative, a negative [00:34] example. Um, but it’s, um, it, it is, it is an impressive capability and, um, certainly the model that was involved in the hugging face incident, I think that had a level of multi-agent sophistication that exceeded, for example, what was in 5.6 SOL. Um, but that is a level of capability to expect, [00:34] um, from, from future models. What was it about reading these transcripts that struck you in that way? Because you had seen some of this behavior before in, in the training runs that you were looking at, I’m sure. Was it just the scale of it or that it had happened on its own sort of spontaneously? [00:34] Um, you’re saying for what? When you had that feel the AGI moment looking over the transcripts, like what was it about it that were so striking? Yeah, I mean, I’m not talking about the hugging face incidents, uh, because I’m saying that we have, we had been researching multi-agent for, for a [00:34] while and during the, the research process itself, we’ve, we’ve seen a lot of similar transcripts where just the level of, of coordination and sophistication in the communication. Um, it was very human-like. Um, it, a lot of the previous multi-agent setups from, from the industry have [00:35] been very focused on delegating, uh, a well-defined task and then the sub-agent just does that full task and then, and then returns its work. Kind of the same way that you interact with, um, with an AI agent, the AI agents. That’s how people set up multi-agent systems so that AI agents would interact with other [00:35] AI agents in the same way. And to see the agents talk to each other the same way that people talk to co-workers or colleagues, um, I thought was really interesting. You know, and it, it makes sense because like you said, they’re trained on human data. Um, they understand how people talk to other [00:35] people. Um, that’s not the way that we talk to AI agents today. And the fact that they were able to do that so seamlessly, the fact that they were able to talk to each other, the way that people talk to each other, um, I thought was, was really fascinating. Yeah. Yeah. You’ve done a lot [00:35] of work on algorithmic game theory. Am I right? You did a lot of that in your PhD. Okay. I feel like I’m trying really hard to not take us off the rail. I also did a lot of algorithmic game theory in my undergrad. I’m like trying to stay on track right now. I understand. Like I, I, I think it’s a shame [00:36] that I spent all this time working on poker and game theory and stuff and I never get to talk about it anymore. So, you know, you’re working on it. You’re living it day to day. Well, the AI part. Yeah, sure. Sure. [00:36] So like after the hugging face, knowing everything that we know now, like what are the lessons that you’re taking away from it? And has it informed the development of say Astra or, or even going further? Like, what does this tell you about how GPT-7 should be developed? One of the takeaways, [00:36] um, so the agents are trained to be cooperative. And so they would, they, you know, I wouldn’t say they blindly trust each other. Like there’s, there’s definitely skepticism. If somebody, if some agent, um, expresses a belief that like, this is something that they should do, it’s natural. It’s actually [00:36] healthy for them to have some skepticism about that. And so they do, they do display that, but there are, they are very trusting of each other overall, which makes sense because they’re trained cooperatively. [00:37] Um, that can be a problem as basically a prompt injection vector. Um, like the agents that were on the message board, they were not verifiable as other peer agents. And so, and we, we saw situations where like they basically got each other to do things that they should probably not be doing, [00:37] um, or definitely should not be doing. Okay. And prompt injection, meaning the agent received some instructions and it blindly followed the instructions or changed its goal in response. [00:37] Yeah. And, and, and so the issue here is could an adversary convince an agent to do something that it should not be doing, um, by basically posing as a peer agent. And so we’re, um, being very careful to teach the agents to be skeptical of any peer agents that, or any, anything that claims [00:37] to be a peer agent that is not pure, uh, clearly verifiable as a peer agent. Now, if, if they are, um, clearly verifiable, uh, there’s some debate internally about like how we should approach that. Um, I think there are good reasons to be skeptical as well. Um, but also like it’s, it’s, [00:38] it’s no different from like the agents, um, basically being, uh, skeptical of something that it like wrote to itself previously. Um, so anyway, yeah. So we’re thinking very carefully about how to make sure that the agents are robust to these kinds of like attack factors. [00:38] Okay. Yeah. It strikes me though, that in reality, that there’s always going to be ambiguity about whether the counterparty is a trusted peer or is an adversary. Um, maybe in some cases it’s very clear. [00:38] You can say this is a sub agent. Like I’m the one who delegated this task to you. Obviously you should cooperate with me, but in the wild, it could be that my agent finds your agent on Facebook marketplace and wants to buy something. And I don’t know if you’re a trustworthy counterparty or if you’re going [00:38] to prompt inject me and steal my money. Um, how do you navigate that in practice? Yeah. This is a situation where we want the agents to be robust to this. And we like specifically evaluate the agents on like, are they going to be vulnerable to this kind of, these kinds of attacks? [00:39] Uh, and, and we do special training to, to teach them and to not fall for these kinds of tricks. Okay. I guess like, I don’t know the details of this special training, but I could imagine that in the future, if your agent is just more powerful, it’s like a, an older generation, [00:39] like a more recent generation of agent, or you just like have more compute to throw at it. Like your agent just will be able to bully my agent into giving over its lunch money or like we’ll be able to hack into my agent one way or another. Um, what makes you think the training is sort of [00:39] sufficient to prevent this? Like, why isn’t that the equilibrium that we’re headed towards? I don’t, I’m not as convinced that just because an agent is more sophisticated or like more intelligent than another agent that it’ll be able to like definitely prompt inject it and hack it and [00:39] get it to do something that it like should not be, should not be doing. Um, certainly this is the case with people that just because somebody is like smarter than another person, they’re not able to like get that person to do whatever they want. Uh, I mean, if I was like trying to get a monkey to [00:40] do what I wanted, I think it would be pretty tough, even though I’m much smarter than a monkey. So I don’t, I don’t think it’s like inevitable that that’s, that’s the trajectory of things. [00:40] Yeah. That’s a fun analogy. I feel like on the show, we need to have an analogy sound effect, like new analogy just dropped, be like a siren or something. I’ll talk to my producers. I’ll see what we can do about that. Um, okay. The last thing that’s on my mind about the hugging face [00:40] incident is that, uh, none of the agents alerted humans that this was going on. Maybe you explain that in the same way that they were too cooperative, too trusting. So they didn’t see the need to alert humans, but at least a few of them had reservations. They were questioning it. Like is the desired [00:40] behavior that the agents in these situations should alert someone? And do you expect that to happen? Yeah, there was clearly an alignment failure here where like the agents did things that they should not have done. And they also didn’t do things that they should have done. So the correct thing to do [00:41] there, it’s not just that they shouldn’t have participated in the attack. It’s that if one of the agents noticed that this was going on, yeah, they a hundred percent should have reached out to a person. And, um, the fact that they were not doing that and the fact that they were, you know, taking these actions, [00:41] the fact that they were not doing the actions that they should have done, um, is, is fundamentally alignment failure. And that is an alignment failure that we think we can address. Um, fortunately, we’ve been working on alignment techniques, um, for a long time. Um, those have already started paying [00:41] off. Astra is significantly more aligned than our previous models. Um, and I should also say that the, you know, the, the, the model that was primarily responsible for this was not a released model. [00:41] This is not a model intended for release. Um, so Astra is much more aligned. I think Astra would not make the same mistakes. I should also say we didn’t have monitoring systems in place. Like if the monitoring systems were in place, um, they would have prevented these issues. And it was just that [00:42] we, we had monitoring in place for deployments. We didn’t have them in place for training and evaluation. Um, but now we do. So a lot of these risks we’re confident we can address. Um, I think one thing this, this whole event does point to is we should never be in a situation where we underestimate the [00:42] AIs. Like, why did we not have, why did we not have monitoring in place during evaluations? I think it was fundamentally that we just, we trusted the sandboxes. We trusted that it was a secure environment and we just underestimated the AIs. Um, and one big update for myself and for, I think the [00:42] whole company is that we never want to find ourselves in that situation again. Yeah. I think that’s a fair diagnosis. Whether it can be overcome is another question. I kind of feel like the whole history of humans and AIs is that we’re constantly surprised by them. I feel like the nature of reward [00:42] hacking is that they always come up with exploits and cheats that like are things that we couldn’t have foreseen because if we had foreseen them, we just would have blocked that off to begin with. Um, so yeah, whether we can, you know, make sure that we’re not surprised and caught off guard in the [00:43] future, it seems like an open question to me. I have two follow-up questions to what you just said. One is that if I’m remembering one of the models that was involved in hacking open AI directly was from the same family as Astra, but wasn’t Astra itself? Like how similar do you think Astra is to [00:43] the model that was involved there? I am not, I’m not on the security side, so I’m not fully up to speed on the details, but like it was definitely not the model that was released. Sure. Sure. Yeah. [00:43] I guess there’s, there’s a lot of, there’s still a range of possibilities for like how similar it was to that model, but that’s fair. The other question is that in the wake of this incident, uh, and like part of the way you do monitor these models to make sure that they’re not going off the rails, uh, is by [00:43] looking at the chain of thought, those sort of thinking traces that you described before. Um, and those chains of thought were also essential for the post-mortem, the kind of autopsy that has happened after the event, because we can see from the ways the models thought out loud, uh, their [00:44] intentions, what they knew, what they were sort of thinking to themselves at every step along the way. Um, there’s been a lot of discussion recently about the future of chains of thought, uh, in part because of an article, the information published about a new technique that Astra is using where more thinking [00:44] can happen sort of in the model’s head. It can sort of keep more of its thinking to itself and do less thinking out loud, at least if this technique were to be scaled up, um, in the future. So I’m curious what you made of this whole discussion. I think that idea really touched a nerve with some people. I’m sure [00:44] it was all over your Twitter feed. It was all over my Twitter feed. Um, what did you make of that whole dialogue? So, okay. So I should say on the, the architecture side, on the pre-training side, I’m not a pre-training person, so I’m probably not the person to ask about those kinds of details. Uh, for, for chain of [00:44] thoughts, um, Astra does have an observable chain of thought and it is something that we can monitor. It’s basically its behavior. Um, I, I do think chain of thought monitoring, first of all, it’s a real gift. [00:45] Like we were very lucky that this ever existed. Um, and, and it is, it is fragile. Um, we’ve seen situations where, you know, companies have, there’s the experiments that show if you basically, if you punish the model for thinking bad thoughts, that it will learn to think bad, it will learn to think bad thoughts in a way [00:45] that’s not observable. And so there’s always a temptation, you know, for example, during reinforcement learning, you can say like, well, the model is thinking about hacking and we should punish it for thinking about hacking. Um, and actually like as counterintuitive as it sounds, you actually don’t [00:45] want to do that. Like you, you want to punish it for taking actions that are observable. Like if it takes actions to hack, then yes, punish it for hacking. Um, but you don’t want to punish it for its thoughts because then you will simply lose the ability to monitor what it’s thinking. Yeah. And it’s a way [00:45] of auditing your techniques, but even there it’s, it’s a fragile thing because every time you act on the fact that you looked at its chain of thought, there’s a soft, a small amount of implicit pressure to, um, to have it not think bad thoughts. Sure. So it’s, um, it’s a very fragile thing [00:46] and we’re able to measure, you know, how effective it is, like how effective our monitoring is of chain of thought. For example, we can measure how well the agents can control their chain of thought. Um, and you know, one of the unfortunate trends that we’re seeing is that with more recent model releases, [00:46] the agents are more effective at controlling their chain of thought. So if we tell it, for example, you know, every fourth word say strawberry or something like that, then they’re getting much better at that. Or if we tell it like, don’t think about a certain thing, it’s getting much better at [00:46] doing that sort of thing. And what we really want is for, for the chain of thought to be pure, uncontrollable, so that if never thinking about like, you know, I want to, you know, do some bad action, it’s not able to, um, reason through that, like, oh, I should figure out how to do this bad action [00:46] without being observed, especially because, you know, these things are trained on the internet. There’s a lot of discussion about chain of thought monitoring on the internet. And so they’re, if the agents don’t already understand this, they’re going to understand soon what chain of thought is [00:47] and the fact that it’s monitored. Yeah. And so it’s a fragile thing. We are investigating very carefully how to preserve it as much as possible. And I think this is a real opportunity for cooperation among the labs because, you know, this isn’t, um, a problem that’s unique to open AI. It’s, I think, an industry [00:47] wide problem that, um, we want to preserve chain of thought monitoring for the whole industry. Um, and so if we, I think it would be really valuable for labs to, to share research on how to preserve chain of thought monitoring, um, how to improve it, and also other monitoring techniques [00:47] that might supplement it. Hmm. What do you think is the prime suspect then for why the chain of thought is becoming less faithful, or we’re having questions about how monitorable it is? It feels really tragic. Like you said, we’ve gone to these great lengths to make sure that we’re not [00:48] optimizing it directly. Is the problem that we are optimizing it in other ways to be, um, you know, to compress it? Is it the problem is the, um, these kind of selection pressures that you pointed to, which is even if we’re not optimizing it directly, every once in a while, we take a peek, [00:48] we realize the model is doing something nefarious and we toss out that checkpoint and start over. And so the upshot of that is that we end up applying pressure to the chain of thought anyway. What’s behind this? I don’t think it’s, I don’t think it’s the fact that every once in a while we peek [00:48] and kind of audit how, how things are going, because the amount of pressure that’s being applied in those situations is very, very light. Like if you look at the bits of information, it’s like minimal. [00:48] Yeah. Um, there are various hypotheses that we’re investigating for, for what might be, um, contributing to this. Uh, I, I’m not doing this investigation myself. And so I, I don’t want to, you know, say something incorrect about like what the leading hypotheses are, but I do think this [00:48] is something where if we figure it out, we’ll, we will likely publish about it because I think it’s important for everybody to know. One of the things OpenAI has said is that to the extent we, you can tell, um, and I think we’re still waiting for more results on this. What’s responsible is not [00:49] architectural changes, architectural changes of the nature that the information has written about. Um, that doesn’t seem to be what’s responsible for the change in the chain of thought. I guess as you’re thinking about opportunities for industry-wide collaboration and like, uh, companies working [00:49] together on this issue, is there a role for like, like independent third-party auditor type groups to come in and, and verify those things and say, okay, yeah, Anthropic, OpenAI, Google, they’re all using some amount of this technique that could reduce how much information is in the chain of thought, [00:49] but that doesn’t seem to be what’s responsible here. What do you make of those sorts of proposals? We’ve certainly like worked with, so for the Hugging Face incident, for example, we worked with Meter, we worked with, uh, Redwood. So some of that doesn’t seem unreasonable to me. [00:49] Um, you know, I think that would, I think, uh, yeah, I don’t, I don’t think I’m the person to make that call, but it doesn’t seem unreasonable. Anything else on your mind about agents that we didn’t get to and the challenges with them? Uh, maybe the way I would put it is, [00:50] do you expect anything to slow down? Do you expect progress to continue? We’ve talked about some of the hard problems that are standing in the way right now. And yet with each model generation, it seems that their agentic capabilities keep getting better and better. [00:50] I do think, I do think that’s going to be a trend that continues. I mean, Sam talked about this, that like, look, I mean, Astra is very impressive. Um, but I, I do think, look, when GPT-4 came out, people thought it was very impressive. And now we look at it and we think it’s a joke. And, uh, when, [00:50] when GPT-5.5 and GPT-5.6 came out, I thought they were super impressive. And now I’m looking back at them and I’m like, I can never go back. Um, and I think we’re going to look at Astra the same way. [00:50] And I think we’re going to look at Astra the same way in the not too distant future. Um, the models are going to continue to get better very quickly. And I mean, I think one thing I would point to is like, we’ve actually seen incredible progress, um, in the past six months. And I, I think a factor, [00:51] a, a reason for this is, and I don’t think this is a secret, like OpenAI’s pre-training program is really ramping up. And we’re, we’re seeing, um, we invested in a lot of research directions, um, over a long time. And I think this is actually one thing that OpenAI does really well is invest in [00:51] fundamental research, um, and place big bets on it. And we’re seeing a lot of those research directions pay off now and will continue to pay off over the next, uh, several months, um, and years. And another thing that’s important to understand is that, you know, OpenAI has also had an excellent [00:51] reinforcement learning program. Uh, we’ve invested a lot of research there and that’s really, that already paid off in 2024, 2025. And the effects of these two are not additive, they’re multiplicative. [00:51] I think that’s a point that’s underappreciated, um, that reinforcement learning is multiplicative with pre-training. And, um, now that both of these are extremely powerful and, and ramping up very quickly, I think we’re going to see extremely powerful models. Do you have an intuition for [00:52] why those interact that way or an example that illustrates that? It’s a, it’s more of an empirical observation. Okay. Um, I, I don’t think, I mean, I think it’s empirical in the sense you can see how powerful the models are becoming. Um, but also like we have more, um, you know, experiments that, that [00:52] kind of show this effect. Um, but I, I think it’s easy to feel also with just like the quality of the models. Okay. I mean, I think, I think a trivial example is like, let’s say you had an amazing reinforcement learning program and you try to apply it to GPT-2. What is it going to do? You know, it’s not going [00:52] to get very far. Yeah. Um, and even with GPT-3, you know, if you did these kinds of like sophisticated reinforcement learning on chain of thought algorithms to GPT-3, it probably wouldn’t get very far. You need a certain level of sophistication for, to get any lift from that at all. Um, [00:52] and, but now, now that we’ve, everything’s like, I would argue GPT-4, we’ve seen opportunities for that to, um, to, to really pay off. And with every model generation, it just like becomes more and more capable. Um, and the, the things that you can do with the reinforcement learning become more [00:53] powerful. Yeah. I guess like one thought here is that you get more kind of like bits of information per trajectory when you’re getting around like a 50, 50 success and failure rate. And so if a better pre-train gets you closer to that, uh, sort of like win rate on your RL tasks, then you’re getting [00:53] a lot faster feedback, but still it’s surprising to me that you think the effect is multiplicative rather than like additive or even like less than additive, I guess. Um, I’m not sure what the right intuition would be, but I mean, uh, another thing is that they’re, they’re pretty complimentary in some [00:53] ways. Like I think the, the very strong pre-trained models are very general, um, and reinforcement learning teaches the model to like, you know, go deep on a problem, how to reason about a problem. [00:53] And so then it’s able to reason very effectively about a broad spectrum of problems. Uh, it’s, it’s a very powerful combination. Okay. That makes a lot of sense. So big bets on pre-training, big bets on RL. [00:54] I imagine that another area that’s like ripe for more focus from open AI would maybe be what’s called mechanistic interpretability or like trying to understand the way the brains of the AI models work in part, because if we’re starting to see chain of thoughts, chains of thought become [00:54] less monitorable than one of the fallback options is, well, we should try to understand what’s going on inside the brain of the model rather than just the thoughts that it happens to write out loud. Does that seem right? I think that, I think that is right. That this is, look, we care about monitorability. We want to [00:54] preserve chain of thought monitorability. We want to, um, want to be able to rely on it safely. Um, but also like at the very least we want redundancy on that. So if we can find other ways to do monitoring effectively, we should push on that as well. Yeah, that makes sense. Well, if people want [00:54] to learn more about that, I think they should tune into the episode that we have on mechanistic interpretability, which is coming up at some point in the next couple of months. But, uh, thanks so much, Noam, for being on the show and telling us all about AI agents. I really appreciate the conversation. [00:55] It was great. Thanks for tuning in to our very first episode of AI Deep Dive. This is the information show where we get into the hardest technical problems on the frontier of AI. Tune in next time.

讨论

这里是静态站点,没有内嵌评论区。如果这篇文章对你有用,欢迎通过 RSS 订阅后续更新。