Lee Robinson《Always-on Agents》讲座英文逐字稿(Stanford CS146S,45 分钟,带时间轴)
这份英文逐字稿是对 Lee Robinson 在 Stanford CS146S 讲课视频完整音轨的转写:从开发者与 AI 协作的四个阶段讲起,接着是模型「变了什么」(工具调用、computer use、文件即记忆、skills、routines),然后是常驻智能体的内部构造(睡/醒、Firecracker 沙箱、thin client / thick server、durable workflow、优先级与批处理)、harness 的设计(单一 send-to-user 工具、动态工具发现、三层信息检索阶梯、helper 子智能体、auto-review 安全),以及上下文工程(prompt caching、cache miss 即 bug、在热缓存窗口内做 compaction、记忆分层)与训练飞轮。按章节、逐段带时间轴,产品名与专有名词做过校正。
演讲人:Lee Robinson(@leerob,SpaceX AI)· 场合:Stanford CS146S(The Modern Software Developer)· 时长:约 45 分钟 来源:X 帖文 Lee Robinson — How do always-on, proactive agents like @Bot work? Watch my lecture at Stanford CS146S(2026-10-09) 中文精读讲义:《Always-on Agents》讲义:Lee Robinson 讲常驻智能体的积木
说明:这份逐字稿是把推文里那段视频的音轨交给本地 Whisper(large-v3-turbo)转写的,不是官方字幕。 口语重复、语气词与断句基本照录,未作润色;只对少量明显的专有名词做了校正(例如 GrokBot、SpaceX AI、OpenRouter、Temporal、Firecracker、Moltbook)。 演讲中的产品名、版本号与数字都来自他本人口述,本文未独立核验。 时间轴对应视频;章节标题沿用他在帖文里给出的章节(How we got here / What changed in the models / Inside an always-on agent / The harness / Context engineering / Where this is going)。
本章节速览:
00:00开场 ·02:35How we got here ·06:13What changed in the models ·10:06Inside an always-on agent ·18:07The harness ·27:26Context engineering ·34:26Where this is going(至45:16结束)
开场
00:00 So as mentioned, I’m Lee. We’re going to talk a little bit today about how this new category of always-on, proactive, persistent agents work and kind of give you some of the understandings of the building blocks for how these are made.
00:13 Just as a, you know, fun statement up here, some of the things we’ll talk about are in progress. We’re working on it at SpaceX AI. And in week one, you all built this very kind of simple harness that helps you understand the core idea, which is that you have a model, it can call tools in a loop, and that’s actually very powerful. It unlocks a lot of new capabilities for the models, but it does come with some drawbacks.
00:38 Notably, you close your laptop and the process stops. It’s no longer running. And I want to talk about how you can take this simple idea, it’s very powerful for how to get more out of models, and extend it to this new type of agent, which is an always-on agent.
00:53 So rather than just running in your terminal, this can run on a server. It can run in the cloud. You can close your computer. It can restart. It can survive crashes, for example.
01:02 This agent has its own computer. So, not just your laptop, but a Linux machine in the cloud, where it has its own desktop files, browser, more, maybe some 3D apps.
01:13 It can save files on the machine. It can understand more about your progress over time, and build up a memory of how you work with it.
01:23 So, unlike a harness in the terminal that is responding to every message that you send, this type of always-on agent is more like a colleague, where it doesn’t always immediately send you a message.
01:34 It knows when to be quiet. Sometimes it might just want to work in the background a little bit, and only come back to you when there’s an important update.
01:40 And it can be triggered more than just you typing. So, it’s not just you prompting it in the terminal.
01:46 Maybe there was a Slack message, or an email, or some other type of event that wakes up this agent to go do work.
01:52 And ideally, the hope with these type of always-on agents is that they can do work over very long periods of time.
02:00 Not just hours or days, but maybe even weeks or months, and that just requires a little bit of a different architecture, which we’ll get into.
02:07 So, we’re going to go decently in depth on how we built this architecture, how we built the harness, kind of what’s got us here today, and where we’re going in the future.
02:18 And we’ll get into some of the specifics of how we’ve engineered this system to be proactive, to be persistent, and to follow instructions very well.
02:27 And, of course, time at the end for questions, if you all want to ask anything about this, or anything outside of what we’ve covered today.
02:34 So, first, how we got here.
How we got here:我们是怎么走到这一步的
02:38 There’s kind of been four eras of developers working with AI so far, I would say.
02:44 The first one was copy-pasting code out of ChatGPT, and putting it into your editor of choice, which feels like, you know, decades ago at this point, but it was really only a few years ago.
02:55 And then, once the model started to get better at being able to have kind of that simplified harness we talked about in the last lecture,
03:05 they were able to edit files, run shell commands, kind of have this very simple interface for how you work with them in the terminal.
03:13 And this really took off because, all of a sudden, the model had this new capability where it could go do more proactive work for you and be more productive.
03:22 Then, this third era was taking the coding agents and where they were kind of simple to start and giving them these dedicated apps.
03:30 And at this point, people were starting to run many agents in parallel.
03:33 Maybe they’re running them in the cloud, not just on their machine.
03:36 But still, at this point, you were kind of the bottleneck in that you were reviewing every command that they were running.
03:42 You wanted to make sure every shell command was just right.
03:45 You’re kind of really watching them and being in the loop on everything.
03:48 And now, with improvements in models, in harnesses, in products, we’re now in this new era where you’re starting to have these always-on agents.
03:56 And they can kind of run ambiently in the background where they have their own computer.
04:00 They have their own memory.
04:01 They learn new skills on the job.
04:03 And they keep working after you go to bed or go do other work.
04:07 A good example of this is we’ve kind of moved away from this world of prompting an agent every single day,
04:14 although we still do that sometimes.
04:15 But you can set up agents where it will kind of just run in the background.
04:19 So maybe in the before times, it was you would prompt an agent, you would wait.
04:23 You’d prompt an agent, you would wait.
04:24 And if it asks you to run a shell command, you’d have to approve or disallow that command.
04:30 And now, instead of kind of queuing commands, you can steer the agent.
04:35 And you might say, “Hey, go check this thread.”
04:37 And then immediately after, “Oh, by the way, actually go do that on Thursday.”
04:40 And prior generations of models would get kind of confused when you would immediately interrupt it and give it some other command.
04:47 But now they kind of understand just like how you might ramble on talking to a friend and they can kind of parse that these two things are related.
04:53 The models and the harnesses can now handle that.
04:56 And if you say, “Hey, go send this email, go send this Slack message,” there are additional safety checks in place, which a lot of companies are calling auto-review, which is essentially a separate AI model that’s checking all of the commands that get run.
05:11 So, you know, kind of funnily enough, that’s actually more secure than the fatigue of having to watch hundreds of shell commands and approve each one.
05:20 We’ve kind of reached the point where if you did a blind test with what the AI model is checking on the commands and what a human is checking on the commands,
05:26 most of the time the model is actually doing a better job.
05:29 Kind of paradoxically, but as the models have gotten better, the interface has gotten quite a bit more simple actually.
05:38 And this type of always-on agent we’re calling a bot.
05:43 But as you’ll notice, there’s actually quite a few of these new types of products popping up right now and gaining a lot of traction.
05:49 And I think the main reason is because we all know how to text.
05:53 You know, everyone knows how to text.
05:54 Everyone is pretty proficient with using their phones.
05:56 And we’ve kind of distilled down this way of working with AI that feels like you’re just texting a friend,
06:02 but it can actually go and do real work for you.
06:04 It’s not just answering questions in a Q&A style setting anymore.
06:07 It’s actually going and doing longer tasks like you would hand off to a coworker, for example.
What changed in the models:模型变了什么
06:13 So there’s been a few changes specifically in the models that have enabled this to happen.
06:19 I’ll go over six, but there’s, you know, quite a bit more.
06:22 The first one is that the models just over the past year or two years have really gotten better at being able to follow instructions over a longer period of time where they don’t get confused or make mistakes.
06:34 And now models can pretty confidently run for hours or days with the right compression and compaction, which we’ll talk about in a little bit.
06:42 Two, models now can essentially hand off to helpers.
06:46 And those helpers are called sub-agents.
06:48 And the reason why a sub-agent is helpful is the model has kind of its working memory of the process that it’s on.
06:55 And that’s going to fill up eventually.
06:58 So it’s nice if it can hand off to these helpers that have a fresh set of working memory.
07:03 And that effectively can give it an unlimited context, basically, where it can go and do work that might be expensive in terms of the tokens it’s using or the tools that it’s calling.
07:12 And then come back to the main agent and report its updates.
07:15 And I’ll have some diagrams of that later.
07:18 Models now not only are just better at calling tools.
07:21 I mean, just a few years ago, they basically couldn’t call tools very well and they would hallucinate all the time.
07:25 It’s funny, you don’t really hear people talk about hallucinations anymore because it’s kind of a solved problem, mostly.
07:31 But not only can they call tools, but you can give them very powerful tools the same way as if you onboarded somebody onto your team.
07:38 So obviously running shell commands and files, but also you give them a full computer, a full Linux computer.
07:44 You allow them to make demo videos and such.
07:47 They can click around on a screen just like a human would and they’re getting increasingly better at doing that.
07:53 Otherwise known as computer use, which basically means that any SAS product or any website that didn’t expose an API or an MCP server can be used headlessly, basically.
08:05 So autonomously through computer use by these models and these bots.
08:09 You can store all of the information in files.
08:12 And the reason you can do that is because it’s actually very easy for the models then to recall and search this information, which we’ll talk about.
08:18 And then finally, you can just kind of show the bot how to do some work.
08:22 And it will kind of turn that into skills or memories or routines or things that it can reuse later.
08:29 So on the file thing, for example, this is kind of a simplified example.
08:32 But you might remember or you might imagine that when you’re working with a bot under the hood, it might look something like this.
08:39 You build up this profile or this understanding of the user and that’s going to be included in every single prompt.
08:45 Maybe you like lowercase text-like messaging and every single message should have that.
08:52 Also, as you’re working with these bots over time, you’re building up kind of like a working log.
08:57 Just like if you were working with a colleague and they were writing down notes every day of the work that they’ve done.
09:02 So that’s all being put into files that the model can then decide to search or use later.
09:06 If you say, hey, what did I do last week?
09:07 It just can go and look that up.
09:09 And the models are very good at using Linux commands, running shell commands to go grep and search and look through these files.
09:17 Similarly, with skills, you are essentially taking your knowledge of how to do a specific task and encoding it into a markdown file.
09:27 The very specific process you need to do, the things that you believe make that successful or unsuccessful.
09:33 And that’s not just your skills.
09:35 There’s a whole ecosystem now of companies and developers writing these skills for how to send a great email or how to build a great API or many different things.
09:45 And for example, in GrokBot, there is a whole system of plugins.
09:48 So basically any tool that you want to use, there’s a plugin you can install which uses skills or servers.
09:53 And then also routines, which if you’re familiar with cron jobs, you know, you’re running some prompt basically on some kind of schedule, which is pretty useful as well.
10:02 So basically everything is computer.
10:04 Part three.
10:05 So let’s go a little bit further into the kind of the mind of this always on agent or going behind the scenes.
Inside an always-on agent:常驻智能体内部长什么样
10:13 What does it mean for the agent to be always on?
10:16 You can kind of read this chart from left to right.
10:19 But it starts out with the server being asleep and also this computer that the agent can use being asleep.
10:26 And so whether you send in a message or some kind of trigger decides to make the agent start up, which we’ll talk about in a second.
10:34 That’s basically loading up and booting up all of the state for this agent and it’s running on a server.
10:40 If it decides that it needs to use a computer, maybe it wants to go use the computer’s browser or, you know, go click around on the computer, it can start up the computer as well.
10:49 And then once you’re done using it, it also has the ability then to basically put those to sleep.
10:53 Why is this important?
10:54 It would be, you know, expensive and also there’s not enough computers in the world to just have those computers running all of the time for everyone as well as a lot of server usage too.
11:04 And so this state can then be persisted to, you know, a database, the conversation logs, the memories, the routines, all of that.
11:13 How does the agent get woken up then?
11:15 Well, obviously you can prompt it directly, but I think what’s increasingly interesting is having other things running in the background that can then decide when you want to go reach out to your bots.
11:25 Maybe that is somebody pinging you on Slack and your bots can live in Slack.
11:29 Maybe that’s a phone call or a meeting that you put your bot in.
11:32 It’s increasingly common now at SpaceX AI where we see bots joining meetings for employees where they could make it to the meeting, but they want their bot to go and just take notes and come back to them, which is kind of funny.
11:45 Maybe you had some routine or some background work that was taking, you know, hours or days to do and then you want to come back and give you a summary.
11:52 That all happens.
11:53 And then we have basically a router that dedupes that and figures out how to put it into the conversation.
11:58 And then it has this heuristic of, do I need to bug the user about this or not?
12:03 Because not everything is actually important.
12:05 Just like if you have a coworker working on something, it would be pretty annoying if every single time they made progress on something, they came and bugged and were like, hey, check it out.
12:12 Like I made progress on this thing.
12:13 You really only need a status update maybe, I don’t know, every day, depending on the task.
12:17 So I talked about how the agent has this computer, this Linux box.
12:23 This is running a firecracker virtual machine.
12:26 And the reason this is really helpful is because it’s very easy then to essentially snapshot the state of the machine and shut it down and start it back up.
12:34 And on this machine, you of course have files, but you can kind of connect into the machine and have a live view of the machine and actually control it yourself if you want.
12:43 The machine can, of course, run commands, which is great if you’re going to run some shell commands.
12:48 Maybe you don’t want it on your personal laptop.
12:49 You want it in this other environment that is, you know, secured and locked down.
12:53 And then another interesting thing is, of course, it has Chrome and it can use a browser, but you can also make it essentially have like virtual desktops.
13:01 So if you’re running five bots, ten bots, and each one of them needs their own browser, you can essentially let them have different screens.
13:09 And all of this, again, is kind of shared infrastructure between these bots where they can go call on these computers.
13:14 So we talked about kind of the triggers that will wake up the agent.
13:19 That’s kind of on the left.
13:20 So different apps or slack messages or emails.
13:22 Then you have this middle piece, which is the server.
13:25 And then on the right is the computer that we just talked about where you can run files and commands and other things.
13:30 So let’s talk a little bit more about the server.
13:32 I already mentioned, you know, commands will come in.
13:35 They get put into this queue.
13:36 And then it calls the purple box, which is the agent loop.
13:39 And this is what we covered last week.
13:41 Like this is a model calling tools in a loop.
13:46 And then when you shut it down, you can kind of save the state to the database or storage.
13:51 So it’s like it’s interesting, I think, to look at this and realize there’s this one kind of small piece that can then have all of these abstractions built on top to create a very powerful product when you give it all of these helpful tools and you connect it to all of these different services.
14:06 The reason the agent runs on the server is a pretty common practice in computing where you want to have kind of a thin client and a thick server.
14:15 And computing has kind of oscillated between these things over the years.
14:19 But, for example, a few good reasons for this is let’s say you need to push an update to the, you know, the agent code.
14:25 It would, you know, kind of suck if you had to push an app store update just to get that code out there versus just pushing a bunch of server updates very quickly.
14:32 Also, sometimes you need to do heavy computation work, you know, like transcription or image generation or something.
14:38 And that’s probably going to be faster on a server than on device depending on what the devices are.
14:44 And, you know, there’s very established ways that we know how to scale servers, especially with variable traffic and variable growth.
14:52 But the most critical thing here is that if you put the agent loop on the virtual machine, then you can’t shut down the virtual machine independently from the agent server.
15:04 So, yeah, you want to save everything on the server so you can always kind of rebuild that virtual machine.
15:11 For example, let’s say you need to do some critical zero day Linux update on the VM.
15:18 Well, it would be great to be able to rebuild that and not, you know, blow away the entire agent loop and have downtime in the application.
15:24 When you hit send, this is maybe a little in the weeds, so I’ll just kind of gloss through it.
15:31 But basically, if you want to dig more into learning more about distributed systems and thinking about, you know, strong versus eventual consistency, which are just fun topics to get into, you know, you send a message and you kind of want the client to immediately identify that and acknowledge it.
15:48 Then when you go into the turn, you’re obviously putting these tools in a loop with the model, saving that to the database.
15:55 And then the clients are essentially putting subscriptions with all of the servers so they know if something changed and they’re listening for these signals.
16:03 And if they get a signal, then they can go and essentially fetch the latest information from the server.
16:07 And there’s some fancy deduping you can do to make sure if, you know, if there were some kind of retries that you would not show, you know, multiple messages in the list.
16:17 And so the, like, underlying infrastructure primitives to think about here is a workflow versus a job queue.
16:25 And the main difference here is with a job queue, let’s say you’re kind of going along on steps, you ask your agent to go and do some kind of complicated task.
16:33 And it’s making its way from step one to step two to step three.
16:36 But then for some reason there was an infrastructure issue, there was some problem.
16:42 With a normal queue, if that crashes and you go to restart it, well, now you have to replay steps one, two, and three, which is obviously that could be pretty problematic.
16:51 So what you really want here is you want a durable workflow, which means that it’s safe for you to actually retry that.
16:57 And when you retry after a failure, it picks right up on step three.
17:01 And this is, like, kind of a solved problem in some ways in distributed systems and in infrastructure.
17:07 So we use an open source tool called Temporal, which is interesting to look into if you’re curious, so that we don’t have to rebuild a lot of this very complicated stuff from scratch.
17:16 Of course, you still have to run the infrastructure for it, but some of the pieces to put together there.
17:20 So let’s say a new message comes in.
17:25 There’s different priorities based on whether I’m messaging my agents or messaging my bots versus if it’s in a group chat or a Slack thread or some background routine, for example.
17:37 So if I say, oh, actually, you know, change the date for this event, we actually want that agent loop to stop.
17:43 And my message is going to have priority over everything else that’s happening, which kind of makes sense.
17:48 For everybody else, you’re kind of putting things into the queue in the back of the line.
17:53 And if the agent gets, you know, three or four messages very quickly, it’s able to kind of batch those together.
17:59 So it’s just one turn.
18:00 You can kind of figure out whatever heuristic you want there.
18:03 And then it kind of goes through and does each turn at a time.
The harness:harness 的设计
18:07 Okay.
18:08 So let’s talk about the harness itself.
18:11 And this is going to kind of take the basic harness that we talked about last week
18:16 and add some new tools, some new concepts to it.
18:20 The biggest difference with the GrokBot harness and probably some of the other tools like this on the market is
18:29 the way we’ve done the thin client fixed server architecture is there’s only one tool that communicates between the client and the server.
18:37 And it’s just send to user.
18:39 And the great thing about this is that it’s actually very architecturally simple where the client is just texting the server in some ways.
18:48 And you’re putting a lot of these heavier tools on the server.
18:51 And there’s many of those tools.
18:52 There’s running shell commands, reading files, but also calling on a cloud agent or recording your screen or doing demos, for example.
18:59 And this, I think, has worked pretty well because now regardless of where you spin up a new client, whether it’s a mobile app, a server, a desktop app, email slack,
19:09 it all just has this one single tool that communicates back and forth with the server.
19:13 On the client though, even though it just has this one tool, when it receives the tool, it can decide how it wants to display that in a rich and interactive way that makes the most sense for the user.
19:24 So if it gets information back from the server that the user is trying to log into something, it can kind of show these unique sign-in forms and connect to tools like 1Password.
19:33 It can go and call on services like Stripe Link or use XMoney or other services to go and allow people to pay for their services.
19:42 If you’re trying to do something that you can’t reverse it, sending an email or a Slack message, it might come back and ask you to edit and approve a draft before you send it.
19:53 And then sometimes it just, you know, you can just leave an emoji reaction if you don’t actually need to send a message.
19:59 So going back to what I was saying with dynamic tool discovery.
20:03 So for any model, there’s a context window, which is kind of its working memory.
20:09 And that’s some limit in tokens.
20:11 And if you just jam every single tool possible into that context window, you’re not going to have a lot of space.
20:19 And the way the tools work is you have kind of the name of the tool.
20:23 You have the schema.
20:24 You have all these different details.
20:26 And if you start installing your Google Drive connector, your Gmail connector, your YouTube connector, like all these different services that use MCP servers, which all have tools, you’re going to start looking at something like the top, where a lot of the blocks are already filled.
20:43 And the problem with that is it doesn’t give the model as much space to actually do the work of the conversation.
20:49 So one kind of context engineering or harness trick that a lot of the popular harnesses now use is it’s called dynamic context.
20:59 Or I think some other folks call it the tool search tool, which is a fun name.
21:04 But basically, you’re only putting the names of the tools into the every message that gets sent to the model.
21:11 And then the model can decide, okay, I actually want to go use this tool.
21:16 I can go read the information from a file.
21:18 It’s kind of files and a file system all the way down for all of these things.
21:22 And that just helps you save on your context usage, which then also helps save on costs, save on usage, just better for everybody.
21:30 In general, the principle here really is to try the cheapest option or the most reliable option first and kind of work your way up the ladder here.
21:39 So the easiest one is just what the agent or what the bot already knows based on the conversation history.
21:46 If it doesn’t, if it’s not in the immediate context, maybe there’s a plugin, an MCP server, some kind of API that it can go and pull, you know, your banking information from Plaid, for example.
21:58 Maybe that doesn’t work and it needs to go do some kind of web search to see, you know, what was the score of the Cubs game last night.
22:06 Maybe if it doesn’t have that, it actually needs to go to the computer and open up the browser and do some searches on there.
22:14 And finally, like maybe it needs to go and actually use the entire desktop, so run scripts or run commands or build things on that Linux computer.
22:21 And only if all of those fail, then go bug the user and ask them to go give input or approve something.
22:27 And this really helps it feel like what it would be like to work with a colleague.
22:30 It’s like they’re not going to bug you for all these different things.
22:33 Ideally, they’re going to try a lot of these options first.
22:37 And on that computer, so the Linux machine that the agent is able to use, there’s kind of three different ways that it can retrieve information from it.
22:47 Of course, it can just take a screenshot of the screen, but there’s more efficient ways of doing this.
22:52 It could also look at the HTML of the page, but you can also look at the accessibility tree, which is a nice way of like cutting down a lot of the token usage.
23:01 And basically to help you find the buttons or links or other elements that you want the browser to click on.
23:07 Of course, if none of that works, you can just click on pixels.
23:10 So some websites, for example, have modals or dialogues that pop up that you need to click.
23:15 And so your model, your agent needs to have that functionality as well.
23:19 Or again, kind of worst case on that ladder is you go back to the user and like, oh, there’s like a two factor auth code or some kind of captcha or something like that.
23:30 So the idea is that kind of putting the pieces together, you have this main agent and it’s kind of like your orchestrator.
23:39 And the main agent, you want to keep the conversation as sparse as possible because it now has the ability to go call on all of these different helpers.
23:50 And the helpers are just sub agents.
23:52 And so this is helpful because if you hand it off to a helper to go run a bunch of shell commands or do a lot of expensive tool calls,
23:59 all of that work that ultimately is ending up in some answer to a question or a task, none of that is bloating the conversation in the context of the main agent.
24:08 It’s only coming back with the result.
24:10 And this is how you get to a product that feels like it has infinite context.
24:15 Obviously, infinite context is not a real thing.
24:18 Every model has a context limit.
24:19 But there are these tricks and workarounds such that, especially when you train models to get better at this,
24:24 that it doesn’t feel like you’re running up against the limits of the context window.
24:28 And you don’t feel a degradation of the model’s quality as you get further along that as well.
24:33 So we have these different helpers, each with different skills.
24:37 Each one of these is its own durable workflow.
24:40 So going back to the cues versus workflows thing, that means if it fails, the video helper crashes,
24:45 something uses up too much memory or something that can restart and keep going without losing progress along the way.
24:52 And then the main agent, it is kind of like a coworker where it’s managing a to-do list.
24:58 You have a few things you can do.
24:59 You can go check on it.
25:00 You can send it a message.
25:01 You can stop the helpers, for example.
25:04 And so a lot of this is modeled after the ideal way you would scale a team of engineers working on a very complicated problem.
25:11 So just a few little interesting tidbits from building this out on the team.
25:17 I think we’ll talk a little bit more about prompt caching here, but ideally you don’t add or remove new tools during turns.
25:26 And just like I mentioned with the slide with all of the boxes, if you have extremely long tool descriptions or schemas,
25:34 that’s going to take up a lot of space in the context window, which is bad for pricing.
25:38 It’s bad for usage.
25:39 Ideally you want to minimize that.
25:40 So there’s plenty of optimizations you can make here on the engineering side to just trim that down.
25:46 Another interesting one is like you might hear people talk about how, oh, like we can have a harness that just has one tool.
25:52 And that one tool is like running shell commands.
25:54 And this is true, and you can model a lot of things this way.
25:57 But if you see the model doing, you know, 90% of its time doing shell commands for this one thing,
26:02 it actually can be more efficient to make a dedicated tool for some of that stuff.
26:06 And it’s a bit more legible for people actually reviewing how the product is working.
26:11 And the last one’s kind of funny too.
26:13 Obviously if you’re writing errors for humans, it’s helpful to make it very descriptive.
26:17 Like, okay, this thing failed, like click on this link or here’s what you need to do next.
26:21 Turns out that’s also very helpful for agents.
26:23 So there is some DX in designing good error messages for agents too.
26:27 And when you’re trusting your bots to go and do work over a long period of time,
26:34 it is very important that you trust them to be secure with your information,
26:39 to be secure with how it operates.
26:41 So even working on its own separate machine, not polluting anything on your machine,
26:46 still you would like it if, you know, a model reviews every shell command before it goes.
26:51 So auto review.
26:52 It would be great if, you know, when these untrusted inputs like an email in your inbox,
26:57 it could be a phishing email or some web hook,
26:59 you don’t want to send that to the model and act like it was a user instruction.
27:03 And there’s plenty of work to do here in the industry on just making sure
27:06 that that is properly protected from prompt injections.
27:10 For these kind of non-reversible things like sending emails, right,
27:16 you want to really ask the user first.
27:18 And then also you can build in some workflows into the product to keep passwords
27:22 and other things outside of the model conversations.
27:25 Okay.
Context engineering:上下文工程
27:26 Context engineering is also known as harness engineering.
27:31 Like I feel like it’s kind of the same thing.
27:32 They’ve evolved over time.
27:33 But basically what are some tips and tricks we can do
27:36 to minimize the amount of context that’s being used in these harnesses,
27:41 which is good for efficiency, for reducing costs, and other reasons.
27:46 I talked about prompt caching.
27:48 This is a good example of why it’s really important.
27:51 So if you think about when you’re working with a coding agent or any agent,
27:55 most of the tokens used are actually ideally cached tokens
28:00 because you’re sending in a conversation and then you’re doing a change,
28:03 you’re adding one more message,
28:05 and then you’re basically resending the entire conversation over again.
28:08 So if you look at the pricing for models, there’s huge discounts of course on if a token is cached.
28:13 And you want to keep them in the cached window as long as possible.
28:17 And so a way to do this is you try to keep the tool definitions, the system prompts,
28:22 all this information as static as possible between different calls
28:27 so that you don’t break the cache and then have to pay for the uncached tokens every single time.
28:32 So if you have that fixed part at the top, then you know when you get down to the bottom
28:36 and you only have the newest message that comes in,
28:38 that’s the only delta that you’re paying the uncached tokens for.
28:42 And there’s a whole art and science to this as well.
28:45 Basically because that main agent in the most ideal sense,
28:50 this is the one you’re having this very, very, very long conversation with
28:53 that’s compacting or compressing, so summarizing the context many, many times over
28:59 without you even noticing sometimes, which we’ll show an example of.
29:03 So you really want to keep that context small.
29:06 Some ways of doing that is anytime there’s a very long output,
29:11 rather than stuffing that in the context, you can just put it in a file.
29:14 And the agents are very good at reading files.
29:17 So this actually works pretty well, which is just like reading a skill
29:20 or reading a memory, for example.
29:22 If you give the model a screenshot or some other very expensive work,
29:28 you just delegate that out to a helper.
29:30 Again, so it goes to the sub-agent and it doesn’t affect the main agent’s context.
29:35 And then also, in an ideal world, you want to kind of be always gardening
29:41 and trimming and removing unnecessary stuff from this conversation,
29:45 from this context window, as you’re going.
29:47 Both for cost, but for also recall and just making sure the model’s performing well.
29:51 So if there’s a cache miss, it’s kind of like a bug in the system, really.
29:57 And if you go all the way back to January of this year, when OpenClaw was really getting popular
30:05 and people were using it inside of subscription plans for model providers,
30:11 this was a very new usage pattern that model providers hadn’t really seen.
30:15 And the harness hadn’t really yet been optimized for this type of workflow.
30:21 So there was a lot of cache misses.
30:23 And when you’re doing something like this, where agents are running all the time,
30:25 that’s obviously problematic.
30:27 So lots of fixes had to go in to make that robust to keeping the prompt cached as much as possible.
30:35 So like a fun trick here is just trying to keep everything in the prompt in order such that
30:43 even if you deploy a new change where you want to test out some new change to the system prompt
30:47 or add a new tool, you can almost hash each section and always keep a fixed order.
30:53 Like let’s say you’re doing an A/B test and you want to put a new piece of the prompt in there.
30:57 You really don’t want that to screw up the prompt order for everybody else.
31:00 Otherwise, you just busted the cache and then that’s problematic.
31:05 Another fun one is when you open the chat and you start typing,
31:09 the server can actually go and prepare the prompt early to basically warm the cache,
31:14 almost like if you’re on a website and you go to hover on a link,
31:17 it can go and prefetch the next page.
31:19 So there’s like tons and tons of little engineering optimizations
31:23 that compound how far these models can go in terms of token usage
31:28 and how smart they can be when you use them for very long periods of time.
31:32 And the key thing is that when you’re working over a very long period of time,
31:36 you essentially have to get very good at compaction or summarization
31:40 because let’s say the model has a 200,000 token context window.
31:44 You’re going and going and going.
31:46 Eventually, it’s going to hit a point where it needs to summarize
31:48 and all compression is lossy.
31:50 So you have to figure out the best algorithm to do the compression
31:54 and not lose important details.
31:56 And there’s, I’m sure, entire PhDs just for this problem.
32:00 Like it’s very, very tough to do, right?
32:02 But just to show an example, like let’s say you’re going back and forth with your bot,
32:06 you have over 100,000 tokens, and then you go and do something else.
32:11 Well, there is this time period where the model providers have like a cache window when the cache is still warm.
32:17 Maybe it’s 10 minutes. Maybe it’s 60 minutes.
32:19 One trick that we do, which I think is really helpful, is while you’re still within the warm cache period,
32:25 you should probably do the compaction right now because you have this long conversation.
32:30 It’s going to be much cheaper.
32:31 So just go ahead, do the compaction.
32:33 And now when the user comes back in 30 minutes or whatever, they’re starting from this summary.
32:39 And the summary now is going to be much cheaper for them to continue adding messages on.
32:43 So lots of work to do there to make the algorithm very high quality for how you do the summarization.
32:50 When you talk about compaction and what information you choose to include in the agent’s memory,
32:56 I kind of think about it in a few different layers.
32:59 The first one is like facts about the user that are very important to be included every single time.
33:06 Going back to earlier when I said, oh, like the user always wants to have lowercase text.
33:11 That probably needs to be in every single message.
33:13 Otherwise, it would be really weird if one message now has uppercase letters.
33:16 If you’re going like full Sam Altman style.
33:19 Then you probably also have some prompts that you want or some facts that you want to include in there.
33:24 And there’s many different ways you can make this algorithm to determine what are the most important facts to keep.
33:31 So ideally, there’s this kind of garbage collection of this fact.
33:34 And like if you think about it, this is kind of how our brains work.
33:37 It’s like this fact right now is important to keep.
33:39 It’s relevant for this conversation today.
33:42 You’re probably not thinking about a soccer match right now.
33:45 And your brain is very good at keeping those ones in context right now and kind of remembering them for later.
33:51 We tried to model some of that here where the harness has, of course, these core facts.
33:57 It also has these dated logs of kind of your past conversation history essentially.
34:03 And then also these kind of scratch pad notes.
34:06 So just like little things that you’ve been working on right now that are designed to be ephemeral and they will go away.
34:11 But they might be relevant just for the next 30 minutes or an hour.
34:14 And everything else can be found through files.
34:17 This is the magic of files of the models being good at running commands is like they can always just go look stuff up.
34:24 And they’re actually very good at doing that.
Where this is going:接下来往哪里走
34:26 So where where is this going?
34:28 There is this flywheel of training the model to get better at the product.
34:34 And this is a high level refresher on kind of the stages of training a large language model.
34:40 First you have pre-training where the model learns a general understanding of the world from large amounts of internet text or other private data.
34:49 Then you have supervised fine tuning where you’re teaching the model how to behave in a certain way.
34:55 Maybe as a chat assistant or as an agent.
34:58 And then finally you have reinforcement learning where it’s kind of like teaching the model how to play games and get better at those games.
35:05 And in the context of these always on agents it’s very helpful for example to train the model to see examples of the harness.
35:14 And to know how to work inside of the harness in SFT.
35:17 So it can understand the tools that you have available.
35:20 It’s helpful to train the model to be better at instruction following during reinforcement learning.
35:25 How to understand the right skill to call when there’s hundreds of skills for example.
35:30 How to click on the right spots when you’re given in a browser and it needs to kind of click around.
35:35 And not only the training itself but also how you serve the model.
35:40 The inference that you do.
35:41 The main agent you’re talking with you know ideally you want it to have very low latency.
35:45 You want it to come back to you right away.
35:47 Give you really fast replies.
35:48 And you can tune the harness and the inference to do that.
35:51 But then when you go hand it off to a helper and you’re asking it to go do some work for a very very long period of time.
35:57 The difference between 10 minutes and 12 minutes is kind of negligible at that point.
36:01 So that allows you to kind of control how you optimize things.
36:05 The way that this kind of flywheel works is that let’s say your agent gets something wrong.
36:12 Ideally then that gets turned into some kind of evaluation.
36:16 So the model you know accidentally called this tool and it wasn’t supposed to.
36:21 Then on the back end you can go and fix the harness.
36:24 You can fix the bug.
36:25 You can fix the prompt.
36:26 Maybe it’s an inference issue.
36:28 You can check and make sure it’s actually fixed in the evals and then shift the change.
36:32 And you know the teams and the products the teams working on these products they’re kind of doing this every single day.
36:38 They’re finding bugs.
36:39 They’re fixing it.
36:40 They’re measuring it with evals.
36:41 That’s kind of this inner loop that you’re always getting better at.
36:43 But then there’s this outer loop which is like every few months or every month or however long you’re training new models.
36:49 And ideally the new models are learning from all of these bugs and these failures.
36:53 And they’re kind of climbing on the evals to get better at the things that you care about.
36:57 And you can drop the new versions into the product and make sure that they’re improving at all these different things.
37:02 And the fun thing about this is then when the new models come out they have new capability jumps.
37:07 It allows you to use the product in new ways.
37:09 So there hasn’t really been a lot of training data for example on using bots in a group chat.
37:15 This is kind of an emergent thing that’s being figured out and it’s important to get it right.
37:19 Obviously if you’ve seen things like Moltbook or the Hugging Face incident like there are ways that that goes poorly.
37:25 So it’s important to be able to train the models how to behave well in a multi-agent system where different bots are talking to each other for example.
37:35 And so ideally you’re always kind of running these inner loops and outer loops as a product team building these type of harnesses at a company working on AI together.
37:45 One of the principles that we had at Cursor and now at SpaceX AI is to delete the product.
37:52 And what we mean by this is you kind of have to internalize and assume that the models today are the worst they’ll ever be.
37:59 Which means that in six months you might need to completely rebuild the UI.
38:03 Like most of the stuff that I’ve talked about today it’s a pretty big update from if I would have given this talk six months ago or definitely a year ago.
38:11 Because we’ve learned a lot of new practical ways of building harnesses and doing context engineering.
38:17 So a good strategy here is you really want to try to build for the next generation of the model.
38:23 Like how do you build as little as possible and get by while making just a really really simple interface to work with these models and allow them to think and act for you.
38:34 So some things that I think will happen in the next six to twelve months we’ll see I think these are decently safe bets.
38:43 But you know right now let’s say you have your agents or your bots working for maybe a few hours maybe a day depending on how tough it is.
38:50 I think it’s very possible that pretty soon your agents will be working for weeks or months.
38:57 And it does bring up some kind of interesting questions in terms of how we train and evaluate those models.
39:02 I think the models are going to get better and better about how they recall on past information in conversations.
39:10 Plenty of different approaches and research here for the type of algorithms used or the representation of the data under the hood.
39:17 Maybe it’s a graph.
39:18 Maybe it’s files.
39:19 Maybe something else.
39:20 I think ideally if you have this digital colleague you give it literally all the tools that you have if you were onboarding a real human.
39:27 And so I think organizations are still warming up to that idea of course.
39:31 So there’s more to do there.
39:34 I think the interfaces will maybe get even easier.
39:37 I think some people will prefer to just talk to their kind of bots or agents in a speech to speech way.
39:44 Maybe they’ll have their like Johnny Ive device that is hopefully really cool and they can like talk to that all day.
39:51 I think the models and the products will get really good at learning skills on the job and then encoding that into files
40:00 and instructions skills and taking that raw intelligence and actually making it useful for whatever whatever task you’re asking it to do.
40:09 And people and companies are going to get much better I think about writing down how all of these kind of bespoke processes work so that they’re legible for agents to use.
40:20 So some interesting problems I think that like could just spark ideas for you all to research or to look into.
40:27 The first one is like going back to this idea of agents that run for a month.
40:30 What do you do when you need to evaluate an agent that runs for a month but the models ship every month.
40:36 You know it’s like you can’t fully evaluate that the model is sufficiently behaving as expected.
40:44 So you either have to figure out how to do a better eval or you know not ship models as fast.
40:48 And I think that’s going to be a problem that the industry has to figure out in the next six months.
40:54 Memory obviously we talked about the different algorithms.
40:57 I think there’s so much to do on the security side for how to make a secure harness.
41:02 How to make the infrastructure around these agents very secure and so they can’t be taken over by bad actors.
41:09 I think we didn’t really talk about it too much but the heuristics around when the agent is silent or when it kind of bugs you.
41:17 I don’t think that we have it perfect yet and I think trying to thread the needle between being proactive and suggesting like oh hey like I saw that you missed class yesterday.
41:27 Like I had a recording there and I took the notes and like here’s the thing like maybe that’s helpful or maybe that’s kind of annoying.
41:34 You know and like trying to figure out the right balance there.
41:37 I also think that it’s a pretty safe bet to assume that at this time next year there will be significantly significantly more tokens flowing in the world.
41:49 Which means that optimizing the inference optimizing the models optimizing the harnesses trying to drive down the tokens to be more efficient will also the token usage to be more efficient.
42:01 I think will also be very important.
42:03 So there’s kind of these new building blocks to think about.
42:06 I believe it was Alex from OpenRouter that I saw on on X.
42:11 He had a really good take on this.
42:12 There’s people saying like oh every app today looks the same.
42:15 You’ve got the sidebar.
42:16 You know you’ve got the agent.
42:18 You’ve got the chat box like everybody’s going the same thing.
42:20 And he was like yeah that’s like in 2005.
42:22 He said everybody’s doing the same thing.
42:24 They’ve got a server.
42:25 They’ve got a database.
42:27 They’ve got some like crud interface to like update and delete things.
42:31 And the reality is this is just kind of the new normal now where every new product is going to have to have some kind of agentic side to it.
42:40 Either integrating or it will become a you know a headless piece of data for another agent to use.
42:47 And so we’ve talked about a lot of these products today.
42:50 Durable workflows.
42:51 Sandboxes.
42:52 Computers in the cloud.
42:53 Memory skills.
42:55 Connecting to other services.
42:57 Doing voice.
42:58 Payments.
42:59 Identity.
43:00 Evals.
43:01 Observability.
43:02 Each one of those boxes is like multiple billions of dollars of VC capital.
43:05 Probably trillions.
43:06 There’s 20 startups for each one of those.
43:09 And there’s so many different things to explore and all of those.
43:12 Because they kind of are the new primitives for the next generation of companies.
43:16 The next generation of products.
43:18 And they all have to integrate and be part of harnesses.
43:20 So there’s definitely a lot to do there.
43:22 A few things to take away and then we’ll have time for questions.
43:26 Obviously you all are engineers.
43:27 You’re thinking about how to write software.
43:29 Where it’s more important than ever to think about the broader system design.
43:33 And the architecture of the code that you’re working on.
43:36 Because the decisions you make there will really compound.
43:40 So I’m still pro looking at code.
43:43 I still look at the code that I’m writing.
43:45 Some people are not looking at the code.
43:46 I still think it’s helpful.
43:47 Especially to understand the architecture that you’re shipping.
43:49 It’s crazy how much that’s changed in just a couple years.
43:52 Secondly, you know, it sounds like you all were kind of messing around with building own versions.
43:56 Or understanding of how the harness worked in last week.
44:00 I encourage you to think about some of the things that I did today.
44:05 And try to apply that.
44:06 And think about how you might build a version for yourself.
44:09 I’m pretty sure there’s like open GrokBot on GitHub somewhere.
44:13 Where somebody’s like, you know, slot forked our code and made a version.
44:18 And you could probably set up like the GUI for that with, you know, a sandbox somewhere.
44:24 And, you know, a VM.
44:27 And like all sorts of different things.
44:29 So that could be interesting if you want to learn a little bit more about how these pieces work.
44:32 And then finally, I would not shy away from just learning a little bit more about how LLMs work.
44:39 You know, we gave a little bit of the TLDR here of how they’re trained.
44:43 But there’s so much more you can go into here without necessarily going deep in ML or deep into the science part of it.
44:49 But I think it’s helpful to understand just like when you’re building a web app.
44:53 It’s helpful to know how databases work.
44:55 Because increasingly, there are so many decisions you make at the product and the harness level.
45:01 That are influenced by the models themselves.
45:03 And having an intuition of how that works I think will make you very well positioned for whatever role you’re operating in.
45:10 And that’s all I have today.
45:13 Thank you.
讨论
用 GitHub 账号留言;评论保存在公开仓库chengshu-blog-discussions的 Discussions 里。也可通过 RSS 订阅后续文章。