{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Chengshu",
  "home_page_url": "https://blog.chengshu.space/",
  "feed_url": "https://blog.chengshu.space/feed.json",
  "description": "Chengshu 的个人写作空间。",
  "language": "zh-CN",
  "authors": [
    {
      "name": "Chengshu"
    }
  ],
  "items": [
    {
      "id": "https://blog.chengshu.space/posts/movez-opus55-04-overnight-critique-ship/",
      "url": "https://blog.chengshu.space/posts/movez-opus55-04-overnight-critique-ship/",
      "title": "用 Opus 5.5 搭动态设计工作室（四）：导演简报、自审与交付",
      "summary": "从一夜导演简报到让 Opus 回看自己的帧，再到多画幅导出、打包成 skill、甚至卖成服务。一句话换来片子，harness 换来工作室。",
      "date_published": "2026-09-29T04:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Opus 5.5",
        "Claude Code",
        "Motion Design",
        "中英对照",
        "X转录"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/movez-opus55-03-engine-springs-sound/",
      "url": "https://blog.chengshu.space/posts/movez-opus55-03-engine-springs-sound/",
      "title": "用 Opus 5.5 搭动态设计工作室（三）：seek(t) 引擎、弹簧与配乐",
      "summary": "路线 A：两页文件搭起 seek(t) 渲染器；闭式弹簧让运动有质量；节拍网格上合成配乐。也给出 Remotion / HyperFrames 的路线 B。",
      "date_published": "2026-09-29T03:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Opus 5.5",
        "Claude Code",
        "Motion Design",
        "中英对照",
        "X转录"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/movez-opus55-02-brand-reference-spec/",
      "url": "https://blog.chengshu.space/posts/movez-opus55-02-brand-reference-spec/",
      "title": "用 Opus 5.5 搭动态设计工作室（二）：品牌、参考与状态清单",
      "summary": "把 showreel 对准你的产品：给 URL、真截图和音乐；用一帧参考锁住风格；再用 XML 式状态清单代替「感觉对了」。提示词从一句话升到可执行的规格。",
      "date_published": "2026-09-29T02:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Opus 5.5",
        "Claude Code",
        "Motion Design",
        "中英对照",
        "X转录"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/movez-opus55-01-intro-pixels-setup-oneliner/",
      "url": "https://blog.chengshu.space/posts/movez-opus55-01-intro-pixels-setup-oneliner/",
      "title": "用 Opus 5.5 搭动态设计工作室（一）：像素、环境与一句话片头",
      "summary": "大多数人用 Opus 5.5 做动态设计，最后都落到同一条片子：渐变底上居中大字、全员淡入、片尾一个 logo。本篇从「模型写的是程序不是视频」讲起，装好工作室，再拆开那句走红的一句话 showreel prompt。",
      "date_published": "2026-09-29T01:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Opus 5.5",
        "Claude Code",
        "Motion Design",
        "中英对照",
        "X转录"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/anatomy-of-collective-communication/",
      "url": "https://blog.chengshu.space/posts/anatomy-of-collective-communication/",
      "title": "《走进 TPU 与 GPU 集群：集合通信的解剖》中英对照全译",
      "summary": "Aleksa Gordić 的长文《Inside TPU and GPU Clusters: The Anatomy of Collective Communication》全文对照翻译：左栏英文原文、右栏对应中文。从 TPU 的 2D/3D torus 与 GPU 的 fat tree 拓扑讲起，把 All-Gather、Reduce-Scatter、All-Reduce、All-to-All 的 ring/tree 实现、代价模型，以及 SHARP 与跨节点分层通信一次讲透。",
      "date_published": "2026-09-28T13:30:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "集合通信",
        "TPU",
        "GPU",
        "All-Gather",
        "All-Reduce",
        "中英对照"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/inside-transformer-life-of-a-token/",
      "url": "https://blog.chengshu.space/posts/inside-transformer-life-of-a-token/",
      "title": "《走进 Transformer：一个 token 的一生》中英对照全译",
      "summary": "Aleksa Gordić 的长文《Inside the Transformer: The Life of a Token》全文对照翻译：左栏英文原文、右栏对应中文。用 Rnj 1.5 的真实结构走一遍前向传播，逐个拆开 RMSNorm、GeGLU MLP、多头注意力、YaRN 与 core attention，并把 KV cache、参数量、FLOPs 和集群规模一次算清。",
      "date_published": "2026-09-28T13:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Transformer",
        "LLM 架构",
        "YaRN",
        "注意力机制",
        "KV cache",
        "中英对照"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/making-deep-learning-go-brrrr/",
      "url": "https://blog.chengshu.space/posts/making-deep-learning-go-brrrr/",
      "title": "《让深度学习跑出 Brrrr》中英对照全译：从第一性原理看清 compute、带宽与 overhead",
      "summary": "Horace He 的经典长文《Making Deep Learning Go Brrrr From First Principles》全文对照翻译：左栏英文原文、右栏中文译文。用「工厂」类比讲清 compute / memory bandwidth / overhead 三种性能 regime，以及 operator fusion、compute intensity、CUDA Graph 这些优化的适用边界。",
      "date_published": "2026-09-28T12:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "深度学习",
        "性能优化",
        "PyTorch",
        "GPU",
        "中英对照",
        "operator fusion"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/vllm-anatomy-05-benchmarking-and-epilogue/",
      "url": "https://blog.chengshu.space/posts/vllm-anatomy-05-benchmarking-and-epilogue/",
      "title": "走进 vLLM（五）：benchmark 与 auto-tuning —— latency 与 throughput 之争",
      "summary": "为什么延迟和吞吐天生互相拉扯？TTFT、ITL、TPOT、E2E、Goodput 到底各测的是什么？roofline 模型怎么解释「算 1 个 token 和算 10 个 token 差不多快」？以及 vLLM 自带的 vllm bench 三个脚本分别在什么场景下用。",
      "date_published": "2026-09-28T02:40:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "vLLM",
        "LLM 推理",
        "推理引擎",
        "benchmark",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/vllm-anatomy-04-scaling-and-serving/",
      "url": "https://blog.chengshu.space/posts/vllm-anatomy-04-scaling-and-serving/",
      "title": "走进 vLLM（四）：从单卡到多机——MultiProcExecutor 与分布式服务",
      "summary": "权重装不下一张卡之后怎么办：MultiProcExecutor 怎么用共享内存消息队列把 8 个 worker 进程编排起来；DP 复制、DPCoordinator 与 DP wave 又怎么把两个 8×H100 节点拼成一套服务；最后完整走一遍一条 curl 请求的生命周期。",
      "date_published": "2026-09-28T02:30:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "vLLM",
        "LLM 推理",
        "推理引擎",
        "分布式",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/vllm-anatomy-03-specdec-and-pd/",
      "url": "https://blog.chengshu.space/posts/vllm-anatomy-03-specdec-and-pd/",
      "title": "走进 vLLM（三）：speculative decoding 与 disaggregated P/D",
      "summary": "一次大模型 forward 最多换回 k+1 个 token，而且分布和逐个采样完全等价——speculative decoding 是怎么做到的？vLLM V1 为什么不用小模型做 draft，而用 n-gram / EAGLE / Medusa？以及为什么 prefill 和 decode 值得被拆到两批实例上跑。",
      "date_published": "2026-09-28T02:20:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "vLLM",
        "LLM 推理",
        "推理引擎",
        "speculative decoding",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/vllm-anatomy-02-chunked-prefill-prefix-caching/",
      "url": "https://blog.chengshu.space/posts/vllm-anatomy-02-chunked-prefill-prefix-caching/",
      "title": "走进 vLLM（二）：chunked prefill、prefix caching 与 guided decoding",
      "summary": "把基础 engine 流程扩出来的第一批高级特性：chunked prefill 怎么让长 prompt 不再独占一步；prefix caching 为什么本质上是「别重算你已经见过的前缀」；guided decoding 又是怎么用一台 FSM 把 logits 里不该出现的 token 直接按成负无穷。",
      "date_published": "2026-09-28T02:10:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "vLLM",
        "LLM 推理",
        "推理引擎",
        "prefix caching",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/vllm-anatomy-01-engine-core/",
      "url": "https://blog.chengshu.space/posts/vllm-anatomy-01-engine-core/",
      "title": "走进 vLLM（一）：LLM Engine 与 Engine Core",
      "summary": "一个现代高吞吐 LLM 推理系统的最小内核长什么样：从一次离线 generate 调用出发，拆开 LLM Engine 的构造函数、generate 函数、Scheduler，以及一次 forward pass 里到底发生了什么——顺带把 paged attention、continuous batching、KV cache block 这些词讲成人话。",
      "date_published": "2026-09-28T02:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "vLLM",
        "LLM 推理",
        "推理引擎",
        "系统设计",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/openship-self-hosted-deploy-platform/",
      "url": "https://blog.chengshu.space/posts/openship-self-hosted-deploy-platform/",
      "title": "Openship 拆解：一个把控制面放进桌面应用的开源部署平台，和 Vercel、Coolify 差在哪",
      "summary": "Openship（oblien/openship）把「push 一下就上线」的体验搬到你自己拥有的机器上：构建跑在你的笔记本，产物走纯 SSH 推到目标机，服务器上不装常驻 agent。本文按仓库和官方文档逐条核对它的六步流水线、三种运行形态，把它和 Vercel/Netlify、Coolify/Dokploy/Dokku/CapRover/Kamal 放进同一张表对比，最后泼几盆冷水。",
      "date_published": "2026-09-28T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Openship",
        "自托管",
        "部署平台",
        "Vercel",
        "Coolify",
        "PaaS"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/anthropic-when-ai-builds-itself/",
      "url": "https://blog.chengshu.space/posts/anthropic-when-ai-builds-itself/",
      "title": "《当 AI 开始建造自己》中英对照全译：逐句翻译、动效复刻与批注",
      "summary": "把 Anthropic Institute 的《When AI builds itself》整篇译成中文并转载：每个段落先英文后中文，复刻了原文的滚动驱动时间线与点阵 hero 动效，并在关键观点处加了批注。",
      "date_published": "2026-09-27T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Anthropic",
        "递归自我改进",
        "中英对照",
        "AI 研究",
        "批注",
        "逐句翻译"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/codex-goal-command/",
      "url": "https://blog.chengshu.space/posts/codex-goal-command/",
      "title": "不达目的誓不罢休：拆解 Codex 的 Goal 命令是怎么实现、又是怎么设计的",
      "summary": "从 codex-rs/ext/goal 源码出发，讲清楚 Codex Goal（长时程目标）的四件事：目标怎么持久化、凭什么自动续跑、为什么不会自己骗自己说做完了，以及 token 预算是怎么计的。",
      "date_published": "2026-09-26T00:00:00.000Z",
      "date_modified": "2026-09-26T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Codex",
        "Agent Harness",
        "源码阅读",
        "长任务"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/iltb-487-ben-thompson-ai-boom-runs-out-of-money/",
      "url": "https://blog.chengshu.space/posts/iltb-487-ben-thompson-ai-boom-runs-out-of-money/",
      "title": "《当 AI 泡沫的钱烧完了》：Ben Thompson 谈大科技、中国与 AI 资本周期",
      "summary": "Ben Thompson 在 Invest Like The Best 第 487 期里讲了 80 分钟：为什么「美国赢下 AI 竞赛」对世界是危险的；为什么这轮 AI 最近在眼前的约束不是算力、电力，而是「钱」——以及 1870 年代铁路的久期错配；为什么消费者 AI 最终只能靠广告；为什么台积电其实是把风险转嫁给了大科技公司；为什么「待在不在前沿」对数字公司来说才是真正的风险；以及英伟达最怕的不是别人造出更好的芯片，而是电力变得便宜。含 8 张双语配图。",
      "date_published": "2026-09-26T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Ben Thompson",
        "Stratechery",
        "聚合理论",
        "AI 泡沫",
        "资本开支",
        "大宗商品",
        "台积电",
        "英伟达",
        "Meta",
        "讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/iltb-487-ben-thompson-ai-boom-runs-out-of-money-transcript/",
      "url": "https://blog.chengshu.space/posts/iltb-487-ben-thompson-ai-boom-runs-out-of-money-transcript/",
      "title": "《当 AI 泡沫的钱烧完了》中英对照逐字稿（80 分钟，带时间轴）",
      "summary": "2026 年 8 月，Ben Thompson 在 Invest Like The Best 第 487 期里与 Patrick O'Shaughnessy 聊了 80 分钟：AI 竞赛与中美、铁路式的久期错配、可验证与不可验证的领域、推理的真实成本、广告与注意力、航运与内存的大宗商品周期、台积电如何把风险转嫁给大科技公司、亚马逊与苹果的位置、微软的 IBM 剧本、Meta 与 TikTok、英伟达与电力。这份逐字稿按 16 个章节逐段给出中英对照。",
      "date_published": "2026-09-26T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Ben Thompson",
        "Stratechery",
        "逐字稿",
        "中英对照",
        "AI 泡沫",
        "台积电",
        "英伟达",
        "Meta"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-13-hands-on/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-13-hands-on/",
      "title": "上手篇：把 V4.1-Flash 跑起来",
      "summary": "从零到发出第一个请求的完整清单：镜像只能选 nightly、环境变量照抄哪几行、启动参数逐条拆解、验证为什么必须发两条请求（正确答案 323）、thinking 默认开着会把小 max_tokens 的请求变成空回复、四个性能旋钮、内核生态速查（FlashInfer/CuTeDSL/DeepGEMM/AITER/MFMA/gfx942/gfx950/CUDA graph/torch.compile）、硬件选型与 KV cache offloading。",
      "date_published": "2026-09-25T02:27:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "上手指南",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-12-pd-disaggregation/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-12-pd-disaggregation/",
      "title": "PD 分离、NIXL 与 vllm-router：prefill 和 decode 为什么分居",
      "summary": "prefill 吃算力、decode 吃带宽，混跑会互相踩脚。官方唯一验证过的进阶布局就是把它们物理隔离：GB200 NVL4 上每角色一台 tray、池内 TP4、KV 经 NIXL 交接、前面由 vllm-router 分发，并把并发限制在 32。还包括为什么这台模型的 KV 小到可以搬，以及开投机解码必须两个池子一起开的原因。",
      "date_published": "2026-09-25T02:26:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "PD 分离",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-11-parallelism-tp-dp-ep/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-11-parallelism-tp-dp-ep/",
      "title": "并行三兄弟：TP / DP / EP(DEP)",
      "summary": "一张卡装不下 511 GB 时必须切，但切法有三种：TP 切层内权重（每层都要通信）、DP 复制模型切数据（不省显存）、EP 把专家分卡（all-to-all）。对比 H200 / MI355X / MI325X / Blackwell 的默认配置，看同一份模型为什么在不同卡上得出完全不同的答案，以及 DEP 一次性打开了哪四样东西。",
      "date_published": "2026-09-25T02:25:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "并行",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-10-speculative-decoding/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-10-speculative-decoding/",
      "title": "猜 5 个再验一遍：投机解码与 DSpark 草稿头",
      "summary": "decode 是带宽瓶颈，搬一次权重只写一个字太浪费。投机解码让草稿器先猜 5 个 token，再由主模型一次验证——DSpark 是三阶段（每阶段 128 专家取 3）、读取第 37–39 层注意力输入的小模型。还包括 AMD 上必须关掉自适应验证的两个具体检查，以及 --max-cudagraph-capture-size 里 128×(1+5) 的由来。",
      "date_published": "2026-09-25T02:24:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "推理加速",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-09-number-formats-memory/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-09-number-formats-memory/",
      "title": "BF16/FP8/MXFP4/MXFP8/UE8M0：数字格式与 511GB 显存账本",
      "summary": "\"MXFP4 量化\"翻译成人话就是一个参数半个字节。把官方 511 GB 的账本逐项算一遍：专家 259.5 GiB、Engram 183.1 GiB，而光\"块 scale\"就要占 21.9 GiB。顺带说明数字格式如何决定硬件能不能跑（MI325X 没有 FP4 MFMA，MXFP4 专家只能走 Triton），以及 GB 与 GiB 之间那 7% 的差异。",
      "date_published": "2026-09-25T02:23:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "量化",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-08-rope-yarn-context/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-08-rope-yarn-context/",
      "title": "1M 上下文怎么来的：RoPE、YaRN 与 theta",
      "summary": "这台模型只在 65,536 token 的训练窗口里见过世界，却提供 1,048,576 的上下文：65,536 × 16 = 1,048,576，那个 16 就是 YaRN 的 factor。从 RoPE 的旋转直觉讲到\"模型内部有两套位置刻度\"——压缩 KV 为什么要用自己的 theta 160,000，以及你能调的其实只有 --max-model-len。",
      "date_published": "2026-09-25T02:22:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "长上下文",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-07-hyper-connections/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-07-hyper-connections/",
      "title": "超连接与 Sinkhorn：残差流为什么要开四条副本",
      "summary": "残差连接从\"系数恒为 1 的一条线\"变成\"4 份并行副本 + 每层自算 pre/post/combine 系数\"，再用 20 次 Sinkhorn 迭代把混合矩阵压成双随机矩阵。顺带解释官方 Prerequisites 里那句看起来毫不相干的 mHC + AITER 说明，以及为什么 AMD 镜像必须够新。",
      "date_published": "2026-09-25T02:21:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "架构",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-06-engram/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-06-engram/",
      "title": "Engram：给模型配一本 384M 行的词组小抄",
      "summary": "这台模型里唯一不参与\"思考\"的大块参数：第 1、14 层各有一张 384M 行 × 256 维的哈希表，用 2/3/4-gram 哈希查表，通过学习的门写进残差流。196.6B 参数、183 GiB 显存，也是唯一适合 offload 到 CPU 内存的部分——但\"输出不变\"不等于\"速度不变\"。",
      "date_published": "2026-09-25T02:20:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "MoE",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-appendix-recipe-zh/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-appendix-recipe-zh/",
      "title": "附录：vLLM DeepSeek-V4.1-Flash 部署说明中文全译",
      "summary": "vLLM Recipes 官方页面 deepseek-ai/DeepSeek-V4.1-Flash 的逐节中文翻译：从 Overview、Context length、Images，到 Prerequisites、Verifying、Performance tuning，以及 MI355X / MI325X / H200 / Blackwell 的并行配置与 KV cache offloading。命令、参数、链接保持原样，是整个系列的事实底稿。",
      "date_published": "2026-09-25T02:05:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "技术翻译"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-05-sparse-attention/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-05-sparse-attention/",
      "title": "两级稀疏注意力：滑窗、压缩 KV latent 与 indexer",
      "summary": "1M 上下文凭什么可行：每层 128 token 滑窗看近处，压缩 KV latent 看远处，indexer 用 2048 块粗筛 + 512 精筛决定\"该看哪里\"。还有两个关键细节——只有第 2、8、14、20 层真正写摘要、其余 36 层共享，以及压缩 KV 为什么要用自己的一套 RoPE theta。",
      "date_published": "2026-09-25T02:04:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "注意力机制",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-04-attention-kv-cache/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-04-attention-kv-cache/",
      "title": "注意力与 KV cache：为什么「长上下文」本质是显存问题",
      "summary": "prefill 吃算力、decode 吃带宽：从 TTFT 讲到 KV cache 为什么随长度和并发线性增长。以及这个模型的反直觉之处——890 字节/token 让上下文几乎不再是显存问题，真正吃显存的是权重和批大小；顺带说清 --max-model-len、--max-num-seqs、--max-num-batched-tokens 该怎么排优先级，KV offloading 为什么是\"用延迟换容量\"。",
      "date_published": "2026-09-25T02:03:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "KV cache",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-03-vision-path/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-03-vision-path/",
      "title": "模型也会看图：ViT、aligner、图像 token 与 Encoder parallel",
      "summary": "一张图片如何从像素变成向量、再插进文本序列：32 层 ViT、patch 14、3× 下采样 aligner、每图 1024 token 上限。顺带讲清两个互斥的部署开关（--language-model-only 与 Encoder parallel），以及为什么验证服务时必须发一条图片请求。",
      "date_published": "2026-09-25T02:02:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "多模态",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-02-moe-routing/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-02-moe-routing/",
      "title": "384 个专家里只叫醒 6 个：MoE 与路由入门",
      "summary": "从\"一家公司有 384 位专家、每来一个问题只准叫 6 个\"讲起：MoE 为什么要切开 FFN，路由器这个\"前台\"到底在做什么，sqrtsoftplus 与 noaux_tc 偏置各自解决什么问题，为什么图像 token 要单独排一队，以及那句最该带走的话——MoE 省的是电费，不是房租。",
      "date_published": "2026-09-25T02:01:00.000Z",
      "date_modified": "2026-09-25T12:35:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "MoE",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/deepseek-deploy-01-model-overview/",
      "url": "https://blog.chengshu.space/posts/deepseek-deploy-01-model-overview/",
      "title": "它到底有多大：552B、196B、8B/16B 三个数字分别是什么",
      "summary": "拆开 DeepSeek-V4.1-Flash 的三个核心数字：552B 主干、196B Engram 记忆、每 prompt token 激活 8B / 每 output token 激活 16B。为什么容量按总量付费、速度按激活量算账，为什么 511 GB 权重只能用 Docker 跑，以及\"1M 上下文\"到底要不要 1M 显存。",
      "date_published": "2026-09-25T02:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "DeepSeek",
        "vLLM",
        "模型部署",
        "MoE",
        "术语科普"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/buffett-1998-university-of-florida/",
      "url": "https://blog.chengshu.space/posts/buffett-1998-university-of-florida/",
      "title": "巴菲特 1998 年佛罗里达大学演讲讲义：「买下同学 10% 的人生」",
      "summary": "1998 年秋天，巴菲特在佛罗里达大学做了一场 82 分钟的演讲：六分钟讲品格，剩下一个多小时全部用来回答提问。这篇讲义按顺序整理了他讲的东西——从「买下同学 10% 的人生」那个思维实验，到长期资本管理公司为什么会让 16 个最聪明的人破产、护城河与心智份额、喜诗糖果 2500 万美元的买价怎么算出来、可口可乐「没有味觉记忆」的秘密、他犯过的错大多是「漏做」、华尔街靠活动赚钱而你靠不活动赚钱、分散投资的两种相反答案，以及最后那个「卵巢彩票」。含 14 张双语配图，全部截自视频对应时刻。",
      "date_published": "2026-09-24T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "巴菲特",
        "投资",
        "护城河",
        "能力圈",
        "长期资本管理",
        "可口可乐",
        "分散投资",
        "讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/buffett-1998-university-of-florida-transcript/",
      "url": "https://blog.chengshu.space/posts/buffett-1998-university-of-florida-transcript/",
      "title": "巴菲特 1998 年佛罗里达大学演讲英文逐字稿（82 分钟，带时间轴）",
      "summary": "这份英文逐字稿是对 1998 年巴菲特佛罗里达大学演讲完整音轨的转写：开场那段品格课，以及随后一个多小时的全部现场问答——日本、长期资本管理公司、工作与薪水、护城河、喜诗糖果、可口可乐、犯过的错、宏观与华尔街、分红与卖出、分散投资、品牌与定价权、下跌的市场，直到最后关于「卵巢彩票」的问题。按句合并、逐段带时间轴，人名与公司名做过校正。",
      "date_published": "2026-09-24T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "巴菲特",
        "投资",
        "逐字稿",
        "长期资本管理",
        "可口可乐",
        "分散投资"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/webcafe-xinci-find-build-promote/",
      "url": "https://blog.chengshu.space/posts/webcafe-xinci-find-build-promote/",
      "title": "「哥飞的朋友们」45 篇新词比赛复盘读完：找需求、做开发、做推广",
      "summary": "把 Web.Cafe「比赛复盘文章」页上全部 45 篇复盘（2025 年度赛到 2026 年 8 月，共 10 届）读完后的汇总。只收录至少被两个人独立提到、且有人靠它拿到名次的做法：平台子域名找词、Trends 词根库、intitle 与 SERP 意图判断、模板与 SOP、MVP 拆分、多语言的取与舍、AdSense 低价值红线、有流量时改页面的代价、外链的持续性与质量、社媒裂变与去源头发评论。附 45 篇分届清单。",
      "date_published": "2026-09-24T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "出海",
        "SEO",
        "新词站",
        "独立开发",
        "复盘",
        "心法",
        "哥飞"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/advanced-evals-hidden-ai-failures/",
      "url": "https://blog.chengshu.space/posts/advanced-evals-hidden-ai-failures/",
      "title": "《Advanced evals》讲义：先找出「值得测的失败」——error discovery 三步法",
      "summary": "Hamel Husain 和 Shreya Shankar 的进阶续篇：做 evals 最该先做、大多数团队却跳过的第一步是 error discovery。为什么「太早写指标」会测错东西；为什么 agent 抓不到判断型失败（criteria drift）；100 条真实 trace 的对照结论；以及用 coding agent 走完三步——从 traces 开始、审阅与标注、把失败模式变成产品优先级。含 7 张双语配图。原文为 Substack 付费文章，本讲义覆盖免费可见部分的全文，付费墙之后的部分已据实标注。",
      "date_published": "2026-09-23T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "evals",
        "AI 产品",
        "error discovery",
        "错误发现",
        "标注",
        "coding agent",
        "讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/advanced-evals-hidden-ai-failures-transcript/",
      "url": "https://blog.chengshu.space/posts/advanced-evals-hidden-ai-failures-transcript/",
      "title": "《Advanced evals》英文原文全文照录（Hamel Husain & Shreya Shankar）",
      "summary": "Lenny's Newsletter 付费长文《Advanced evals: How to find (and fix) hidden AI failures in your product》的英文全文照录，保留作者原有小标题体系与 7 张配图。原文为 Substack 付费文章，本次抓取可读到 Step 3 开头，付费墙之后的内容未获取，文中已注明。",
      "date_published": "2026-09-23T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "evals",
        "AI 产品",
        "error discovery",
        "英文原文",
        "Hamel Husain"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/ai-native-services-100b-opportunity/",
      "url": "https://blog.chengshu.space/posts/ai-native-services-100b-opportunity/",
      "title": "《AI-native services: a $100B opportunity》讲义：卖「完成的工作」，而不是卖工具",
      "summary": "Greg Isenberg 的判断：AI-native services 是现在最值得做的生意，背后约 1000 亿美元。本文逐段翻译他的完整指南——那笔每年 12 万美元、一直被「人」锁住的预算为什么打开了；为什么模型变强对服务生意是利好而不是利空；unit / intake / engine / rulebook / review layer / delivery / pricing / distribution 八个零件分别是什么；以及「已经外包 + 可核对」这张机会地图，和一份 2026 年具体可做的清单。6 张配图是原文原图 + 英文原句 + 中文翻译。",
      "date_published": "2026-09-21T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "X长文",
        "AI-native services",
        "AI 服务化",
        "产品化服务",
        "定价",
        "一人公司",
        "讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/building-a-harness-with-jev/",
      "url": "https://blog.chengshu.space/posts/building-a-harness-with-jev/",
      "title": "用 Jev 搭 Agent Harness：面向决策而非生成文本的新模型",
      "summary": "Jev 是为决策而非生成文本优化的 System One 模型；本文译自 Sydney Runkle / Hunter Lovell，讲如何用它搭 agent harness（路由与 Auto Mode）。",
      "date_published": "2026-09-20T00:00:00.000Z",
      "date_modified": "2026-09-20T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "X转录",
        "Jev",
        "Agent",
        "LangChain",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/the-information-ai-deep-dive-noam-brown/",
      "url": "https://blog.chengshu.space/posts/the-information-ai-deep-dive-noam-brown/",
      "title": "《AI Deep Dive》第 1 集讲义：Noam Brown 谈智能体、多智能体，与「让 AI 改进 AI」",
      "summary": "The Information 新节目 AI Deep Dive 首集精读：Noam Brown（OpenAI）讲清楚了什么是智能体、推理为什么是可靠性的前提、90/10 效应下人的注意力去了哪里、研究品味为什么仍是差距；以及 Hugging Face 事件里，一群被隔离训练的智能体如何找到漏洞互通、协调越权行动，为什么「没有人类被告知」本身就是对齐失败，思维链监控这份「礼物」为什么脆弱，和为什么预训练 × 强化学习是乘法而不是加法。22 张配图截自视频对应时刻，图中拼有当时的英文原句与中文翻译。",
      "date_published": "2026-09-20T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "AI Deep Dive",
        "Noam Brown",
        "The Information",
        "AI Agent",
        "多智能体",
        "递归式自我改进",
        "讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/the-information-ai-deep-dive-noam-brown-transcript/",
      "url": "https://blog.chengshu.space/posts/the-information-ai-deep-dive-noam-brown-transcript/",
      "title": "《AI Deep Dive》第 1 集英文逐字稿（Noam Brown，55:13，带时间轴）",
      "summary": "The Information 节目 AI Deep Dive 首集的完整英文逐字稿：Noam Brown 谈智能体、推理、强化学习、研究品味、多智能体、Hugging Face 事件与思维链监控。这期直播回放没有任何字幕轨，逐字稿由本地 Whisper 转写整理，按句合并、逐条带时间轴。",
      "date_published": "2026-09-20T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "AI Deep Dive",
        "Noam Brown",
        "The Information",
        "逐字稿",
        "AI Agent"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/patrick-winston-how-to-speak/",
      "url": "https://blog.chengshu.space/posts/patrick-winston-how-to-speak/",
      "title": "《How to Speak》讲义：Patrick Winston 的 60 分钟演讲方法课，逐条拆开",
      "summary": "MIT 的经典演讲课《How to Speak》全解析：开场先给「赋能承诺」，想法要围篱笆，告知用黑板和道具、展示才用幻灯片；求职报告前 5 分钟要交付愿景与成果；最后一张幻灯片写 Contributions，而不是「谢谢」或「Questions?」。本文按 Winston 的讲课顺序逐节精读，19 张配图均截自视频对应时刻，并把当时的完整观点句（英文原句＋中文翻译）拼进图里。",
      "date_published": "2026-09-14T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "How to Speak",
        "演讲方法",
        "表达",
        "MIT",
        "Patrick Winston",
        "讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/patrick-winston-how-to-speak-transcript/",
      "url": "https://blog.chengshu.space/posts/patrick-winston-how-to-speak-transcript/",
      "title": "《How to Speak》英文逐字稿（Patrick Winston / MIT，1:03:42，带时间轴）",
      "summary": "MIT 传奇讲座《How to Speak》的完整英文逐字稿，按句合并、逐条带时间轴，保留了现场对话、笑声与掌声标记；专有名词（right-hand screw rule、Bartos Theater、Delores Etter 等）按上下文校正。这份逐字稿由 YouTube 自动字幕整理，原讲座以 CC BY-NC-SA 许可发布。",
      "date_published": "2026-09-14T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "How to Speak",
        "演讲方法",
        "MIT",
        "逐字稿",
        "Patrick Winston"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/dongxi-why-kv-cache-not-q/",
      "url": "https://blog.chengshu.space/posts/dongxi-why-kv-cache-not-q/",
      "title": "LLM 为什么缓存 K 和 V，却通常不缓存 Q？",
      "summary": "为什么 attention 通常缓存 K/V 而不缓存 Q：prefill/decode、causal attention，以及显存代价。转载自 Dongxi 东锡 NLP。",
      "date_published": "2026-09-13T00:00:00.000Z",
      "date_modified": "2026-09-13T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "X转录",
        "LLM",
        "KV Cache",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/karpathy-lets-build-gpt-1-bigram-baseline/",
      "url": "https://blog.chengshu.space/posts/karpathy-lets-build-gpt-1-bigram-baseline/",
      "title": "Karpathy《Let's build GPT》第一篇：语言模型、字符级 tokenizer 与 bigram 基线",
      "summary": "Karpathy《Let's build GPT》精读系列第一篇。这一篇把地基打好：语言模型为什么是「逐 token 预测下一个」；1MB 的 tiny Shakespeare 怎么变成 65 个字符的字符级 tokenizer；一段 9 个字符里为什么藏着 8 个训练样本；4×8 的 batch 里 32 个样本为什么互不通信；然后用一个 vocab×vocab 的查表（bigram 语言模型）跑通「取 batch → 算损失 → 反传 → 更新」的完整闭环，并亲手算出初始损失 4.87 比均匀分布的理论下界 4.17 还差意味着什么。文中 13 张配图均截自视频对应时刻，并把当时的完整观点句（英文原句＋中文翻译）拼进图里。",
      "date_published": "2026-09-13T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Karpathy",
        "GPT",
        "Transformer",
        "nanoGPT",
        "从零实现",
        "系列教程"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/karpathy-lets-build-gpt-2-self-attention/",
      "url": "https://blog.chengshu.space/posts/karpathy-lets-build-gpt-2-self-attention/",
      "title": "Karpathy《Let's build GPT》第二篇：自注意力——从「加权平均的数学技巧」到 Q / K / V",
      "summary": "Karpathy《Let's build GPT》精读系列第二篇，也是整门课最核心的一段。先用一个玩具例子（B,T,C = 4,8,2）把「用下三角矩阵乘法做加权聚合」这个数学技巧练熟：全 1 矩阵左乘等于按行求和、tril 是「未来不许说话」的闸门、masked_fill(-inf)+softmax 让权重从常数变成亲和度。然后把这个技巧搬进真模型，推出单头自注意力：每个 token 发出 query（我在找什么）和 key（我包含什么），q@k^T 得到亲和度、除以 √head_size 控制方差、掩码、softmax、最后 out = wei @ v。并逐条讲清四个要点：注意力是通信机制、它没有空间概念（所以要位置编码）、batch 之间永不通信、encoder block 与 decoder block 只差删掉掩码那一行。文中 14 张配图均截自视频对应时刻，并把当时的完整观点句（英文原句＋中文翻译）拼进图里。",
      "date_published": "2026-09-13T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Karpathy",
        "GPT",
        "Transformer",
        "nanoGPT",
        "从零实现",
        "系列教程"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/karpathy-lets-build-gpt-3-transformer-block/",
      "url": "https://blog.chengshu.space/posts/karpathy-lets-build-gpt-3-transformer-block/",
      "title": "Karpathy《Let's build GPT》第三篇：多头注意力、前馈、残差与 LayerNorm——拼出完整的 Transformer Block",
      "summary": "Karpathy《Let's build GPT》精读系列第三篇：把 Transformer Block 剩下的零件全部装上，并首次跑出真正的 GPT 结构。多头注意力（多条并行通信频道，输出拼接后投影回 n_embd）→ 前馈网络（逐 token 独立的 Linear-ReLU-Linear，中间放大 4 倍，负责「计算」）→ 残差连接（加法把梯度等量分给两路，形成通往输入的梯度高速公路）→ LayerNorm（归一化「行」而非「列」，只改一个维度数字；本课用 pre-norm）→ dropout，最后把模型放大到 n_embd=384、6 头、6 层、dropout 0.2。验证损失沿 2.4 → 2.28 → 2.24 → 2.08 → 2.06 → 1.48 一路下降，每一个数字对应一个具体零件。文中 13 张配图均截自视频对应时刻，并把当时的完整观点句（英文原句＋中文翻译）拼进图里。",
      "date_published": "2026-09-13T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Karpathy",
        "GPT",
        "Transformer",
        "nanoGPT",
        "从零实现",
        "系列教程"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/karpathy-lets-build-gpt-4-nanogpt-and-scale/",
      "url": "https://blog.chengshu.space/posts/karpathy-lets-build-gpt-4-nanogpt-and-scale/",
      "title": "Karpathy《Let's build GPT》第四篇：nanoGPT 走读、encoder vs decoder，以及从 1000 万参数到 GPT-3 的距离",
      "summary": "Karpathy《Let's build GPT》精读系列第四篇（收官）：把论文架构图讲完（encoder、cross-attention 与 decoder-only 的关系），走读 nanoGPT（多头合并成四维批处理、MLP 用 GELU、checkpoint 与 DDP 等工程件），对照从 1000 万参数 / 30 万 token 到 GPT-3 的 175B / 300B token——结构几乎相同，规模差约一百万倍。最后讲清 ChatGPT 的两阶段：预训练只给你一个「文档补全器」，助手化还需要 SFT → 奖励模型 → PPO 三步对齐。文末另附「从 BERT 到 GPT：同一个 Transformer，三种用法」对照表（本文补充，专门写给做过 BERT 微调、没有大模型实战经验的人）。文中 10 张配图均截自视频对应时刻，并把当时的完整观点句（英文原句＋中文翻译）拼进图里。",
      "date_published": "2026-09-13T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Karpathy",
        "GPT",
        "Transformer",
        "nanoGPT",
        "从零实现",
        "系列教程"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/paperroute-blender-gpt-astra-full-game/",
      "url": "https://blog.chengshu.space/posts/paperroute-blender-gpt-astra-full-game/",
      "title": "4 天做一个「真的能玩」的完整游戏：Blender + GPT 5.6 Astra 做 PaperRoute 的全流程",
      "summary": "Emm Tee 用 GPT 5.6 Astra 写代码、Blender 做模型，四天把 Paperboy 重做成可玩的 PaperRoute。本文逐段翻译他这份六步 runbook（机制先行、美术后置、两条流分开跑、Blender-as-code、review 渲染回路、分支实验），并站在开发者视角补齐它缺口的部分：这类项目在工程上应该怎么走——CI 门禁、性能预算、资产许可、关卡数据驱动，以及一份可执行的流程清单。",
      "date_published": "2026-09-13T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "X长文",
        "游戏开发",
        "AI 编程",
        "Blender",
        "GPT 5.6 Astra",
        "开发流程",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/google-1-hour-agentic-engineering/",
      "url": "https://blog.chengshu.space/posts/google-1-hour-agentic-engineering/",
      "title": "Google 用 1 小时把 Agent 讲清楚了：Agent → Memory → Loop → MCP → Multi-Agent",
      "summary": "这是一门被拼成 1 小时合辑的 Google ADK 课程：先用一个「规划—校验—重写」的写博客 Agent 讲清 Agent 是什么，再把记忆拆成 session、数据库会话加用户画像、Memory Bank 三层，接着用入职协调 Agent 讲长时间运行 Agent 的三个必要条件，最后落地 MCP Server 与「把 Agent 当工具挂给另一个 Agent」的多智能体结构。",
      "date_published": "2026-09-12T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "X转录",
        "AI Agent",
        "MCP",
        "Google ADK",
        "多智能体",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/xiu-mo-lun-pin-wei-de-biao-zhun/",
      "url": "https://blog.chengshu.space/posts/xiu-mo-lun-pin-wei-de-biao-zhun/",
      "title": "休谟《论品味的标准》：趣味真的无可争辩吗？",
      "summary": "1757 年休谟写下《Of the Standard of Taste》：如果趣味人人不同，好作品凭什么成立？他既不同意「一切情感都对」的怀疑论，也不同意先验规则，而是把问题换成「谁的感觉器官是健全的」，并给出真正批评家的五个条件——精微、练习、比较、免于偏见、健全理智。",
      "date_published": "2026-09-12T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "哲学",
        "休谟",
        "美学",
        "读书笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/zen-me-yong-ppt-rang-guan-zhong-ji-zhu-ni/",
      "url": "https://blog.chengshu.space/posts/zen-me-yong-ppt-rang-guan-zhong-ji-zhu-ni/",
      "title": "怎么用 PPT 让观众记住你？瑞哥那的「选题—框架—设计—讲解」全流程",
      "summary": "小红书博主瑞哥那拆解自己那套可复用的 PPT 方法：先分清信息主导与主讲人主导；选题先看受众状态、再找「我能讲别人不能讲」的独家素材；框架用标题定调并主动回答「为什么是我」；设计交给 AI 的只有风格、素材、字体三个元素；讲解时 PPT 补充你，而不是你补充 PPT。",
      "date_published": "2026-09-12T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "小红书转录",
        "PPT",
        "演讲",
        "AI",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/huangyun-editing-jobs-content-team/",
      "url": "https://blog.chengshu.space/posts/huangyun-editing-jobs-content-team/",
      "title": "为什么纯执行型剪辑岗位正承压：从“做视频”转向“懂业务”",
      "summary": "张元认为，企业账号真正要“垂直”的不是内容品类，而是目标人群：围绕同一群客户的多种焦虑选题，仍然能吸引同一批人。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "date_modified": "2026-09-12T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "X转录",
        "内容",
        "AI",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-01-overview-tokenization/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-01-overview-tokenization/",
      "title": "Stanford CS336 第一讲精读：从零构建语言模型，为什么先讲分词？",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第一讲完整讲义：这门课为什么存在、苦涩的教训的正确读法（accuracy = efficiency × resources）、从 Shannon 到 Transformer 的历史坐标、从零构建语言模型的完整流水线，以及字符级/字节级/词级分词为什么都不行、BPE 到底怎么训练出来的。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "语言模型",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-02-pytorch-resource-accounting/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-02-pytorch-resource-accounting/",
      "title": "Stanford CS336 第二讲精读：PyTorch 与资源核算——训一个大模型到底要多少显存和算力？",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第二讲完整讲义：用两道「餐巾纸算术」题讲清训练的资源账——6 × 参数量 × token 数的算力从哪来、显存为什么是 16 字节/参数、float32 / float16 / bfloat16 / fp8 怎么选、PyTorch 张量的 storage 与 stride 为什么重要、MFU 怎么算，以及从参数初始化、优化器一路到激活值、梯度、优化器状态的完整内存账本，最后落到混合精度训练。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "PyTorch",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-03-architectures-hyperparameters/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-03-architectures-hyperparameters/",
      "title": "Stanford CS336 第三讲精读：Transformer 架构与超参数，哪些改动是真的重要？",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第三讲完整讲义：主讲 Tatsunori Hashimoto 把 2017–2025 的架构演化当成一份实验数据来读——pre-norm 是唯一的铁律，RMSNorm 与删掉 bias 背后是访存而非 FLOPs，SwiGLU 为什么赢了，位置编码如何收敛到 RoPE，d_ff、头维度、宽深比、词表大小的共识与例外，以及 z-loss、QK norm、MQA/GQA 与稀疏注意力这些稳定性与推理侧的关键改动。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "Transformer架构",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-04-mixture-of-experts/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-04-mixture-of-experts/",
      "title": "Stanford CS336 第四讲精读：混合专家（MoE），以及 DeepSeek V3 是怎么搭起来的",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第四讲完整讲义：MoE 到底是什么（其实和「专家分领域」毫无关系）、为什么在 FLOPs 对齐下它稳定地赢、token choice top-K 路由为什么成为唯一收敛的答案、细粒度专家与共享专家、负载均衡的 F·P 损失与 DeepSeek V3 的「无辅助损失」偏置、专家并行与 token dropping，以及把 DeepSeek V1/V2/V3 一路拆到 MLA 与多 token 预测。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "MoE",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-05-gpus/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-05-gpus/",
      "title": "Stanford CS336 第五讲精读：GPU 硬件、访存瓶颈，以及 Flash Attention 是怎么被逼出来的",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第五讲完整讲义：为什么这一讲只讲单卡硬件、CPU 优化延迟而 GPU 优化吞吐、SM/warp/block 的执行模型与物理内存层级、同样重要的 TPU 支线，以及那张「波浪形」矩阵乘法性能图的三重谜底（roofline、tiling 对齐、wave quantization）——低精度、算子融合、重计算、访存合并、tiling 这五件提速法宝最终如何拼成 Flash Attention。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "GPU",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-06-kernels-triton/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-06-kernels-triton/",
      "title": "Stanford CS336 第六讲精读：Kernels 与 Triton——从 PyTorch 一路降到 PTX",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第六讲完整讲义：为什么 benchmark 之前必须 warm-up 加 cuda.synchronize、torch.profiler 与 Nsight Systems 里 CPU 为什么总跑在 GPU 前面、一行 print 如何毁掉整条流水线、cutlass kernel 名字里的 tile 尺寸，以及「仓库—工厂」的算子融合比喻如何变成三个真实的 GELU 实现——手写 CUDA、Triton、torch.compile——最后用 Triton 写一个融合 softmax 收尾。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "Triton",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-07-parallelism-1/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-07-parallelism-1/",
      "title": "Stanford CS336 第七讲精读：并行（上）——ZeRO/FSDP、流水线并行与张量并行",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第七讲完整讲义：单卡为什么不够用、8 卡一个盒子的网络层级、必须背下来的 all-reduce ≡ reduce-scatter + all-gather、数据并行的显存账单（16 字节每参数、5 份权重）与 ZeRO 三阶段（stage 1 在带宽受限下几乎免费，stage 3 就是 FSDP）、batch size 为什么是一种被花掉的资源、流水线并行的气泡与 zero-bubble 技巧、张量并行的切法与 8 卡经验法则、激活内存公式与序列并行，最后落到 3D/4D 并行的经验法则与 Megatron、DeepSeek、Llama 3 的真实配置。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "并行训练",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-08-parallelism-2/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-08-parallelism-2/",
      "title": "Stanford CS336 第八讲精读：并行（下）——把集合通信、数据/张量/流水线并行写成代码",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第八讲完整讲义：上一讲讲概念，这一讲把概念落成代码。从 1980 年代的集合通信原语讲起（broadcast/scatter/gather/reduce/all-gather/reduce-scatter 与 all-reduce ≡ reduce-scatter + all-gather），再看 PCIe/以太网与 NVLink/NVSwitch 两代硬件（H100：18 条 NVLink 4.0、总带宽 900 GB/s）、NCCL 与 torch.distributed 的分工，然后逐行实现三个 bare-bones 训练脚本——数据并行（在 SGD 里插一行 all-reduce）、张量并行（沿宽度切、一路 all-gather）、流水线并行（按层切、micro-batch 与点对点通信），最后用 JAX/Levanter 的十行代码对照，并总结「重计算 / 存内存 / 存到别的 GPU 再通信」这条贯穿全课的主线。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "并行训练",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-09-scaling-laws-1/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-09-scaling-laws-1/",
      "title": "Stanford CS336 第九讲精读：缩放定律（上）——从 1993 年的学习曲线，到 Chinchilla 的三种拟合方法",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第九讲完整讲义，主讲 Tatsunori Hashimoto。这一讲先把缩放定律放回统计学习理论的历史里（VC 维的 1/√m、非参数速率的 n^(−β/(2β+1))、1993 年 Bell Labs 那篇几乎没有引用的学习曲线论文、Hestness 2017 的三阶段曲线），再用两个玩具例子说明'幂律为什么是自然的'（估计均值得斜率 1、非参数回归得斜率 −1/d 与内在维度），然后逐一拆解模型工程里最实用的几件事——架构/优化器选择、深度与宽度、batch size 与临界 batch size、学习率与 muP——最后落到'更多数据还是更大模型'：联合缩放定律与 Chinchilla 的三种方法（最小包络、IsoFLOP、联合拟合），以及 Epoch AI 如何通过'从图里取证读点'复现方法三、发现原论文只是曲线拟合的残差没零均值。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "缩放定律",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-10-inference/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-10-inference/",
      "title": "Stanford CS336 第十讲精读：推理——为什么 generation 是访存受限的，以及 KV cache 的五种瘦身法",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第十讲完整讲义，主讲 Percy Liang。这一讲把'推理'当成一门系统工程课来讲：先给出一张算力账（算术强度、accelerator intensity、prefill 与 generation 的两阶段模型），说明为什么训练的瓶颈是算力而推理的瓶颈是显存带宽；再用 Llama-2 13B + H100 把延迟、吞吐与 batch size 的三角关系算成具体数字；然后用整场后半段讲 KV cache 的五种瘦身法（GQA、MLA、跨层注意力 CLA、局部注意力、以及它们的组合拳），并跳出 Transformer 看状态空间模型、线性注意力与扩散模型；最后收在两类工程手段上——有损的量化与剪枝、无损的投机解码，以及服务系统里的连续批处理、选择性批处理与 PagedAttention。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "LLM推理",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-11-scaling-laws-2/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-11-scaling-laws-2/",
      "title": "Stanford CS336 第十一讲精读：缩放定律（下）——真实世界的 scaling 配方，与 muP 的两条谱条件",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第十一讲完整讲义，主讲 Percy Liang。这是 scaling laws 两讲中的第二讲，也是更偏案例与细节的一讲。上半场把'真实模型是怎么用 scaling law 的'拆成四个案例：Cerebras-GPT 与 MiniCPM 如何用 muP 把超参从规模里解耦、MiniCPM 如何用 WSD（warm-up / stable / decay）学习率把 Chinchilla 复现从 n² 次训练压到近似一次、DeepSeek LLM 如何不靠 muP 直接网格搜索最优 batch 与学习率、以及近一年 Llama 3（约 39:1）、Hunyuan-1（96:1）、MiniMax-01 这些公开 scaling 研究给出了什么。下半场用四十分钟深潜 muP：从两条'谱条件'（初始化时激活为 Θ(1)、走一步梯度后激活变化仍为 Θ(1)）出发，在一个深线性网络上推导出 1/√fan-in 的初始化和逐层学习率 fan-out/fan-in（SGD）或 1/fan-in（Adam），再对照大规模消融实验看它什么时候有效、什么时候失效。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "缩放定律",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-12-evaluation/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-12-evaluation/",
      "title": "Stanford CS336 第十二讲精读：评估——从 perplexity 到 Agent 与安全基准，一场『评估危机』里的方法论",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第十二讲完整讲义，主讲 Percy Liang。这一讲把『评估』拆成四件事——输入、如何调用模型、如何评判输出、如何解读结果——然后带着这条线索走遍今天的评测版图：perplexity 与数据污染、MMLU / MMLU-Pro / GPQA / Humanity's Last Exam 这四代知识型基准、Chatbot Arena / IFEval / AlpacaEval / WildBench 这类指令遵循评测、SWE-bench / Cybench / MLE-bench 与 ARC-AGI 这些 Agent 与推理基准，以及 HarmBench / AIR-Bench / 越狱与『能力 vs 倾向』的安全评测。最后落到两个最容易被忽略的问题：评测离真实用例有多远（quizzing 还是 asking），以及『有效性』——训练集污染、数据质量与『我们到底在评测方法还是评测系统』。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "模型评估",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-13-data-1/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-13-data-1/",
      "title": "Stanford CS336 第十三讲精读：Data 1——数据不会从天上掉下来，从 BERT 的 BooksCorpus 到 Nemotron-CC 的完整数据工程史",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第十三讲完整讲义，主讲 Percy Liang。这一讲把『预训练数据』当成一门工程来教：开场先给出一句暴论——数据是决定语言模型好坏最重要的东西，然后从 BERT 的 BooksCorpus 与 Wikipedia 讲起，走过 Common Crawl、CCNet、C4、WebText、GPT-3、The Pile、MassiveText、LLaMA、RefinedWeb、FineWeb、Dolma、DCLM/DataComp，一直讲到 Nemotron-CC 的『改写 + 任务化』合成路线；中间穿插数据投毒、影子图书馆、代码与问答数据；最后用大段时间讲版权——许可、Creative Commons、合理使用的四个要素、以及『合法也可能被服务条款挡住』。收尾是那句关键的一课：数据不会从天上掉下来，你得动手去弄。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "数据治理",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-14-data-2/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-14-data-2/",
      "title": "Stanford CS336 第十四讲精读：Data 2——过滤与去重的算法机制，从 KenLM、fastText、重要性重采样到布隆过滤器与 MinHash LSH",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第十四讲完整讲义，主讲 Percy Liang。上一讲是数据集的编年史，这一讲转向机制：先把『过滤』抽象成『给定小份高质量目标数据 T 和大量原始数据 R，找出与 T 相似的子集 T′』，再依次讲三种实现——n-gram 语言模型（KenLM / Kneser-Ney）、fastText 线性分类器（词袋 + hashing trick）、以及重要性重采样（DSIR）；接着用语言识别、质量过滤、毒性过滤三个案例说明同一套机器怎么复用；后半程进入去重，先分清精确重复与近重复，讲哈希分组与布隆过滤器（含假阳性率的推导与最优哈希函数个数 k≈(m/n)ln2），再讲 Jaccard 相似度、MinHash 的碰撞概率证明、以及 LSH 的『与/或』band 结构如何把相似度概率曲线锐化成一条近乎阶跃的 sigmoid。收尾是那句总结：哈希之所以重要，是因为它把成对的相似与碰撞变成了一元函数，从而让线性时间成为可能。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "数据治理",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-15-alignment-sft-rlhf/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-15-alignment-sft-rlhf/",
      "title": "Stanford CS336 第十五讲精读：Alignment——人类怎么把自己的偏好装进模型，从 SFT 到 RLHF",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第十五讲完整讲义，主讲 Tatsunori Hashimoto。这一讲回答一个问题：GPT-3 已经很强了，但那支从 GPT-3 指向 ChatGPT 的箭到底是怎么射出去的？前半程拆三种完全不同的 SFT 数据（FLAN 的 NLP 任务聚合、OpenAssistant 的人类众包、Alpaca 的模型生成），讲为什么『完全正确且非常详细』的数据反而会教模型编造引用、安全微调如何在拒绝与过度拒绝之间走钢丝、mid-training 如何把指令数据塞回预训练；后半程进入 RLHF：成对偏好数据怎么收、标注员是谁为什么重要、生成-验证差距，最后把 InstructGPT 的目标函数、Bradley-Terry 奖励模型、PPO 的三次渐近尝试与 DPO 的完整推导一步步走完。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "对齐",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-16-alignment-rl-1/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-16-alignment-rl-1/",
      "title": "Stanford CS336 第十六讲精读：Alignment——从 PPO 到 GRPO，再到 R1 / Kimi / Qwen3 的可验证奖励 RL",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第十六讲完整讲义，主讲 Tatsunori Hashimoto。这一讲先把 RLHF 收尾（DPO 的梯度形状、SimPO 与长度归一化、过优化与校准度崩塌），然后正式转向 RL from verifiable rewards：从最朴素的策略梯度推到 PPO、再推到删掉 value model 与 GAE 的 GRPO；接着逐条拆解 GRPO 的两个理论瑕疵（除以标准差、按长度归一化）以及 Dr. GRPO 的修正；最后用三个案例——DeepSeek-R1、Kimi k1.5、Qwen3——对比同一套思路下不同的配方：奖励怎么设、数据怎么筛、思维链长度怎么控、thinking 模式怎么融合。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "强化学习",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-cs336-lecture-17-alignment-rl-2/",
      "url": "https://blog.chengshu.space/posts/stanford-cs336-lecture-17-alignment-rl-2/",
      "title": "Stanford CS336 第十七讲精读：Alignment——把策略梯度和 GRPO 的机制拆到底",
      "summary": "斯坦福 CS336（Language Modeling from Scratch, Spring 2025）第十七讲完整讲义，主讲 Percy Liang，也是全课最后一讲。这一讲不引入新内容，而是把第十六讲 RL from verifiable rewards 的算法机制拆开重讲：从状态/动作/奖励与结果奖励的设定出发，推朴素策略梯度为什么是「被奖励加权的 SFT」、为什么稀疏奖励会让梯度消失；用两个状态的玩具例子讲清方差问题与基线（baseline）的等价变换，再接到优势函数；然后把 GRPO 的代码从头走一遍——排序任务与部分分奖励、compute_deltas 的三种选择、冻结参数与 no_grad、裁剪的重要性比率、低方差 KL 估计，最后用一次真实的小实验说明「损失曲线为什么会骗人」。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "CS336",
        "强化学习",
        "课程讲义"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-ai-supercycle-illustrated/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-ai-supercycle-illustrated/",
      "title": "斯坦福 MS&E435 第一讲图解：AI 的钱到底被谁赚走了？",
      "summary": "Stanford MS&E435 开篇课图解版：用一张「倒三角」讲清 AI 价值链——云是 App > Infra > Semi，AI 却是 Semi 3000 亿 > Infra 750 亿 > App 600 亿；以及 75% 新增收入流向芯片层、AWS 用掉 8 年、消费级 AI 每用户年收入只有约 10 美元为什么可能靠广告解锁。",
      "date_published": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "经济学",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/what-is-vllm/",
      "url": "https://blog.chengshu.space/posts/what-is-vllm/",
      "title": "What is vLLM? Efficient AI Inference for Large Language Models",
      "summary": "IBM Technology 视频笔记：vLLM 用 PagedAttention 分页管理 KV cache、用 continuous batching 填满 GPU，论文口径吞吐约为 Hugging Face Transformers / TGI 的 24 倍。",
      "date_published": "2026-09-10T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "vLLM",
        "LLM推理",
        "PagedAttention"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/how-we-built-grok-bot/",
      "url": "https://blog.chengshu.space/posts/how-we-built-grok-bot/",
      "title": "How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)",
      "summary": "Lenny's Podcast 笔记：SpaceXAI 的 Roman Ugarte 讲述四周做出 Grok Bot——隔离小团队、云-first、每个 bot 一台电脑，以及从 has 转向 can 的产品哲学。",
      "date_published": "2026-09-08T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Grok Bot",
        "Agent",
        "产品"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/grok-bot-delegatable-team-report/",
      "url": "https://blog.chengshu.space/posts/grok-bot-delegatable-team-report/",
      "title": "Grok Bot 综合研究报告：把多智能体做成「可委派的同事团队」",
      "summary": "Grok Bot 真正重做的不是多 Agent 并行，而是让非技术用户不必当 Agent 管理员——用同事、岗位、交接与例程把多智能体产品化。",
      "date_published": "2026-09-05T00:00:00.000Z",
      "date_modified": "2026-09-05T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "Grok Bot",
        "Agent",
        "产品设计",
        "多智能体"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/sam-altman-building-openai/",
      "url": "https://blog.chengshu.space/posts/sam-altman-building-openai/",
      "title": "Sam Altman：构建 OpenAI 与押注不可能",
      "summary": "Altman 认为，AI 能力提升得很快，但经济与人的行为具有巨大惯性；这种惯性会让转型更慢，也可能让社会适应过程更平滑。今天缺少的未必是底层技术，而是类似 iPhone 的产品时刻。",
      "date_published": "2026-08-23T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "OpenAI",
        "AI",
        "创业",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-gpu-economy/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-gpu-economy/",
      "title": "Stanford MS&E435：GPU 经济 | The GPU Economy",
      "summary": "Stanford MS&E435 课堂对谈笔记：Brad Gerstner 与 Sunny Madra 讨论 GPU/推理成本、token 工厂与 AI 商业模式。",
      "date_published": "2026-07-23T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "GPU",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-ai-in-life-sciences/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-ai-in-life-sciences/",
      "title": "Stanford MS&E435：AI in Life Sciences——从分子设计到自动化湿实验室",
      "summary": "Chai 把自己定位为“分子的计算机辅助设计套件”，目标是减少抗体等药物设计中的反复试验，长期甚至希望从计算机直接生成接近可用的候选分子。",
      "date_published": "2026-07-17T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "生命科学",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-ai-supercycle-economics/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-ai-supercycle-economics/",
      "title": "Stanford MS&E435：生成式 AI 的经济学 | Economics of the AI Supercycle",
      "summary": "Stanford MS&E435 开篇课笔记：Apoorv Agrawal 用 cloud vs AI 价值三角形，讨论生成式 AI 的价值流向、推理成本与应用层翻转。",
      "date_published": "2026-07-17T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "经济学",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-enterprise-ai-saas/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-enterprise-ai-saas/",
      "title": "Stanford MS&E435：企业 AI 与 SaaS | Enterprise AI",
      "summary": "Stanford MS&E435 课堂对谈笔记：Databricks CEO Ali Ghodsi 谈企业 AI 落地、上下文瓶颈、流程重构与 SaaS 护城河。",
      "date_published": "2026-07-13T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "SaaS",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-applications-coding-ai/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-applications-coding-ai/",
      "title": "Stanford MS&E435：Coding AI 与 Agent Cloud | Guillermo Rauch",
      "summary": "Stanford MS&E435 课堂对谈笔记：Vercel 创始人 Guillermo Rauch 谈 Coding AI、Agent Cloud，以及软件免费之后部署、组合与治理的价值。",
      "date_published": "2026-06-23T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "Coding",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-building-ai-factories/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-building-ai-factories/",
      "title": "Stanford MS&E435：建设 AI 工厂 | Building AI Factories",
      "summary": "Stanford MS&E435 课堂对谈笔记：Crusoe CEO Chase Lochmiller 谈吉瓦级 AI 数据中心与从电子到 token 的工厂经济学。",
      "date_published": "2026-06-17T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "基础设施",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-applications-applied-ai/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-applications-applied-ai/",
      "title": "Stanford MS&E435：Applications, Applied AI——从 Baseten 看推理经济",
      "summary": "Tuhin 将推理视为 AI 产品交付价值时持续发生的成本项。Baseten 的定位，是把模型优化、部署、跨云容量、可靠性、可观测性和安全能力打包成生产平台。",
      "date_published": "2026-06-05T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "应用",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-infrastructure-capstone-case/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-infrastructure-capstone-case/",
      "title": "Stanford MS&E435：AI 基础设施与前沿实验室 | Sachin Katti",
      "summary": "Katti 把 AI 算力定义为一整套工业系统，而不只是 GPU：芯片、内存、网络、电力、制冷、数据中心、发电、配电和土地必须在同一时间到位。真正困难的工作往往从签完合同之后才开始。",
      "date_published": "2026-05-27T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "基础设施",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/stanford-mse435-enterprise-internal-knowledge/",
      "url": "https://blog.chengshu.space/posts/stanford-mse435-enterprise-internal-knowledge/",
      "title": "Stanford MS&E435：企业内部知识与专用模型 | Yash Patil",
      "summary": "Stanford MS&E435 课堂对谈笔记：Applied Compute CEO Yash Patil 谈企业内部知识如何变成专用模型，以及 Evals、RL 环境与持续学习。",
      "date_published": "2026-05-22T00:00:00.000Z",
      "date_modified": "2026-09-10T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "youtube转录",
        "Stanford",
        "AI",
        "企业知识",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/huo-shan-yin-qing-force-dong-ji-da-hui-shou-ji/",
      "url": "https://blog.chengshu.space/posts/huo-shan-yin-qing-force-dong-ji-da-hui-shou-ji/",
      "title": "火山引擎 Force 冬季大会手记",
      "summary": "",
      "date_published": "2025-12-18T00:00:00.000Z",
      "date_modified": "2025-12-18T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": []
    },
    {
      "id": "https://blog.chengshu.space/posts/ai-du-li-kai-fa-zhe-huo-dong-bei-wang/",
      "url": "https://blog.chengshu.space/posts/ai-du-li-kai-fa-zhe-huo-dong-bei-wang/",
      "title": "AI独立开发者活动备忘",
      "summary": "日拱一卒，难在日拱",
      "date_published": "2025-12-15T00:00:00.000Z",
      "date_modified": "2025-12-15T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": []
    },
    {
      "id": "https://blog.chengshu.space/posts/yi-ci-hei-ke-song-de-fu-pan/",
      "url": "https://blog.chengshu.space/posts/yi-ci-hei-ke-song-de-fu-pan/",
      "title": "一次黑客松的复盘",
      "summary": "那些我以为懂、其实没懂的产品常识",
      "date_published": "2025-12-15T00:00:00.000Z",
      "date_modified": "2025-12-14T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "思考",
        "建站"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/wei-shen-me-ni-de-shen-du-hao-wen-gen-ben-mei-ren-kan/",
      "url": "https://blog.chengshu.space/posts/wei-shen-me-ni-de-shen-du-hao-wen-gen-ben-mei-ren-kan/",
      "title": "为什么你的深度好文，根本没人看？",
      "summary": "Paul Graham用“引力场模型”讲透了内容创作的残酷真相。",
      "date_published": "2025-11-18T00:00:00.000Z",
      "date_modified": "2025-11-18T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "ai generated"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/quan-jing-guan-huan-jia-xue/",
      "url": "https://blog.chengshu.space/posts/quan-jing-guan-huan-jia-xue/",
      "title": "《权经》与官宦家学：一套关于权力运作的历史叙事",
      "summary": "视频以《权经》为线索，把“圣人之说”与“官宦家学”区分开：前者提供公开的道德语言，后者被作者描述为处理现实利益、组织与权力关系的实践知识。",
      "date_published": "2025-09-28T00:00:00.000Z",
      "date_modified": "2026-09-11T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "bilibili转录",
        "历史",
        "权力",
        "笔记"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/example-1/",
      "url": "https://blog.chengshu.space/posts/example-1/",
      "title": "2025年微短剧行业深度趋势报告",
      "summary": "微短剧现状 AI 调研",
      "date_published": "2025-01-20T00:00:00.000Z",
      "date_modified": "2025-11-16T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "ai generated",
        "manus",
        "短剧"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/go_talk/",
      "url": "https://blog.chengshu.space/posts/go_talk/",
      "title": "围棋的启发2",
      "summary": "时间又过了三个多月， 来到了 8k 水平， 开始比较关注厚势了。",
      "date_published": "2021-12-05T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "talk",
        "go"
      ]
    },
    {
      "id": "https://blog.chengshu.space/posts/go/",
      "url": "https://blog.chengshu.space/posts/go/",
      "title": "围棋的启发",
      "summary": "从2020年8月份开始学习围棋, 断断续续玩了一年了, 水平可能有了10k, 这个进步的速度其实是比较慢的, 究其原因还是自己太懒, 不愿意去下苦功夫研究定式和死活, 当成了一个纯粹工作之余的消遣活动",
      "date_published": "2021-08-22T00:00:00.000Z",
      "authors": [
        {
          "name": "Chengshu"
        }
      ],
      "tags": [
        "history",
        "go"
      ]
    }
  ]
}