Anthropic 深读¶
Anthropic 公开 4 大研究 team:Alignment / Interpretability / Societal Impacts / Frontier Red Team。加上 agent 应用层和工程基础设施,本章按以下顺序展开:
- 工程深读 — 怎么训出 Claude(pretraining infra / JAX 栈 / 数据 / serving)
- 机制可解释性 — SAE / circuit analysis / attribution graphs(Interpretability team)
- 对齐研究 — CAI / Alignment Faking / Reward Hacking / Sycophancy(Alignment team)
- Agents 工程 — Computer Use / MCP / Tool Use(agent 应用层)
- Safety + Frontier Red Team — Constitutional Classifiers / RSP / 多层防御(Red Team)
- 社会影响研究 — Economic Index / Project Vend / Project Deal(Societal Impacts team)
跟其他实验室相比 Anthropic 的特色:
- 公开度居中,但 alignment / interp 研究 paper 数量遥遥领先
- 4 大 team 有专门 societal impacts 研究(OpenAI / DeepSeek 都没有这条线)
- 工程层面 JAX + 多硬件栈是独特路径
↑ 上级 · K. 前沿实验室深读