第 10 周 · Tether 源码课
Projection 与 Persistence:事实、视图和缓存
学完你能做到
- 区分 authoritative event log、projection state 与 checkpoint cache
- 解释 ProjectionDefinition 的 state/view schema、init、apply、StateOf 与 stateVersion
- 说明 online fold、full replay 和 checkpoint + tail 为什么应得相同结果
- 用 SessionProjectionRegistryTests 验证一致 cut 与 cache failure 边界
课程进度
- 第 1 周
- 第 2 周
- 第 3 周
- 第 4 周
- 第 5 周
- 第 6 周
- 第 7 周
- 第 8 周
- 第 9 周
- 第 10 周
- 第 11 周
- 第 12 周
- 第 13 周
- 第 14 周
- 第 15 周
- 第 16 周
一句话先懂
Persistence 保住事实,Projection 把事实折叠成好读的答案,Checkpoint 只让重算少走一段路。
日志是 authority;projection 与 cache 即使全部删除,也应能从完整日志恢复出同样的 snapshot。
如果每次显示“共 382 条 event、7 个 turn、当前标题”都从 seq 0 重放,答案可靠但代价会增长。Projection 把重放规则变成可注册的纯 fold,checkpoint 再保存中间进度;两者都不能反过来覆盖原始 log。
先看大图
一个类比:同一份流水账生成余额与月报
银行流水是一笔笔不可变交易;余额、月度支出和分类图表都能从流水折叠出来。每天保存一次余额能加快第二天计算,但那张便签丢了,重新加总流水即可。若便签与流水冲突,应丢便签,不应改流水迎合它。
映射到 Tether:
- Session events 是流水。
ProjectionDefinition<TState,TView>是“怎样加总”的规则;host-only unit 只有 state。ProjectionSnapshot是某个asOfSeq的余额表。- checkpoint 是带
stateVersion与seq的中间便签。 - persistence Provider 保管可恢复的 event log bytes。
类比边界:Projection state 是带 schema 的 typed fold 值,client view 是另一份 schema;Provider lease、Coordinator write-behind、torn-tail repair 与 cache identity 也远比便签复杂。类比只说明 authority 与 derivation 的方向。
一项 Projection 的 state 与 view
src/Tether.Session.Projection/ProjectionTypes.cs 把内部 fold 与 client view 分开:
| 字段 | 作用 | 纪律 |
|---|---|---|
Key | 全局识别这一项 projection | 兼容定义才共享;不兼容 fail loud |
StateVersion + state schema | state shape 与校验 | 仅 version 相同不够;畸形 checkpoint 必须丢弃重放 |
Init / Apply | 空日志与纯转移 | 无关 event 必须返回同一个 state 实例 |
viewSchema / View | client-visible unit 的 wire 值 | host-only unit 省略;snapshot 不含它们 |
StateOf | 同进程 borrowed host state | 不 clone、不进 API/downlink |
为什么“无变化返回同一 reference”?registry 用 reference equality 判断是否触发 change feed。host-only 变化也不会进入 snapshot。
Registry 实现与 checkpoint store 在 Tether.Session.Projection.Runtime:消费者只引用 Service Definition。cell 绑定 exact ISession 实例。cache identity 对齐 header 的 version/createdAt/cwd/composition/profile。
三条得到同一答案的路
Online fold
SessionProjectionRegistry 订阅 committed session/event,按 exact ISession 实例推进 cell。Snapshot(session) 只读取 client-visible view,返回一致 cut 与 AsOfSeq。
Full replay
late registration 或 cold start 没有 cell 时,从 session.Events 的 seq 0 调用 Init,再逐条 Apply。它较慢,却是验证其他路径的基准。
Checkpoint + tail
checkpoint row 记录 Ver / Seq / Val。若 key 与 version 匹配,restore 从该 seq 后的 log tail 继续 fold;不匹配就必须回到更早位置或从 0 重算。RestoreFloor 计算所有 registered units 都能安全恢复所需的最早 tail 起点。
Persistence 与 Projection 不互相代替
ISessionPersistence 是 Coordinator 提供给 consumers 的协调面:
- batch 必须连续,first seq 必须等于 stored next seq;
AttachAsync把 exact live session 接上 write-behind;FlushAsync等待排空,退出或需要 durability fence 时必须调用;InspectAsync只读检查 torn tail,LoadAsync才持久修复;- 它不发布 live Agent,也不把 projection cache 当事实来源。
SessionPersistenceCoordinatorPlugin 提供这条 seam,并消费 composition 中唯一的 ISessionEventStore。base/headless/CLI 默认把该 store 绑定到 JSONL;SQLite patch 只替换 session-store row。两个 Provider 都只拥有物理 codec、原子 batch、range/list、repair fence 与 lease,不订阅 Session event,也不构造 Session/Agent。
首个 tag 前 physical schema 没有兼容承诺;教材只依赖 provider-neutral 契约,不要求学生背 JSONL frame 或 SQLite table layout。
动手实验:三条路线必须相等
实验目标:用测试中的 test/count projection 比较 online snapshot、late full replay 与 checkpoint + tail。
1. 预测
测试里的 Count 定义对每一条 event加 1。一个 session 追加 turn/start 与 turn/end 后:
AsOfSeq是 1 还是 2?- count view 是 1 还是 2?
- late register 后 full replay 会不会少掉注册之前的 events?
再预测:若 checkpoint 的 StateVersion 与 definition 不同,系统应信旧 Val,还是回到 log?
2. 运行
dotnet test \
tests/Tether.Session.Projection.Tests/Tether.Session.Projection.Tests.csproj \
--filter FullyQualifiedName~SessionProjectionRegistryTests
3. 观察
在 tests/Tether.Session.Projection.Tests/SessionProjectionRegistryTests.cs 对照四个场景:
- online snapshot:
AsOfSeq == 1,count 是 2;seq 从 0 开始。 - late registration:第一次读取会 full replay,因此与 online 相等。
- restore floor:usable rows 加 log tail 后与 online fold 相等。
- cache + selected store:online、cold cache + tail、full log 三条路线相等;alien identity 的 cache 被拒绝。
4. 解释
“cache write failure stays behind the log”表示 checkpoint 写失败不能让 committed session append 失败。最坏结果只是下次多 replay 一段;如果为了 cache 成功而回滚 log,优化层就反过来控制了 authority。
一张选择表
| 需求 | 应读取什么 | 原因 |
|---|---|---|
| 证明某个事件是否发生 | durable session log | 事实 authority |
| 显示当前统计/标题 | projection snapshot | 已折叠且有 asOfSeq |
| 加快 cold read | version-matched checkpoint + log tail | 结果仍由 log 验证 |
| 退出前确保 bytes 落盘 | FlushAsync | write-behind 需要显式 fence |
| 怀疑崩溃留下 torn tail,但不想改盘 | InspectAsync | 只读逻辑视图 |
检查理解
1. checkpoint 文件删除后,session 是否损坏?
查看答案
不应损坏。checkpoint 是 discardable fold shortcut;只要 authoritative log 完整,就能从 seq 0 重建 projection。性能会下降,语义不应改变。
2. projection 的 StateVersion 为什么不能拿 SQLite schema version 代替?
查看答案
它们描述不同边界:StateVersion 表示某个 projection key 的 state shape;SQLite schema version 表示 provider 的物理布局;SessionFormat 又表示 event vocabulary。任一变化不必强迫另外两者一起变化。
3. write-behind 已 attach,为什么还需要 FlushAsync?
查看答案
attach 让提交后的 events 异步排队落盘,不承诺某个具体时点已经 durable。FlushAsync 才是调用方可等待的 durability fence。
本周带走
- durable event log 是 authority,projection state 与 checkpoint 都是可重建派生物。
- online、full replay、checkpoint + tail 必须在同一 asOfSeq 给出相同 snapshot。
- cache failure 最多损失性能;不能越过 log 改变事实语义。
下一周沿 Agent 的 request 走到 LLM stream,观察网络 frame、typed update 与 session commit 的边界。详细契约可查会话投影和会话持久化。