13 · 多 Agent、Job、Goal 与 Schedule:四类不同的持续工作¶
不要把所有“继续干活”都叫自主 Agent。DSH 分开:子 Agent 拥有自己的 Session;Job 跟踪后台 producer 的状态和 output;Goal 在同一 Session 中自动提交下一轮目标输入;Schedule 由宿主定时向指定 Session 投递 followup;实验 Agent Teams 在 continuable children 上加任务板和邮箱。这些机制可组合,但结果、取消和持久化保证不同。
13.1 Subagent seam 不绑定一种 provider¶
ctx.subagents 注册 SubagentProvider。one-shot start 请求含 parent、prompt、signal 和可选 agentOptions/outputSchema/maxDepth/toolFilter/persona;服务检查 provider capabilities,一项不支持就明确 UNSUPPORTED_CAPABILITY,而非悄悄忽略。provider 可在进程内创建 DSH Agent,也可委派给外部产品;能力声明与取消义务决定如何消费,而不是看到名字像 Agent 就推断可 fork/resume。请求与能力,start/capability 检查。
agentOptions 是 route/effort/tokens override,不是任意新子树 config;persona 在 in-process child 中遮蔽 deployment prefix;toolFilter 在 unpublished setup 阶段 restrict,使提示词与执行可见性一致;outputSchema 是 object-rooted supported JSON schema,结果不满足就 error,不保证 structured 必然存在。
13.2 spawn 与 fork 的真实上下文差异¶
spawn-in-process 不带 parent conversation seed,inheritsParentContext=false。child 仍可从全局部署获得一般 persona/capabilities,并从 parent 继承 workspace/route 等指定 metadata;“fresh”不能解释成整个产品依赖树完全隔离。
fork-in-process 取父最后 turn/end inclusive 的事件前缀,排除当前开放 turn(尤其当前 delegation tool call),未完成任何 turn 则 fresh。prefix 在创建时截一次,continuable child 冷恢复以后继续自己的 durable history,不重新取父最新历史。spawn,fork。
这是与通用 SessionStore.fork 切任意有效 seq 并补 open-tail 的区别。你写 reviewer Agent 可 fresh 防止继承作者的叙述偏见,写接手编码的 child 可 fork 保留已完成背景;独立审查是否成立还要看任务、源码、证据和修订复验,不能只看 fresh flag。
13.3 深度预算跨恢复保持¶
delegationDepthOf 取 max(session.header.delegationDepth ?? 0, options.subagentDepth ?? 0),runtime 可加深不能降低;runtime 值校验 safe integer/nonnegative/no negative zero。childDepth=parentDepth+1,maxDepth 是 absolute cap,不是“这个父还剩几个孩子”,childDepth 不能超过 cap。深度全部算法。
默认 maxDepth=1,默认 maxActiveSubagents=8(volatile Config)。provider-managed 表示外部 provider 自己管深度,不是把本地深度检查无限放宽后还声称递归受控。自己的多代理产品应设任务数量/层级/模型调用预算,并避免把 restarted child 当成深度0。
13.4 one-shot:一个 run,但 result 与 release 都要成功¶
provider.start 返回 SubagentRun,result 表示最终 child outcome,dispose 负责释放。signal 启动前 abort 要清理 partial resources 后 reject,run 已发布后要取消剩余工作。subagent/start/end 以 runId 成对发布,listener 失败 contained;end 不可仅从“dispose 正常结束”推断 completed,要从孩子实际 Session 的 turn/error/refusal/cancel 结局来。
output 取最后非空 assistant message;没有这类消息可回 accumulated assistant text 或[];stopReason 非 completed 时 output 可能 partial。外部 diagnostic 必须保留在助手内容外,且作者约定不能泄露工具输入、文件、环境和原始 protocol。结果字段,生命周期观察器。
背景 one-shot 用 settleRun 映射 Job:completed 带最终 text;无 diagnostic 的 localaborted 是 killed;remote diagnosed abort、error、max-tokens、refusal 为 failed。不论 result 如何先 await 它再 dispose;dispose 失败也 failed,双失败保留两段 detail。settleRun 全文。
13.5 continuable:稳定 Session,不是每次消息创建 Job¶
continuation manager 预留 childId,provider.prepareContinuable 只返回 detached seed data,不交付 Agent/result/dispose/resume 能力。manager 拥有所有后续 materialize、scopedcomposition、inbox、归属和 release。必须有 persistence 与 session-query,否则报 PERSISTENCE_UNAVAILABLE/CONTINUATION_UNAVAILABLE。创建与冷恢复入口。
startContinuable 返回 childId 与 accepted inbox messageId,不等孩子一轮结束。调用方取消仅拥有 pre-acceptance window;accepted 之后 child 由 manager 管理。continuable 没有“每条消息的 result Promise”或对应 Job,不能把 one-shot 等待方式直接套过来。
sendMessage 只授权相邻 parent/directcontinuablechild,sender attribution 从 exactliveAgent 构造;running target steer 到下一 step,idle targetwake,冷 child 从 durable descriptor 恢复。hostprotocol 的 source 透传走私有 symbol seam,普通模型不能冒充宿主来源。公开消息语义。
interrupt 是 fire-and-return 发 cancel;仅保留未 claimpending、activation 和 descendants,claimed 输入不自动重排;whenidle 之后新 wake 重启停着的 FIFO。target 不存在/one-shot 未知/无 manager 是 acceptedno-op,live target 错误 authority 则 UNAUTHORIZED。drain 先关 admission,再停止 descendant,await 已接收 materialization 并 child-firstrelease。取消不是删除 durable child,也不等于把它所有后代杀掉。interrupt/drain。
13.6 Jobs 是进程内资源追踪,不是 durable workflow 引擎¶
LocalJobRegistry 把 lifecycle、boundedoutputring、modelcursor 留在内存,注册记录可超出 producer/controllerFiber 寿命;Agent 或 service 卸载会 cancel 其 livework 并 awaitcompliantproducer。进程重启不会凭这张 Map 恢复所有外部工作。provider 状态。
start 先 resolveowner 为 exactliveAgent,确认 owner scope 有 controller(如 tool-jobs),再验证 kind/label/outputLimit 与 active count。默认每 owner10 个 running+stopping;live ring256KiB、settled16KiB、pullpoll150ms 可配置。id 在 spec.run starter 之前发放,starter throw 不留注册记录;starter 可先 appendoutput,commit 后 producerState.job 绑定 record,再 publishregistered,再启动 pump。
hooks.done reject 是 producer 契约违例,normalizedfailed;pump 最后 drain 后才 settle 和 trim。view 是 freshcopy,ringread 返回 copies,不给工具裸 mutablejob。ownedjob 只有相同 Sessionid caller 可读取/kill;unownedbucket 共享,不要把 publiccaller-lessread 当 superuser。
read 移动 modelcursor,仅第一次 settledread 交 result,terminalread 之后 trim 到 settledcap;readAt 不移动 cursor,用于观察 UI。ringeviction 会有 lossy 标志,settlement 临时保留 unconsumedbytes 并不能保证 output 无限无损。kill 先调用 cancel,成功后 stopping,重复 reason 最后 writerwin;cancelthrow 不会假装状态已变 killed。wait timeout 表示“等待结束但 job 还活着”,不是 kill;remove 只允许 terminal。start/read/kill 核心。
13.7 Goal:持久 phase 与临时 activation 分离¶
GoalSnapshot 有 id/revision/objective/phase/maxGoalRounds,GoalRef 为 compare-and-set。phase=active/paused/blocked/complete 是 durable;activation=armed/disarmed 是 process-local。create 新 revision1 且 armed;恢复 Agent、重新挂 driver 都 disarm,因此磁盘里的 active 不自动授予重启后继续耗费模型/工具的权限。状态定义,service 初始化。
create 拒绝已有非 completegoal;edit 至少改 objective/cap 且保持 phase;pause 只 active→paused/disarmed;resume 允许 active/paused/blocked 但不能 alreadyactivearmed,必须预算还有剩余,revision+1/armed;complete 接受 active/paused/blocked;block 只 active 且带 code/message;clear 写 tombstone,不删 history。每次 mutation 旧 ref 被拒绝,避免异步 UI 拿旧 revision 误改新 goal。所有变更。
当前 goal/change 是 log-only durable 事件,完整 snapshot/change metadata 通过 strict decoder/fold 解释;不能直接将它硬当作 user/message。当前 GoalMessageSource.round 要求正整数;旧 round-zero 输入在 strict fold 中拒绝,driver 的清除分支保留防御处理不代表当前 API 还产生 round-zero。roundsStarted 在 positive goal-round user/message 实际进入时增长,排队或 claimed 不算已开始回合。purefold。
maxGoalRounds 默认256,是“自动 goalround”上限,不等于总 tokens、总 walltime 或 step 数。Pi 课程的 token budget 不应照搬成 DSH 的同名 API。要控制成本还需 providerusage、外部预算插件和 host 监督。
13.8 goal-round-driver 如何避免幽灵续轮¶
readyToDrive 要求 driverFiberACTIVE、exactAgent 仍 live、Agentidle、无 competingprompt 且不 stopping。状态按 Agent 对象而非裸 Sessionid 索引,旧 lifecycle 不能误驱动新的同 idAgent。
mutation/round 结束后先 sessions.flushcheckpoint,再重新检查条件;flushfaildisarm。当前 attempt 在 queued/claimed/admitted 三个阶段维护身份 goalId/revision/round/messageId 和完整 content。只有一次 reservation 可存在,requestDrive 把触发 coalesce 到单 agentserializedrun。driver 读/预约。
goalactivearmed 且有 budget 时 followup 生成 positivegoalrounduserinput;达到 cap 则 blocked(round-limit)。其他 queuedhumanprompt 标 competing 并把 queuedroundstale;agent/pre-step 先验证 reservation 再 next,再验证 next 的 async 期间没有 revision/authority 变化。无效时 restoreOtherClaimed 避免误丢其他 context,reject 自己的 round;成功返回 {...decision,startsRequestSeries:true},保留下游 rewrittenmessages 和其他声明。admission 双重检查。
hostpause 在 running 且 currentInitiator 不是 agent 自身时 cancel({user},keepInbox:true);model 工具内 pause 则本 turn 正常收尾。取消/错误/max-tokens 会暂停或 disarm,不能因为 phase 还 active 就马上再 wake 无限循环。driver 卸载先 stop/reservationstale/cancelneededround/drainrun,再移除 listeners,以保留 stepfence 到 quiescence。
13.9 Schedule:确认投递,不等于完成工作¶
ScheduleRuntime 保存至多一个 timer,requestDrive 清旧 timer 并 serializedmanagement/delivery 事务;dispose 关闭 timer 并等待已 admitteddelivery。MAX_TIMER_DELAY_MS=2147483647,远期定时需分段 recompute,timer.unref 不单独保活进程。runtime 全文。
drive 取到期 active tasks;recurring 同 Session 可 batch,resume/resolveAgent 的 await 之后再检查 wallclock,以免时钟回拨把 future 任务提前发送。followup 同步 durableinboxsplice,再 sessions.flush;必须返回 true 才承认 persistenceack;之后才 committaskreceipt/status/history 和 recurringnexttime。
这确认“Session 输入投递”,不是“模型已经成功处理提醒”,也不是 external 副作用 exactly-once。inboxflush 之后 taskreceiptcommit 失败存在跨两个存储事实的窗口,需要按产品语义考虑重复/核对;源码不提供一个涵盖它们的 distributedtransaction。失败任务在当前 scan 的 failedset 排除自动短周期 retry,后续 explicitrecompute 可重试;不是永久取消也不是保证 backoff 重试。投递与失败路径。
13.10 Experimental Agent Teams:已发布、opt-in、预稳定¶
TeamId 是顶层 LeadSessionid;continuablesubagent 是 teammate。成员 durablephaseprovisioning/active/failed,runtimeview 还含 running/inactive;任务 pending/in_progress/completed/deleted,每 mutationrevision+1。leadlog 存全 team/member、team/task、queued/deliveredfacts,目标 Session 存带 team-messageidentity 的实际消息。完整数据词汇。
TeamJournal 按 rootid 串行 read-check-append-and-flush;失败 tail 用 then(success/failure)收敛,后续事务不因前一次 reject 永久卡住。projectionfailure 时 failclosed,不用破损任务状态继续执行。journal。
TaskBoard.claim 要求 pending、ready、owner 未被别人占;complete 要求 in_progress 且 owner/lead 授权;release/reopen 去 owner;reassign 仅 lead;delete 有 dependent 则拒绝。dependencies 校验存在、未 deleted、不重复、不 self,再完整 DAGcyclecheck。expectedRevision 过期则 STALE_REVISION;无互斥 CAS 的 read-then-write 会让两个 teammate 同时 claim,所以队列/版本都不可省。任务板,graph 校验。
writeScopes 为 advisory 冲突诊断,不是 filesystem 锁或 sandbox 授权;roster/membervisibility 不是拿到任意 Sessionsecret 的权限。不要把实验 feature 描述成生产分布式事务平台。
13.11 Team mailbox 的跨 Session 确认顺序¶
send 在 roottransaction 中验证 exactmembership、target、自发给自己、pendingcap 和 sender-framedUTF8size;先 appendAndFlushqueuedfact,再注册 target-localdispatch,维持并发 sender 的 durableorder。每 messageinflightset 防同进程重复,每 targetdispatchtail 串行 admission。
如果 target 已记录 identity,先 flush 再 ack,不再次 steer。冷 target 先读 persistentownsuffix 判 dedup,uncertainty keepsqueued;recoverFor 按 Leadpendingfacts 重投。targetuser/messageobserver 也异步 checkpointack,所以 receipt 可迟于 sendaccepted。accepted/queued 不等于目标“答复已完成”,失效或失败仍在 durablemailbox 等待恢复。mailbox 全流程。
去重由 target 已接收 identity 及日志事实支撑,不覆盖模型接收以后执行的所有外部副作用。root 与 target 分别 flush 也不是全局事务。实际测试需覆盖 crash 窗口和队列顺序,不应仅验证工具返回 accepted。
13.12 为自己的 Agent 选择最少必要机制¶
想并行拿一个审查结论:one-shot 子 Agent,要求 result+dispose 成功、parent 自己的预算/整合;想保持一个长期专家:continuable child,并设计 activation/cold-resume、消息授权和 durability;想跑耗时进程:Job+可取消 producer+控制工具;想按目标自主继续:Goal+rounddriver+roundbudget;想到期提醒:Schedule 投递和 receipt,不把它当答案验收;想多个 peer 共同 claim 任务:实验 Teams+revisionDAG/mailbox,并额外解决 writeconflict 和 humanreview。
本章所述主流程来自连续源码;外部 Codex/ACP/SSH transport 实际联网、多进程崩溃压力、真实模型以及实验 Teams 生产长期运行不属于已完成的离线验收。各 adapter 配置和版本差异还需对应章节,而不能由 in-process 行为推断。