
一轮 assistant 消息可以带多个内容块。写死的一轮里,模型先写一句文字,后面跟着两个 tool_use,一个查天气,一个查汇率,用户的问题要求两件事一起给。循环跑完却只动了一个工具:第二个调用留在消息里,没人执行,也没有对应的 tool_result,循环自己报告成功。两个调用只落地了一个,另外一半的意图断在循环里,外面看不出来。
桩响应里两个 tool_use 都算数
先把这一轮的 assistant 消息钉死:
== 桩响应:第一轮 assistant 消息(写死,drive 最小 agent 循环) ==
role: assistant
block[0] text "要两个数据,一起查。"
block[1] tool_use id=toolu_w1 name=get_weather input={"city":"北京"}
block[2] tool_use id=toolu_f1 name=get_fx_rate input={"base":"USD","quote":"CNY"}
user 提问:"查一下北京今天的天气,再查一下美元兑人民币的汇率,用一句话总结。"
模型在同一轮里怎么安排这两个块,没有实测;这里用写死的桩响应驱动一个最小循环,只盯循环那一侧拿到消息后怎么处理。取调用那行通常写成:
const call = msg.content.find(b => b.type === 'tool_use'); // 取第一个
content.find 和 content[0] 在这个消息上返回同一块,都是 get_weather。第二个块只有用 filter、for...of 或按下标遍历才拿得到。
只取第一个之后
把循环连跑两种桩。桩模型的行为也写死:一种发现汇率没到,就把同一个计划再发一遍;另一种把那件事直接略过。
== 配置:只取第一个 tool_use / 桩模型:跳过 ==
round 1 assistant 文本:要两个数据,一起查。
tool_use:get_weather(toolu_w1), get_fx_rate(toolu_f1)
执行:get_weather -> {"city":"北京","cond":"晴","temp_c":24}
丢弃(未执行):get_fx_rate(toolu_f1)
round 2 assistant 文本:北京今天晴,24 度。汇率那部分略过。
tool_use:无
历史 role 序列:user assistant tool assistant
未配对的 tool_use:toolu_f1(有调用,卡在历史里没有对应 tool_result)
执行工具总次数:1(get_weather)
结束方式:第 2 轮模型不再发 tool_use,循环正常结束
最终一句话:北京今天晴,24 度。汇率那部分略过。
含天气=true 含汇率=false
两轮就收尾,循环认为功课做完了。最终那句话少了一整个汇率,剩下的答案看着挺完整,缺了什么都不报错。这是一类静默的坏法:循环成功结束,因为模型把第二件事自己跳过了。
另一种桩更难看:
== 配置:只取第一个 tool_use / 桩模型:追问 ==
round 1 assistant 文本:要两个数据,一起查。
tool_use:get_weather(toolu_w1), get_fx_rate(toolu_f1)
执行:get_weather -> {"city":"北京","cond":"晴","temp_c":24}
丢弃(未执行):get_fx_rate(toolu_f1)
round 2 assistant 文本:汇率还没拿到,再要一次。
tool_use:get_weather(toolu_w1), get_fx_rate(toolu_f1)
执行:get_weather -> {"city":"北京","cond":"晴","temp_c":24}
丢弃(未执行):get_fx_rate(toolu_f1)
round 3 assistant 文本:汇率还没拿到,再要一次。
tool_use:get_weather(toolu_w1), get_fx_rate(toolu_f1)
执行:get_weather -> {"city":"北京","cond":"晴","temp_c":24}
丢弃(未执行):get_fx_rate(toolu_f1)
round 4 assistant 文本:汇率还没拿到,再要一次。
tool_use:get_weather(toolu_w1), get_fx_rate(toolu_f1)
执行:get_weather -> {"city":"北京","cond":"晴","temp_c":24}
丢弃(未执行):get_fx_rate(toolu_f1)
round 5 assistant 文本:汇率还没拿到,再要一次。
tool_use:get_weather(toolu_w1), get_fx_rate(toolu_f1)
执行:get_weather -> {"city":"北京","cond":"晴","temp_c":24}
丢弃(未执行):get_fx_rate(toolu_f1)
round 6 assistant 文本:汇率还没拿到,再要一次。
tool_use:get_weather(toolu_w1), get_fx_rate(toolu_f1)
执行:get_weather -> {"city":"北京","cond":"晴","temp_c":24}
丢弃(未执行):get_fx_rate(toolu_f1)
历史 role 序列:user assistant tool assistant tool assistant tool assistant tool assistant tool assistant tool
未配对的 tool_use:toolu_f1(有调用,卡在历史里没有对应 tool_result)
执行工具总次数:6(get_weather, get_weather, get_weather, get_weather, get_weather, get_weather)
结束方式:到 6 轮上限仍未收敛
最终一句话:汇率还没拿到,再要一次。
含天气=false 含汇率=false
天气在每一轮都被重新查一遍,六轮下来查了六次,汇率一次没跑。模型每一轮发出同一个计划,循环每一轮只认头一个块,两边对着转,谁也推不动。这就是另一类现象:模型反复追问同一件事,而追问的那件事早就在它自己发的消息里。
循环最后收敛成什么样,取决于桩模型怎么应对那半份返回。桩把两条路都写死了:一条当作没发过,把整个计划重发一遍;另一条绕开缺的那件事,用手里的数据收尾。真实的模型会走哪条,没有实测,能确定的是循环把两条路都走到了,一条走到轮数上限,一条两轮收尾。
丢掉的那块,历史里没有下文
两种下场的共同点在那个被丢掉的块。toolu_f1 是模型发出的调用,历史里却没有任何一条 tool_result 认领它。要求「每个 tool_use 都必须有配对的 tool_result」的服务端,会直接回 400 把整条请求拒掉。这条服务端行为没联网实测,是从消息结构的配对要求推出来的结论;本机这段桩能坐实的只有配对缺失本身。
要提前发现这种断链,可以在消息入历史前对一遍两个集合:这一轮所有 tool_use 的 id 收一边,历史上所有 tool_result 的 tool_use_id 收另一边,两边的差集非空,就说明有调用没人接。差集里出现的 id,就是被循环跳过的那些块。
取全部之后
把取块那行换成 filter,两个调用都执行:
== 配置:取全部 tool_use / 桩模型:完整 ==
round 1 assistant 文本:要两个数据,一起查。
tool_use:get_weather(toolu_w1), get_fx_rate(toolu_f1)
执行:get_weather -> {"city":"北京","cond":"晴","temp_c":24}
执行:get_fx_rate -> {"base":"USD","quote":"CNY","rate":7.12}
round 2 assistant 文本:北京今天晴,24 度;美元兑人民币 7.12。
tool_use:无
历史 role 序列:user assistant tool tool assistant
未配对的 tool_use:无
执行工具总次数:2(get_weather, get_fx_rate)
结束方式:第 2 轮模型不再发 tool_use,循环正常结束
最终一句话:北京今天晴,24 度;美元兑人民币 7.12。
含天气=true 含汇率=true
和两种只取第一个的写法摆在一起:
== 两种写法对照 ==
取全部 :2 轮,执行 2 次工具,未配对 0,最终含天气=true 含汇率=true
只取第一个(追问桩):6 轮,执行 6 次工具,未配对 1,最终含天气=false 含汇率=false
只取第一个(跳过桩):2 轮,执行 1 次工具,未配对 1,最终含天气=true 含汇率=false
后面两行是同一个循环代码,差别只在桩模型怎么应对缺的返回:轮数一个 2 一个 6,最终答案都缺一块。第一行两个数都对上了。
改在取块的那一行
取块那行改成拿全部:
const calls = msg.content.filter(b => b.type === 'tool_use');
for (const call of calls) {
const result = runTool(call);
messages.push({ role: 'tool', content: [{ type: 'tool_result', tool_use_id: call.id, content: result }] });
}
每个 tool_use 都要落到一条 tool_result,id 从调用块上取,顺序无所谓,一条都不能少。如果这一轮本来就只打算跑一个,得把消息里剩下的 tool_use 块删干净再进历史,别让它们悬在那儿。
一条 assistant 消息里 tool_use 的个数不固定,可能是零个、一个,也可能像这一轮一样好几个。取块那行按「这个块在不在」写成单数,就会在多块的那一轮漏掉后面的。
取舍上我不做「只跑第一个、其余略过」的降级:静默丢一个调用,比直接报错难查得多。报错会停在出问题的那一步,静默丢块只会在几轮之后交一句缺斤少两的答案,中间没有一处提醒。宁可循环里多一层循环,也不要一份对不上的历史。