Table of contents
Open Table of contents
- 1. Tool Call 不是 Execution
- 2. Tool Runtime 在 Harness 中的位置
- 3. Tool Registry:能力目录
- 4. Tool Schema:不是越复杂越好
- 5. Tool Schema 的常见坏味道
- 6. Validation:模型输出必须被当成不可信输入
- 7. Unknown Tool:也应该成为 Observation
- 8. Effect Metadata:Tool 不能只按名字治理
- 9. Effect Metadata 为什么重要
- 10. Permission:Tool 能调用,不等于能执行
- 11. Scheduler:不要默认所有 Tool Call 都 asyncio.gather
- 12. Tool 并发的基本原则
- 13. Resource Conflict
- 14. 推荐 Scheduler 思维
- 15. Idempotency:是否允许自动 Retry 的关键
- 16. 幂等性示例
- 17. Idempotency Key
- 18. Execution Journal:避免“不知道执行没执行”
- 19. Execution Runtime:真正执行发生在哪里
- 20. 为什么要隔离 Agent 主进程
- 21. Timeout、Cancellation、Failure 必须区分
- 22. Cancellation 必须向下传播
- 23. Process Cancellation 的暗坑
- 24. Error Taxonomy:所有错误不能只有 Exception
- 25. Error Result 应该告诉模型什么
- 26. Tool Error 与 Engine Error 再次分离
- 27. Result Normalization:Tool 输出不能直接裸返回
- 28. 为什么需要 summary 与 output 分开
- 29. Raw Result、Normalized Result、Context Projection 三层分离
- 30. Artifact:大结果不应该全部塞进 ToolResult 文本
- 31. Tool-specific Output Policy
- 32. Tool Output 与 Context 的边界
- 33. Tool Versioning
- 34. Tool Capability Discovery 与 Dynamic Exposure
- 35. Tool Name 与语义应稳定
- 36. Tool Retry Policy
- 37. Retry Budget
- 38. Backoff 属于 Runtime,不应该交给模型
- 39. Runtime Recovery vs Model Recovery
- 40. Tool Hooks 与生命周期
- 41. Tool Runtime Trace
- 42. Tool Runtime Metrics
- 43. Tool Runtime Failure Taxonomy
- 44. Tool Runtime 与 MCP 的关系
- 45. Tool Runtime 与 Sandbox 的关系
- 46. Tool Runtime 与 Editing Runtime 的关系
- 47. 推荐工业级 Tool Runtime 架构
- 48. 推荐 Tool Runtime 伪代码
- 49. 推荐实践实验
- 50. 本章必答的 22 个问题
- 51. 与下一篇 Editing Runtime 的衔接
- 52. 一句话总纲
1. Tool Call 不是 Execution
模型返回:
{
"name": "bash",
"arguments": {
"command": "pytest tests/test_auth.py"
}
}
这只代表:
模型希望执行这个动作。
它并不代表:
Tool 存在
参数正确
权限允许
执行安全
可以并发
环境可用
执行成功
结果适合直接进入 Context
因此必须经过独立的 Runtime Pipeline。
flowchart TD
A["Model Tool Call"]
B["Tool Lookup"]
C["Schema Validation"]
D["Semantic / Policy Validation"]
E["Permission Gate"]
F["Effect Classification"]
G["Scheduler"]
H["Execution Runtime"]
I["Sandbox / Resource Boundary"]
J["Raw Result"]
K["Result Normalization"]
L["Retention / Context Projection"]
M["Tool Observation"]
A --> B
B --> C
C --> D
D --> E
E --> F
F --> G
G --> H
H --> I
I --> J
J --> K
K --> L
L --> MTool Call 是 Intent,Tool Runtime 才负责把 Intent 变成 Effect。
2. Tool Runtime 在 Harness 中的位置
从整个 Harness 看:
Model
↓
Agent Loop
↓
Tool Call
↓
Tool Runtime
↓
Environment
↓
Tool Result
↓
Context Manager
↓
Model
因此 Tool Runtime 是:
模型世界与真实执行环境之间的边界层。
它连接两个性质完全不同的世界:
| 模型侧 | 执行侧 |
|---|---|
| 概率输出 | 确定执行 |
| JSON / Structured Call | OS / API / FS / Process |
| 可能幻觉 Tool | 实际注册 Tool |
| 参数可能错误 | 参数必须合法 |
| 不理解真实权限 | 必须遵循 Policy |
| 可以随意重试 | 副作用可能不可重复 |
| 输出只关心语义 | Runtime 必须记录事实 |
3. Tool Registry:能力目录
Tool Runtime 首先需要一个明确的 Tool Registry。
例如:
class ToolRegistry:
def register(self, tool): ...
def get(self, name): ...
def list_tools(self): ...
def get_schema(self, name): ...
Tool Registry 不应该只保存函数引用。
推荐 Tool Metadata:
@dataclass
class ToolDefinition:
name: str
description: str
input_schema: dict
effect_type: str
timeout_seconds: float
requires_approval: bool
sandbox_profile: str | None
idempotency: str
concurrency_group: str | None
这使 Runtime 能在真正调用 Tool 前做治理决策。
4. Tool Schema:不是越复杂越好
Tool Schema 解决:
模型应该以什么结构表达动作。
例如:
{
"name": "read_file",
"description": "Read a range of lines from a file.",
"parameters": {
"path": "string",
"start_line": "integer",
"end_line": "integer"
}
}
一个好的 Schema 应该:
语义清晰
字段少而明确
避免互斥参数组合
默认值可预测
错误容易修正
约束显式
5. Tool Schema 的常见坏味道
5.1 一个 Tool 承担太多行为
例如:
file_operation(
action = read/write/delete/move/copy/search/patch
)
模型需要先选择:
Tool
再选择:
Action
错误空间更大。
通常:
read_file
edit_file
delete_file
更容易治理。
5.2 参数过度自由
例如:
bash(command: str)
虽然强大,但治理困难。
因此通常需要:
Permission
Sandbox
Command Policy
Timeout
Output Limit
作为补偿。
5.3 返回结果结构不稳定
同一个 Tool 有时返回:
string
有时:
dict
有时:
None
会让 Runtime 和 Context Projection 复杂化。
应统一进入:
ToolResult
6. Validation:模型输出必须被当成不可信输入
推荐至少两层校验。
6.1 Schema Validation
检查:
字段是否存在
类型是否正确
枚举是否合法
必填字段
长度
数值范围
例如:
validation_error = schema.validate(call.arguments)
if validation_error:
return ToolResult(
status="error",
error_type="invalid_arguments",
output=str(validation_error),
)
6.2 Semantic Validation
Schema 合法,不代表语义合理。
例如:
read_file(path="/etc/passwd")
Schema 完全正确,但可能违反 Workspace Policy。
又例如:
start_line = 1000
end_line = 10
类型正确,但语义错误。
因此还要检查:
Path Policy
Workspace Boundary
Argument Relationship
Resource Existence
Command Policy
Business Constraint
7. Unknown Tool:也应该成为 Observation
模型可能产生:
tool = "read_directory_recursive"
但 Runtime 没有这个 Tool。
错误做法:
raise KeyError
正确做法:
ToolResult
status = ERROR
error_type = UNKNOWN_TOOL
available_tools = [...]
让模型能够调整策略。
例如:
Unknown tool 'read_directory_recursive'.
Available alternatives: glob, grep, read_file.
Tool Runtime 的原则不是“避免一切错误”,而是把可恢复错误转成高质量 Observation。
8. Effect Metadata:Tool 不能只按名字治理
Runtime 必须知道:
这个 Tool 执行后会产生什么副作用?
推荐分类:
PURE
READ_ONLY
LOCAL_WRITE
PROCESS_EXECUTION
NETWORK_READ
NETWORK_WRITE
EXTERNAL_SIDE_EFFECT
DESTRUCTIVE
例如:
| Tool | Effect |
|---|---|
read_file | READ_ONLY |
grep | READ_ONLY |
edit_file | LOCAL_WRITE |
bash | PROCESS_EXECUTION |
web_search | NETWORK_READ |
send_email | EXTERNAL_SIDE_EFFECT |
delete_resource | DESTRUCTIVE |
9. Effect Metadata 为什么重要
它会影响:
Permission
Concurrency
Retry
Sandbox
Audit
Idempotency
因此 Tool Scheduler 不应该只知道:
call.name
而应该知道:
call.effect_type
10. Permission:Tool 能调用,不等于能执行
推荐区分:
Capability
↓
Policy
↓
Approval
↓
Execution
模型看到 Tool Schema,只说明:
Agent 拥有这种 Capability。
真正执行前还必须经过 Permission Engine。
flowchart TD
A["Validated Tool Call"]
B["Policy Evaluation"]
C{"Auto Allow?"}
D["Execute"]
E{"Need Human Approval?"}
F["Request Approval"]
G{"Approved?"}
H["Denied ToolResult"]
A --> B
B --> C
C -- "Yes" --> D
C -- "No" --> E
E -- "Yes" --> F
F --> G
G -- "Yes" --> D
G -- "No" --> H
E -- "Policy Deny" --> HPermission Denied 不是 Agent Crash,而是:
Observation
模型可以选择其他方案。
11. Scheduler:不要默认所有 Tool Call 都 asyncio.gather
模型一次可能返回:
read_file(A)
read_file(B)
grep(C)
edit_file(A)
教学实现可能:
await asyncio.gather(...)
但这很危险。
Runtime 需要判断:
是否存在数据依赖?
是否操作同一个 Resource?
是否有写冲突?
是否有外部副作用?
12. Tool 并发的基本原则
可以并发:
read_file(A)
read_file(B)
通常可以并发:
grep(pattern1)
grep(pattern2)
不应该盲目并发:
edit_file(A)
edit_file(A)
不应该盲目并发:
git checkout branch
edit_file(...)
也不能随意并发:
create_resource
delete_resource
13. Resource Conflict
一个更好的 Scheduler 可以基于:
read_set
write_set
external_effect
例如:
ToolCall 1
read_set = {A}
write_set = {}
ToolCall 2
read_set = {B}
write_set = {}
ToolCall 3
read_set = {A}
write_set = {A}
可以判断:
1 和 2 可并发
1 和 3 存在冲突
14. 推荐 Scheduler 思维
flowchart TD
A["Validated Tool Calls"]
B["Classify Effects"]
C["Extract Resource Sets"]
D["Build Dependency / Conflict Graph"]
E["Create Execution Batches"]
F["Parallel Read-only Batch"]
G["Serialized Write / Side-effect Batch"]
H["Collect Results"]
A --> B
B --> C
C --> D
D --> E
E --> F
E --> G
F --> H
G --> H不一定第一次就实现完整依赖图,但要建立:
并发是 Runtime 决策,而不是模型输出 Tool Calls 后的默认行为。
15. Idempotency:是否允许自动 Retry 的关键
考虑:
Tool 执行成功
↓
Runtime 在收到结果前崩溃
↓
恢复
能不能重新执行?
取决于 Tool 的幂等性。
推荐分类:
PURE
IDEMPOTENT
CONDITIONALLY_IDEMPOTENT
NON_IDEMPOTENT
16. 幂等性示例
PURE
read_file
grep
多次执行通常不会产生副作用。
IDEMPOTENT
例如:
set_config(key="x", value="1")
重复执行结果相同。
CONDITIONALLY_IDEMPOTENT
例如:
edit_file
如果依赖:
expected file hash
则可安全重放。
如果盲目按字符串替换,可能不幂等。
NON_IDEMPOTENT
例如:
send_email
create_order
append_comment
create_issue
重复执行会产生重复副作用。
17. Idempotency Key
对于外部副作用,可以引入:
idempotency_key
operation_id
例如:
send_email(
...,
idempotency_key="session123-call456"
)
如果底层服务支持,就可以防止重复提交。
18. Execution Journal:避免“不知道执行没执行”
上一章 Session 中提到:
ToolExecutionStarted
↓
Crash
↓
No ToolExecutionCompleted
Tool Runtime 应维护执行日志:
operation_id
call_id
tool
arguments_hash
started_at
finished_at
status
external_reference
恢复时:
查 Journal
↓
判断是否已经产生副作用
↓
决定 Retry / Reconcile / Ask Human
Tool Retry 的前提不是“上次没有 Result”,而是“我们确定重试是安全的”。
19. Execution Runtime:真正执行发生在哪里
Tool Runtime 可以有不同执行后端:
In-process Function
Subprocess
Sandbox Process
Container
Remote Worker
External API
Browser Runtime
MCP Server
因此推荐:
Tool Definition
↓
Execution Runtime Adapter
而不是所有 Tool 都直接运行在 Agent 主进程。
20. 为什么要隔离 Agent 主进程
如果 Tool 直接:
subprocess.run(...)
在主进程环境执行,一旦发生:
死循环
OOM
文件破坏
环境污染
进程树泄漏
可能拖垮整个 Agent。
因此执行层应逐步支持:
Timeout
Cancellation
Process Group
Working Directory
Environment Filtering
Resource Limit
Sandbox
21. Timeout、Cancellation、Failure 必须区分
这三种状态语义不同。
Timeout
系统执行预算耗尽
例如:
command > 30s
Cancellation
外部主动要求停止
例如:
User Stop
Steering
Session Suspend
Parent Task Cancel
Failure
工具自己执行完成,但结果失败
例如:
pytest exit code 1
推荐显式状态:
SUCCESS
ERROR
TIMEOUT
CANCELLED
DENIED
SANDBOX_VIOLATION
UNKNOWN_OUTCOME
22. Cancellation 必须向下传播
用户点击 Stop 时不能只:
停止下一轮模型调用
如果当前 Tool 还在运行:
pytest
npm install
browser task
subprocess
它必须收到取消信号。
推荐:
flowchart TD
A["Cancellation Requested"]
B["Mark Session Cancelling"]
C["Cancel Pending Tool Calls"]
D["Signal Running Runtime"]
E["Terminate / Kill Process Group"]
F["Persist Cancel Events"]
G["Return CANCELLED Observation"]
H["Session Safe Stop"]
A --> B
B --> C
C --> D
D --> E
E --> F
F --> G
G --> H23. Process Cancellation 的暗坑
如果只杀:
parent process
可能留下:
child process
grandchild process
server
watcher
因此执行层通常需要考虑:
Process Group
Job Object
Container
Sandbox Runtime
确保取消能覆盖整个执行树。
24. Error Taxonomy:所有错误不能只有 Exception
推荐至少分类:
UNKNOWN_TOOL
INVALID_ARGUMENTS
SEMANTIC_VALIDATION_ERROR
PERMISSION_DENIED
TIMEOUT
CANCELLED
SANDBOX_VIOLATION
PROCESS_ERROR
NETWORK_ERROR
RESOURCE_NOT_FOUND
OUTPUT_LIMIT
RUNTIME_EXCEPTION
UNKNOWN_OUTCOME
为什么重要?
因为模型对不同错误的恢复策略不同。
25. Error Result 应该告诉模型什么
例如 File Not Found:
差:
Error.
好:
RESOURCE_NOT_FOUND
Tool: read_file
Path: src/foo.py
The requested file does not exist.
Consider using glob or grep to locate the correct path.
例如 Permission Denied:
PERMISSION_DENIED
The operation was blocked by policy.
Do not repeat the same command unchanged.
Choose a read-only or less destructive alternative.
26. Tool Error 与 Engine Error 再次分离
Tool Error:
pytest fails
file missing
permission denied
timeout
应该:
ToolResult
↓
Context
↓
Model
Engine Error:
Tool Registry corrupt
Session Store unavailable
Scheduler internal invariant broken
Execution Runtime unavailable
才可能:
Retry Runtime
Suspend Session
Fatal Error
27. Result Normalization:Tool 输出不能直接裸返回
不同 Tool 的原始输出:
str
bytes
dict
process result
HTTP response
file handle
exception
都应该进入统一模型。
推荐:
@dataclass
class ToolResult:
call_id: str
status: str
summary: str
output: str | None
error_type: str | None = None
exit_code: int | None = None
duration_ms: int | None = None
truncated: bool = False
metadata: dict = field(default_factory=dict)
28. 为什么需要 summary 与 output 分开
例如 Shell 输出 50K:
summary:
pytest failed: 3 failed, 81 passed
output:
[truncated detailed logs]
ContextManager 可以优先保留:
summary
必要时再读取:
full artifact
这是 Tool Runtime 与 Context Management 的重要接口。
29. Raw Result、Normalized Result、Context Projection 三层分离
推荐:
flowchart TD
A["Raw Tool Result<br/>完整执行事实"]
B["Normalized Tool Result<br/>统一状态 / 元数据 / Artifact"]
C["Context Projection<br/>适合模型阅读的摘要"]
D["Model Observation"]
A --> B
B --> C
C --> D这三层不要混在一起。
原因:
Raw Result 用于审计
Normalized Result 用于 Runtime
Context Projection 用于推理
30. Artifact:大结果不应该全部塞进 ToolResult 文本
例如:
大型 Diff
完整 Build Log
网页 HTML
PDF
数据库导出
二进制文件
应该存为 Artifact:
artifact_id
type
size
location
hash
ToolResult 只返回:
摘要
Artifact Reference
如何继续读取
例如:
Build output stored as artifact build-log-123.
Summary: 4 compiler errors.
Use read_artifact_range(...) for details.
31. Tool-specific Output Policy
不同工具需要不同 Normalizer。
例如:
Bash
保留:
command
exit_code
stderr / stdout summary
head-tail output
Tests
保留:
passed
failed
failed test names
key traceback
grep
保留:
top matches
file count
match count
read_file
保留:
line range
file version/hash
content
truncation marker
git diff
保留:
changed files
hunks
stats
因此架构可以:
Tool
↓
ResultFormatter
↓
Artifact Policy
↓
Context Projection
32. Tool Output 与 Context 的边界
Tool Runtime 应该:
尽量保留完整执行事实
ContextManager 应该:
选择模型下一步真正需要的部分
所以不要在 Runtime 层为了节省 Token:
直接永久丢掉完整日志
更好的方式:
Raw Result → Artifact Store
Normalized Result → Session
Projection → Context
33. Tool Versioning
长期 Harness 中 Tool 会升级。
例如:
edit_file v1
edit_file v2
Schema 或行为可能不同。
因此 Tool Call / Event 最好记录:
tool_name
tool_version
schema_version
否则 Replay 时:
旧 Session
可能被新 Tool 行为错误解释。
34. Tool Capability Discovery 与 Dynamic Exposure
上一章提到 Tool Schema 本身也占 Context。
因此:
Registry 有 100 个 Tool
不代表每轮都要暴露 100 个。
可以:
Global Registry
↓
Capability Selection
↓
Current Tool Set
↓
Model
例如 Coding 阶段只暴露:
read
grep
edit
bash
lsp
部署阶段再暴露:
deploy
cloud
release
35. Tool Name 与语义应稳定
如果同一功能今天叫:
read_file
明天变成:
filesystem_read
会影响:
Prompt Cache
模型 Tool 选择习惯
历史 Trace
Eval
所以生产 Tool API 应像真正 API 一样:
重视兼容性。
36. Tool Retry Policy
不要统一:
retry_on_exception = True
应该按错误和 Effect 决策。
例如:
| 情况 | 自动 Retry |
|---|---|
| Read-only + Network Timeout | 通常可以 |
| Read-only + 5xx | 可以有限重试 |
| Invalid Arguments | 不应该,交给模型修正 |
| Permission Denied | 不应该重复 |
| Non-idempotent External Write | 默认不应自动重试 |
| Tool Process Crash | 视幂等性 |
| Rate Limit | 可退避重试 |
37. Retry Budget
所有自动 Retry 都应该有预算:
max_attempts
max_elapsed_time
backoff
jitter
并记录:
attempt
last_error
next_retry
避免 Runtime 自己产生死循环。
38. Backoff 属于 Runtime,不应该交给模型
如果模型 API 或 Tool Backend 返回:
429
503
Temporary Network Error
Harness 可以在 Runtime 内:
retry with backoff
不需要每次都把暂态错误交给模型。
但如果:
重试预算耗尽
再转成 Observation 或 Runtime Failure。
这就是:
Runtime 自愈和 Model 自愈的边界。
39. Runtime Recovery vs Model Recovery
Runtime Recovery
适合:
网络抖动
暂态 5xx
进程启动失败一次
Rate Limit
目标:
保持语义不变,透明恢复
Model Recovery
适合:
参数错误
文件不存在
测试失败
权限拒绝
命令本身不成立
目标:
模型根据新 Observation 改变策略
不要把所有失败都交给模型“反思”,也不要把所有失败都 Runtime 自动重试。
40. Tool Hooks 与生命周期
Tool Runtime 很适合提供生命周期钩子:
BeforeValidate
AfterValidate
BeforePermission
BeforeExecute
AfterExecute
OnError
BeforeProject
AfterProject
用途:
Policy
Audit
Metrics
Secret Redaction
Formatter
Lint
Custom Guard
但 Hook 不应该破坏核心不变量:
不能绕过 Permission
不能伪造 Execution Result
不能静默吞掉关键错误
41. Tool Runtime Trace
一次 Tool Call 至少应该能追踪:
call_id
tool_name
tool_version
arguments_hash
effect_type
permission_decision
scheduler_batch
runtime
sandbox_profile
start_time
end_time
status
retry_count
exit_code
output_size
artifact_ids
error_type
这样才能回答:
为什么这个 Tool 慢?
为什么它被拒绝?
为什么它执行了两次?
为什么 Agent 认为它失败?
42. Tool Runtime Metrics
推荐:
| 指标 | 作用 |
|---|---|
| Tool Success Rate | 基础可靠性 |
| Validation Failure Rate | Schema 设计质量 |
| Permission Denied Rate | 权限策略影响 |
| Timeout Rate | 执行预算合理性 |
| Cancellation Latency | Stop 响应能力 |
| Retry Rate | Backend 稳定性 |
| Duplicate Execution Rate | 幂等问题 |
| Avg Tool Latency | Tool 性能 |
| Output Size | Context 压力 |
| Truncation Rate | Result Policy |
| Recovery Rate | Tool 错误后 Agent 是否恢复 |
43. Tool Runtime Failure Taxonomy
为了 Eval,建议进一步区分:
MODEL_TOOL_SELECTION_ERROR
UNKNOWN_TOOL
ARGUMENT_SCHEMA_ERROR
SEMANTIC_ARGUMENT_ERROR
POLICY_DENIED
USER_DENIED
SCHEDULING_CONFLICT
EXECUTION_TIMEOUT
EXECUTION_CANCELLED
SANDBOX_VIOLATION
TOOL_PROCESS_ERROR
EXTERNAL_API_ERROR
UNKNOWN_OUTCOME
RESULT_PARSE_ERROR
OUTPUT_OVERFLOW
这样才能知道:
问题出在模型,还是 Tool Runtime。
44. Tool Runtime 与 MCP 的关系
MCP 可以标准化:
Tool Discovery
Schema
Invocation
Result
但 MCP 不会自动解决:
Permission
Sandbox
Retry
Idempotency
Concurrency Conflict
Session Journal
Context Retention
因此:
MCP 是 Tool 接入协议,不是完整 Tool Runtime。
Harness 仍然需要自己的治理层。
45. Tool Runtime 与 Sandbox 的关系
Tool Runtime 决定:
是否执行
如何调度
如何重试
如何观察
Sandbox 决定:
执行时能访问什么资源
因此:
Permission
≠
Tool Runtime
≠
Sandbox
三者职责不同。
46. Tool Runtime 与 Editing Runtime 的关系
Coding Agent 中:
read
grep
edit
patch
bash
lsp
test
都是 Tool。
但 Editing Runtime 还需要解决更高层问题:
Read-before-Edit
Patch Correctness
File Version
Diff
Conflict
LSP Validation
Test Feedback
所以:
下一篇 Editing Runtime 是 Tool Runtime 在 Coding Agent 场景下的专业化实现。
47. 推荐工业级 Tool Runtime 架构
flowchart TD
A["Agent Loop"]
B["Tool Registry"]
C["Validator"]
D["Policy / Permission"]
E["Effect Analyzer"]
F["Scheduler"]
G["Execution Journal"]
H["Execution Runtime"]
I["Sandbox"]
J["Result Normalizer"]
K["Artifact Store"]
L["Session Event Log"]
M["Context Projection"]
N["Model Observation"]
A --> B
B --> C
C --> D
D --> E
E --> F
F --> G
G --> H
H --> I
I --> J
J --> K
J --> L
K --> M
L --> M
M --> N这张图可以作为本章最终架构模型。
48. 推荐 Tool Runtime 伪代码
class ToolRuntime:
def __init__(
self,
registry,
validator,
permission_engine,
scheduler,
execution_backend,
journal,
artifact_store,
):
self.registry = registry
self.validator = validator
self.permission = permission_engine
self.scheduler = scheduler
self.execution = execution_backend
self.journal = journal
self.artifacts = artifact_store
async def execute_calls(self, calls, session):
planned = []
for call in calls:
tool = self.registry.get(call.name)
if tool is None:
planned.append(
self.error_result(
call,
"UNKNOWN_TOOL",
f"Unknown tool: {call.name}",
)
)
continue
validation = self.validator.validate(tool, call.arguments)
if not validation.ok:
planned.append(
self.error_result(
call,
"INVALID_ARGUMENTS",
validation.message,
)
)
continue
permission = await self.permission.evaluate(
tool=tool,
arguments=call.arguments,
session=session,
)
if permission.denied:
planned.append(
self.error_result(
call,
"PERMISSION_DENIED",
permission.reason,
)
)
continue
planned.append(
ExecutableCall(
call=call,
tool=tool,
effect=tool.effect_metadata,
permission=permission,
)
)
plan = self.scheduler.plan(planned)
results = []
for batch in plan.batches:
batch_results = await self._run_batch(batch, session)
results.extend(batch_results)
return results
async def _run_batch(self, batch, session):
return await asyncio.gather(
*[
self._execute_one(item, session)
for item in batch
]
)
async def _execute_one(self, item, session):
operation_id = self.journal.create_operation(
call=item.call,
tool=item.tool,
)
await session.append_event(
"ToolExecutionStarted",
{
"operation_id": operation_id,
"call_id": item.call.call_id,
"tool": item.tool.name,
},
)
try:
raw = await self.execution.run(
tool=item.tool,
arguments=item.call.arguments,
timeout=item.tool.timeout_seconds,
sandbox_profile=item.tool.sandbox_profile,
)
except asyncio.TimeoutError:
result = ToolResult(
call_id=item.call.call_id,
status="TIMEOUT",
error_type="EXECUTION_TIMEOUT",
summary="Tool exceeded execution time budget.",
)
except asyncio.CancelledError:
result = ToolResult(
call_id=item.call.call_id,
status="CANCELLED",
error_type="EXECUTION_CANCELLED",
summary="Tool execution was cancelled.",
)
except Exception as exc:
result = ToolResult(
call_id=item.call.call_id,
status="ERROR",
error_type=type(exc).__name__,
summary=str(exc),
)
else:
result = item.tool.result_normalizer.normalize(raw)
self.journal.finish_operation(
operation_id=operation_id,
result=result,
)
await session.append_event(
"ToolExecutionCompleted",
{
"operation_id": operation_id,
"call_id": item.call.call_id,
"status": result.status,
},
)
return result
这仍然是教学伪代码。
实际系统还需要处理:
Pending Approval
Unknown Outcome
Retry Policy
Remote Worker
Artifact Persistence
Process Tree Cancellation
Secret Redaction
Partial Result49. 推荐实践实验
实验 1:Unknown Tool 自愈
让模型调用不存在的:
find_symbol
Runtime 返回:
UNKNOWN_TOOL
并给出:
grep
lsp
作为替代。
观察模型能否调整。
实验 2:Schema Validation
故意生成:
read_file(
start_line=100,
end_line=10
)
实现:
Semantic Validation
而不是执行后报错。
实验 3:并发读
一次让模型:
read A
read B
read C
对比:
串行
vs
并发
的 Wall Time。
实验 4:写冲突
同时提交:
edit A
edit A
Scheduler 必须:
串行
或拒绝冲突。
实验 5:Timeout
创建:
sleep 60
Tool timeout:
5s
验证:
TIMEOUT
是否能被模型正确理解。
实验 6:Cancellation
运行长命令。
用户点击 Stop。
确认:
父进程
子进程
全部停止。
实验 7:Crash + Idempotency
模拟:
ToolExecutionStarted
↓
实际 Side Effect 成功
↓
Process Crash
重启后验证:
不会盲目重复执行
实验 8:Artifact
让 Tool 生成:
100K build log
完整内容保存 Artifact。
Context 只注入:
Summary + Artifact Ref
50. 本章必答的 22 个问题
- 为什么 Tool Call 不等于 Tool Execution?
- Tool Registry 为什么不应该只存函数引用?
- Tool Schema 的主要设计目标是什么?
- Schema Validation 与 Semantic Validation 有什么区别?
- Unknown Tool 为什么应该成为 Observation?
- Effect Metadata 解决什么问题?
- 为什么 Permission 与 Capability 要分离?
- 为什么 Tool Calls 不能默认全部并发?
- read_set / write_set 有什么价值?
- Tool 幂等性为什么决定 Retry 策略?
- 什么是 Unknown Outcome?
- Execution Journal 有什么价值?
- Timeout、Cancellation、Failure 有什么区别?
- 为什么 Cancellation 必须向子进程传播?
- Tool Error 与 Engine Error 如何区分?
- Raw Result、Normalized Result、Context Projection 有什么区别?
- 为什么大结果应该进入 Artifact Store?
- Tool-specific Result Policy 为什么优于统一字符串截断?
- Runtime Recovery 与 Model Recovery 如何分工?
- 为什么所有 Retry 都必须有 Budget?
- MCP 为什么不等于 Tool Runtime?
- Tool Runtime 与 Sandbox 的职责边界是什么?
51. 与下一篇 Editing Runtime 的衔接
现在已经建立:
Model Tool Call
↓
Validation
↓
Permission
↓
Scheduling
↓
Execution
↓
Result
但 Coding Agent 还有一个更加专业的问题:
“修改代码”并不是一个普通 Tool Call。
它涉及:
文件版本
定位
Read-before-Edit
Patch
Diff
冲突
LSP
Build
Tests
回滚
所以下一篇:
Coding Agent Editing Runtime:Read、Search、Edit、Diff、LSP 与 Test
会重点研究:
Repo Navigation
Read / Search Strategy
Patch vs Overwrite
File Version / Preconditions
Diff-based Editing
LSP Feedback
Build / Test Loop
Edit Conflict
Verification
这也是从“通用 Agent Runtime”正式进入“Coding Agent 专业 Runtime”的关键一步。
52. 一句话总纲
工业级 Tool Runtime 的本质,是把模型产生的概率性动作意图转换成受 Schema、Policy、Permission、Effect、Concurrency、Idempotency、Timeout、Cancellation 与 Sandbox 约束的确定执行,并把完整执行事实归一化成模型可以继续推理的 Observation;它是 Agent 从“会思考”跨越到“可靠行动”的核心边界层。