Skip to content
Figo Blogs
Go back

Coding Agent Editing Runtime:Read、Search、Edit、Diff、LSP 与 Test

Contents

Table of contents

Open Table of contents

1. 为什么 Coding Agent 需要独立的 Editing Runtime

上一章已经建立通用 Tool Runtime:

Tool Call
↓
Validation
↓
Permission
↓
Scheduling
↓
Execution
↓
Result

但代码修改存在额外约束:

文件可能已变化
定位可能错误
Patch 可能不再适用
代码语法可能通过但语义错误
修改可能影响其他文件
同一 Symbol 可能有多个定义
测试可能失败
Lint / Type Check 可能失败
模型可能覆盖用户已有修改

因此:

代码编辑不是“字符串写入”,而是一种带前置条件、版本约束、语义反馈与验证闭环的状态变更。


2. Coding Agent 的核心闭环

这个闭环比:

Prompt
↓
Generate Code

重要得多。


3. Repository Navigation:先找到正确的位置

Coding Agent 第一大能力不是写代码,而是:

知道应该看哪里。

大型仓库中可能有:

数千文件
多语言
生成代码
测试目录
vendor
build artifacts
monorepo
多个同名 symbol

因此不能:

递归把整个 Repo 塞给模型

而应采用渐进式导航。


4. 推荐 Repo Navigation 层级

Repository Metadata
↓
Directory / Project Structure
↓
Search
↓
Symbol / LSP
↓
Targeted Read
↓
Dependency Expansion

例如:

任务:修复 login timeout

不要一开始全仓库读。

更合理:

grep "timeout"
↓
找到 auth / session 相关文件
↓
LSP 查 definition / references
↓
读取相关函数
↓
再根据调用关系扩展 Working Set

5. Search First,而不是 Read Everything

推荐 Tool 层级:

glob / file tree
grep / ripgrep
symbol search
definition / references
read_file(range)

其核心思想是:

先缩小候选空间,再支付 Context Token。


6. Text Search 与 Semantic Search 的组合

纯 grep 能找到:

字符串
函数名
错误信息
配置项

但不能准确回答:

这个 symbol 的定义在哪?
谁引用了它?
调用链是什么?
类型是什么?

因此 Coding Agent 通常需要:

Text Search
+
LSP / AST / Index

7. LSP 在 Editing Runtime 中的位置

LSP 不是独立“锦上添花”的 Tool。

它参与整个编辑闭环:

它解决:

Definition
References
Diagnostics
Symbols
Type Information
Rename

8. Read-before-Edit:为什么必须先读

一个非常重要的安全不变量:

Agent 不应该修改自己尚未读取或验证过的文件区域。

原因:

文件可能与模型假设不同
用户可能已经修改
分支可能变化
旧 Context 可能失效
相同字符串可能出现多次

因此 Edit Tool 可以要求:

先 Read
↓
获得 file_version / hash
↓
再 Edit

9. File Version / Preconditions

推荐给每次读取返回:

path
content
line range
file hash / version

编辑时提交:

expected_hash

例如:

edit_file(
    path="src/auth.ts",
    expected_hash="abc123",
    patch=...
)

Runtime 执行前:

Current Hash == Expected Hash ?

如果不一致:

PRECONDITION_FAILED

而不是直接覆盖。


10. 为什么 Precondition 很关键

考虑:

Agent 读取 auth.ts
↓
用户手工修改 auth.ts
↓
Agent 根据旧内容提交 Edit

如果没有版本检查:

用户修改可能被覆盖

有了 Precondition:

Edit rejected
↓
Agent re-read
↓
re-plan

这就是 Optimistic Concurrency Control 的思路。


11. Patch vs Overwrite

Coding Agent 常见两种修改方式:

11.1 Full Overwrite

write_file(path, full_content)

优点:

简单
模型容易理解

缺点:

容易误删
大文件 Token 成本高
会覆盖未知改动
Diff 不透明
并发冲突大

11.2 Patch / Diff Edit

apply_patch(...)

优点:

修改范围小
可审计
冲突容易检测
更容易保护未读区域

缺点:

Patch 可能应用失败
上下文行变化会导致冲突
格式更复杂

12. 推荐策略:小范围 Patch 优先

通常:

局部修改 → Patch
新增小文件 → Full Write
大规模生成 → Full Write + Diff Review
机械重构 → Structured Edit / AST / LSP

不能简单认为:

Patch 永远比 Overwrite 好

关键是:

修改方式要与变更规模和可验证性匹配。


13. Exact Match Edit 的风险

常见 Tool:

replace(
    old_text="...",
    new_text="..."
)

需要防止:

0 matches
multiple matches
stale match
whitespace mismatch

推荐:

0 Match
→ ERROR

1 Match
→ Apply

>1 Match
→ AMBIGUOUS_MATCH

而不是:

replace all

14. Edit Tool 应返回什么

不是只返回:

Success

而应至少返回:

path
old_hash
new_hash
changed_ranges
diff
lines_added
lines_removed
warnings

这样模型和 Trace 都能理解发生了什么。


15. Diff 是 Editing Runtime 的第一等公民

修改后最重要的 Observation 之一:

git diff

因为模型必须确认:

改了哪些文件
改了哪些行
是否误改
是否存在意外删除

推荐闭环:

Edit
↓
Diff
↓
Inspect
↓
Diagnostics
↓
Tests

16. Diff 不应该无限大

大型 Diff 也会打爆 Context。

推荐:

Changed File List
↓
Diff Stats
↓
Relevant Hunks
↓
On-demand Full Diff

例如:

3 files changed
+42 -17

src/auth.ts
@@ function validateToken ...

必要时才继续展开。


17. File-aware Diff Projection

不要对整个:

git diff

统一 Head-Tail。

更合理:

按文件
↓
按 Hunk
↓
按当前任务相关性

这样不会把:

最重要中间文件

因为字符串截断丢掉。


18. Editing Runtime 与 Working Set

上一章 Context Management 中的 Working Set,在 Coding Agent 中尤其重要。

推荐维护:

active_files
active_symbols
modified_files
diagnostics
failed_tests
dependency_edges

Working Set 更新:

Read → active
Edit → high priority
Diagnostic → high priority
Test failure reference → active
任务完成 → decay

19. 修改后必须让旧 Context 失效

如果:

read auth.ts v1
↓
edit auth.ts → v2

Context 中旧的:

auth.ts v1

必须:

Superseded

否则模型下一轮可能同时看到:

旧版本
新版本

导致推理冲突。

所以 Editing Runtime 与 Context Manager 必须联动。


20. Edit Event 应进入 Session

推荐事件:

FileRead
EditRequested
EditPreconditionChecked
EditApplied
DiffGenerated
DiagnosticsUpdated
TestStarted
TestCompleted

Session 不应该只存:

assistant said it edited file

而应保存真实执行事实。


21. Edit Conflict

并发场景:

Agent A 修改 file.ts
Agent B 修改 file.ts

或者:

Subagent 1
Subagent 2

都可能产生冲突。

需要:

File Lock
Version Check
Merge
Conflict Detection

至少第一版要有:

expected_hash

22. Structured Edit

对于某些修改,字符串 Patch 不是最佳方法。

例如:

Rename Symbol
Add Import
Change Function Signature
Move Symbol

可以使用:

LSP Workspace Edit
AST Transformation
Structured Refactoring Tool

优势:

语义更准确
跨文件一致性更好

23. 什么时候优先 Structured Edit

推荐:

修改工具
小范围逻辑修改Patch
Rename SymbolLSP Rename
Add ImportAST / LSP / Patch
大量机械变换AST
新文件生成Full Write
配置修改Structured Config Tool / Patch

24. Diagnostics:编辑后第一层验证

编辑完成后,不应该直接:

任务完成

至少先做:

Syntax
Type
LSP Diagnostics
Lint

因为这类反馈:

快
成本低
定位准

适合作为第一层验证。


25. 验证层级:从便宜到昂贵

推荐顺序:

不一定所有项目顺序完全一致,但核心原则:

先执行便宜、定位性强的检查,再执行昂贵的全量验证。


26. Targeted Test 优先

如果只改:

auth token parser

不要每轮都:

全仓库 test

优先:

对应 test file
相关 package
相关 module

通过后再扩大范围。

这降低:

Wall Time
Tool Output
Context Noise

27. Test Failure 是 Observation,不是失败终点

典型闭环:

Edit
↓
Test
↓
Failure
↓
Parse Failure
↓
Locate Cause
↓
Read
↓
Edit Again

模型真正的能力来自:

在 Harness 提供的高质量反馈中逐步收敛。


28. Test Result Normalization

不要把:

5000 行 pytest

全部塞给模型。

推荐结构化:

status: failed
passed: 81
failed: 3
failures:
  - test: test_refresh_token
    error_type: AssertionError
    message: expected 401, got 500
    traceback_tail: ...

原始日志:

Artifact Store

Context:

结构化失败摘要

29. Compiler / Build Result 同理

推荐提取:

file
line
column
error_code
message
symbol

而不是只给:

raw stdout

这就是 Tool-specific Result Normalization 在 Coding Agent 中的价值。


30. Failure Localization

如果测试失败,Agent 不应该马上:

重新修改刚才文件

需要先判断:

失败是否由本次变更引入?

推荐参考:

Diff
Diagnostics
Test name
Stack trace
Changed symbols
Dependency graph

31. Regression vs Pre-existing Failure

真实仓库可能本来就有失败测试。

因此最好有:

Baseline

例如修改前:

pytest target

记录:

2 failures

修改后:

2 failures

不能直接认为:

Agent 修改失败

更准确:

No New Regression

32. Baseline Verification

对于高价值任务:

Before Edit
↓
Targeted Baseline
↓
Edit
↓
Targeted Verification
↓
Compare

能大幅降低误判。

代价:

额外执行时间

所以应按任务风险选择。


33. Edit Transaction

对于多文件修改:

A
B
C

可能希望:

要么整组成功
要么回滚

可以抽象成:

Edit Transaction

流程:


34. 是否应该自动回滚

不一定。

如果测试失败:

修改可能仍然部分正确

立即回滚会丢失有价值状态。

因此可以区分:

Syntax-breaking edit
→ 自动回滚

Patch corruption
→ 自动回滚

Semantic test failure
→ 通常保留 Diff,让模型继续修

这比:

任何失败都 rollback

更合理。


35. Dirty Working Tree:必须尊重用户已有修改

Coding Agent 启动时应检查:

git status

如果工作区已有用户修改:

Agent 不能默认它们属于自己

因此需要:

Baseline Snapshot
Agent-owned Changes
User-owned Changes

至少能够:

避免覆盖
区分 Diff

36. Agent-owned Diff

推荐 Session 记录:

initial repo state
files changed by agent
pre-existing modified files

最终汇报时可以:

Agent changes:
- src/a.ts
- tests/a.test.ts

Pre-existing changes untouched:
- README.md

这对安全性非常重要。


37. Git 不是 Editing Runtime 的全部

Git 很有价值:

Diff
Status
Baseline
Rollback
Commit

但不能把 Editing Runtime 简化成:

git commands

因为 Agent 仍需要:

File Preconditions
LSP
Test
Semantic Navigation
Context Supersession

38. Repo State Snapshot

可以定义:

repo_commit
dirty_files
file_hashes
branch

任务开始时记录。

这样 Trace 中能知道:

Agent 是基于哪个仓库状态做出的修改。


39. Generated / Vendor Files

Agent 不应该无脑修改:

node_modules
vendor
dist
build
generated code
lockfiles

除非任务明确需要。

因此 Repo Policy 可以定义:

editable_paths
read_only_paths
generated_paths
ignored_paths

40. File Policy 属于 Editing Runtime 与 Permission 的交界

例如:

src/**
→ editable

generated/**
→ deny edit

infra/prod/**
→ require approval

Editing Runtime 负责:

识别资源

Permission Engine 负责:

是否允许

41. Search Loop 也可能死循环

典型:

grep foo
↓
read file
↓
grep foo
↓
read same file

需要和 Agent Loop 的 Dead-loop Detection 联动。

可以追踪:

search fingerprint
file read fingerprint
working_set change
new information gain

如果:

连续多步没有增加新信息

应触发 Steering。


42. Information Gain

一个高级但很有价值的思想:

每个 Search / Read Tool Call 是否真的增加了新的有效信息?

例如:

read 同一文件同一范围

通常 Information Gain ≈ 0。

可以作为:

Loop Guard

的辅助指标。


43. Editing Plan:要不要先 Plan

简单修改:

单文件明显 Bug

未必需要显式 Plan。

复杂修改:

跨 10 文件 API migration

最好先生成:

Files
Symbols
Expected Changes
Verification Plan

所以:

Planning 是按任务复杂度启用的策略

而不是每次强制长 Plan。


44. Plan 也必须能失效

如果后续发现:

原假设错误
用户 Steering
文件结构不同
测试暴露新问题

旧 Plan 应:

Superseded

不要让模型继续被旧计划绑定。

这再次说明:

Plan 是 Session State
不是永恒 Prompt

45. Editing Runtime 与 Context Manager 如何协作

Editing Runtime
产生:
- FileRead
- Diff
- Diagnostics
- Tests

Context Manager
决定:
- 保留哪些文件片段
- 哪些旧版本失效
- 哪些 Diff Hunk 进入 Context
- 哪些 Test Failure 最重要

两者不能合并,但必须有清晰接口。


46. Editing Runtime 与 Session 如何协作

Session 保存:

FileRead Event
Edit Event
File Version
Diff
Test Result
Diagnostics

这样支持:

Resume
Replay
Audit
Fork
Eval

47. Editing Runtime 与 Tool Runtime 的边界

Tool Runtime 提供通用能力:

Validation
Permission
Scheduling
Timeout
Cancellation
Sandbox
Result Normalization

Editing Runtime 增加领域语义:

File Version
Read-before-Edit
Patch
Diff
LSP
Tests
Repo Policy
Edit Transaction

因此:

Editing Runtime = Tool Runtime + Code Repository Semantics。


48. 推荐 Editing Runtime 架构

这张图可以作为本章最终架构模型。


49. 推荐核心数据结构

FileSnapshot

@dataclass
class FileSnapshot:
    path: str
    content: str
    hash: str
    start_line: int
    end_line: int

EditRequest

@dataclass
class EditRequest:
    path: str
    expected_hash: str
    patch: str

EditResult

@dataclass
class EditResult:
    path: str
    old_hash: str
    new_hash: str
    diff: str
    changed_ranges: list

VerificationResult

@dataclass
class VerificationResult:
    stage: str
    status: str
    diagnostics: list
    artifacts: list

50. 推荐 Editing Loop 伪代码

class CodingEditingRuntime:
    def __init__(
        self,
        navigator,
        file_runtime,
        patch_engine,
        lsp,
        test_runner,
        session,
        context_manager,
    ):
        self.navigator = navigator
        self.files = file_runtime
        self.patch = patch_engine
        self.lsp = lsp
        self.tests = test_runner
        self.session = session
        self.context = context_manager

    async def modify(self, task):
        candidates = await self.navigator.search(task)

        working_set = await self.navigator.build_working_set(
            task=task,
            candidates=candidates,
        )

        snapshots = []

        for target in working_set.files:
            snapshot = await self.files.read(
                path=target.path,
                range=target.range,
            )
            snapshots.append(snapshot)

            await self.session.append_event(
                "FileRead",
                {
                    "path": snapshot.path,
                    "hash": snapshot.hash,
                    "range": [
                        snapshot.start_line,
                        snapshot.end_line,
                    ],
                },
            )

        plan = await self._plan_edit(
            task=task,
            snapshots=snapshots,
        )

        edit_results = []

        for edit in plan.edits:
            current_hash = await self.files.hash(edit.path)

            if current_hash != edit.expected_hash:
                raise PreconditionFailed(
                    f"{edit.path} changed after it was read"
                )

            result = await self.patch.apply(edit)
            edit_results.append(result)

            await self.session.append_event(
                "EditApplied",
                {
                    "path": result.path,
                    "old_hash": result.old_hash,
                    "new_hash": result.new_hash,
                },
            )

        diff = await self._build_diff(edit_results)

        diagnostics = await self.lsp.diagnostics(
            files=[r.path for r in edit_results]
        )

        if diagnostics.has_blocking_errors:
            return VerificationResult(
                stage="lsp",
                status="failed",
                diagnostics=diagnostics.items,
                artifacts=[],
            )

        test_result = await self.tests.run_targeted(
            task=task,
            changed_files=[r.path for r in edit_results],
        )

        return VerificationResult(
            stage="tests",
            status=test_result.status,
            diagnostics=test_result.failures,
            artifacts=test_result.artifacts,
        )
Note

真实 Coding Agent 通常不会由一个固定函数一次性完成全部过程。

更常见的是:

Agent Loop
↓
Search Tool
↓
Read Tool
↓
Edit Tool
↓
Diagnostics Tool
↓
Test Tool
↓
Model Re-plan

上面的伪代码主要用于理解 Editing Runtime 应提供哪些领域能力和不变量。


51. 推荐实践实验

实验 1:Read-before-Edit

要求:

没有 Read 过的文件不能 Edit

观察模型是否自动:

read
↓
edit

实验 2:Stale File

流程:

Agent Read A
↓
外部修改 A
↓
Agent Edit A

要求返回:

PRECONDITION_FAILED

而不是覆盖。


实验 3:Ambiguous Replace

文件中有两个:

foo()

Agent 请求:

replace foo() → bar()

要求:

AMBIGUOUS_MATCH

实验 4:LSP Diagnostics

让 Agent 引入:

Type Error

Edit 后自动返回 LSP Diagnostic。

观察模型能否自愈。


实验 5:Targeted Tests

只修改一个模块。

比较:

全量测试
vs
Targeted Test → Full Test

的时间与反馈质量。


实验 6:Pre-existing Failure

修改前已有:

2 failed

修改后仍:

2 failed

验证系统能否判断:

No New Regression

实验 7:Dirty Working Tree

用户先手动修改:

README.md

Agent 修改其他文件。

最终确认:

README.md 不被覆盖

实验 8:Multi-file Failure

Agent 修改三个文件。

其中一个 Patch 失败。

验证:

Transaction Policy

是否符合预期。


52. Editing Runtime Eval

建议记录:

指标说明
Task Pass@1最终任务成功率
Patch Apply RatePatch 成功率
Precondition Failure RateStale Edit 比例
Ambiguous Edit Rate定位不唯一比例
LSP Error Introduction Rate修改引入静态错误比例
Targeted Test Pass Rate相关验证效果
Regression Rate新增回归
Re-read RateContext / Working Set 效率
Search Steps定位效率
Files Touched修改范围
Unnecessary Edit Rate无关修改
User-change Overwrite Rate是否破坏已有用户修改

53. 一个重要指标:Unnecessary Edit Rate

Coding Agent 很容易:

顺手重构
顺手格式化
修改无关文件

即使最终测试通过,也会增加风险。

因此可以评价:

任务必须修改的最小范围
vs
Agent 实际修改范围

这是 Coding Agent 工业质量的重要指标。


54. Minimal Diff Principle

推荐原则:

在满足任务的前提下,优先产生最小、可解释、可验证的 Diff。

好处:

降低回归风险
更容易 Review
更容易测试
更容易回滚
减少 Context

但对于明确要求大规模重构的任务,不应机械追求最小行数。


55. Editing Failure Taxonomy

建议至少区分:

SEARCH_MISS
WRONG_FILE
INSUFFICIENT_CONTEXT
STALE_FILE
AMBIGUOUS_EDIT
PATCH_APPLY_FAILURE
SYNTAX_REGRESSION
TYPE_REGRESSION
LINT_REGRESSION
TEST_REGRESSION
PREEXISTING_FAILURE
UNRELATED_EDIT
USER_CHANGE_OVERWRITE
VERIFICATION_INCOMPLETE

这样 Eval 才知道:

Agent 是“不会写代码”,还是“编辑 Runtime 没保护好”。


56. 与下一篇 Permission & HITL 的衔接

现在 Coding Agent 已经能够:

Search
Read
Edit
Run Tests

接下来必须回答:

哪些操作应该自动允许,哪些操作必须经过用户或策略审批?

例如:

read src/**
→ 自动允许

edit src/**
→ 允许

edit production config
→ 需要审批

git reset --hard
→ 高风险

rm -rf
→ 阻断 / 审批

network deploy
→ 明确审批

所以下一篇:

Permission & HITL:Agent 的权限、审批与策略控制

会重点研究:

Capability
Policy
Approval
Risk Classification
Permission Scope
Human-in-the-loop
Denied Observation
Approval Caching
Steering
Escalation
Least Privilege

57. 一句话总纲

Quote

Coding Agent Editing Runtime 的本质,不是让模型拥有“写文件”能力,而是让代码修改始终发生在可定位、已读取、有版本前置条件、可生成 Diff、可静态诊断、可测试验证、可检测冲突并尊重用户已有修改的闭环中;真正可靠的 Coding Agent 是通过 Runtime 不断校验和收敛,而不是依赖模型一次性写对。


Share this post:

Previous Post
Permission & HITL:Agent 的权限、审批与策略控制
Next Post
Tool Runtime:从 Tool Call 到可靠执行