有一类 bug 特别折磨人:某个变量的值莫名其妙变了,但你翻遍代码也找不到哪里改了它。可能是数组越界写到了隔壁,可能是野指针,可能是另一个线程。断点帮不上忙,因为你根本不知道该在哪一行下断点。这时要用的是观察点(watchpoint):不是「程序执行到某一行时停」,而是「某块内存被修改时停」。

这是「C++ 调试实战」系列的第 5 篇。前面讲的断点是盯着「代码」,观察点是盯着「数据」。

一、一个「账对不上」的 bug

下面是一个记账程序:Ledger 里有三个账户余额,外加一个审计总额 audit_total,它应该始终等于三个余额之和。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
// watch.cpp
#include <cstdio>

struct Ledger {
int balances[3];
int audit_total; // 应当始终等于 balances 之和
};

void deposit(Ledger& l, int idx, int amount) {
l.balances[idx] += amount;
l.audit_total += amount;
}

void apply_fee(Ledger& l, int n) {
// Bug:本该是 i < n,多循环了一次
for (int i = 0; i <= n; ++i) {
l.balances[i] -= 1;
}
}

int main() {
Ledger ledger{{100, 200, 300}, 600};
deposit(ledger, 0, 50);
apply_fee(ledger, 3);
std::printf("balances = %d %d %d, audit = %d\n", ledger.balances[0],
ledger.balances[1], ledger.balances[2], ledger.audit_total);
return 0;
}
行号 代码
9–10 deposit():改余额、改审计总额
15–16 apply_fee() 的循环:每个账户扣 1 元手续费
21–23 main():初始化、存款、扣费
1
2
g++ -std=c++20 -g -O0 watch.cpp -o watch && ./watch
# balances = 149 199 299, audit = 649

三个余额之和是 149 + 199 + 299 = 647,审计总额应该是 600 + 50 - 3 = 647,可程序输出的是 649。bug 已经写在注释里了(i <= n 多循环一次,balances[3] 越界,正好写到了紧挨着的 audit_total 上),但假设你不知道,而且这是一个几万行的项目,audit_total 被几十个地方读写,你该怎么找?

答案:让调试器盯住 audit_total,谁改它就停在谁那里。

二、watch:变量被写入时停下

先停在 main 里、ledger 初始化完成之后(第 22 行),然后设置观察点:

1
2
3
4
5
6
7
8
(gdb) break watch.cpp:22
Breakpoint 1 at 0x958: file watch.cpp, line 22.
(gdb) run

Breakpoint 1, main () at watch.cpp:22
22 deposit(ledger, 0, 50);
(gdb) watch ledger.audit_total
Hardware watchpoint 2: ledger.audit_total

Hardware watchpoint 表示这是一个硬件观察点(第六节解释)。继续运行:

1
2
3
4
5
6
7
8
(gdb) continue

Hardware watchpoint 2: ledger.audit_total

Old value = 600
New value = 650
deposit (l=..., idx=0, amount=50) at watch.cpp:11
11 }

第一次停下:deposit 把它从 600 改成了 650,这是正常的存款逻辑。再继续:

1
2
3
4
5
6
7
8
(gdb) continue

Hardware watchpoint 2: ledger.audit_total

Old value = 650
New value = 649
apply_fee (l=..., n=3) at watch.cpp:15
15 for (int i = 0; i <= n; ++i) {

抓到了:apply_fee 把审计总额从 650 改成了 649,可这个函数根本不应该碰 audit_total。看看现场:

1
2
3
4
5
6
7
8
9
10
11
12
13
(gdb) bt
#0 apply_fee (l=..., n=3) at watch.cpp:15
#1 0x0000aaaaaaaa0974 in main () at watch.cpp:23
(gdb) print i
$1 = 3
(gdb) print n
$2 = 3
(gdb) print l
$3 = (Ledger &) @0xfffffffffb18: {balances = {149, 199, 299}, audit_total = 649}
(gdb) print &l.balances[i]
$4 = (int *) 0xfffffffffb24
(gdb) print &l.audit_total
$5 = (int *) 0xfffffffffb24

i == 3,而 balances 只有 3 个元素(下标 0~2)。最后两行把证据摆得清清楚楚:&l.balances[3] 和 &l.audit_total 是同一个地址。越界写入正好落在了结构体的下一个成员上。

有两个细节要注意:

  • 停下的位置是「写入之后」。观察点是在那条写内存的指令执行完之后触发的,所以显示的是下一条要执行的语句——这里是循环头的第 15 行(++i 和判断),而真正的「凶手」是它前面刚执行完的第 16 行 l.balances[i] -= 1;。
  • Old value / New value 直接告诉你改之前和改之后的值,不用自己去比。

LLDB 的输出(macOS 上结果一样):

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
(lldb) watchpoint set variable ledger.audit_total
Watchpoint created: Watchpoint 1: addr = 0x16fdfdd7c size = 4 state = enabled type = m
watchpoint spec = 'ledger.audit_total'
...
(lldb) continue

Watchpoint 1 hit:
old value: 600
new value: 650
Process 5544 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = watchpoint 1
frame #0: 0x0000000100000524 watch.mac`deposit(l=0x000000016fdfdd70, idx=0, amount=50) at watch.cpp:11:1
-> 11 }
^
12
(lldb) continue

Watchpoint 1 hit:
old value: 650
new value: 649
Process 5544 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = watchpoint 1
frame #0: 0x0000000100000568 watch.mac`apply_fee(l=0x000000016fdfdd70, n=3) at watch.cpp:17:5
-> 17 }
^
18 }
(lldb) v i n
(int) i = 3
(int) n = 3
(lldb) p &l.balances[i]
(int *) 0x000000016fdfdd7c
(lldb) p &l.audit_total
(int *) 0x000000016fdfdd7c

watchpoint set variable 可以缩写为 w s v ledger.audit_total。输出里的 type = m 表示 modify 类型:值被改变时触发。LLDB 同样停在「写入之后」,Clang 生成的行号让它停在了第 17 行的 } 上。

三、观察点的作用域

上面的观察点设置在 main 里的局部变量 ledger 上。当 main 返回、ledger 的生命周期结束时,GDB 会自动删除它:

1
2
3
4
(gdb) continue

Watchpoint 2 deleted because the program has left the block in
which its expression is valid.

如果你是在一个子函数里、通过引用或指针设置观察点,这个问题就更明显了。比如停在 deposit 里,watch l.audit_total——l 是 deposit 的参数,deposit 一返回观察点就被删了,根本等不到 apply_fee 去改它。

解决办法是 watch -l(-location):先把表达式求值成一个地址,然后盯住这个地址本身,不再关心表达式里的变量是否还在作用域内:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
(gdb) break deposit
(gdb) run
...
(gdb) watch -l l.audit_total
Hardware watchpoint 2: -location l.audit_total
(gdb) continue

Hardware watchpoint 2: -location l.audit_total

Old value = 600
New value = 650
deposit (l=..., idx=0, amount=50) at watch.cpp:11
11 }
(gdb) continue

Hardware watchpoint 2: -location l.audit_total

Old value = 650
New value = 649
apply_fee (l=..., n=3) at watch.cpp:15
15 for (int i = 0; i <= n; ++i) {

离开 deposit 之后观察点依然有效,在 apply_fee 里成功抓到了越界写入。

LLDB 的 watchpoint set variable 本身就是按地址观察的,另外也可以显式地按地址设置:

1
(lldb) watchpoint set expression -- &l.audit_total

经验:只要被观察的对象不是当前函数的局部变量,就用 watch -l。 堆上的对象(new 出来的、容器里的元素)尤其如此。

四、条件观察点

观察点也可以加条件,只在满足条件时停。比如只关心 balances[1] 什么时候跌破 200:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
(gdb) break apply_fee
(gdb) run
...
(gdb) watch l.balances[1] if l.balances[1] < 200
Hardware watchpoint 2: l.balances[1]
(gdb) continue

Hardware watchpoint 2: l.balances[1]

Old value = 200
New value = 199
apply_fee (l=..., n=3) at watch.cpp:15
15 for (int i = 0; i <= n; ++i) {
(gdb) print i
$1 = 1

LLDB 先设置观察点,再用 watchpoint modify -c 加条件。条件里也可以用程序里的其他变量,比如只在循环第 2 轮(i == 1)时停:

1
2
3
4
5
6
7
8
9
10
(lldb) watchpoint set variable -w read_write l.balances[1]
(lldb) watchpoint modify -c 'i == 1'
(lldb) continue
...
* thread #1, queue = 'com.apple.main-thread', stop reason = watchpoint 1
frame #0: 0x0000000100000560 watch.mac`apply_fee(l=0x000000016fdfdd70, n=3) at watch.cpp:16:23
-> 16 l.balances[i] -= 1;
^
(lldb) v i
(int) i = 1

watchpoint modify 不写编号时,作用于最后创建的那个观察点。这里用的是 -w read_write(读和写都触发,下一节讲),所以在第 16 行 -= 读取旧值时就停下了,还没来得及写。

五、读观察点与访问观察点

除了「写入时停」,还可以「读取时停」:

类型 GDB LLDB 触发时机
写观察点 watch expr watchpoint set variable var(默认 -w modify) 值被修改。GDB 和 LLDB 的默认行为都是只在值真的变了时报告;LLDB 用 -w write 可以在「写入了相同的值」时也停
读观察点 rwatch expr watchpoint set variable -w read var 值被读取
访问观察点 awatch expr watchpoint set variable -w read_write var 读或写

读观察点在一些场景里很有用,比如「这个已经释放的对象,还有谁在读它?」。在我们的例子里,audit_total 按说只有 main 最后打印时才会读,看看 rwatch 会抓到什么:

1
2
3
4
5
6
7
8
9
10
11
12
(gdb) break apply_fee
(gdb) run
...
(gdb) rwatch l.audit_total
Hardware read watchpoint 2: l.audit_total
(gdb) continue

Hardware read watchpoint 2: l.audit_total

Value = 650
0x0000aaaaaaaa08d4 in apply_fee (l=..., n=3) at watch.cpp:16
16 l.balances[i] -= 1;

apply_fee 在第 16 行读取了 audit_total——l.balances[3] -= 1 要先读出旧值再减 1,读的正是 audit_total 那块内存。同一个 bug,读观察点比写观察点还早一步抓到它。

六、硬件观察点与软件观察点

6.1 硬件观察点

现代 CPU 都有专门的调试寄存器,可以设置「某个地址被访问时触发异常」,完全由硬件检测,程序全速运行、没有任何性能损失。这就是 Hardware watchpoint。

但调试寄存器数量很少,x86_64 和常见的 ARM64 处理器都只有 4 个左右,每个能监视的长度也有限(通常最多 8 字节)。LLDB 会直接告诉你:

1
2
(lldb) watchpoint list
Number of supported hardware watchpoints: 4

超出限制时会报错。在本文的 Docker 环境里,同时设了一个 watch、一个 rwatch、一个 awatch,continue 时就失败了:

1
2
3
4
5
6
(gdb) continue
Warning:
Could not insert hardware watchpoint 3.
Could not insert hardware watchpoint 4.
Could not insert hardware breakpoints:
You may have requested too many hardware breakpoints/watchpoints.

(容器跑在虚拟机里,能用的调试寄存器可能比物理机更少。)遇到这种情况就删掉不需要的观察点,一次只盯一两个关键变量。

6.2 软件观察点

硬件资源用完、或者要观察的数据太大(比如整个结构体、一个大数组)时,GDB 会退回软件观察点:每执行一条指令就停下来检查一次值有没有变。这会让程序慢成百上千倍,只适合很短的代码段。

可以用 set can-use-hw-watchpoints 0 强制 GDB 使用软件观察点(一般用不上,只在怀疑硬件观察点有问题时用来对照)。

6.3 观察什么,才能用上硬件观察点

  • 观察单个标量(int、指针、double),而不是整个对象。想知道 std::string 什么时候被改,观察它内部的长度字段或数据指针,而不是整个 string;
  • 表达式里不要带函数调用,比如 watch v.size() 是不行的,调试器无法把它映射成一个固定地址;
  • 优先用 watch -l,按地址观察。

七、管理观察点

观察点和断点共用一套编号,GDB 里用管理断点的命令管理它们:

1
2
3
4
5
6
(gdb) info watchpoints
Num Type Disp Enb Address What
2 hw watchpoint keep y l.balances[0]
3 read watchpoint keep y l.audit_total
4 acc watchpoint keep y l.balances[1]
stop only if l.balances[1] < 200

info watchpoints 只列观察点,info breakpoints 会连断点一起列出。disable、enable、delete 用法和断点完全一样。

LLDB 有独立的一套:

操作 LLDB
列出 watchpoint list
禁用 / 启用 watchpoint disable 1 / watchpoint enable 1
删除 watchpoint delete 1
加条件 watchpoint modify -c '条件' 1
忽略前 N 次 watchpoint ignore -i N 1

八、小结

需求 GDB LLDB
值被写入时停 watch x watchpoint set variable x(w s v x)
按地址观察(跨作用域) watch -l expr watchpoint set expression -- &expr
值被读取时停 rwatch x w s v -w read x
读或写都停 awatch x w s v -w read_write x
条件观察点 watch x if 条件 watchpoint modify -c '条件'
列出观察点 info watchpoints watchpoint list

观察点最适合的场景:

  • 变量被意外修改,不知道是谁改的;
  • 怀疑数组越界、缓冲区溢出踩坏了相邻数据;
  • 追踪一个值在复杂逻辑里是什么时候、被谁改成错误值的。

它的局限是只能盯住少数几个小变量。如果你面对的是「程序到处乱写内存、一跑就崩」,更好的工具是下一篇要讲的 AddressSanitizer。

C++ 调试实战系列第 5 篇完。下一篇:崩溃现场分析——段错误、core dump 与 AddressSanitizer。