单线程程序出了问题,大不了一步一步走。多线程程序就麻烦了:断点在哪个线程上命中?单步时别的线程在干什么?程序卡住不动了,到底是谁在等谁?而且多线程 bug 常常「一调试就消失」——调试器改变了线程执行的节奏。这一篇讲多线程调试的基本功:查看和切换线程、给所有线程打印调用栈、attach 到一个卡死的进程定位死锁,以及用 ThreadSanitizer 抓数据竞争。

这是「C++ 调试实战」系列的第 7 篇。本篇的所有命令都建立在前几篇的基础上:线程切换之后,每个线程里的调用栈、变量,查看方式和单线程完全一样。

编译多线程程序时,GCC 需要链接线程库。较新的 glibc 已经把 pthread 并入了 libc,直接编译就行;老系统上需要加 -pthread。

一、示例一:两个线程累加同一个计数器

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
// race.cpp
#include <cstdio>
#include <thread>

int counter = 0; // 没有任何同步

void add_many() {
for (int i = 0; i < 100000; ++i) {
++counter;
}
}

int main() {
std::thread t1(add_many);
std::thread t2(add_many);
t1.join();
t2.join();
std::printf("counter = %d (expect 200000)\n", counter);
return 0;
}

两个线程各加 10 万次,期望结果是 20 万。运行两次:

1
2
3
4
5
$ g++ -std=c++20 -g -O0 race.cpp -o race.linux
$ ./race.linux
counter = 121549 (expect 200000)
$ ./race.linux
counter = 121236 (expect 200000)

(和前几篇一样,Linux 版可执行文件带 .linux 后缀,macOS 版带 .mac 后缀,下面输出里的线程名就是程序名。)

每次结果都不一样,而且都远小于 20 万。先别急着修,用它来熟悉多线程调试命令。

二、查看与切换线程

2.1 断点在哪个线程命中

1
2
3
4
5
6
7
8
(gdb) break add_many
Breakpoint 1 at 0xdac: file race.cpp, line 7.
(gdb) run
[New Thread 0xfffff7a3f180 (LWP 17)]
[Switching to Thread 0xfffff7a3f180 (LWP 17)]

Thread 2 "race.linux" hit Breakpoint 1, add_many () at race.cpp:7
7 for (int i = 0; i < 100000; ++i) {

和单线程程序相比,多了几条信息:

  • [New Thread ... (LWP 17)]:新线程创建了。LWP(Light Weight Process)是 Linux 内核里的线程 ID;
  • [Switching to Thread ...]:断点命中的线程和之前选中的不同,GDB 自动切换过去;
  • Thread 2 "race.linux" hit Breakpoint 1:命中断点的是 2 号线程(GDB 内部编号),线程名是 race.linux(默认就是程序名)。

2.2 info threads:列出所有线程

1
2
3
4
(gdb) info threads
Id Target Id Frame
1 Thread 0xfffff7fecc00 (LWP 15) "race.linux" __GI___clone3 () at ../sysdeps/unix/sysv/linux/aarch64/clone3.S:60
* 2 Thread 0xfffff7a3f180 (LWP 17) "race.linux" add_many () at race.cpp:7
列 含义
* 当前选中的线程
Id GDB 的线程编号,切换线程时用它
Target Id 系统层面的线程标识:pthread_t 的值、LWP 号、线程名
Frame 这个线程当前停在哪个函数

1 号线程是主线程,此时正在 clone3 里——它正在创建第二个工作线程。

要点:GDB 默认是「全停」模式(all-stop),任何一个线程停下来,所有线程都会停下来。 所以你看到的是整个进程在同一时刻的快照。

继续运行,第二个线程也命中了断点:

1
2
3
4
5
6
7
8
9
10
11
(gdb) continue
[New Thread 0xfffff722f180 (LWP 18)]
[Switching to Thread 0xfffff722f180 (LWP 18)]

Thread 3 "race.linux" hit Breakpoint 1, add_many () at race.cpp:7
7 for (int i = 0; i < 100000; ++i) {
(gdb) info threads
Id Target Id Frame
1 Thread 0xfffff7fecc00 (LWP 15) "race.linux" __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/aarch64/syscall_cancel.S:50
2 Thread 0xfffff7a3f180 (LWP 17) "race.linux" add_many () at race.cpp:7
* 3 Thread 0xfffff722f180 (LWP 18) "race.linux" add_many () at race.cpp:7

现在有 3 个线程:主线程已经进入 t1.join() 在等待(停在 __syscall_cancel_arch,这是等待类系统调用的入口),两个工作线程都在 add_many 的第 7 行。

2.3 thread N:切换线程

1
2
3
(gdb) thread 2
[Switching to thread 2 (Thread 0xffff8421f180 (LWP 10))]
#0 futex_wait (futex_word=0xaaaae4850020 <mtx_a>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126

(这段输出来自下面的死锁示例。)切换之后,bt、frame、print、info locals 都是针对这个线程的。

LLDB 的对应命令:

操作 GDB LLDB
列出线程 info threads thread list
切换线程 thread 2 thread select 2
所有线程的调用栈 thread apply all bt bt all(thread backtrace all)
在指定线程执行命令 thread apply 2 3 print i ——

2.4 只在某个线程上停

断点可以限定线程,只有指定线程执行到这里时才停:

1
(gdb) break race.cpp:8 thread 2

LLDB 用 breakpoint set -f race.cpp -l 8 -t <线程ID>,或者用 -T <线程名> 按线程名过滤。给线程起名(Linux 上用 pthread_setname_np)能让多线程调试轻松很多,info threads 里一眼就能看出哪个线程是干什么的。

2.5 单步时别让其他线程乱跑

在一个线程里 next,默认情况下其他线程也会跟着运行。如果你正在单步观察共享变量,它可能在你两次单步之间被别的线程改掉。GDB 可以锁住调度器:

1
(gdb) set scheduler-locking step
取值 效果
off 所有线程都自由运行
step 单步(next / step)时只运行当前线程;continue 时所有线程都运行
on 任何时候都只运行当前线程

注意 on 模式下如果当前线程在等一个锁、而持有锁的线程被锁住不能运行,程序就「卡死」了——这是你自己造成的,不是程序的 bug。

LLDB 的单步命令有 --run-mode 选项:thread step-over --run-mode this-thread。

三、示例二:程序卡住了——定位死锁

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
// threads.cpp
#include <chrono>
#include <cstdio>
#include <mutex>
#include <thread>

std::mutex mtx_a;
std::mutex mtx_b;

void worker1() {
std::lock_guard<std::mutex> la(mtx_a);
std::this_thread::sleep_for(std::chrono::milliseconds(100));
std::lock_guard<std::mutex> lb(mtx_b); // 等 mtx_b
std::puts("worker1 done");
}

void worker2() {
std::lock_guard<std::mutex> lb(mtx_b);
std::this_thread::sleep_for(std::chrono::milliseconds(100));
std::lock_guard<std::mutex> la(mtx_a); // 等 mtx_a —— 与 worker1 顺序相反
std::puts("worker2 done");
}

int main() {
std::thread t1(worker1);
std::thread t2(worker2);
t1.join();
t2.join();
return 0;
}
行号 代码
10 / 12 worker1:先锁 mtx_a,再锁 mtx_b
17 / 19 worker2:先锁 mtx_b,再锁 mtx_a
26 t1.join();

运行之后,程序什么也不输出,也不退出,就这么卡着。这是经典的死锁:worker1 拿着 mtx_a 等 mtx_b,worker2 拿着 mtx_b 等 mtx_a,谁也等不到。(sleep_for 是为了让两个线程几乎必然各自拿到第一把锁,稳定复现死锁。)

3.1 两种进入调试器的方式

方式 1:在调试器里运行,卡住后按 Ctrl+C。 GDB 和 LLDB 都会中断程序,停在当前位置,然后就可以查看了。

方式 2:attach 到已经在运行的进程。 这是实际工作中更常见的情况:服务已经跑了很久,突然不响应了。你不能重启它(重启就丢失现场了),而是把调试器「挂」上去:

1
2
3
$ ./threads.linux &       # 模拟一个已经卡住的进程
$ pgrep -f threads.linux # 找到它的 PID(本文实测为 7)
$ gdb -p 7 # attach 上去

LLDB 是 lldb -p <PID>,或者按进程名 lldb -n threads.mac。

attach 需要权限,两个常见问题:

  • Linux 上报 ptrace: Operation not permitted:很多发行版开启了 Yama 安全模块,/proc/sys/kernel/yama/ptrace_scope 为 1 时,只允许调试自己启动的子进程。临时解决办法是用 sudo gdb -p <PID>;Docker 里则需要 --cap-add=SYS_PTRACE(第 00 篇);
  • macOS 上:attach 到别的进程需要开发者权限,首次使用时系统会弹窗要求授权,或者需要 sudo。系统自带的程序受系统完整性保护(SIP),不能 attach。

attach 成功后,进程会被暂停;调试结束时用 detach 让进程继续运行,或者直接 quit(GDB 会询问是否 detach)。

3.2 看所有线程在干什么

attach 之后,第一个命令就是 info threads:

1
2
3
4
5
(gdb) info threads
Id Target Id Frame
* 1 Thread 0xffffb280fc00 (LWP 7) "threads.linux" __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/aarch64/syscall_cancel.S:50
2 Thread 0xffffb1a4f180 (LWP 10) "threads.linux" futex_wait (futex_word=0xaaaabb230020 <mtx_a>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126
3 Thread 0xffffb225f180 (LWP 9) "threads.linux" futex_wait (futex_word=0xaaaabb230050 <mtx_b>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126

GDB 很贴心地把地址翻译成了符号:线程 2 在等 <mtx_a>,线程 3 在等 <mtx_b>。futex_wait 是 Linux 上线程阻塞等待锁的底层系统调用。

再用 thread apply all bt 看每个线程完整的调用栈(省略了标准库内部的部分帧):

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
(gdb) thread apply all bt

Thread 3 (Thread 0xffffb225f180 (LWP 9) "threads.linux"):
#0 futex_wait (futex_word=0xaaaabb230050 <mtx_b>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126
#1 __GI___lll_lock_wait (futex=futex@entry=0xaaaabb230050 <mtx_b>, private=private@entry=0) at ./nptl/lowlevellock.c:49
...
#5 std::mutex::lock (this=0xaaaabb230050 <mtx_b>) at /usr/include/c++/15/bits/std_mutex.h:115
#6 0x0000aaaabb212758 in std::lock_guard<std::mutex>::lock_guard (this=0xffffb225e7b0, __m=...) at /usr/include/c++/15/bits/std_mutex.h:252
#7 0x0000aaaabb210fd0 in worker1 () at threads.cpp:12
...

Thread 2 (Thread 0xffffb1a4f180 (LWP 10) "threads.linux"):
#0 futex_wait (futex_word=0xaaaabb230020 <mtx_a>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126
#1 __GI___lll_lock_wait (futex=futex@entry=0xaaaabb230020 <mtx_a>, private=private@entry=0) at ./nptl/lowlevellock.c:49
...
#5 std::mutex::lock (this=0xaaaabb230020 <mtx_a>) at /usr/include/c++/15/bits/std_mutex.h:115
#6 0x0000aaaabb212758 in std::lock_guard<std::mutex>::lock_guard (this=0xffffb1a4e7b0, __m=...) at /usr/include/c++/15/bits/std_mutex.h:252
#7 0x0000aaaabb2110d0 in worker2 () at threads.cpp:19
...

Thread 1 (Thread 0xffffb280fc00 (LWP 7) "threads.linux"):
#0 __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/aarch64/syscall_cancel.S:50
...
#5 0x0000ffffb23bddb8 [PAC] in __pthread_clockjoin_ex (threadid=281473670574464, thread_return=0x0, clockid=0, abstime=0x0, cancel=<optimized out>) at ./nptl/pthread_join_common.c:67
#6 0x0000ffffb2627944 [PAC] in std::thread::join() () from /usr/lib/aarch64-linux-gnu/libstdc++.so.6
#7 0x0000aaaabb2111bc [PAC] in main () at threads.cpp:26

输出里的 [PAC] 是 ARM64 的指针认证(Pointer Authentication)标记,表示这个返回地址是被签名过的,在 x86_64 上不会出现。

用第 04 篇的读栈方法——找第一个属于自己代码的帧:

线程 自己代码的帧 在做什么
3(LWP 9) #7 worker1 () at threads.cpp:12 在第 12 行等 mtx_b
2(LWP 10) #7 worker2 () at threads.cpp:19 在第 19 行等 mtx_a
1(LWP 7) #7 main () at threads.cpp:26 在 t1.join() 等线程 1 结束

3.3 找到锁的持有者

知道了谁在等哪把锁,还要知道每把锁被谁拿着。在 Linux 上,std::mutex 底层是 pthread_mutex_t,它的内部结构里记录了持有者的 LWP:

1
2
3
4
5
6
7
8
9
10
(gdb) thread 2
[Switching to thread 2 (Thread 0xffff8421f180 (LWP 10))]
#0 futex_wait (futex_word=0xaaaae4850020 <mtx_a>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126
(gdb) frame 7
#7 0x0000aaaae48310d0 in worker2 () at threads.cpp:19
19 std::lock_guard<std::mutex> la(mtx_a); // 等 mtx_a —— 与 worker1 顺序相反
(gdb) print mtx_a._M_mutex.__data.__owner
$1 = 9
(gdb) print mtx_b._M_mutex.__data.__owner
$2 = 10

(这次 attach 的是重新运行的进程,所以地址和上面不同,LWP 编号恰好一致。)

把线索串起来:

1
2
线程 3(LWP 9,worker1)  ──持有──▶  mtx_a  ◀──等待──  线程 2(LWP 10,worker2)
线程 3(LWP 9,worker1) ──等待──▶ mtx_b ◀──持有── 线程 2(LWP 10,worker2)

一个完美的环:循环等待。这就是死锁的铁证。

_M_mutex.__data.__owner 是 libstdc++ 和 glibc 的内部实现细节,不同平台不一样(macOS 的 libc++ 就没有这个字段)。但思路是通用的:先看每个线程在等什么(调用栈),再想办法确认谁持有什么。实在找不到持有者字段,就在每个线程里看它的调用栈上有哪些 lock_guard 已经构造完成。

3.4 LLDB 版

在 macOS 上 attach:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
(lldb) process attach --pid 7106
Process 7106 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP
...
(lldb) thread list
Process 7106 stopped
* thread #1: tid = 0x72b0c, 0x0000000188831af8 libsystem_kernel.dylib`__ulock_wait + 8, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP
thread #2: tid = 0x72b2d, 0x00000001888329dc libsystem_kernel.dylib`__psynch_mutexwait + 8
thread #3: tid = 0x72b2e, 0x00000001888329dc libsystem_kernel.dylib`__psynch_mutexwait + 8
(lldb) bt all
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP
* frame #0: 0x0000000188831af8 libsystem_kernel.dylib`__ulock_wait + 8
frame #1: 0x0000000188876114 libsystem_pthread.dylib`_pthread_join + 616
frame #2: 0x00000001887aae24 libc++.1.dylib`std::__1::thread::join() + 36
frame #3: 0x000000010058896c threads.mac`main at threads.cpp:26:8
frame #4: 0x00000001884a84e4 dyld`start + 6992
thread #2
frame #0: 0x00000001888329dc libsystem_kernel.dylib`__psynch_mutexwait + 8
...
frame #4: 0x0000000100588a60 threads.mac`std::__1::lock_guard<std::__1::mutex>::lock_guard[abi:nqe220106](this=0x000000016f8fef10, __m=0x0000000100590040) at lock_guard.h:32:10
...
﹍ frame #6: 0x0000000100588670 threads.mac`worker1() at threads.cpp:12:33
thread #3
frame #0: 0x00000001888329dc libsystem_kernel.dylib`__psynch_mutexwait + 8
...
frame #4: 0x0000000100588a60 threads.mac`std::__1::lock_guard<std::__1::mutex>::lock_guard[abi:nqe220106](this=0x000000016f98af10, __m=0x0000000100590000) at lock_guard.h:31:19
...
﹍ frame #6: 0x00000001005888c4 threads.mac`worker2() at threads.cpp:19:33

macOS 上等待互斥锁的系统调用叫 __psynch_mutexwait。线程 2 在 worker1 第 12 行等地址 0x...590040 的锁,线程 3 在 worker2 第 19 行等 0x...590000 的锁(两个全局 mutex 相邻存放,前者是 mtx_b,后者是 mtx_a,可以用 p &mtx_a 确认)。结论和 Linux 上完全一致。

﹍ 和 ﹉ 符号表示 LLDB 把中间几个标准库的帧折叠了(第 7、8 帧)。切到自己代码的帧看源码:

1
2
3
4
5
6
7
8
9
10
(lldb) thread select 2
(lldb) frame select 6
frame #6: 0x0000000100588670 threads.mac`worker1() at threads.cpp:12:33
9 void worker1() {
10 std::lock_guard<std::mutex> la(mtx_a);
11 std::this_thread::sleep_for(std::chrono::milliseconds(100));
-> 12 std::lock_guard<std::mutex> lb(mtx_b); // 等 mtx_b
^
13 std::puts("worker1 done");
14 }

3.5 修复

死锁的根源是两个线程加锁顺序相反。修复方法:

  • 所有线程按统一的顺序加锁(都先 mtx_a 再 mtx_b);
  • 或者用 C++17 的 std::scoped_lock,一次性锁住多把锁,它内部用死锁避免算法:
1
2
3
4
void worker1() {
std::scoped_lock lock(mtx_a, mtx_b); // 同时锁两把,不会死锁
std::puts("worker1 done");
}

四、数据竞争:ThreadSanitizer

回到第一个例子 race.cpp。计数器结果不对的原因是 ++counter 不是原子操作,它实际上是「读、加一、写回」三步,两个线程交错执行时会丢失更新。

这类 bug 用调试器很难抓:你在 ++counter 上下断点,线程就被暂停了,时序完全被改变;而且它不崩溃、不卡死,只是结果「偶尔不对」。专门的工具是 ThreadSanitizer(TSan):

1
2
g++ -std=c++20 -g -O0 -fsanitize=thread race.cpp -o race.tsan
./race.tsan
1
2
3
4
5
6
7
8
9
10
11
12
13
14
==================
WARNING: ThreadSanitizer: data race (pid=12)
Read of size 4 at 0xaaaae1b7001c by thread T2:
#0 add_many() /work/race.cpp:8 (race.tsan+0x1098)
...
Previous write of size 4 at 0xaaaae1b7001c by thread T1:
#0 add_many() /work/race.cpp:8 (race.tsan+0x10b4)
...
Location is global 'counter' of size 4 at 0xaaaae1b7001c (race.tsan+0x2001c)
...
SUMMARY: ThreadSanitizer: data race /work/race.cpp:8 in add_many()
==================
...
counter = 200000 (expect 200000)

(为了篇幅省略了 BuildId 和标准库的栈帧。)报告的结构和 ASan 很像:

部分 含义
Read of size 4 ... by thread T2 线程 T2 在 race.cpp:8 读了这块内存
Previous write ... by thread T1 此前线程 T1 在 race.cpp:8 写过同一块内存,两次访问之间没有任何同步
Location is global 'counter' 出问题的变量是全局变量 counter

这正是 ++counter 那一行。有意思的是,这次运行打印出了正确的 200000——TSan 的插桩改变了程序的执行节奏,竞争碰巧没有造成丢失更新。但 TSan 照样报告了问题:它检测的是「两个线程访问同一内存且没有同步」这个事实本身,不依赖于这次运行是否真的算错了。这正是它比「多跑几次看结果」可靠的地方。

修复:把 counter 改成 std::atomic<int>,或者用 std::mutex 保护。

TSan 的注意事项:

  • 不能和 ASan 同时使用;
  • 运行开销比 ASan 大,程序可能慢 5~15 倍,内存占用也更多;
  • 同样只能发现实际执行到的竞争,所以测试要覆盖到并发路径。

五、多线程调试的几条经验

经验 说明
程序卡住,先抓全体线程的栈 thread apply all bt(LLDB:bt all)是定位死锁、卡死的第一步;可以把输出保存下来慢慢分析
不要急着重启 卡死的进程就是最好的现场,先 attach,看完再说
给线程起名字 info threads 里一眼能认出每个线程是干什么的
调试器会改变时序 断点、单步都会让线程停下来,竞争类 bug 可能因此消失。这类问题优先用 TSan
单步时锁住调度器 想专注观察一个线程时,set scheduler-locking step
崩溃在别的线程 多线程程序崩溃时,GDB / LLDB 会自动切到出问题的线程,直接 bt 即可;core dump 里也包含所有线程

一个实用的小脚本:不用进入交互界面,直接抓取某个进程所有线程的调用栈并保存下来:

1
gdb -p <PID> -batch -ex "thread apply all bt" > stacks.txt

-batch 表示执行完命令就退出(并 detach,进程继续运行)。线上服务偶发卡顿时,隔几秒抓一次,对比几份栈,就能看出线程长时间卡在哪里。

六、小结

需求 GDB LLDB
列出线程 info threads thread list
切换线程 thread N thread select N
所有线程的调用栈 thread apply all bt bt all
指定线程的断点 break 位置 thread N br set ... -t 线程ID
单步时只跑当前线程 set scheduler-locking step thread step-over --run-mode this-thread
attach 到进程 gdb -p PID lldb -p PID / lldb -n 进程名
脱离进程 detach detach
检测数据竞争 -fsanitize=thread -fsanitize=thread

C++ 调试实战系列第 7 篇完。下一篇:IDE 图形化调试与速查表——VS Code 调试配置,以及 GDB / LLDB 命令对照总表。