单线程程序出了问题,大不了一步一步走。多线程程序就麻烦了:断点在哪个线程上命中?单步时别的线程在干什么?程序卡住不动了,到底是谁在等谁?而且多线程 bug 常常「一调试就消失」——调试器改变了线程执行的节奏。这一篇讲多线程调试的基本功:查看和切换线程、给所有线程打印调用栈、attach 到一个卡死的进程定位死锁,以及用 ThreadSanitizer 抓数据竞争。
这是「C++ 调试实战」系列的第 7 篇。本篇的所有命令都建立在前几篇的基础上:线程切换之后,每个线程里的调用栈 、变量 ,查看方式和单线程完全一样。
编译多线程程序时,GCC 需要链接线程库。较新的 glibc 已经把 pthread 并入了 libc,直接编译就行;老系统上需要加 -pthread。
一、示例一:两个线程累加同一个计数器 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 #include <cstdio> #include <thread> int counter = 0 ; void add_many () { for (int i = 0 ; i < 100000 ; ++i) { ++counter; } } int main () { std::thread t1 (add_many) ; std::thread t2 (add_many) ; t1. join (); t2. join (); std::printf ("counter = %d (expect 200000)\n" , counter); return 0 ; }
两个线程各加 10 万次,期望结果是 20 万。运行两次:
1 2 3 4 5 $ g++ -std=c++20 -g -O0 race.cpp -o race.linux $ ./race.linux counter = 121549 (expect 200000) $ ./race.linux counter = 121236 (expect 200000)
(和前几篇一样,Linux 版可执行文件带 .linux 后缀,macOS 版带 .mac 后缀,下面输出里的线程名就是程序名。)
每次结果都不一样,而且都远小于 20 万。先别急着修,用它来熟悉多线程调试命令。
二、查看与切换线程 2.1 断点在哪个线程命中 1 2 3 4 5 6 7 8 (gdb) break add_many Breakpoint 1 at 0xdac: file race.cpp, line 7. (gdb) run [New Thread 0xfffff7a3f180 (LWP 17)] [Switching to Thread 0xfffff7a3f180 (LWP 17)] Thread 2 "race.linux" hit Breakpoint 1, add_many () at race.cpp:7 7 for (int i = 0; i < 100000; ++i) {
和单线程程序相比,多了几条信息:
[New Thread ... (LWP 17)]:新线程创建了。LWP(Light Weight Process)是 Linux 内核里的线程 ID;
[Switching to Thread ...]:断点命中的线程和之前选中的不同,GDB 自动切换过去;
Thread 2 "race.linux" hit Breakpoint 1:命中断点的是 2 号线程 (GDB 内部编号),线程名是 race.linux(默认就是程序名)。
2.2 info threads:列出所有线程 1 2 3 4 (gdb) info threads Id Target Id Frame 1 Thread 0xfffff7fecc00 (LWP 15) "race.linux" __GI___clone3 () at ../sysdeps/unix/sysv/linux/aarch64/clone3.S:60 * 2 Thread 0xfffff7a3f180 (LWP 17) "race.linux" add_many () at race.cpp:7
列
含义
*
当前选中的线程
Id
GDB 的线程编号,切换线程时用它
Target Id
系统层面的线程标识:pthread_t 的值、LWP 号、线程名
Frame
这个线程当前停在哪个函数
1 号线程是主线程,此时正在 clone3 里——它正在创建第二个工作线程。
要点:GDB 默认是「全停」模式(all-stop),任何一个线程停下来,所有线程都会停下来。 所以你看到的是整个进程在同一时刻的快照。
继续运行,第二个线程也命中了断点:
1 2 3 4 5 6 7 8 9 10 11 (gdb) continue [New Thread 0xfffff722f180 (LWP 18)] [Switching to Thread 0xfffff722f180 (LWP 18)] Thread 3 "race.linux" hit Breakpoint 1, add_many () at race.cpp:7 7 for (int i = 0; i < 100000; ++i) { (gdb) info threads Id Target Id Frame 1 Thread 0xfffff7fecc00 (LWP 15) "race.linux" __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/aarch64/syscall_cancel.S:50 2 Thread 0xfffff7a3f180 (LWP 17) "race.linux" add_many () at race.cpp:7 * 3 Thread 0xfffff722f180 (LWP 18) "race.linux" add_many () at race.cpp:7
现在有 3 个线程:主线程已经进入 t1.join() 在等待(停在 __syscall_cancel_arch,这是等待类系统调用的入口),两个工作线程都在 add_many 的第 7 行。
2.3 thread N:切换线程 1 2 3 (gdb) thread 2 [Switching to thread 2 (Thread 0xffff8421f180 (LWP 10))] #0 futex_wait (futex_word=0xaaaae4850020 <mtx_a>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126
(这段输出来自下面的死锁示例。)切换之后,bt、frame、print、info locals 都是针对这个线程的。
LLDB 的对应命令:
操作
GDB
LLDB
列出线程
info threads
thread list
切换线程
thread 2
thread select 2
所有线程的调用栈
thread apply all bt
bt all(thread backtrace all)
在指定线程执行命令
thread apply 2 3 print i
——
2.4 只在某个线程上停 断点可以限定线程,只有指定线程执行到这里时才停:
1 (gdb) break race.cpp:8 thread 2
LLDB 用 breakpoint set -f race.cpp -l 8 -t <线程ID>,或者用 -T <线程名> 按线程名过滤。给线程起名(Linux 上用 pthread_setname_np)能让多线程调试轻松很多,info threads 里一眼就能看出哪个线程是干什么的。
2.5 单步时别让其他线程乱跑 在一个线程里 next,默认情况下其他线程也会跟着运行 。如果你正在单步观察共享变量,它可能在你两次单步之间被别的线程改掉。GDB 可以锁住调度器:
1 (gdb) set scheduler-locking step
取值
效果
off
所有线程都自由运行
step
单步(next / step)时只运行当前线程;continue 时所有线程都运行
on
任何时候都只运行当前线程
注意 on 模式下如果当前线程在等一个锁、而持有锁的线程被锁住不能运行,程序就「卡死」了——这是你自己造成的,不是程序的 bug。
LLDB 的单步命令有 --run-mode 选项:thread step-over --run-mode this-thread。
三、示例二:程序卡住了——定位死锁 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 #include <chrono> #include <cstdio> #include <mutex> #include <thread> std::mutex mtx_a; std::mutex mtx_b; void worker1 () { std::lock_guard<std::mutex> la (mtx_a) ; std::this_thread::sleep_for (std::chrono::milliseconds (100 )); std::lock_guard<std::mutex> lb (mtx_b) ; std::puts ("worker1 done" ); } void worker2 () { std::lock_guard<std::mutex> lb (mtx_b) ; std::this_thread::sleep_for (std::chrono::milliseconds (100 )); std::lock_guard<std::mutex> la (mtx_a) ; std::puts ("worker2 done" ); } int main () { std::thread t1 (worker1) ; std::thread t2 (worker2) ; t1. join (); t2. join (); return 0 ; }
行号
代码
10 / 12
worker1:先锁 mtx_a,再锁 mtx_b
17 / 19
worker2:先锁 mtx_b,再锁 mtx_a
26
t1.join();
运行之后,程序什么也不输出,也不退出,就这么卡着。这是经典的死锁 :worker1 拿着 mtx_a 等 mtx_b,worker2 拿着 mtx_b 等 mtx_a,谁也等不到。(sleep_for 是为了让两个线程几乎必然各自拿到第一把锁,稳定复现死锁。)
3.1 两种进入调试器的方式 方式 1:在调试器里运行,卡住后按 Ctrl+C。 GDB 和 LLDB 都会中断程序,停在当前位置,然后就可以查看了。
方式 2:attach 到已经在运行的进程。 这是实际工作中更常见的情况:服务已经跑了很久,突然不响应了。你不能重启它(重启就丢失现场了),而是把调试器「挂」上去:
1 2 3 $ ./threads.linux & $ pgrep -f threads.linux $ gdb -p 7
LLDB 是 lldb -p <PID>,或者按进程名 lldb -n threads.mac。
attach 需要权限,两个常见问题:
Linux 上报 ptrace: Operation not permitted :很多发行版开启了 Yama 安全模块,/proc/sys/kernel/yama/ptrace_scope 为 1 时,只允许调试自己启动的子进程。临时解决办法是用 sudo gdb -p <PID>;Docker 里则需要 --cap-add=SYS_PTRACE(第 00 篇);
macOS 上 :attach 到别的进程需要开发者权限,首次使用时系统会弹窗要求授权,或者需要 sudo。系统自带的程序受系统完整性保护(SIP),不能 attach。
attach 成功后,进程会被暂停;调试结束时用 detach 让进程继续运行,或者直接 quit(GDB 会询问是否 detach)。
3.2 看所有线程在干什么 attach 之后,第一个命令就是 info threads:
1 2 3 4 5 (gdb) info threads Id Target Id Frame * 1 Thread 0xffffb280fc00 (LWP 7) "threads.linux" __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/aarch64/syscall_cancel.S:50 2 Thread 0xffffb1a4f180 (LWP 10) "threads.linux" futex_wait (futex_word=0xaaaabb230020 <mtx_a>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126 3 Thread 0xffffb225f180 (LWP 9) "threads.linux" futex_wait (futex_word=0xaaaabb230050 <mtx_b>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126
GDB 很贴心地把地址翻译成了符号:线程 2 在等 <mtx_a>,线程 3 在等 <mtx_b>。futex_wait 是 Linux 上线程阻塞等待锁的底层系统调用。
再用 thread apply all bt 看每个线程完整的调用栈(省略了标准库内部的部分帧):
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 (gdb) thread apply all bt Thread 3 (Thread 0xffffb225f180 (LWP 9) "threads.linux"): #0 futex_wait (futex_word=0xaaaabb230050 <mtx_b>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126 #1 __GI___lll_lock_wait (futex=futex@entry=0xaaaabb230050 <mtx_b>, private=private@entry=0) at ./nptl/lowlevellock.c:49 ... #5 std::mutex::lock (this=0xaaaabb230050 <mtx_b>) at /usr/include/c++/15/bits/std_mutex.h:115 #6 0x0000aaaabb212758 in std::lock_guard<std::mutex>::lock_guard (this=0xffffb225e7b0, __m=...) at /usr/include/c++/15/bits/std_mutex.h:252 #7 0x0000aaaabb210fd0 in worker1 () at threads.cpp:12 ... Thread 2 (Thread 0xffffb1a4f180 (LWP 10) "threads.linux"): #0 futex_wait (futex_word=0xaaaabb230020 <mtx_a>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126 #1 __GI___lll_lock_wait (futex=futex@entry=0xaaaabb230020 <mtx_a>, private=private@entry=0) at ./nptl/lowlevellock.c:49 ... #5 std::mutex::lock (this=0xaaaabb230020 <mtx_a>) at /usr/include/c++/15/bits/std_mutex.h:115 #6 0x0000aaaabb212758 in std::lock_guard<std::mutex>::lock_guard (this=0xffffb1a4e7b0, __m=...) at /usr/include/c++/15/bits/std_mutex.h:252 #7 0x0000aaaabb2110d0 in worker2 () at threads.cpp:19 ... Thread 1 (Thread 0xffffb280fc00 (LWP 7) "threads.linux"): #0 __syscall_cancel_arch () at ../sysdeps/unix/sysv/linux/aarch64/syscall_cancel.S:50 ... #5 0x0000ffffb23bddb8 [PAC] in __pthread_clockjoin_ex (threadid=281473670574464, thread_return=0x0, clockid=0, abstime=0x0, cancel=<optimized out>) at ./nptl/pthread_join_common.c:67 #6 0x0000ffffb2627944 [PAC] in std::thread::join() () from /usr/lib/aarch64-linux-gnu/libstdc++.so.6 #7 0x0000aaaabb2111bc [PAC] in main () at threads.cpp:26
输出里的 [PAC] 是 ARM64 的指针认证(Pointer Authentication)标记,表示这个返回地址是被签名过的,在 x86_64 上不会出现。
用第 04 篇的读栈方法——找第一个属于自己代码的帧:
线程
自己代码的帧
在做什么
3(LWP 9)
#7 worker1 () at threads.cpp:12
在第 12 行等 mtx_b
2(LWP 10)
#7 worker2 () at threads.cpp:19
在第 19 行等 mtx_a
1(LWP 7)
#7 main () at threads.cpp:26
在 t1.join() 等线程 1 结束
3.3 找到锁的持有者 知道了谁在等 哪把锁,还要知道每把锁被谁拿着 。在 Linux 上,std::mutex 底层是 pthread_mutex_t,它的内部结构里记录了持有者的 LWP:
1 2 3 4 5 6 7 8 9 10 (gdb) thread 2 [Switching to thread 2 (Thread 0xffff8421f180 (LWP 10))] #0 futex_wait (futex_word=0xaaaae4850020 <mtx_a>, expected=2, private=0) at ../sysdeps/nptl/futex-internal.h:126 (gdb) frame 7 #7 0x0000aaaae48310d0 in worker2 () at threads.cpp:19 19 std::lock_guard<std::mutex> la(mtx_a); // 等 mtx_a —— 与 worker1 顺序相反 (gdb) print mtx_a._M_mutex.__data.__owner $1 = 9 (gdb) print mtx_b._M_mutex.__data.__owner $2 = 10
(这次 attach 的是重新运行的进程,所以地址和上面不同,LWP 编号恰好一致。)
把线索串起来:
1 2 线程 3(LWP 9,worker1) ──持有──▶ mtx_a ◀──等待── 线程 2(LWP 10,worker2) 线程 3(LWP 9,worker1) ──等待──▶ mtx_b ◀──持有── 线程 2(LWP 10,worker2)
一个完美的环:循环等待 。这就是死锁的铁证。
_M_mutex.__data.__owner 是 libstdc++ 和 glibc 的内部实现细节,不同平台不一样(macOS 的 libc++ 就没有这个字段)。但思路是通用的:先看每个线程在等什么 (调用栈),再想办法确认谁持有什么 。实在找不到持有者字段,就在每个线程里看它的调用栈上有哪些 lock_guard 已经构造完成。
3.4 LLDB 版 在 macOS 上 attach:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 (lldb) process attach --pid 7106 Process 7106 stopped * thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP ... (lldb) thread list Process 7106 stopped * thread #1: tid = 0x72b0c, 0x0000000188831af8 libsystem_kernel.dylib`__ulock_wait + 8, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP thread #2: tid = 0x72b2d, 0x00000001888329dc libsystem_kernel.dylib`__psynch_mutexwait + 8 thread #3: tid = 0x72b2e, 0x00000001888329dc libsystem_kernel.dylib`__psynch_mutexwait + 8 (lldb) bt all * thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGSTOP * frame #0: 0x0000000188831af8 libsystem_kernel.dylib`__ulock_wait + 8 frame #1: 0x0000000188876114 libsystem_pthread.dylib`_pthread_join + 616 frame #2: 0x00000001887aae24 libc++.1.dylib`std::__1::thread::join() + 36 frame #3: 0x000000010058896c threads.mac`main at threads.cpp:26:8 frame #4: 0x00000001884a84e4 dyld`start + 6992 thread #2 frame #0: 0x00000001888329dc libsystem_kernel.dylib`__psynch_mutexwait + 8 ... frame #4: 0x0000000100588a60 threads.mac`std::__1::lock_guard<std::__1::mutex>::lock_guard[abi:nqe220106](this=0x000000016f8fef10, __m=0x0000000100590040) at lock_guard.h:32:10 ... ﹍ frame #6: 0x0000000100588670 threads.mac`worker1() at threads.cpp:12:33 thread #3 frame #0: 0x00000001888329dc libsystem_kernel.dylib`__psynch_mutexwait + 8 ... frame #4: 0x0000000100588a60 threads.mac`std::__1::lock_guard<std::__1::mutex>::lock_guard[abi:nqe220106](this=0x000000016f98af10, __m=0x0000000100590000) at lock_guard.h:31:19 ... ﹍ frame #6: 0x00000001005888c4 threads.mac`worker2() at threads.cpp:19:33
macOS 上等待互斥锁的系统调用叫 __psynch_mutexwait。线程 2 在 worker1 第 12 行等地址 0x...590040 的锁,线程 3 在 worker2 第 19 行等 0x...590000 的锁(两个全局 mutex 相邻存放,前者是 mtx_b,后者是 mtx_a,可以用 p &mtx_a 确认)。结论和 Linux 上完全一致。
﹍ 和 ﹉ 符号表示 LLDB 把中间几个标准库的帧折叠了(第 7、8 帧)。切到自己代码的帧看源码:
1 2 3 4 5 6 7 8 9 10 (lldb) thread select 2 (lldb) frame select 6 frame #6: 0x0000000100588670 threads.mac`worker1() at threads.cpp:12:33 9 void worker1() { 10 std::lock_guard<std::mutex> la(mtx_a); 11 std::this_thread::sleep_for(std::chrono::milliseconds(100)); -> 12 std::lock_guard<std::mutex> lb(mtx_b); // 等 mtx_b ^ 13 std::puts("worker1 done"); 14 }
3.5 修复 死锁的根源是两个线程加锁顺序相反 。修复方法:
所有线程按统一的顺序 加锁(都先 mtx_a 再 mtx_b);
或者用 C++17 的 std::scoped_lock,一次性锁住多把锁,它内部用死锁避免算法:
1 2 3 4 void worker1 () { std::scoped_lock lock (mtx_a, mtx_b) ; std::puts ("worker1 done" ); }
四、数据竞争:ThreadSanitizer 回到第一个例子 race.cpp。计数器结果不对的原因是 ++counter 不是原子操作,它实际上是「读、加一、写回」三步,两个线程交错执行时会丢失更新。
这类 bug 用调试器很难抓:你在 ++counter 上下断点,线程就被暂停了,时序完全被改变;而且它不崩溃、不卡死,只是结果「偶尔不对」。专门的工具是 ThreadSanitizer(TSan) :
1 2 g++ -std=c++20 -g -O0 -fsanitize=thread race.cpp -o race.tsan ./race.tsan
1 2 3 4 5 6 7 8 9 10 11 12 13 14 ================== WARNING: ThreadSanitizer: data race (pid=12) Read of size 4 at 0xaaaae1b7001c by thread T2: #0 add_many() /work/race.cpp:8 (race.tsan+0x1098) ... Previous write of size 4 at 0xaaaae1b7001c by thread T1: #0 add_many() /work/race.cpp:8 (race.tsan+0x10b4) ... Location is global 'counter' of size 4 at 0xaaaae1b7001c (race.tsan+0x2001c) ... SUMMARY: ThreadSanitizer: data race /work/race.cpp:8 in add_many() ================== ... counter = 200000 (expect 200000)
(为了篇幅省略了 BuildId 和标准库的栈帧。)报告的结构和 ASan 很像:
部分
含义
Read of size 4 ... by thread T2
线程 T2 在 race.cpp:8 读 了这块内存
Previous write ... by thread T1
此前线程 T1 在 race.cpp:8 写 过同一块内存,两次访问之间没有任何同步
Location is global 'counter'
出问题的变量是全局变量 counter
这正是 ++counter 那一行。有意思的是,这次运行打印出了正确的 200000——TSan 的插桩改变了程序的执行节奏,竞争碰巧没有造成丢失更新。但 TSan 照样报告了问题 :它检测的是「两个线程访问同一内存且没有同步」这个事实本身,不依赖于这次运行是否真的算错了。这正是它比「多跑几次看结果」可靠的地方。
修复:把 counter 改成 std::atomic<int>,或者用 std::mutex 保护。
TSan 的注意事项:
不能和 ASan 同时使用;
运行开销比 ASan 大,程序可能慢 5~15 倍,内存占用也更多;
同样只能发现实际执行到 的竞争,所以测试要覆盖到并发路径。
五、多线程调试的几条经验
经验
说明
程序卡住,先抓全体线程的栈
thread apply all bt(LLDB:bt all)是定位死锁、卡死的第一步;可以把输出保存下来慢慢分析
不要急着重启
卡死的进程就是最好的现场,先 attach,看完再说
给线程起名字
info threads 里一眼能认出每个线程是干什么的
调试器会改变时序
断点、单步都会让线程停下来,竞争类 bug 可能因此消失。这类问题优先用 TSan
单步时锁住调度器
想专注观察一个线程时,set scheduler-locking step
崩溃在别的线程
多线程程序崩溃时,GDB / LLDB 会自动切到出问题的线程,直接 bt 即可;core dump 里也包含所有线程
一个实用的小脚本:不用进入交互界面,直接抓取某个进程所有线程的调用栈并保存下来:
1 gdb -p <PID> -batch -ex "thread apply all bt" > stacks.txt
-batch 表示执行完命令就退出(并 detach,进程继续运行)。线上服务偶发卡顿时,隔几秒抓一次,对比几份栈,就能看出线程长时间卡在哪里。
六、小结
需求
GDB
LLDB
列出线程
info threads
thread list
切换线程
thread N
thread select N
所有线程的调用栈
thread apply all bt
bt all
指定线程的断点
break 位置 thread N
br set ... -t 线程ID
单步时只跑当前线程
set scheduler-locking step
thread step-over --run-mode this-thread
attach 到进程
gdb -p PID
lldb -p PID / lldb -n 进程名
脱离进程
detach
detach
检测数据竞争
-fsanitize=thread
-fsanitize=thread
C++ 调试实战系列第 7 篇完。下一篇:IDE 图形化调试与速查表——VS Code 调试配置,以及 GDB / LLDB 命令对照总表 。