State
On the Linux experimental server 1.30.164014, SIGTERM starts a normal-looking shutdown (the script module and the game are destroyed within about 2 seconds), but the process never exits. Its #Main thread then busy-loops in user space at about 100 % of one core, without a single syscall, until it is killed with SIGKILL. The stable server 1.29.163709 run the same way logs
--- Termination successfully completed --- and exits within 2 seconds.
Cause (found with gdb and confirmed by experiment): after the game is destroyed, the main thread calls
getc(stdin) in a loop, a console-command loop that only ends when the line quit is read. If stdin is not a terminal and is at end-of-file /dev/null, a closed pipe, a container without -i, a systemd service), getc returns EOF immediately and forever (glibc keeps EOF sticky, so not even a read syscall happens): the thread spins at 100 % CPU and never reaches the exit. With stdin an open pipe the thread blocks at 0 % CPU until quit is written, and then the server exits. The 1.29 server does not enter this loop on shutdown.
This makes a normal killsystemctl stoppodman stop of the experimental server take the full stop timeout (we use 120 s) and end in SIGKILL.
Reproduced without any mod, not as PID 1, not in a container, on a plain DayZServer started from a shell.
Server (affected): Steam app 1042420 "DayZ Server Exp", build id 25319221, Version 1.30.164014.27
Server (works): Steam app 223350 "DayZServer", build id 24570360, Version 1.29.163709
OS: Debian 13 (trixie), Linux 7.2.8, x86_64, 2 vCPU of an AMD Ryzen 7 7840HS, 16 GB RAM, btrfs |
Run as an unprivileged user, from a shell, no container, no mods, no -servermod-mod
Mission: dayzOffline.chernarusplus (the one shipped in the server's mpmissions/) |
Command line (identical for both builds apart from the ports):
./DayZServer -config=serverDZ.cfg -port=44800 -profiles=prof -dologs -nosplash -nopause
serverDZ.cfg is the stock file with a hostname, a password, maxPlayers = 10, steamQueryPort = 44810
and template = "dayzOffline.chernarusplus". No player ever connects.
Experimental, SIGTERM sent at 03:06:31:
03:06:32.077 SCRIPT : ~DayZGame()
03:06:32 ENGINE : Destroying game
03:06:33.394 SCRIPT : Cleaning up script module globals 'Game'
03:06:33.394 SCRIPT (E): Leaked 'BunkerBroadcastManager' script instance (1x)!
03:06:33.394 SCRIPT (E): ==== Total Leaks (1x)! ====
(nothing more, ever)Stable, SIGTERM sent at 03:09:52:
03:09:53 ENGINE : Destroying game
03:09:54 --- Termination successfully completed ---
(process exited, status 255)Side remark: the stable server exits with status 255 after a clean, requested termination.
Measured on the vanilla native run, 12 s after SIGTERM:
#Main thread: state R (running). 495 CPU ticks in 5 s (about 99 % of a core). **0 voluntary context switches in 5 s**, compared with 80 before the signal.
strace -f -p <pid> for 8 s (all 18 remaining threads attached, 1875 lines): **the main thread makes no syscall at all**. It is spinning in user space.
All other threads are idle: DebugThreadWatc, enfResourceLoad, rvFileServer*, bankSignatureCh, enfWork*, CFileWriterThre, SteamEngineWatc, CHTTP*, CJobMgr::m_Work, CNet Encrypt wait in futex; IOCP Thread waits in epoll_wait.
- One Steam thread IPC:CSteamEngin…) keeps polling: epoll_wait(…, 45–50 ms), then read(eventfd) returning
EAGAIN, then poll(fd, 0), repeatedly, plus periodic access("/sys/kernel/debug/tracing/trace_marker") returning EACCES. That is a normal idle Steam IPC loop, not a blocked one.
No signal is received or sent in the trace, no file is written and no network syscall sendtorecvfromconnect) appears in it.
So the shutdown logic finishes its visible work and the main thread then sits in a user-space loop. Not a
deadlock: see the gdb section for the loop.
SIGTERM)gdb -p <pid> -batch -ex "thread 1" -ex "bt 40" (the server binary is stripped, so only libc symbols resolve; the offsets are in the main executable, BuildID 1cc8e9950544725a327c26cc779da501…):
#0 0x00007ff8f4ba86ab in __uflow () from /lib/x86_64-linux-gnu/libc.so.6
#1 0x00007ff8f4ba2fe8 in getc () from /lib/x86_64-linux-gnu/libc.so.6
#2 0x0000000000c10344 in ?? ()
#3 0x0000000000bbee3f in ?? ()
#4 0x00000000005faad5 in ?? ()
#5 0x00000000012a57f3 in ?? ()
#6 0x00000000012aff0a in ?? ()
#7 0x0000000000d20ddf in ?? ()
#8 0x00000000012b01e5 in ?? ()
#9 0x000000000043f044 in ?? ()
#10 0x00007ff8f4b45ca8 in ?? () from /lib/x86_64-linux-gnu/libc.so.6
#11 0x00007ff8f4b45d65 in __libc_start_main () from /lib/x86_64-linux-gnu/libc.so.6
#12 0x00000000004dee7a in ?? ()Three captures 2 s apart are all inside getc, always called from 0xc10344 (in the later captures gdb names that frame db (), a symbol that is left in the binary). All other threads wait in pthread_cond_timedwait or epoll_wait
Same run, only stdin differs. After SIGTERM and 12 s:
Stdin: /dev/null (EOF), result: state R, 274 CPU ticks in 3 s, no syscall, never exits
Stdin: open pipe, no data, result: blocked in anon_pipe_read, 0 CPU ticks, waits
Stdin: open pipe, after writing an empty line, still blocked in anon_pipe_read , waits
Stdin: open pipe, after writing quit : process exits
So the shutdown ends in a command loop that reads stdin and leaves on quit, and a stdin at EOF turns it into a busy loop. Suggested fix: do not enter the stdin command loop during a signal-initiated shutdown (as 1.29 does not), or stop reading when getc returns EOF.
Operators stop servers with SIGTERM (systemd, podman, docker, kill). On 1.30 experimental this always ends in the supervisor's SIGKILL after its timeout (here 2 minutes), and a script that waits for the process to exit hangs. The server also does not log any persistence save during this shutdown.
Workaround for operators: run the experimental server with a stdin that stays open and write quit after SIGTERM, or kill it with SIGKILL once Destroying game has been logged.
Start the experimental server as above and wait until script_*.log shows the mission scripts (about 1 minute), then 10 more seconds.
kill -TERM <pid>.
Watch: the process does not exit. ps shows it running at 100 % CPU.
The process exits (as the 1.29 server does).
Actual: it keeps running until SIGKILL.
Activity
The hang had a different cause than I reported. I had said the server read a quit command on stdin. The 1.30 experimental server raises Assertion failed ... Script is leaking! at shutdown and waits for an answer to (A)bort (R)etry (I)gnore on stdin. The vanilla BunkerBroadcastManager leak triggers it - quit only worked because its third letter is i.
You are not signed in. Please sign in to see more details and to reply.
State