Skip to content

systemd 服务管理 ​

systemd 是现代 Linux 的初始化系统和服务管理器。它替代了 SysV init,几乎所有主流发行版(Ubuntu 15.04+、Debian 8+、RHEL 7+)都使用它。掌握 systemd = 掌握服务的生命周期管理。

核心概念 ​

systemd 把一切抽象为 Unit,不同类型用不同后缀:

Unit 类型用途示例
.service系统服务nginx.service
.timer定时任务(替代 cron)logrotate.timer
.socket套接字激活docker.socket
.mount挂载点管理home.mount
.target运行级别/状态组multi-user.target
.device设备管理dev-sda.device

systemctl 命令速查 ​

bash
# === 服务生命周期 ===
sudo systemctl start nginx          # 启动
sudo systemctl stop nginx           # 停止
sudo systemctl restart nginx        # 重启(stop + start)
sudo systemctl reload nginx         # 重载配置(不重启,发 SIGHUP)
sudo systemctl reload-or-restart nginx   # 支持 reload 就 reload,否则 restart

# === 状态查询 ===
systemctl status nginx              # 服务状态 + 最近 10 行日志
systemctl is-active nginx           # 是否在运行
systemctl is-enabled nginx          # 是否开机自启
systemctl is-failed nginx           # 是否启动失败

# === 开机自启 ===
sudo systemctl enable nginx         # 设为开机自启
sudo systemctl disable nginx        # 取消开机自启
sudo systemctl enable --now nginx   # 设开机自启 + 立即启动

# === 全局操作 ===
systemctl list-units                # 列出所有活动的 unit
systemctl list-units --all          # 包括未活动的
systemctl list-units --type=service # 只看 service 类型
systemctl list-unit-files           # 所有已安装的 unit 文件
systemctl daemon-reload             # 修改 unit 文件后重新加载
systemctl list-timers               # 查看所有定时器

编写自己的 systemd 服务 ​

以 Python Web 应用为例,创建 /etc/systemd/system/myapp.service:

ini
[Unit]
Description=My Python Web Application
After=network.target postgresql.service
# After: 在这些服务之后启动(不保证它们完全就绪!)

[Service]
Type=simple
User=myapp
Group=myapp
WorkingDirectory=/opt/myapp
EnvironmentFile=/opt/myapp/.env
Environment="FLASK_ENV=production"
ExecStart=/opt/myapp/venv/bin/gunicorn -w 4 -b 0.0.0.0:8000 app:app
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
RestartSec=5
# 安全加固
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=strict
ProtectHome=yes
ReadWritePaths=/opt/myapp/data /var/log/myapp

[Install]
WantedBy=multi-user.target
bash
# === 前置步骤 ===
# 1. 创建专用服务用户(不允许登录)
sudo useradd -r -s /bin/false myapp

# 2. 创建数据和日志目录
sudo mkdir -p /opt/myapp/data /var/log/myapp

# 3. 安装 gunicorn(在虚拟环境中)
python3 -m venv /opt/myapp/venv
/opt/myapp/venv/bin/pip install gunicorn

# 4. 写入环境变量文件
cat > /opt/myapp/.env << 'EOF'
DB_HOST=localhost
DB_PORT=5432
FLASK_ENV=production
EOF

# 5. 设置目录权限
sudo chown -R myapp:myapp /opt/myapp /var/log/myapp

# === 部署流程 ===
sudo cp myapp.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now myapp
sudo systemctl status myapp

Unit 配置详解 ​

指令含义推荐值
Typesimple(默认)/ forking / oneshot / notifysimple 适用于大多数程序
User / Group以哪个用户运行非 root,专用服务用户
EnvironmentFile从文件加载环境变量敏感信息用 .env 文件
ExecStart启动命令(必须用绝对路径)—
ExecReload重载命令(systemctl reload 触发)—
Restartno / on-success / on-failure / on-abnormal / on-watchdog / on-abort / alwayson-failure(异常退出或被杀才重启)
RestartSec重启前等待秒数推荐 3-10 秒(默认 100ms)
TimeoutStopSec停止超时90s(默认),长任务增大
WorkingDirectory工作目录设为项目根目录
PrivateTmp独立的 /tmp 目录yes

Restart 策略选择 ​

策略触发条件
no永不自动重启(默认)
on-success进程正常退出(退出码 0)且未被信号终止时重启
on-failure进程异常退出(退出码非 0)或被信号杀死时重启【推荐】
on-abnormal进程因信号异常终止时重启(如 SIGABRT、SIGSEGV),不包含看门狗超时
on-watchdog看门狗超时时重启(需配合 WatchdogSec= 使用)
on-abort进程因未捕获信号(非 SIGHUP/SIGINT/SIGTERM/SIGPIPE)退出时重启
always任何情况都重启,包括正常退出

Restart=always 需配合启动频率限制

Restart=always 会让正常退出的进程也被无限重启。更关键的隐患是 systemd 默认有启动频率限制(DefaultStartLimitBurst=5, DefaultStartLimitIntervalSec=10s),即 10 秒内启动超过 5 次,服务会被彻底停止。如果期望快速重启失败的服务,需显式设置:

ini
StartLimitBurst=10
StartLimitIntervalSec=10

journalctl 查日志 ​

bash
# === 基础用法 ===
journalctl -u nginx                  # 查看 nginx 服务日志
journalctl -u nginx -f               # 实时跟踪(tail -f 模式)
journalctl -u nginx --since today    # 今天的日志
journalctl -u nginx --since "2026-07-26 10:00" --until "2026-07-26 12:00"

# === 高级过滤 ===
# 按优先级
journalctl -p err                     # 只看 ERROR 及以上
journalctl -p warning                 # WARNING + ERROR + CRIT

# 按启动记录(-b)
journalctl -b                         # 本次启动以来的日志
journalctl -b -1                      # 上次启动
journalctl --list-boots               # 列出所有启动记录

# 内核日志
journalctl -k                         # 等同于 dmesg

# 组合使用
journalctl -u nginx -p err --since "1 hour ago"    # 最近1小时 nginx 的错误
journalctl -u nginx -u postgresql -f               # 同时看两个服务

Timer 定时器(cron 替代方案) ​

bash
# 创建 service(实际要运行的命令)
# /etc/systemd/system/log-cleanup.service
[Unit]
Description=Clean up old logs

[Service]
Type=oneshot
ExecStart=/usr/local/bin/cleanup-logs.sh

# 创建 timer(调度规则)
# /etc/systemd/system/log-cleanup.timer
[Unit]
Description=Daily log cleanup

[Timer]
OnCalendar=daily
# OnCalendar=*-*-* 03:00:00    # 每天凌晨3点(二选一,只能有一个 OnCalendar= 生效)
# OnBootSec=5min               # 开机5分钟后(可与 OnCalendar= 并存)
Persistent=true                 # 错过时间窗口后补执行

[Install]
WantedBy=timers.target

# 启用
sudo systemctl enable --now log-cleanup.timer
systemctl list-timers log-cleanup.timer

常见坑与排故 ​

坑 1:Type=forking 配错 ​

bash
# forking: 程序启动后 fork 到后台(如 nginx/php-fpm 传统模式)
# 必须配合 PIDFile 使用,否则 systemd 不知道谁是主进程

# 错误写法:
Type=forking
ExecStart=/usr/local/nginx/sbin/nginx    # 缺少 PIDFile!

# 正确写法:
Type=forking
PIDFile=/run/nginx.pid
ExecStart=/usr/local/nginx/sbin/nginx

Type 选择指南

  • 你的程序是前台运行的(如 gunicorn/npm start)→ Type=simple
  • 你的程序自己 daemonize(如传统 nginx)→ Type=forking + PIDFile
  • 你的程序支持 sd_notify → Type=notify
  • 一次性任务 → Type=oneshot

坑 2:环境变量不生效 ​

bash
# 问题:在 .bashrc 或 /etc/environment 设了变量,但 systemd 服务读不到

# 原因:systemd 有自己的环境,不继承用户 shell

# 解决:
# 方法1:在 unit 文件中声明
Environment="DB_HOST=localhost"
Environment="DB_PORT=5432"

# 方法2:用 EnvironmentFile(推荐)
EnvironmentFile=/opt/myapp/.env
# .env 文件格式:KEY=value(不支持 export 语法)
# ⚠️ systemd v254+ 会自动去除引号,旧版本不会
# DB_HOST=localhost        ✓
# export DB_HOST=localhost ✗
# DB_HOST="localhost"      ✓ v254+ 自动去引号;旧版会把引号当值的一部分

坑 3:daemon-reload 忘了执行 ​

bash
# 每次修改 unit 文件后必须执行,否则不生效!
sudo systemctl daemon-reload
sudo systemctl restart myapp

坑 4:服务启动就退出,状态显示 "inactive (dead)" ​

bash
# 调试步骤:
# 1. 看完整日志
journalctl -u myapp --since "1 min ago"

# 2. 手动跑一遍 ExecStart 命令,看报什么错
/opt/myapp/venv/bin/gunicorn -w 4 app:app

# 3. 检查 WorkingDirectory 是否存在、User 是否有权限
sudo -u myapp ls /opt/myapp

# 4. 检查是否 80% 因为路径不对或环境变量缺失

生产环境经验 ​

  • 不要用 root 跑服务:每个服务建专用用户,最小权限原则
  • Restart=on-failure 而非 always:正常退出的程序不应被无限重启。若用 always,务必设置 StartLimitBurst 和 StartLimitIntervalSec 防止速率限制导致服务被彻底停掉
  • TimeoutStopSec 适当增大:数据库等慢关服务设 120s+
  • 日志持久化:/etc/systemd/journald.conf 设 Storage=persistent
  • 限制日志大小:SystemMaxUse=500M 防止 /var/log/journal 撑爆磁盘
  • 用 systemd-analyze 诊断启动:systemd-analyze blame 看哪个服务启动慢
加载练习题中...

有问题或补充?欢迎留言