主题
systemd 服务管理
systemd 是现代 Linux 的初始化系统和服务管理器。它替代了 SysV init,几乎所有主流发行版(Ubuntu 15.04+、Debian 8+、RHEL 7+)都使用它。掌握 systemd = 掌握服务的生命周期管理。
核心概念
systemd 把一切抽象为 Unit,不同类型用不同后缀:
| Unit 类型 | 用途 | 示例 |
|---|---|---|
.service | 系统服务 | nginx.service |
.timer | 定时任务(替代 cron) | logrotate.timer |
.socket | 套接字激活 | docker.socket |
.mount | 挂载点管理 | home.mount |
.target | 运行级别/状态组 | multi-user.target |
.device | 设备管理 | dev-sda.device |
systemctl 命令速查
bash
# === 服务生命周期 ===
sudo systemctl start nginx # 启动
sudo systemctl stop nginx # 停止
sudo systemctl restart nginx # 重启(stop + start)
sudo systemctl reload nginx # 重载配置(不重启,发 SIGHUP)
sudo systemctl reload-or-restart nginx # 支持 reload 就 reload,否则 restart
# === 状态查询 ===
systemctl status nginx # 服务状态 + 最近 10 行日志
systemctl is-active nginx # 是否在运行
systemctl is-enabled nginx # 是否开机自启
systemctl is-failed nginx # 是否启动失败
# === 开机自启 ===
sudo systemctl enable nginx # 设为开机自启
sudo systemctl disable nginx # 取消开机自启
sudo systemctl enable --now nginx # 设开机自启 + 立即启动
# === 全局操作 ===
systemctl list-units # 列出所有活动的 unit
systemctl list-units --all # 包括未活动的
systemctl list-units --type=service # 只看 service 类型
systemctl list-unit-files # 所有已安装的 unit 文件
systemctl daemon-reload # 修改 unit 文件后重新加载
systemctl list-timers # 查看所有定时器编写自己的 systemd 服务
以 Python Web 应用为例,创建 /etc/systemd/system/myapp.service:
ini
[Unit]
Description=My Python Web Application
After=network.target postgresql.service
# After: 在这些服务之后启动(不保证它们完全就绪!)
[Service]
Type=simple
User=myapp
Group=myapp
WorkingDirectory=/opt/myapp
EnvironmentFile=/opt/myapp/.env
Environment="FLASK_ENV=production"
ExecStart=/opt/myapp/venv/bin/gunicorn -w 4 -b 0.0.0.0:8000 app:app
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
RestartSec=5
# 安全加固
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=strict
ProtectHome=yes
ReadWritePaths=/opt/myapp/data /var/log/myapp
[Install]
WantedBy=multi-user.targetbash
# === 前置步骤 ===
# 1. 创建专用服务用户(不允许登录)
sudo useradd -r -s /bin/false myapp
# 2. 创建数据和日志目录
sudo mkdir -p /opt/myapp/data /var/log/myapp
# 3. 安装 gunicorn(在虚拟环境中)
python3 -m venv /opt/myapp/venv
/opt/myapp/venv/bin/pip install gunicorn
# 4. 写入环境变量文件
cat > /opt/myapp/.env << 'EOF'
DB_HOST=localhost
DB_PORT=5432
FLASK_ENV=production
EOF
# 5. 设置目录权限
sudo chown -R myapp:myapp /opt/myapp /var/log/myapp
# === 部署流程 ===
sudo cp myapp.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now myapp
sudo systemctl status myappUnit 配置详解
| 指令 | 含义 | 推荐值 |
|---|---|---|
Type | simple(默认)/ forking / oneshot / notify | simple 适用于大多数程序 |
User / Group | 以哪个用户运行 | 非 root,专用服务用户 |
EnvironmentFile | 从文件加载环境变量 | 敏感信息用 .env 文件 |
ExecStart | 启动命令(必须用绝对路径) | — |
ExecReload | 重载命令(systemctl reload 触发) | — |
Restart | no / on-success / on-failure / on-abnormal / on-watchdog / on-abort / always | on-failure(异常退出或被杀才重启) |
RestartSec | 重启前等待秒数 | 推荐 3-10 秒(默认 100ms) |
TimeoutStopSec | 停止超时 | 90s(默认),长任务增大 |
WorkingDirectory | 工作目录 | 设为项目根目录 |
PrivateTmp | 独立的 /tmp 目录 | yes |
Restart 策略选择
| 策略 | 触发条件 |
|---|---|
no | 永不自动重启(默认) |
on-success | 进程正常退出(退出码 0)且未被信号终止时重启 |
on-failure | 进程异常退出(退出码非 0)或被信号杀死时重启【推荐】 |
on-abnormal | 进程因信号异常终止时重启(如 SIGABRT、SIGSEGV),不包含看门狗超时 |
on-watchdog | 看门狗超时时重启(需配合 WatchdogSec= 使用) |
on-abort | 进程因未捕获信号(非 SIGHUP/SIGINT/SIGTERM/SIGPIPE)退出时重启 |
always | 任何情况都重启,包括正常退出 |
Restart=always 需配合启动频率限制
Restart=always 会让正常退出的进程也被无限重启。更关键的隐患是 systemd 默认有启动频率限制(DefaultStartLimitBurst=5, DefaultStartLimitIntervalSec=10s),即 10 秒内启动超过 5 次,服务会被彻底停止。如果期望快速重启失败的服务,需显式设置:
ini
StartLimitBurst=10
StartLimitIntervalSec=10journalctl 查日志
bash
# === 基础用法 ===
journalctl -u nginx # 查看 nginx 服务日志
journalctl -u nginx -f # 实时跟踪(tail -f 模式)
journalctl -u nginx --since today # 今天的日志
journalctl -u nginx --since "2026-07-26 10:00" --until "2026-07-26 12:00"
# === 高级过滤 ===
# 按优先级
journalctl -p err # 只看 ERROR 及以上
journalctl -p warning # WARNING + ERROR + CRIT
# 按启动记录(-b)
journalctl -b # 本次启动以来的日志
journalctl -b -1 # 上次启动
journalctl --list-boots # 列出所有启动记录
# 内核日志
journalctl -k # 等同于 dmesg
# 组合使用
journalctl -u nginx -p err --since "1 hour ago" # 最近1小时 nginx 的错误
journalctl -u nginx -u postgresql -f # 同时看两个服务Timer 定时器(cron 替代方案)
bash
# 创建 service(实际要运行的命令)
# /etc/systemd/system/log-cleanup.service
[Unit]
Description=Clean up old logs
[Service]
Type=oneshot
ExecStart=/usr/local/bin/cleanup-logs.sh
# 创建 timer(调度规则)
# /etc/systemd/system/log-cleanup.timer
[Unit]
Description=Daily log cleanup
[Timer]
OnCalendar=daily
# OnCalendar=*-*-* 03:00:00 # 每天凌晨3点(二选一,只能有一个 OnCalendar= 生效)
# OnBootSec=5min # 开机5分钟后(可与 OnCalendar= 并存)
Persistent=true # 错过时间窗口后补执行
[Install]
WantedBy=timers.target
# 启用
sudo systemctl enable --now log-cleanup.timer
systemctl list-timers log-cleanup.timer常见坑与排故
坑 1:Type=forking 配错
bash
# forking: 程序启动后 fork 到后台(如 nginx/php-fpm 传统模式)
# 必须配合 PIDFile 使用,否则 systemd 不知道谁是主进程
# 错误写法:
Type=forking
ExecStart=/usr/local/nginx/sbin/nginx # 缺少 PIDFile!
# 正确写法:
Type=forking
PIDFile=/run/nginx.pid
ExecStart=/usr/local/nginx/sbin/nginxType 选择指南
- 你的程序是前台运行的(如 gunicorn/npm start)→
Type=simple - 你的程序自己 daemonize(如传统 nginx)→
Type=forking+PIDFile - 你的程序支持 sd_notify →
Type=notify - 一次性任务 →
Type=oneshot
坑 2:环境变量不生效
bash
# 问题:在 .bashrc 或 /etc/environment 设了变量,但 systemd 服务读不到
# 原因:systemd 有自己的环境,不继承用户 shell
# 解决:
# 方法1:在 unit 文件中声明
Environment="DB_HOST=localhost"
Environment="DB_PORT=5432"
# 方法2:用 EnvironmentFile(推荐)
EnvironmentFile=/opt/myapp/.env
# .env 文件格式:KEY=value(不支持 export 语法)
# ⚠️ systemd v254+ 会自动去除引号,旧版本不会
# DB_HOST=localhost ✓
# export DB_HOST=localhost ✗
# DB_HOST="localhost" ✓ v254+ 自动去引号;旧版会把引号当值的一部分坑 3:daemon-reload 忘了执行
bash
# 每次修改 unit 文件后必须执行,否则不生效!
sudo systemctl daemon-reload
sudo systemctl restart myapp坑 4:服务启动就退出,状态显示 "inactive (dead)"
bash
# 调试步骤:
# 1. 看完整日志
journalctl -u myapp --since "1 min ago"
# 2. 手动跑一遍 ExecStart 命令,看报什么错
/opt/myapp/venv/bin/gunicorn -w 4 app:app
# 3. 检查 WorkingDirectory 是否存在、User 是否有权限
sudo -u myapp ls /opt/myapp
# 4. 检查是否 80% 因为路径不对或环境变量缺失生产环境经验
- 不要用 root 跑服务:每个服务建专用用户,最小权限原则
- Restart=on-failure 而非 always:正常退出的程序不应被无限重启。若用 always,务必设置
StartLimitBurst和StartLimitIntervalSec防止速率限制导致服务被彻底停掉 - TimeoutStopSec 适当增大:数据库等慢关服务设 120s+
- 日志持久化:
/etc/systemd/journald.conf设Storage=persistent - 限制日志大小:
SystemMaxUse=500M防止/var/log/journal撑爆磁盘 - 用 systemd-analyze 诊断启动:
systemd-analyze blame看哪个服务启动慢
加载练习题中...