主题
故障排查
排故命令清单
遇到问题时的标准排查顺序:
bash
# 第一级:快速检查
openclaw status
# 第二级:Gateway 详细状态
openclaw gateway status
# 第三级:实时日志
openclaw logs --follow
# 第四级:深度诊断
openclaw doctor
# 第五级:通道连通性探测
openclaw channels status --probe健康信号
如果以下输出都正常,大概率不是 OpenClaw 本身的问题:
openclaw gateway status显示Runtime: running、Connectivity probe: okopenclaw doctor报告无阻塞性配置/服务问题openclaw channels status --probe显示各通道works或audit ok
查看日志
bash
# 实时查看 Gateway 日志
openclaw logs --follow
# 查看最近 100 条日志
openclaw logs --limit 100
# 日志文件位置(直接查看)
ls ~/.openclaw/logs/
# 或查看 Gateway 当前日志文件
# 日志路径可通过 openclaw gateway status 查看更新后故障
OpenClaw 更新后出现 Gateway 不启动、通道无响应、模型 401 错误的排查流程:
bash
openclaw status --all
openclaw update status --json
openclaw gateway status --deep
openclaw doctor --fix
openclaw gateway restart重点关注:
openclaw status中是否有Update restart待处理或失败- 插件加载失败:
plugin load failed: dependency tree corrupted→ 运行openclaw doctor --fix - Provider 401:
openclaw doctor --fix检查并清理过期的 Agent OAuth 认证
常见报错和解决
端口被占
bash
# 症状:Error: listen EADDRINUSE :::18789
# 查看谁在占用
lsof -i :18789
# 或
ss -tlnp | grep 18789
# 解决:停掉占用进程,或更改端口
openclaw config set gateway.port 18790
openclaw gateway restartAPI Key 无效
bash
# 症状:401 Unauthorized、403 Forbidden
# 检查环境变量
echo $DEEPSEEK_API_KEY
echo $ANTHROPIC_API_KEY
# 检查 .env 文件
cat ~/.openclaw/.env
# 查看 Gateway 日志确认具体 Provider
openclaw logs --follow | grep -E "401|403|auth|unauthorized"
# 重新配置
openclaw onboard --auth-choice deepseek-api-key模型超时
bash
# 症状:Request timeout、Gateway timeout
# 排查
openclaw models status # 检查模型连通性
openclaw logs --follow | grep timeout # 搜索超时日志
# 原因和解决
# 1. 网络问题 → 检查服务器外网连通性
curl -I https://api.deepseek.com
# 2. 模型繁忙 → 配置 fallback
openclaw config set agents.defaults.model.fallbacks '["openai/gpt-5.4"]'
# 3. 超时太短 → 调大超时配置
openclaw config set agents.defaults.timeoutMs 120000Ollama 本地模型连接失败
bash
# 检查 Ollama 是否运行
curl http://127.0.0.1:11434/v1/models
# 如果失败,启动 Ollama
ollama serve
# 检查模型是否已拉取
ollama list
# 如果模型不存在
ollama pull qwen3:14bSSL/HTTPS 问题
bash
# 症状:certificate verify failed、SSL error
# 排查
curl -v https://api.deepseek.com/v1/models
# 自签名证书/内网环境
# 设置环境变量跳过证书验证(不推荐生产环境)
export NODE_TLS_REJECT_UNAUTHORIZED=0
# 或配置自定义 CA 证书
export NODE_EXTRA_CA_CERTS=/path/to/ca-cert.pem分割安装(Split Brain)
更新后 Gateway 服务意外停止,表现为旧版 openclaw 二进制无法加载新版 openclaw.json:
bash
# 检查当前二进制版本
which openclaw
openclaw --version
# 检查配置最后写入版本
openclaw config get meta.lastTouchedVersion
# 修复 PATH 指向新版
# 然后重新安装 Gateway 服务
openclaw gateway install --force
openclaw gateway restart协议不匹配(协议回滚后)
降级后出现 protocol mismatch:
bash
openclaw --version
which -a openclaw
openclaw gateway status --deep
openclaw doctor --deep修复方法:
- 停止或重启旧版 OpenClaw 客户端进程(
gateway status --deep可看到 PID) - 重启嵌入 OpenClaw 的应用(dashboard、编辑器插件等)
- 重新运行
openclaw gateway status --deep确认旧进程已消失
通道收不到消息排查流程
mermaid
graph TD
A[收不到消息] --> B{通道状态正常?}
B -->|否| C[openclaw channels status --probe]
B -->|是| D{DM/群策略正确?}
D -->|否| E[检查 dmPolicy/groupPolicy 配置]
D -->|是| F{发送者在白名单?}
F -->|否| G[添加 allowFrom 或批准配对]
F -->|是| H{群聊需 @?}
H -->|是未@| I[@ 机器人重试]
H -->|否| J[查看实时日志]
J --> K[openclaw logs --follow]
K --> L[搜索关键词:收到消息/error/blocked]通道排查清单
bash
# 1. 检查通道配置
openclaw config get channels.feishu
# 2. 检查通道连通性
openclaw channels status --probe --channel feishu
# 3. 检查配对待审批
openclaw pairing list feishu
# 4. 发送测试消息同时查看日志
# 终端 A:
openclaw logs --follow
# 终端 B:从平台发送消息给 Bot
# 5. 检查 DM 策略和 allowFrom
openclaw config get channels.feishu.dmPolicy
openclaw config get channels.feishu.allowFrom经验:排除法
从链路前端往后排查,是最效率的排故策略:
发送者 → 平台服务器 → 通道连接 → Gateway → Agent → 模型 API → 工具执行
① ② ③ ④ ⑤ ⑥ ⑦| 环节 | 检查方法 |
|---|---|
| ① 发送者 | 确认消息已发出、发送者身份正确 |
| ② 平台服务器 | 飞书开放平台后台查看消息推送状态 |
| ③ 通道连接 | openclaw channels status --probe |
| ④ Gateway | openclaw gateway status、openclaw doctor |
| ⑤ Agent | 看日志中 Agent 是否收到消息并开始处理 |
| ⑥ 模型 API | openclaw models status、检查余额 |
| ⑦ 工具执行 | 看日志中 tool call 是否成功和报错信息 |
深度诊断命令
bash
# 综合健康检查
openclaw doctor --deep
# 带自动修复
openclaw doctor --fix
# 导出诊断报告
openclaw doctor --json > /tmp/openclaw-diagnosis.json
# 测试模型推理
openclaw infer model run --model deepseek/deepseek-v4-flash --prompt "hi" --json强制恢复
极端情况下的恢复操作:
bash
# 强制恢复方法
# 方案 1:重置 dev 环境(仅 --dev 模式有效)
openclaw --dev gateway --reset
# 方案 2:重新安装 Gateway 服务
openclaw gateway uninstall
openclaw gateway install --force
openclaw gateway start
# 方案 3:清理缓存后重装
rm -rf ~/.openclaw/cache/
openclaw doctor --repair
# 降级恢复(仅在必须降级时使用,需确认数据安全)
OPENCLAW_ALLOW_OLDER_BINARY_DESTRUCTIVE_ACTIONS=1 openclaw gateway install --force强制操作风险
--force 和 reset 操作可能丢失会话状态和缓存数据。仅在常规排故无效且确认可接受数据丢失时使用。
下一步
掌握了故障排查后,继续学习 实战案例 →
加载练习题中...
Gateway 配置故障(实战经验)
配置变更导致 Gateway 无法启动
症状: 修改 openclaw.json 后重启 Gateway,systemctl --user status 显示 failed,端口不监听,自己 session 消失。
根因分析:
- Schema 必填字段缺失 — 新增 provider 时只写了
baseUrl和apiKey,漏了models数组。models是 schema 必填字段,少一个都会导致配置校验失败。 - Agent 跑在 Gateway 进程内 — 从当前 session 发
systemctl restart等于 kill 自己。
修复步骤:
bash
# 1. 先备份当前配置(铁律)
cp openclaw.json openclaw.json.bak.$(date +%Y%m%d_%H%M%S)
# 2. 验证 JSON 合法性
python3 -c "import json; json.load(open('/root/.openclaw/openclaw.json'))"
# 3. 检查 provider schema(models 数组是否完整)
python3 -c "
import json
c = json.load(open('/root/.openclaw/openclaw.json'))
for k, v in c.get('models', {}).get('providers', {}).items():
has_models = 'models' in v and len(v.get('models', [])) > 0
print(f'{k}: models={"OK" if has_models else "MISSING"} ')
"
# 4. 修复配置后,让用户手动重启 Gateway(不要从自己 session 发重启)
systemctl --user restart openclaw-gateway.service教训:
- 改
openclaw.json前必须备份 - 改 provider 前先查 schema,
models是必填数组 - 重启 Gateway 前考虑自己是否跑在 Gateway 上
- 优先用官方内置方案(
localprovider),不自己造轮子
Memory 索引故障
症状: openclaw memory status 显示 0/65 files · 0 chunks、Dirty: yes、Vector search: paused。
常见原因和修复:
bash
# 1. 检查状态
openclaw memory status --agent main
# 2. 清理残留的 reindex 锁(上次索引异常中断留下的)
ls -la ~/.openclaw/agents/main/agent/*.reindex-lock*
ps aux | grep openclaw-memory # 确认没有活跃的索引进程
rm -f ~/.openclaw/agents/main/agent/*.reindex-lock*
# 3. 重建索引
openclaw memory index --force --agent main
# 4. 验证
openclaw memory status --agent main # 应显示 65/65 files, Dirty: no🎯 本章要点
- 排故先检查 Gateway 进程状态和端口 18789 是否监听
- 改配置前备份 + 验证 JSON + 查 schema
- 重启 Gateway 前确认自己不在 Gateway 进程里
- Memory 索引 Dirty 时先清残留锁再
--force重建