主题
📁 文件操作
程序跑完数据就没了?把数据存到文件里,下次启动还能用。这一章学 Python 怎么读写文件。
打开文件 — open()
open() 是读写文件的大门,两个关键参数:文件路径 和 模式。
python
# 模式一览
# 'r' 只读(默认),文件不存在报错
# 'w' 只写,文件存在则清空,不存在则创建
# 'a' 追加,文件存在则在末尾写,不存在则创建
# 'x' 新建,文件已存在则报错(安全创建)
# 'b' 二进制模式(配合 r/w/a 用,如 'rb'、'wb')
# 't' 文本模式(默认,配合 r/w/a)
# '+' 读写模式(如 'r+'、'w+')with — 自动关门
python
# ❌ 不推荐:手动关
f = open("test.txt", "w")
f.write("Hello")
f.close() # 容易忘!
# ✅ 推荐:with 自动关,即使出异常也会关
with open("test.txt", "w", encoding="utf-8") as f:
f.write("Hello, World!")永远用 with!
with 是一个上下文管理器,离开 with 块时自动调用 f.close()。不用担心忘记关闭文件。一定要指定 encoding="utf-8",否则可能出乱码。
读取文件
python
# 一次性读全部
with open("data.txt", "r", encoding="utf-8") as f:
content = f.read() # 整个文件一个字符串
print(content)
# 逐行读(小文件)
with open("data.txt", "r", encoding="utf-8") as f:
lines = f.readlines() # 返回列表,每行一个元素
for line in lines:
print(line.strip()) # strip() 去掉末尾换行符
# 逐行读(大文件 — 推荐!)
with open("huge.log", "r", encoding="utf-8") as f:
for line in f: # 一行一行地读,不占内存
if "ERROR" in line:
print(line.strip())大文件怎么读?
文件几百 MB 甚至上 GB 时,read() 会爆内存。用 for line in f 逐行迭代,Python 只缓存一行,内存占用极低。
按块读取
python
with open("large.bin", "rb") as f:
while chunk := f.read(8192): # 每次读 8KB
process(chunk)写入文件
python
lines = ["第一行\n", "第二行\n", "第三行\n"]
with open("output.txt", "w", encoding="utf-8") as f:
# 方式一:逐条写
for line in lines:
f.write(line)
# 方式二:批量写
f.writelines(lines)write vs writelines
write() 写一个字符串,writelines() 写一个字符串列表但不会自动加换行符,需要自己加 \n。
JSON — 结构化数据存储
JSON 是数据交换的事实标准,Python 内置 json 模块:
python
import json
data = {
"name": "小雅",
"age": 25,
"hobbies": ["摄影", "编程", "旅行"],
"is_student": False
}
# 字典/列表 → JSON 字符串
json_str = json.dumps(data, indent=2, ensure_ascii=False)
print(json_str)
# 字典/列表 → JSON 文件
with open("data.json", "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
# JSON 字符串 → 字典/列表
parsed = json.loads(json_str)
print(parsed["name"]) # 小雅
# JSON 文件 → 字典/列表
with open("data.json", "r", encoding="utf-8") as f:
loaded = json.load(f)
print(loaded["hobbies"]) # ['摄影', '编程', '旅行']ensure_ascii=False 很重要
不加这个参数,中文会被转成 \u5c0f\u96c5 这样谁也看不懂的东西。写了 JSON 永远带上 ensure_ascii=False。
CSV — 表格数据处理
python
import csv
# 写入 CSV
with open("users.csv", "w", newline="", encoding="utf-8-sig") as f:
writer = csv.writer(f)
writer.writerow(["姓名", "年龄", "城市"]) # 表头
writer.writerow(["小明", 25, "北京"])
writer.writerow(["小红", 28, "上海"])
# 读取 CSV
with open("users.csv", "r", encoding="utf-8-sig") as f:
reader = csv.reader(f)
for row in reader:
print(row) # ['姓名', '年龄', '城市'] ...
# DictReader — 用字典方式读,更直观
with open("users.csv", "r", encoding="utf-8-sig") as f:
reader = csv.DictReader(f)
for row in reader:
print(f"{row['姓名']}住在{row['城市']}")
# DictWriter — 用字典方式写
fieldnames = ["姓名", "年龄", "城市"]
with open("users2.csv", "w", newline="", encoding="utf-8-sig") as f:
writer = csv.DictWriter(f, fieldnames=fieldnames)
writer.writeheader()
writer.writerow({"姓名": "小雅", "年龄": 25, "城市": "杭州"})encoding="utf-8-sig" 是什么?
utf-8-sig 会在文件开头加 BOM(Byte Order Mark),让 Excel 正确识别 UTF-8 中文。普通文本用 utf-8 就行,CSV 要给 Excel 用就加 -sig。
pathlib — 优雅的路径操作
告别 os.path.join() 的噩梦,用 / 拼接路径:
python
from pathlib import Path
# 创建 Path 对象
data_dir = Path("data")
file_path = data_dir / "users.csv" # 用 / 拼接!
# 基本操作
print(file_path.name) # users.csv(文件名)
print(file_path.stem) # users(不含扩展名)
print(file_path.suffix) # .csv(扩展名)
print(file_path.parent) # data(父目录)
# 检查与创建
if not data_dir.exists():
data_dir.mkdir(parents=True) # 递归创建目录
print(file_path.is_file()) # 是文件?
print(file_path.is_dir()) # 是目录?
# 读写
content = file_path.read_text(encoding="utf-8") # 一次读全部文本
file_path.write_text("Hello", encoding="utf-8") # 写文本
data = file_path.read_bytes() # 读二进制
# 遍历目录
for py_file in Path(".").glob("*.py"): # 所有 .py 文件
print(py_file)
for py_file in Path(".").rglob("*.py"): # 递归(含子目录)
print(py_file)实战:批量处理日志 + 合并 CSV
python
# /opt/kb/docs/python/file_io.py
"""实战:批量处理日志文件、合并 CSV"""
from pathlib import Path
import csv
import json
def scan_errors(log_dir="logs"):
"""扫描日志目录下所有 .log 文件,收集错误行"""
errors = []
for log_file in Path(log_dir).glob("*.log"):
with open(log_file, "r", encoding="utf-8") as f:
for i, line in enumerate(f, 1):
if "ERROR" in line:
errors.append({
"file": log_file.name,
"line": i,
"content": line.strip()
})
return errors
def merge_csv(output="merged.csv", *csv_files):
"""合并多个 CSV 文件"""
all_rows = []
header = None
for filepath in csv_files:
with open(filepath, "r", encoding="utf-8-sig") as f:
reader = csv.reader(f)
file_header = next(reader)
if header is None:
header = file_header
all_rows.extend(reader)
with open(output, "w", newline="", encoding="utf-8-sig") as f:
writer = csv.writer(f)
writer.writerow(header)
writer.writerows(all_rows)
print(f"合并完成:{len(all_rows)} 行 → {output}")
if __name__ == "__main__":
# 示例用法
errs = scan_errors()
print(f"发现 {len(errs)} 条错误")
# merge_csv("all_users.csv", "users1.csv", "users2.csv")常见坑
坑1:忘记用 with
不用 with,出了异常文件就没关。程序长时间跑可能耗尽文件句柄,导致 Too many open files 错误。
坑2:编码不匹配
Windows 的记事本默认保存为 GBK 编码。用 utf-8 读 GBK 文件会乱码或报错:UnicodeDecodeError。解决方法:
- 不确定编码时用
encoding="gbk"试试 - 或用
errors="ignore"跳过无法解码的字节(不推荐) - 最佳实践:统一用 UTF-8
坑3:CSV 写入没加 newline=""
Windows 上 open("file.csv", "w") 会导致行之间多一个空行。加上 newline="" 就正常了。
练习题
- 读一个文本文件,统计每个单词出现次数,输出 TOP 10
- 创建 JSON 文件存 5 个联系人的姓名和电话,然后读出来打印
- 写程序扫描一个目录,找出所有
.jpg文件并列出来(用pathlib) - 有 3 个 CSV 文件(分别记账 1 月/2 月/3 月),合并成一个并计算总支出
- 写一个日志分析脚本:找出今天日志中所有的 ERROR 和 WARNING
📁 文件操作学会了,下一步学 异常处理,让程序遇到错误也能优雅应对。
🎯 本章要点
- with 是上下文管理器,无论 try 块里是否抛异常
- with + open + read 是最标准的读文件模式。
- 'x' = exclusive creation
加载练习题中...