Skip to content

📁 文件操作 ​

程序跑完数据就没了?把数据存到文件里,下次启动还能用。这一章学 Python 怎么读写文件。

打开文件 — open() ​

open() 是读写文件的大门,两个关键参数:文件路径 和 模式。

python
# 模式一览
# 'r'  只读(默认),文件不存在报错
# 'w'  只写,文件存在则清空,不存在则创建
# 'a'  追加,文件存在则在末尾写,不存在则创建
# 'x'  新建,文件已存在则报错(安全创建)
# 'b'  二进制模式(配合 r/w/a 用,如 'rb'、'wb')
# 't'  文本模式(默认,配合 r/w/a)
# '+' 读写模式(如 'r+'、'w+')

with — 自动关门 ​

python
# ❌ 不推荐:手动关
f = open("test.txt", "w")
f.write("Hello")
f.close()  # 容易忘!

# ✅ 推荐:with 自动关,即使出异常也会关
with open("test.txt", "w", encoding="utf-8") as f:
    f.write("Hello, World!")

永远用 with!

with 是一个上下文管理器,离开 with 块时自动调用 f.close()。不用担心忘记关闭文件。一定要指定 encoding="utf-8",否则可能出乱码。

读取文件 ​

python
# 一次性读全部
with open("data.txt", "r", encoding="utf-8") as f:
    content = f.read()       # 整个文件一个字符串
    print(content)

# 逐行读(小文件)
with open("data.txt", "r", encoding="utf-8") as f:
    lines = f.readlines()    # 返回列表,每行一个元素
    for line in lines:
        print(line.strip())  # strip() 去掉末尾换行符

# 逐行读(大文件 — 推荐!)
with open("huge.log", "r", encoding="utf-8") as f:
    for line in f:           # 一行一行地读,不占内存
        if "ERROR" in line:
            print(line.strip())

大文件怎么读?

文件几百 MB 甚至上 GB 时,read() 会爆内存。用 for line in f 逐行迭代,Python 只缓存一行,内存占用极低。

按块读取 ​

python
with open("large.bin", "rb") as f:
    while chunk := f.read(8192):  # 每次读 8KB
        process(chunk)

写入文件 ​

python
lines = ["第一行\n", "第二行\n", "第三行\n"]

with open("output.txt", "w", encoding="utf-8") as f:
    # 方式一:逐条写
    for line in lines:
        f.write(line)

    # 方式二:批量写
    f.writelines(lines)

write vs writelines

write() 写一个字符串,writelines() 写一个字符串列表但不会自动加换行符,需要自己加 \n。

JSON — 结构化数据存储 ​

JSON 是数据交换的事实标准,Python 内置 json 模块:

python
import json

data = {
    "name": "小雅",
    "age": 25,
    "hobbies": ["摄影", "编程", "旅行"],
    "is_student": False
}

# 字典/列表 → JSON 字符串
json_str = json.dumps(data, indent=2, ensure_ascii=False)
print(json_str)

# 字典/列表 → JSON 文件
with open("data.json", "w", encoding="utf-8") as f:
    json.dump(data, f, indent=2, ensure_ascii=False)

# JSON 字符串 → 字典/列表
parsed = json.loads(json_str)
print(parsed["name"])  # 小雅

# JSON 文件 → 字典/列表
with open("data.json", "r", encoding="utf-8") as f:
    loaded = json.load(f)
    print(loaded["hobbies"])  # ['摄影', '编程', '旅行']

ensure_ascii=False 很重要

不加这个参数,中文会被转成 \u5c0f\u96c5 这样谁也看不懂的东西。写了 JSON 永远带上 ensure_ascii=False。

CSV — 表格数据处理 ​

python
import csv

# 写入 CSV
with open("users.csv", "w", newline="", encoding="utf-8-sig") as f:
    writer = csv.writer(f)
    writer.writerow(["姓名", "年龄", "城市"])  # 表头
    writer.writerow(["小明", 25, "北京"])
    writer.writerow(["小红", 28, "上海"])

# 读取 CSV
with open("users.csv", "r", encoding="utf-8-sig") as f:
    reader = csv.reader(f)
    for row in reader:
        print(row)  # ['姓名', '年龄', '城市'] ...

# DictReader — 用字典方式读,更直观
with open("users.csv", "r", encoding="utf-8-sig") as f:
    reader = csv.DictReader(f)
    for row in reader:
        print(f"{row['姓名']}住在{row['城市']}")

# DictWriter — 用字典方式写
fieldnames = ["姓名", "年龄", "城市"]
with open("users2.csv", "w", newline="", encoding="utf-8-sig") as f:
    writer = csv.DictWriter(f, fieldnames=fieldnames)
    writer.writeheader()
    writer.writerow({"姓名": "小雅", "年龄": 25, "城市": "杭州"})

encoding="utf-8-sig" 是什么?

utf-8-sig 会在文件开头加 BOM(Byte Order Mark),让 Excel 正确识别 UTF-8 中文。普通文本用 utf-8 就行,CSV 要给 Excel 用就加 -sig。

pathlib — 优雅的路径操作 ​

告别 os.path.join() 的噩梦,用 / 拼接路径:

python
from pathlib import Path

# 创建 Path 对象
data_dir = Path("data")
file_path = data_dir / "users.csv"  # 用 / 拼接!

# 基本操作
print(file_path.name)       # users.csv(文件名)
print(file_path.stem)       # users(不含扩展名)
print(file_path.suffix)     # .csv(扩展名)
print(file_path.parent)     # data(父目录)

# 检查与创建
if not data_dir.exists():
    data_dir.mkdir(parents=True)  # 递归创建目录

print(file_path.is_file())   # 是文件?
print(file_path.is_dir())    # 是目录?

# 读写
content = file_path.read_text(encoding="utf-8")       # 一次读全部文本
file_path.write_text("Hello", encoding="utf-8")       # 写文本
data = file_path.read_bytes()                         # 读二进制

# 遍历目录
for py_file in Path(".").glob("*.py"):  # 所有 .py 文件
    print(py_file)

for py_file in Path(".").rglob("*.py"):  # 递归(含子目录)
    print(py_file)

实战:批量处理日志 + 合并 CSV ​

python
# /opt/kb/docs/python/file_io.py
"""实战:批量处理日志文件、合并 CSV"""

from pathlib import Path
import csv
import json

def scan_errors(log_dir="logs"):
    """扫描日志目录下所有 .log 文件,收集错误行"""
    errors = []
    for log_file in Path(log_dir).glob("*.log"):
        with open(log_file, "r", encoding="utf-8") as f:
            for i, line in enumerate(f, 1):
                if "ERROR" in line:
                    errors.append({
                        "file": log_file.name,
                        "line": i,
                        "content": line.strip()
                    })
    return errors

def merge_csv(output="merged.csv", *csv_files):
    """合并多个 CSV 文件"""
    all_rows = []
    header = None

    for filepath in csv_files:
        with open(filepath, "r", encoding="utf-8-sig") as f:
            reader = csv.reader(f)
            file_header = next(reader)
            if header is None:
                header = file_header
            all_rows.extend(reader)

    with open(output, "w", newline="", encoding="utf-8-sig") as f:
        writer = csv.writer(f)
        writer.writerow(header)
        writer.writerows(all_rows)

    print(f"合并完成:{len(all_rows)} 行 → {output}")

if __name__ == "__main__":
    # 示例用法
    errs = scan_errors()
    print(f"发现 {len(errs)} 条错误")
    # merge_csv("all_users.csv", "users1.csv", "users2.csv")

常见坑 ​

坑1:忘记用 with

不用 with,出了异常文件就没关。程序长时间跑可能耗尽文件句柄,导致 Too many open files 错误。

坑2:编码不匹配

Windows 的记事本默认保存为 GBK 编码。用 utf-8 读 GBK 文件会乱码或报错:UnicodeDecodeError。解决方法:

  • 不确定编码时用 encoding="gbk" 试试
  • 或用 errors="ignore" 跳过无法解码的字节(不推荐)
  • 最佳实践:统一用 UTF-8

坑3:CSV 写入没加 newline=""

Windows 上 open("file.csv", "w") 会导致行之间多一个空行。加上 newline="" 就正常了。

练习题 ​

  1. 读一个文本文件,统计每个单词出现次数,输出 TOP 10
  2. 创建 JSON 文件存 5 个联系人的姓名和电话,然后读出来打印
  3. 写程序扫描一个目录,找出所有 .jpg 文件并列出来(用 pathlib)
  4. 有 3 个 CSV 文件(分别记账 1 月/2 月/3 月),合并成一个并计算总支出
  5. 写一个日志分析脚本:找出今天日志中所有的 ERROR 和 WARNING

📁 文件操作学会了,下一步学 异常处理,让程序遇到错误也能优雅应对。


🎯 本章要点 ​

  • with 是上下文管理器,无论 try 块里是否抛异常
  • with + open + read 是最标准的读文件模式。
  • 'x' = exclusive creation
加载练习题中...