Python中的pathlib库使用详解 前言在pathlib出现之前处理路径靠的是os.path里的一堆函数os.path.join(a, b)、os.path.splitext(p)、os.path.basename(p)……它们是把路径当作字符串来操作的于是代码里到处是拼接、切片、判断分隔符。pathlibPEP 428Python 3.4 起换了一个视角把路径当作对象来操作。拼接用/运算符取父目录用.parent判断存在用.exists()读写内容用.read_text()/.write_text()。常见的误解有两个。第一以为pathlib会做文件系统检查——其实Path的构造不做任何 I/OPath(不存在的/x.txt)完全合法只有调用.exists()、.read_text()这类方法时才真正访问磁盘。第二以为PurePath和Path差不多——PurePath只做纯字符串运算不碰文件系统适合在无法访问磁盘的场景或需要显式区分PurePosixPath/PureWindowsPath时使用。本文覆盖Path的构造、解析、判断、遍历、读写与常见陷阱并与os.path做对照方便判断什么时候该用哪个。一、构造与解析路径就是对象Path有两大族PurePath系列只做字符串层面的路径运算PurePosixPath、PureWindowsPathPath系列继承它们并加上真正的文件系统操作。可以这样理解PurePath负责长什么样Path负责能不能用。# 适用于 Python 3.8from pathlib import PurePosixPath, PureWindowsPathfrom pathlib import Path# 构造不做任何磁盘访问p Path(data) / reports / 2024 / q1.csvprint(p) # data/reports/2024/q1.csvWindows 上显示反斜杠# 解析print(p.name) # q1.csv 最后一段print(p.stem) # q1 去掉最后一个后缀print(p.suffix) # .csvprint(p.suffixes) # [.csv]print(p.parent) # data/reports/2024print(p.parts) # (data, reports, 2024, q1.csv)print(p.anchor) # 相对路径没有锚点# 跨平台字符串运算不需要真实文件print(PureWindowsPath(rC:\Users\me\file.txt).drive) # C:print(PurePosixPath(/usr/local/bin).parts) # (/, usr, local, bin)几个必须记牢的解析细节。.suffix只返回最后一个点之后的部分Path(archive.tar.gz).suffix是.gz要拿全部后缀得用.suffixes返回[.tar, .gz]。.parents是一个可索引的序列.parents[0]是直接父目录也可以切片或负索引Python 3.10 起支持切片和负索引。.name返回最后一段对文件是文件名对目录是目录名.stem是.name去掉后缀的结果。二、改名类方法with_name / with_suffix / with_stem这几个方法返回新的 Path 对象不修改原对象也不触碰磁盘。# 适用于 Python 3.9from pathlib import Pathp Path(/tmp/report.csv)print(p.with_name(summary.csv)) # /tmp/summary.csvprint(p.with_suffix(.txt)) # /tmp/report.txtprint(p.with_stem(final)) # /tmp/final.csv with_stem 需要 3.9with_stem()是Python 3.9 新增的3.9 之前只能用p.with_name(final p.suffix)模拟。with_suffix()有一个易错点如果传入的字符串不以点开头会直接拼上去而不报错p.with_suffix(txt)得到reporttxt。三、真正的文件系统操作到这里才开始接触磁盘。判断类方法.exists()、.is_file()、.is_dir()、.is_symlink()、.stat()。目录类.mkdir(parentsFalse, exist_okFalse)、.iterdir()、.glob(pattern)、.rglob(pattern)、.unlink()、.rmdir()、.touch()。移动改名.rename(target)、.replace(target)。# 适用于 Python 3.8from pathlib import Pathroot Path(workspace)root.mkdir(parentsTrue, exist_okTrue) # 递归建目录已存在也不报错(root / a.txt).write_text(hello, encodingutf-8)(root / b.txt).write_text(world, encodingutf-8)(root / sub).mkdir(exist_okTrue)(root / sub / c.md).write_text(note, encodingutf-8)# 只列一层不递归for item in root.iterdir():print(item.name, 目录 if item.is_dir() else 文件)# 递归匹配** 表示任意层级for md in root.rglob(*.md):print(markdown:, md)# 只找当前层for txt in root.glob(*.txt):print(text:, txt)mkdir()的parentsTrue会连父目录一起创建相当于os.makedirsexist_okTrue让目录已存在不再抛异常。.rglob(pattern)等价于.glob(**/ pattern)。遍历目录树还可以用Path.walk()它从Python 3.12开始提供用法与os.walk()类似返回(root, dirs, files)三元组但有一处不同当follow_symlinks为假时Path.walk()会把指向目录的符号链接归到filenames而不是dirnames。任务pathlib写法os.path写法拼接a / bos.path.join(a, b)取文件名p.nameos.path.basename(p)取目录p.parentos.path.dirname(p)取后缀p.suffixos.path.splitext(p)[1]判断存在p.exists()os.path.exists(p)判断是文件p.is_file()os.path.isfile(p)建目录p.mkdir(parentsTrue, exist_okTrue)os.makedirs(p, exist_okTrue)遍历树p.rglob(*)/p.walk()os.walk(p)四、读写内容与绝对路径Path直接封装了常用的文本 / 字节读写省去open()的样板。它们与open()一样默认编码依赖平台所以照样要显式写encoding。# 适用于 Python 3.8from pathlib import Pathp Path(notes.txt)p.write_text(你好\n, encodingutf-8)print(p.read_text(encodingutf-8), end)b Path(blob.bin)b.write_bytes(b\x00\x01\x02)print(b.read_bytes())# 需要更细粒度控制读一部分、遍历行时还是用 openwith p.open(r, encodingutf-8) as f:for line in f:print(line.rstrip(\n))read_text(encodingNone, errorsNone, newlineNone)与write_text(data, encodingNone, errorsNone, newlineNone)的参数与open()对应。write_bytes/read_bytes则用于二进制。绝对路径相关的三个方法要注意区别Path.cwd()返回当前工作目录类方法Path.home()返回用户主目录类方法p.resolve()返回解析掉符号链接和..之后的绝对路径它会访问文件系统而p.absolute()只是把相对路径前面接上 cwd、不做规范化会保留..段。判断路径包含关系用.is_relative_to(other)Python 3.9 新增。常见坑点❌ 以为Path(x)会检查文件是否存在✅ 构造不做任何 I/O只有.exists()/.read_text()这类方法才真正访问磁盘❌ 用Path(a/b) c或对 Path 做字符串拼接✅Path不能直接加字符串会抛TypeError用/运算符或p / c❌Path(archive.tar.gz).suffix期望拿到.tar.gz✅ 只返回最后一个后缀.gz全部后缀用.suffixes❌p.with_suffix(txt)忘了点✅ 参数要带点写p.with_suffix(.txt)否则会直接拼出reporttxt❌read_text()/write_text()不写encoding✅ 与open()一致默认编码平台相关一律显式传encodingutf-8❌ 以为p.absolute()会去掉..✅ 它只接上 cwd、不做规范化需要规范化用p.resolve()会访问磁盘、解析符号链接❌ 在循环里p.exists()又p.is_dir()又p.stat()导致多次系统调用✅ 需要多次属性时一次p.stat()拿全或用os.scandir()一次遍历带出类型信息❌p.mkdir()建多级目录报FileNotFoundError✅ 多级目录要p.mkdir(parentsTrue, exist_okTrue)总结类别方法是否访问磁盘构造 / 运算Path(...)、a / b、parents、stem否改名返回新对象with_name、with_suffix、with_stem否判断exists、is_file、is_dir是读写read_text、write_text、read_bytes是遍历iterdir、glob、rglob、walk是pathlib带来的最大改变不是少打几个函数名而是把路径从一堆字符串操作里解放出来变成可以安全传递、组合、拼接的对象。新代码优先用Path只有在需要与老接口对接、或明确要用os.scandir()的性能特性时才回退到os.path。