Python re库怎么用?超全正则表达式实战教程

Python `re` 模块完全指南:从入门到精通

在 Python 的世界里,处理文本数据是开发者的日常。无论是清理脏数据、提取网页信息,还是验证用户输入,正则表达式(Regular Expression) 都是最强大的工具之一。而 Python 内置的 `re` 模块,则是你驾驭正则表达式的核心武器。 许多初学者在面对 `re` 模块时,常被其复杂的符号和多样的函数所劝退。本文将带你系统地梳理 `re` 模块的使用方法,从基础概念到高级技巧,助你成为文本处理的高手。

一、 为什么选择 `re` 模块?

虽然 Python 字符串本身也提供了一些简单的查找和替换方法(如 `find`, `replace`),但 `re` 模块提供了:
  • 强大的模式匹配能力:支持字符类、量词、分组、前瞻断言等复杂逻辑。
  • 高效的性能:底层由 C 语言实现,处理大规模文本时速度极快。
  • 统一的接口:所有操作都围绕“模式”展开,逻辑清晰。

二、 核心概念:正则表达式基础

在使用 `re` 之前,你需要了解几个基本符号:
符号 含义 示例
`.` 匹配除换行符外的任意单个字符 `a.c` 匹配 `abc`, `a1c`
`^` 匹配字符串开头 `^Hello` 匹配以 Hello 开头的字符串
`$` 匹配字符串结尾 `World$` 匹配以 World 结尾的字符串
`` 匹配前一个字符 0 次或多次 `abc` 匹配 `ac`, `abc`, `abbc`
`+` 匹配前一个字符 1 次或多次 `ab+c` 匹配 `abc`, `abbc`
`?` 匹配前一个字符 0 次或 1 次 `ab?c` 匹配 `ac`, `abc`
`{n}` 匹配前一个字符恰好 n 次 `a{3}` 匹配 `aaa`
`[]` 字符集合,匹配其中任意一个字符 `[aeiou]` 匹配任意元音字母
`d` 匹配任意数字,等价于 `[0-9]` `d+` 匹配一串数字
`w` 匹配字母、数字、下划线,等价于 `[a-zA-Z0-9_]` `w+` 匹配单词
`s` 匹配空白字符(空格、制表符等) `s+` 匹配一个或多个空格

三、 `re` 模块常用函数详解

`re` 模块提供了多个函数,每个函数适用于不同的场景。

1. `re.search(pattern, string, flags=0)`

用途:在字符串中搜索第一个匹配项。 返回值:如果找到匹配,返回 `Match` 对象;否则返回 `None`。 ```python import re text = "Contact us at support@example.com or sales@example.com" match = re.search(r'S+@S+', text) if match: print("First email found:", match.group()) # 输出: support@example.com ``` ? 注意:`search` 只返回第一个匹配项,即使后面还有更多匹配。

2. `re.match(pattern, string, flags=0)`

用途:从字符串开头开始匹配。 返回值:如果开头匹配成功,返回 `Match` 对象;否则返回 `None`。 ```python text = "Hello World" match = re.match(r'Hello', text) if match: print("Matched at start:", match.group()) # 输出: Hello ``` ⚠️ 区别:`match` 只检查字符串开头,而 `search` 检查整个字符串。

3. `re.findall(pattern, string, flags=0)`

用途:找出字符串中所有匹配项。 返回值:返回一个列表,包含所有匹配的子串。 ```python text = "My phone is 123-456-7890 and my work is 987-654-3210" phones = re.findall(r'd{3}-d{3}-d{4}', text) print(phones) # 输出: ['123-456-7890', '987-654-3210'] ```

4. `re.sub(pattern, repl, string, count=0, flags=0)`

用途:替换字符串中匹配的部分。 参数:
  • `pattern`:正则表达式
  • `repl`:替换后的字符串或函数
  • `string`:原始字符串
  • `count`:最大替换次数(0 表示全部替换)
```python text = "Hello, World! Hello, Python!" new_text = re.sub(r'Hello', 'Hi', text, count=1) print(new_text) # 输出: Hi, World! Hello, Python! ```

5. `re.split(pattern, string, maxsplit=0, flags=0)`

用途:根据匹配项分割字符串。 返回值:返回一个列表。 ```python text = "apple,banana;cherry|date" fruits = re.split(r'[,;|]', text) print(fruits) # 输出: ['apple', 'banana', 'cherry', 'date'] ```

四、 高级技巧:分组与捕获

正则表达式中的括号 `()` 用于分组,不仅可以逻辑分组,还可以捕获匹配的内容。

1. 捕获组(Capturing Groups)

使用 `group()` 方法访问捕获的内容。 ```python text = "2023-10-05" match = re.search(r'(d{4})-(d{2})-(d{2})', text) if match: print("Full match:", match.group(0)) # 2023-10-05 print("Year:", match.group(1)) # 2023 print("Month:", match.group(2)) # 10 print("Day:", match.group(3)) # 05 ```

2. 命名组(Named Groups)

使用 `(?Ppattern)` 语法,使代码更易读。 ```python text = "User: John Doe, Age: 30" match = re.search(r'User: (?Pw+ w+), Age: (?Pd+)', text) if match: print("Name:", match.group('name')) # John Doe print("Age:", match.group('age')) # 30 ```

3. 非捕获组(Non-capturing Groups)

使用 `(?:pattern)`,仅用于逻辑分组,不捕获内容,提高效率。 ```python

匹配 "abc" 或 "abd",但不捕获 "a" 和 "c/d"

text = "abc" match = re.search(r'ab(?:c|d)', text) ```

五、 标志位(Flags)的使用

`re` 模块支持多种标志位,用于调整匹配行为。
标志 含义 示例
`re.IGNORECASE` 忽略大小写 `re.search(r'hello', 'HELLO', re.I)`
`re.MULTILINE` 多行模式,`^` 和 `$` 匹配每行的开头/结尾 `re.search(r'^start', 'startnend', re.M)`
`re.DOTALL` 使 `.` 匹配包括换行符在内的所有字符 `re.search(r'.', 'anb', re.S)`
`re.VERBOSE` 允许在正则表达式中加入注释和空白,提高可读性 `re.search(r''' ... # comment ... ''', text, re.X)`

示例:使用 `re.VERBOSE` 提高可读性

```python import re pattern = re.compile(r""" ^ # 字符串开头 d{4} # 年份
  • # 连字符
d{2} # 月份
  • # 连字符
d{2} # 日期 $ # 字符串结尾 """, re.VERBOSE) if pattern.match("2023-10-05"): print("Valid date format") ```

六、 最佳实践与常见陷阱

1. 始终使用原始字符串(Raw Strings)

在 Python 中,反斜杠 `` 是转义字符。为了避免混淆,始终在正则表达式前加 `r`。 ```python

推荐

re.search(r'd+', text)

不推荐:需要双重转义

re.search('\d+', text) ```

2. 预编译正则表达式

如果重复使用同一个正则表达式,使用 `re.compile()` 预编译,可提升性能。 ```python import re

预编译

email_pattern = re.compile(r'S+@S+')

多次使用

result1 = email_pattern.search("abc@def.com") result2 = email_pattern.findall("x@y.com z@w.com") ```

3. 避免贪婪匹配导致的性能问题

正则表达式默认是贪婪的,会尽可能多地匹配。在某些情况下,这会导致“灾难性回溯”(Catastrophic Backtracking),使程序卡死。 ```python

危险:可能引发性能问题

re.search(r'(a+)+b', 'a' 10000 + 'c')

推荐:使用非贪婪匹配或原子组

re.search(r'(a+?)+b', 'a' 10000 + 'c') ```

4. 验证输入时,使用 `fullmatch()`

如果你希望整个字符串完全匹配某个模式,使用 `re.fullmatch()` 而不是 `re.match()` 或 `re.search()`。 ```python

验证邮箱格式(简化版)

email = "user@example.com" if re.fullmatch(r'w+@w+.w+', email): print("Valid email") else: print("Invalid email") ```

七、 实战案例:从 HTML 中提取数据

假设我们需要从一个简单的 HTML 字符串中提取所有链接: ```python import re html = """ Example Python """

提取 href 属性值

links = re.findall(r'href=["'](.?)["']', html) print(links)

输出: ['https://example.com', 'https://python.org']

``` `re` 模块是 Python 文本处理的利器,但其学习曲线较为陡峭。建议初学者: 1. 从简单开始:先掌握 `search`, `findall`, `sub` 三个核心函数。 2. 使用在线工具调试:如 [regex101.com](https://regex101.com/),可以实时测试正则表达式并查看匹配过程。 3. 阅读官方文档:Python 的 `re` 模块文档非常详细,是最佳的参考资料。 通过不断练习和积累经验,你将能够轻松应对各种复杂的文本处理任务。祝你编程愉快!