Python re库怎么用?超全正则表达式实战教程 Python `re` 模块完全指南:从入门到精通
在 Python 的世界里,处理文本数据是开发者的日常。无论是清理脏数据、提取网页信息,还是验证用户输入,正则表达式(Regular Expression) 都是最强大的工具之一。而 Python 内置的 `re` 模块,则是你驾驭正则表达式的核心武器。 许多初学者在面对 `re` 模块时,常被其复杂的符号和多样的函数所劝退。本文将带你系统地梳理 `re` 模块的使用方法,从基础概念到高级技巧,助你成为文本处理的高手。
一、 为什么选择 `re` 模块?
虽然 Python 字符串本身也提供了一些简单的查找和替换方法(如 `find`, `replace`),但 `re` 模块提供了:
- 强大的模式匹配能力:支持字符类、量词、分组、前瞻断言等复杂逻辑。
- 高效的性能:底层由 C 语言实现,处理大规模文本时速度极快。
- 统一的接口:所有操作都围绕“模式”展开,逻辑清晰。
二、 核心概念:正则表达式基础
在使用 `re` 之前,你需要了解几个基本符号:
| 符号 | 含义 | 示例 |
| `.` | 匹配除换行符外的任意单个字符 | `a.c` 匹配 `abc`, `a1c` |
| `^` | 匹配字符串开头 | `^Hello` 匹配以 Hello 开头的字符串 |
| `$` | 匹配字符串结尾 | `World$` 匹配以 World 结尾的字符串 |
| `` | 匹配前一个字符 0 次或多次 | `abc` 匹配 `ac`, `abc`, `abbc` |
| `+` | 匹配前一个字符 1 次或多次 | `ab+c` 匹配 `abc`, `abbc` |
| `?` | 匹配前一个字符 0 次或 1 次 | `ab?c` 匹配 `ac`, `abc` |
| `{n}` | 匹配前一个字符恰好 n 次 | `a{3}` 匹配 `aaa` |
| `[]` | 字符集合,匹配其中任意一个字符 | `[aeiou]` 匹配任意元音字母 |
| `d` | 匹配任意数字,等价于 `[0-9]` | `d+` 匹配一串数字 |
| `w` | 匹配字母、数字、下划线,等价于 `[a-zA-Z0-9_]` | `w+` 匹配单词 |
| `s` | 匹配空白字符(空格、制表符等) | `s+` 匹配一个或多个空格 |
三、 `re` 模块常用函数详解
`re` 模块提供了多个函数,每个函数适用于不同的场景。
1. `re.search(pattern, string, flags=0)`
用途:在字符串中搜索第一个匹配项。 返回值:如果找到匹配,返回 `Match` 对象;否则返回 `None`。 ```python import re text = "Contact us at support@example.com or sales@example.com" match = re.search(r'S+@S+', text) if match: print("First email found:", match.group()) # 输出: support@example.com ``` ? 注意:`search` 只返回第一个匹配项,即使后面还有更多匹配。
2. `re.match(pattern, string, flags=0)`
用途:从字符串开头开始匹配。 返回值:如果开头匹配成功,返回 `Match` 对象;否则返回 `None`。 ```python text = "Hello World" match = re.match(r'Hello', text) if match: print("Matched at start:", match.group()) # 输出: Hello ``` ⚠️ 区别:`match` 只检查字符串开头,而 `search` 检查整个字符串。
3. `re.findall(pattern, string, flags=0)`
用途:找出字符串中所有匹配项。 返回值:返回一个列表,包含所有匹配的子串。 ```python text = "My phone is 123-456-7890 and my work is 987-654-3210" phones = re.findall(r'd{3}-d{3}-d{4}', text) print(phones) # 输出: ['123-456-7890', '987-654-3210'] ```
4. `re.sub(pattern, repl, string, count=0, flags=0)`
用途:替换字符串中匹配的部分。 参数:
- `pattern`:正则表达式
- `repl`:替换后的字符串或函数
- `string`:原始字符串
- `count`:最大替换次数(0 表示全部替换)
```python text = "Hello, World! Hello, Python!" new_text = re.sub(r'Hello', 'Hi', text, count=1) print(new_text) # 输出: Hi, World! Hello, Python! ```
5. `re.split(pattern, string, maxsplit=0, flags=0)`
用途:根据匹配项分割字符串。 返回值:返回一个列表。 ```python text = "apple,banana;cherry|date" fruits = re.split(r'[,;|]', text) print(fruits) # 输出: ['apple', 'banana', 'cherry', 'date'] ```
四、 高级技巧:分组与捕获
正则表达式中的括号 `()` 用于分组,不仅可以逻辑分组,还可以捕获匹配的内容。
1. 捕获组(Capturing Groups)
使用 `group()` 方法访问捕获的内容。 ```python text = "2023-10-05" match = re.search(r'(d{4})-(d{2})-(d{2})', text) if match: print("Full match:", match.group(0)) # 2023-10-05 print("Year:", match.group(1)) # 2023 print("Month:", match.group(2)) # 10 print("Day:", match.group(3)) # 05 ```
2. 命名组(Named Groups)
使用 `(?P
pattern)` 语法,使代码更易读。 ```python text = "User: John Doe, Age: 30" match = re.search(r'User: (?Pw+ w+), Age: (?Pd+)', text) if match: print("Name:", match.group('name')) # John Doe print("Age:", match.group('age')) # 30 ``` 3. 非捕获组(Non-capturing Groups)
使用 `(?:pattern)`,仅用于逻辑分组,不捕获内容,提高效率。 ```python 匹配 "abc" 或 "abd",但不捕获 "a" 和 "c/d"
text = "abc" match = re.search(r'ab(?:c|d)', text) ``` 五、 标志位(Flags)的使用
`re` 模块支持多种标志位,用于调整匹配行为。 | 标志 | 含义 | 示例 |
| `re.IGNORECASE` | 忽略大小写 | `re.search(r'hello', 'HELLO', re.I)` |
| `re.MULTILINE` | 多行模式,`^` 和 `$` 匹配每行的开头/结尾 | `re.search(r'^start', 'startnend', re.M)` |
| `re.DOTALL` | 使 `.` 匹配包括换行符在内的所有字符 | `re.search(r'.', 'anb', re.S)` |
| `re.VERBOSE` | 允许在正则表达式中加入注释和空白,提高可读性 | `re.search(r''' ... # comment ... ''', text, re.X)` |
示例:使用 `re.VERBOSE` 提高可读性
```python import re pattern = re.compile(r""" ^ # 字符串开头 d{4} # 年份 d{2} # 月份 d{2} # 日期 $ # 字符串结尾 """, re.VERBOSE) if pattern.match("2023-10-05"): print("Valid date format") ``` 六、 最佳实践与常见陷阱
1. 始终使用原始字符串(Raw Strings)
在 Python 中,反斜杠 `` 是转义字符。为了避免混淆,始终在正则表达式前加 `r`。 ```python 推荐
re.search(r'd+', text) 不推荐:需要双重转义
re.search('\d+', text) ``` 2. 预编译正则表达式
如果重复使用同一个正则表达式,使用 `re.compile()` 预编译,可提升性能。 ```python import re 预编译
email_pattern = re.compile(r'S+@S+') 多次使用
result1 = email_pattern.search("abc@def.com") result2 = email_pattern.findall("x@y.com z@w.com") ``` 3. 避免贪婪匹配导致的性能问题
正则表达式默认是贪婪的,会尽可能多地匹配。在某些情况下,这会导致“灾难性回溯”(Catastrophic Backtracking),使程序卡死。 ```python 危险:可能引发性能问题
re.search(r'(a+)+b', 'a' 10000 + 'c') 推荐:使用非贪婪匹配或原子组
re.search(r'(a+?)+b', 'a' 10000 + 'c') ``` 4. 验证输入时,使用 `fullmatch()`
如果你希望整个字符串完全匹配某个模式,使用 `re.fullmatch()` 而不是 `re.match()` 或 `re.search()`。 ```python 验证邮箱格式(简化版)
email = "user@example.com" if re.fullmatch(r'w+@w+.w+', email): print("Valid email") else: print("Invalid email") ``` 七、 实战案例:从 HTML 中提取数据
假设我们需要从一个简单的 HTML 字符串中提取所有链接: ```python import re html = """ Example Python
""" 提取 href 属性值
links = re.findall(r'href=["'](.?)["']', html) print(links) 输出: ['https://example.com', 'https://python.org']
``` `re` 模块是 Python 文本处理的利器,但其学习曲线较为陡峭。建议初学者: 1. 从简单开始:先掌握 `search`, `findall`, `sub` 三个核心函数。 2. 使用在线工具调试:如 [regex101.com](https://regex101.com/),可以实时测试正则表达式并查看匹配过程。 3. 阅读官方文档:Python 的 `re` 模块文档非常详细,是最佳的参考资料。 通过不断练习和积累经验,你将能够轻松应对各种复杂的文本处理任务。祝你编程愉快!
免责声明:本文内容来源于公开网络、企业供稿或其他合规渠道,仅用于信息交流与学习参考,不构成任何形式的商业建议或结论。若涉及版权、出处或权利争议,请联系我们将在核实后及时处理。