Python迭代器与生成器

时游大约 4 分钟

迭代器与生成器

迭代器协议

迭代器(Iterator)是访问集合元素的一种方式:实现 __iter__() 和 __next__() 两个方法的对象就是迭代器。迭代器记住遍历位置,依次取值,不能回退。

# iter():从可迭代对象获取迭代器
# next():获取下一个值,取完后再取会抛出StopIteration异常
words = ['cat', 'window', 'defenestrate']
it = iter(words)

print(next(it))  # cat
print(next(it))  # window
print(next(it))  # defenestrate
# print(next(it))  # StopIteration:元素已取完

for 循环的本质

for 语句内部就是「调用 iter() 拿迭代器 + 反复调用 next() + 捕获 StopIteration 退出」,任何实现了迭代器协议的对象都能被 for 遍历。

# 这两段代码等价
for word in ['cat', 'window']:
    print(word)

it = iter(['cat', 'window'])
while True:
    try:
        word = next(it)
    except StopIteration:
        break
    print(word)

自定义迭代器

# 实现__iter__和__next__即可让自己的类支持for遍历
class Countdown:
    def __init__(self, start):
        self.current = start

    def __iter__(self):
        return self  # 迭代器返回自身

    def __next__(self):
        if self.current <= 0:
            raise StopIteration  # 遍历结束的信号
        self.current -= 1
        return self.current + 1


for num in Countdown(3):
    print(num)  # 3 2 1

生成器函数(yield)

生成器(Generator)是写迭代器最简便的方式:函数中使用 yield 关键字,调用时不执行函数体,而是返回一个生成器对象;每次 next() 执行到 yield 暂停并交出值,下次从暂停处继续。

def countdown(start):
    print("开始倒数")
    while start > 0:
        yield start  # 执行到这里暂停,把值交出去
        start -= 1  # 下次next()从这里继续


gen = countdown(3)
print(next(gen))  # 开始倒数 -> 3(打印发生在第一次next时,而非调用countdown时)
print(next(gen))  # 2
print(next(gen))  # 1
# print(next(gen))  # StopIteration

# 生成器是迭代器,可直接for遍历
for num in countdown(3):
    print(num)  # 3 2 1

生成器的典型应用

1. 惰性求值:用多少算多少

# 斐波那契数列:用生成器可以表达无限序列,取前n个时才计算
def fibonacci():
    a, b = 0, 1
    while True:
        yield a
        a, b = b, a + b


fib = fibonacci()
print([next(fib) for _ in range(8)])  # [0, 1, 1, 2, 3, 5, 8, 13]

2. 处理大文件:逐行读取,不占内存

# 日志文件可能有几个G,read()一次性读入会撑爆内存
def read_large_file(path):
    with open(path, encoding='utf-8') as f:
        for line in f:  # 文件对象本身就是迭代器,逐行读取
            if 'ERROR' in line:
                yield line.strip()


for line in read_large_file('app.log'):
    print(line)  # 只打印错误行,内存中始终只有一行

3. 生成器管道:流水线处理

def read_numbers(path):
    with open(path, encoding='utf-8') as f:
        for line in f:
            yield int(line)

def filter_even(nums):
    for n in nums:
        if n % 2 == 0:
            yield n

def square(nums):
    for n in nums:
        yield n * n

# 惰性流水线:数据像流水一样逐个穿过每个环节
pipeline = square(filter_even(read_numbers('numbers.txt')))

生成器表达式

语法与列表推导式相同,把方括号换成圆括号。区别:列表推导式立即生成完整列表,生成器表达式按需生成、省内存。

squares_list = [x ** 2 for x in range(10)]    # 立即生成完整列表
squares_gen = (x ** 2 for x in range(10))     # 生成器,几乎不占内存

print(sum(squares_gen))  # 285,生成器可直接传入聚合函数
# 生成器只能消费一次,用完即空
print(list(squares_gen))  # []

yield from:委托给子生成器

def chain(*iterables):
    for it in iterables:
        # yield from:把子迭代器的值逐个yield出来,等价于 for x in it: yield x
        yield from it


print(list(chain([1, 2], 'ab', (True, False))))  # [1, 2, 'a', 'b', True, False]

itertools:迭代工具库

标准库 itertools 提供了一批高效处理迭代器的函数,返回的都是惰性迭代器。

import itertools

# count(10):从10开始的无限计数器 10 11 12 ...
# cycle('AB'):无限循环 A B A B ...
# repeat('x', 3):重复3次 x x x

print(list(itertools.combinations('ABC', 2)))   # [('A','B'), ('A','C'), ('B','C')] 组合
print(list(itertools.permutations('AB')))       # [('A','B'), ('B','A')] 排列
print(list(itertools.product('AB', repeat=2)))  # 笛卡尔积
print(list(itertools.groupby('AABBBCC')))       # 分组:[('A',...), ('B',...), ('C',...)]

迭代器 vs 可迭代对象

概念判断方法说明
可迭代对象 Iterable实现 __iter__list、str、dict、文件对象等,可被 for 遍历
迭代器 Iterator实现 __iter__ + __next__可被 next() 取值,是一次性消费品
生成器 Generator函数含 yield自动实现了迭代器协议,是创建迭代器最简便的方式

记忆点:迭代器一定是可迭代对象,可迭代对象不一定是迭代器;list 可以反复遍历,是因为 for 每次都调用 iter() 创建了新的迭代器。

上次编辑于:
贡献者: 15327360835
Loading...