如何使用asyncio从async for循环中生成?

2024-03-28 08:53:12 发布

您现在位置:Python中文网/ 问答频道 /正文

我试图编写一个简单的异步数据批处理生成器,但在理解如何从异步for循环中生成时遇到了困难。在这里,我写了一个简单的课程来说明我的想法:

import asyncio
from typing import List

class AsyncSimpleIterator:
    def __init__(self, data: List[str], batch_size=None):
        self.data = data
        self.batch_size = batch_size
        self.doc2index = self.get_doc_ids()

    def get_doc_ids(self):
        return list(range(len(self.data)))

    async def get_batch_data(self, doc_ids):
        print("get_batch_data() running")
        page = [self.data[j] for j in doc_ids]
        return page

    async def get_docs(self, batch_size):
        print("get_docs() running")

        _batch_size = self.batch_size or batch_size
        batches = [self.doc2index[i:i + _batch_size] for i in
                   range(0, len(self.doc2index), _batch_size)]

        for _, doc_ids in enumerate(batches):
            docs = await self.get_batch_data(doc_ids)
            yield docs, doc_ids

    async def main(self):
        print("main() running")
        async for res in self.get_docs(batch_size=2):
            print(res)  # how to yield instead of print?

    def gen_batches(self):
        # how to get results of self.main() here?
        loop = asyncio.get_event_loop()
        loop.run_until_complete(self.main())
        loop.close()


 DATA = ["Hello, world!"] * 4
 iterator = AsyncSimpleIterator(DATA)
 iterator.gen_batches()

所以,我的问题是,如何从main()生成一个结果,并将其收集到gen_batches()中?在

当我在main()内打印结果时,我得到以下输出:

^{pr2}$

Tags: inselfidsdocsfordatasizeget
2条回答

基于@user4815162342的工作解决方案对原始问题的回答:

import asyncio
from typing import List


class AsyncSimpleIterator:

def __init__(self, data: List[str], batch_size=None):
    self.data = data
    self.batch_size = batch_size
    self.doc2index = self.get_doc_ids()

def get_doc_ids(self):
    return list(range(len(self.data)))

async def get_batch_data(self, doc_ids):
    print("get_batch_data() running")
    page = [self.data[j] for j in doc_ids]
    return page

async def get_docs(self, batch_size):
    print("get_docs() running")

    _batch_size = self.batch_size or batch_size
    batches = [self.doc2index[i:i + _batch_size] for i in
               range(0, len(self.doc2index), _batch_size)]

    for _, doc_ids in enumerate(batches):
        docs = await self.get_batch_data(doc_ids)
        yield docs, doc_ids

def gen_batches(self):
    loop = asyncio.get_event_loop()

    async def collect():
        return [j async for j in self.get_docs(batch_size=2)]

    items = loop.run_until_complete(collect())
    loop.close()
    return items


DATA = ["Hello, world!"] * 4
iterator = AsyncSimpleIterator(DATA)
result = iterator.gen_batches()
print(result)

I'm trying to write a simple asynchronous data batch generator, but having troubles with understanding how to yield from an async for loop

async for获得的结果与常规收益类似,只是它还必须由async for或等效物收集。例如,get_docs中的yield使其成为异步生成器。如果将print(res)替换为main()中的yield res,那么{}也将成为一个异步生成器。在

the generator in main() should exhaust in gen_batches(), so I can gather all results in gen_batches()

要收集由异步生成器生成的值(例如用print(res)替换为yield res)的{},可以使用helper协同例程:

def gen_batches(self):
    loop = asyncio.get_event_loop()
    async def collect():
        return [item async for item in self.main()]
    items = loop.run_until_complete(collect())
    loop.close()
    return items

collect()助手使用了PEP 530异步理解,可以将其视为更显式的语法糖分:

^{pr2}$

相关问题 更多 >