โมเดลและแพลตฟอร์ม AI

การเรียก API LLM แบบ Asynchronous ใน Python: คู่มือที่ครอบคลุม

mm
เพิ่ม Unite.AI ลงในแหล่งข้อมูลที่คุณต้องการบน Google

ในฐานะนักพัฒนาและนักวิทยาศาสตร์ข้อมูล เรามักพบว่าตัวเองต้องการโต้ตอบกับโมเดลที่ทรงพลังเหล่านี้ผ่าน API อย่างไรก็ตาม เมื่อแอปพลิเคชันของเรามีความซับซ้อนและขนาดใหญ่ขึ้น ความต้องการสำหรับการโต้ตอบ API ที่มีประสิทธิภาพและทำงานได้ดีจึงกลายเป็นสิ่งจำเป็น นี่คือที่ที่การเขียนโปรแกรมแบบ Asynchronous ส่องแสง โดยทำให้เราสามารถเพิ่มประสิทธิภาพสูงสุดและลดความล่าช้าเมื่อทำงานกับ API LLM

ในคู่มือที่ครอบคลุมนี้ เราจะสำรวจโลกของการเรียก API LLM แบบ Asynchronous ใน Python เราจะครอบคลุมทุกอย่างตั้งแต่พื้นฐานของการเขียนโปรแกรมแบบ Asynchronous ไปจนถึงเทคนิคขั้นสูงสำหรับการจัดการ 워์กโฟลว์ที่ซับซ้อน เมื่อถึงจุดสิ้นสุดของบทความนี้ คุณจะมีความเข้าใจที่ดีเกี่ยวกับวิธีการใช้การเขียนโปรแกรมแบบ Asynchronous เพื่อเพิ่มประสิทธิภาพของแอปพลิเคชันที่ใช้ LLM

ก่อนที่เราจะเจาะลึกถึงรายละเอียดของการเรียก API LLM แบบ Asynchronous มาทำความเข้าใจพื้นฐานของการเขียนโปรแกรมแบบ Asynchronous กันก่อน

การเขียนโปรแกรมแบบ Asynchronous ช่วยให้สามารถดำเนินการหลายอย่างพร้อมกันโดยไม่บล็อกเธรดหลักของการดำเนินการ ใน Python สิ่งนี้สามารถทำได้โดยใช้โมดูล asyncio ซึ่งให้เฟรมเวิร์กสำหรับการเขียนโค้ดที่ใช้โคโรทีน การวนซ้ำเหตุการณ์ และฟิวเจอร์

แนวคิดหลัก:

  • โคโรทีน: ฟังก์ชันที่ถูกกำหนดด้วย async def ที่สามารถหยุดและเริ่มต้นใหม่ได้
  • การวนซ้ำเหตุการณ์: กลไกการดำเนินการที่สำคัญที่จัดการและดำเนินการงานแบบ Asynchronous
  • วัตถุAwaitable: วัตถุที่สามารถใช้กับคำสั่ง await (โคโรทีน การทำงาน ฟิวเจอร์)

ต่อไปนี้เป็นตัวอย่างง่ายๆ เพื่ออธิบายแนวคิดเหล่านี้:

import asyncio

<p>async def greet(name):
await asyncio.sleep(1) # จำลองการดำเนินการ I/O</p>

<p>async def main():
await asyncio.gather(
greet(&quot;Alice&quot;),
greet(&quot;Bob&quot;),
greet(&quot;Charlie&quot;)
)</p>

asyncio.run(main())

ในตัวอย่างนี้ เรามีฟังก์ชันแบบ Asynchronous ชื่อ greet ที่จำลองการดำเนินการ I/O ด้วย asyncio.sleep() ฟังก์ชัน main ใช้ asyncio.gather() เพื่อดำเนินการการ поздорованияหลายครั้งพร้อมกัน แม้ว่าจะมีการหน่วงเวลา 1 วินาที แต่การ поздорованияทั้งสามครั้งจะเสร็จสิ้นในเวลาประมาณ 1 วินาที ซึ่งแสดงถึงพลังของการดำเนินการแบบ Asynchronous

ความจำเป็นของ Asynchronous ในการเรียก API LLM

เมื่อทำงานกับ API LLM เรามักพบสถานการณ์ที่ต้องทำการเรียก API หลายครั้ง ไม่ว่าจะเป็นแบบลำดับหรือแบบขนาน การเขียนโค้ดแบบสynchronous สามารถนำไปสู่ปัญหาการทำงานที่ช้าได้ โดยเฉพาะอย่างยิ่งเมื่อทำงานกับการดำเนินการล่าช้าที่สูง เช่น การร้องขอเครือข่ายไปยังบริการ LLM

ลองพิจารณาสถานการณ์ที่เราต้องสร้างสรุปสำหรับบทความ 100 เรื่องโดยใช้ API LLM ด้วยวิธีการแบบสynchronous การเรียก API แต่ละครั้งจะถูกบล็อกจนกว่าจะได้รับการตอบกลับ ซึ่งอาจใช้เวลาหลายนาทีในการดำเนินการすべて การใช้วิธีการแบบ Asynchronous ช่วยให้เราสามารถเริ่มการเรียก API หลายครั้งพร้อมกัน ซึ่งลดเวลาในการดำเนินการโดยรวมลงอย่างมาก

การกำหนดค่าสภาพแวดล้อมของคุณ

เพื่อเริ่มต้นการเรียก API LLM แบบ Asynchronous คุณจะต้องกำหนดค่าสภาพแวดล้อม Python ของคุณด้วยไลบรารีที่จำเป็น ต่อไปนี้คือสิ่งที่คุณต้องการ:

  • Python 3.7 หรือสูงกว่า (สำหรับการสนับสนุน asyncio แบบเนทีฟ)
  • aiohttp: ไลบรารีไคลเอ็นต์ HTTP แบบ Asynchronous
  • openai: ไลบรารีไคลเอ็นต์ Python ของ OpenAI (หากคุณใช้โมเดล GPT ของ OpenAI)
  • langchain: เฟรมเวิร์กสำหรับการสร้างแอปพลิเคชันด้วย LLM (ไม่จำเป็น แต่แนะนำสำหรับ 워์กโฟลว์ที่ซับซ้อน)

คุณสามารถติดตั้งความพึ่งพาเหล่านี้โดยใช้ pip:


<p>pip install aiohttp openai langchain
&lt;div class=&quot;relative flex flex-col rounded-lg&quot;&gt;

การเรียก API LLM แบบ Asynchronous พื้นฐานด้วย asyncio และ aiohttp

มาทำการเรียก API LLM แบบ Asynchronous ที่เรียบง่ายโดยใช้ aiohttp กัน เราจะใช้ API GPT-3.5 ของ OpenAI เป็นตัวอย่าง แต่แนวคิดเหล่านี้สามารถใช้กับ API LLM อื่นๆ ได้

import asyncio
import aiohttp
from openai import AsyncOpenAI

<p>async def generate_text(prompt, client):
response = await client.chat.completions.create(
model=&quot;gpt-3.5-turbo&quot;,
messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: prompt}]
)
return response.choices[0].message.content</p>

<p>async def main():
prompts = [
&quot;อธิบายการคำนวณแบบควอนตัมในคำพูดที่เรียบง่าย.&quot;,
&quot;เขียนฮะอิกูเกี่ยวกับปัญญาประดิษฐ์.&quot;,
&quot;อธิบายกระบวนการของการ фотอสินเทซิส.&quot;
]</p>

<p>async with AsyncOpenAI() as client:
tasks = [generate_text(prompt, client) for prompt in prompts]
results = await asyncio.gather(*tasks)</p>

<p>for prompt, result in zip(prompts, results):
print(f&quot;คำถาม: {prompt}\nคำตอบ: {result}\n&quot;)</p>

asyncio.run(main())

ในตัวอย่างนี้ เรามีฟังก์ชันแบบ Asynchronous ชื่อ generate_text ที่ทำการเรียก API LLM โดยใช้ AsyncOpenAI ฟังก์ชัน main สร้างงานสำหรับคำถามหลายข้อและใช้ asyncio.gather() เพื่อดำเนินการพร้อมกัน

แนวทางนี้ช่วยให้เราสามารถส่งคำถามหลายข้อไปยัง API LLM ในเวลาเดียวกัน ซึ่งลดเวลาทั้งหมดที่ต้องใช้ในการประมวลผลคำถามทั้งหมด

เทคนิคขั้นสูง: การแบตช์และการควบคุมConcurrency

ในขณะที่ตัวอย่างก่อนหน้านี้แสดงให้เห็นถึงพื้นฐานของการเรียก API LLM แบบ Asynchronous แอปพลิเคชันในโลกแห่งความเป็นจริงมักต้องการแนวทางที่ซับซ้อนมากขึ้น มาทำความรู้จักกับเทคนิคสองอย่างที่สำคัญ: การแบตช์คำถามและการควบคุมConcurrency

การแบตช์คำถาม: เมื่อต้องจัดการกับจำนวนคำถามที่มาก การแบตช์คำถามเป็นกลุ่มแทนที่จะส่งคำถามแต่ละข้อแยกกันสามารถช่วยปรับปรุงประสิทธิภาพได้

import asyncio
from openai import AsyncOpenAI

<p>async def process_batch(batch, client):
responses = await asyncio.gather(*[
client.chat.completions.create(
model=&quot;gpt-3.5-turbo&quot;,
messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: prompt}]
) for prompt in batch
])
return [response.choices[0].message.content for response in responses]</p>

<p>async def main():
prompts = [f&quot;บอกข้อมูลเกี่ยวกับตัวเลข {i}&quot; for i in range(100)]
batch_size = 10</p>

<p>async with AsyncOpenAI() as client:
results = []
for i in range(0, len(prompts), batch_size):
batch = prompts[i:i+batch_size]
batch_results = await process_batch(batch, client)
results.extend(batch_results)</p>

<p>for prompt, result in zip(prompts, results):
print(f&quot;คำถาม: {prompt}\nคำตอบ: {result}\n&quot;)</p>

asyncio.run(main())

การควบคุมConcurrency: แม้ว่าการเขียนโปรแกรมแบบ Asynchronous ช่วยให้สามารถดำเนินการหลายอย่างพร้อมกัน แต่ก็สำคัญที่จะต้องควบคุมระดับของConcurrency เพื่อหลีกเลี่ยงการให้ API ซึ่งสามารถทำได้โดยใช้ asyncio.Semaphore

import asyncio
from openai import AsyncOpenAI

<p>async def generate_text(prompt, client, semaphore):
async with semaphore:
response = await client.chat.completions.create(
model=&quot;gpt-3.5-turbo&quot;,
messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: prompt}]
)
return response.choices[0].message.content</p>

<p>async def main():
prompts = [f&quot;บอกข้อมูลเกี่ยวกับตัวเลข {i}&quot; for i in range(100)]
max_concurrent_requests = 5
semaphore = asyncio.Semaphore(max_concurrent_requests)</p>

<p>async with AsyncOpenAI() as client:
tasks = [generate_text(prompt, client, semaphore) for prompt in prompts]
results = await asyncio.gather(*tasks)</p>

<p>for prompt, result in zip(prompts, results):
print(f&quot;คำถาม: {prompt}\nคำตอบ: {result}\n&quot;)</p>

asyncio.run(main())

ในตัวอย่างนี้ เราใช้เซมาฟอร์ในการจำกัดจำนวนคำขอพร้อมกันไว้ที่ 5 เพื่อให้แน่ใจว่าเราไม่ทำให้ API ท่วมท้น

การรับมือข้อผิดพลาดและการทำซ้ำใน API LLM แบบ Asynchronous

เมื่อทำงานกับ API ภายนอก มันสำคัญที่จะนำการรับมือข้อผิดพลาดและการทำซ้ำมาใช้เพื่อให้แน่ใจว่าแอปพลิเคชันของคุณมีความทนทานต่อข้อผิดพลาดและสามารถฟื้นตัวได้จากความล้มเหลว

import asyncio
import random
from openai import AsyncOpenAI
from tenacity import retry, stop_after_attempt, wait_exponential

class APIError(Exception):
pass

<p>@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=4, max=10))
async def generate_text_with_retry(prompt, client):
try:
response = await client.chat.completions.create(
model=&quot;gpt-3.5-turbo&quot;,
messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: prompt}]
)
return response.choices[0].message.content
except Exception as e:
print(f&quot;ข้อผิดพลาดเกิดขึ้น: {e}&quot;)
raise APIError(&quot;ไม่สามารถสร้างข้อความได้&quot;)</p>

<p>async def process_prompt(prompt, client, semaphore):
async with semaphore:
try:
result = await generate_text_with_retry(prompt, client)
return prompt, result
except APIError:
return prompt, &quot;ไม่สามารถสร้างคำตอบหลังจากหลายครั้ง&quot;</p>

<p>async def main():
prompts = [f&quot;บอกข้อมูลเกี่ยวกับตัวเลข {i}&quot; for i in range(20)]
max_concurrent_requests = 5
semaphore = asyncio.Semaphore(max_concurrent_requests)</p>

<p>async with AsyncOpenAI() as client:
tasks = [process_prompt(prompt, client, semaphore) for prompt in prompts]
results = await asyncio.gather(*tasks)</p>

<p>for prompt, result in results:
print(f&quot;คำถาม: {prompt}\nคำตอบ: {result}\n&quot;)</p>

asyncio.run(main())

ในตัวอย่างนี้ เราเพิ่มการรับมือข้อผิดพลาดและการทำซ้ำเข้าไปในโค้ด โดยใช้ไลบรารี tenacity เพื่อทำซ้ำการเรียก API หลังจากระยะเวลาที่กำหนด

การเพิ่มประสิทธิภาพ: การสตรีมคำตอบ

สำหรับการสร้างเนื้อหาที่ยาว การสตรีมคำตอบสามารถปรับปรุงประสิทธิภาพที่รับรู้ของแอปพลิเคชันของคุณได้ โดยไม่ต้องรอจนกว่าจะได้รับคำตอบทั้งหมด คุณสามารถประมวลผลและแสดงส่วนของข้อความที่พร้อมแล้ว

import asyncio
from openai import AsyncOpenAI

<p>async def stream_text(prompt, client):
stream = await client.chat.completions.create(
model=&quot;gpt-3.5-turbo&quot;,
messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: prompt}],
stream=True
)</p>

<p>full_response = &quot;&quot;
async for chunk in stream:
if chunk.choices[0].delta.content is not None:
content = chunk.choices[0].delta.content
full_response += content
print(content, end=&#039;&#039;, flush=True)</p>

<p>print(&quot;\n&quot;)
return full_response</p>

<p>async def main():
prompt = &quot;เขียนเรื่องสั้นเกี่ยวกับนักวิทยาศาสตร์ที่เดินทางข้ามเวลา&quot;</p>

<p>async with AsyncOpenAI() as client:
result = await stream_text(prompt, client)</p>

<p>print(f&quot;คำตอบทั้งหมด:\n{result}&quot;)</p>

asyncio.run(main())

ในตัวอย่างนี้ เราแสดงให้เห็นวิธีการสตรีมคำตอบจาก API โดยพิมพ์ส่วนของข้อความที่พร้อมแล้ว ซึ่งเป็นประโยชน์อย่างยิ่งสำหรับแอปพลิเคชันที่ต้องการให้ผู้ใช้เห็นผลลัพธ์แบบเรียลไทม์

การสร้าง 워์กโฟลว์แบบ Asynchronous ด้วย LangChain

สำหรับแอปพลิเคชันที่ใช้ LLM ที่ซับซ้อน เฟรมเวิร์ก LangChain ให้การสร้างสูงสำหรับการเชื่อมต่อการเรียก LLM หลายครั้งและรวมเครื่องมืออื่นๆ มาทำให้เราสามารถสร้าง 워์กโฟลว์ที่ซับซ้อนได้

ตัวอย่างนี้แสดงให้เห็นวิธีการใช้ LangChain สำหรับการสร้าง 워์กโฟลว์ที่มีการสตรีมและการดำเนินการแบบ Asynchronous

import asyncio
from langchain.llms import OpenAI
from langchain.prompts import PromptTemplate
from langchain.chains import LLMChain
from langchain.callbacks.manager import AsyncCallbackManager
from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler

<p>async def generate_story(topic):
llm = OpenAI(temperature=0.7, streaming=True, callback_manager=AsyncCallbackManager([StreamingStdOutCallbackHandler()]))
prompt = PromptTemplate(
input_variables=[&quot;topic&quot;],
template=&quot;เขียนเรื่องสั้นเกี่ยวกับ {topic}&quot;
)
chain = LLMChain(llm=llm, prompt=prompt)
return await chain.arun(topic=topic)</p>

<p>async def main():
topics = [&quot;ป่ามหัศจรรย์&quot;, &quot;เมืองในอนาคต&quot;, &quot;อารยธรรมใต้น้ำ&quot;]
tasks = [generate_story(topic) for topic in topics]
stories = await asyncio.gather(*tasks)</p>

<p>for topic, story in zip(topics, stories):
print(f&quot;\nหัวข้อ: {topic}\nเรื่องราว: {story}\n{&#039;=&#039;*50}\n&quot;)</p>

asyncio.run(main())

การให้บริการแอปพลิเคชัน LLM แบบ Asynchronous ด้วย FastAPI

เพื่อให้แอปพลิเคชัน LLM แบบ Asynchronous ของคุณพร้อมให้บริการเป็นเว็บเซอร์วิส FastAPI เป็นตัวเลือกที่ดีเนื่องจากการสนับสนุนการดำเนินการแบบ Asynchronous โดยธรรมชาติ

from fastapi import FastAPI, BackgroundTasks
from pydantic import BaseModel
from openai import AsyncOpenAI

app = FastAPI()
client = AsyncOpenAI()

<p>class GenerationRequest(BaseModel):
prompt: str</p>

<p>class GenerationResponse(BaseModel):
generated_text: str</p>

<p>@app.post(&quot;/generate&quot;, response_model=GenerationResponse)
async def generate_text(request: GenerationRequest, background_tasks: BackgroundTasks):
response = await client.chat.completions.create(
model=&quot;gpt-3.5-turbo&quot;,
messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: request.prompt}]
)
generated_text = response.choices[0].message.content</p>

<p># จำลองการประมวลผลเพิ่มเติมในพื้นหลัง
background_tasks.add_task(log_generation, request.prompt, generated_text)</p>

<p>return GenerationResponse(generated_text=generated_text)</p>

<p>async def log_generation(prompt: str, generated_text: str):
# จำลองการบันทึกหรือการประมวลผลเพิ่มเติม
await asyncio.sleep(2)
print(f&quot;บันทึก: คำถาม &#039;{prompt}&#039; สร้างข้อความที่มีความยาว {len(generated_text)}&quot;)</p>

<p>if __name__ == &quot;__main__&quot;:
import uvicorn
uvicorn.run(app, host=&quot;0.0.0.0&quot;, port=8000)

แอปพลิเคชัน FastAPI นี้สร้างจุดสิ้นสุด /generate ที่รับคำถามและคืนคำตอบที่สร้างขึ้น นอกจากนี้ยังแสดงให้เห็นวิธีการใช้ภารกิจพื้นหลังสำหรับการประมวลผลเพิ่มเติมโดยไม่บล็อกคำตอบ

แนวทางปฏิบัติที่ดีที่สุดและข้อผิดพลาดทั่วไป

เมื่อทำงานกับการเรียก API LLM แบบ Asynchronous ควรปฏิบัติตามแนวทางปฏิบัติที่ดีต่อไปนี้:

  1. ใช้การเชื่อมต่อแบบพูล: เมื่อทำการเรียกหลายครั้ง ให้ใช้การเชื่อมต่อแบบพูลเพื่อลดการโอเวอร์เฮด
  2. นำการรับมือข้อผิดพลาดมาใช้: พิจารณาปัญหาเครือข่าย ข้อผิดพลาด API และคำตอบที่ไม่คาดคิด
  3. เคารพขีดจำกัดอัตรา: ใช้เซมาฟอร์หรือกลไกการควบคุมConcurrencyอื่นๆ เพื่อหลีกเลี่ยงการให้ API ท่วมท้น
  4. ติดตามและบันทึก: นำการบันทึกที่ครอบคลุมมาใช้เพื่อติดตามประสิทธิภาพและระบุปัญหา
  5. ใช้การสตรีมสำหรับเนื้อหาที่ยาว: ช่วยให้ผู้ใช้ได้รับผลลัพธ์แบบเรียลไทม์และช่วยให้สามารถประมวลผลส่วนของข้อความที่พร้อมแล้ว

ฉันใช้เวลาที่ผ่านมา 5 ปีในการศึกษาสิ่งที่น่าสนใจเกี่ยวกับ Machine Learning และ Deep Learning ความเชี่ยวชาญและความหลงใหลของฉันทำให้ฉันเข้าร่วมในโครงการพัฒนาซอฟต์แวร์มากกว่า 50 โครงการที่มีความหลากหลาย โดยมุ่งเน้นไปที่ AI/ML ความอยากรู้อยากเห็นของฉันยังทำให้ฉันสนใจในด้าน Natural Language Processing ซึ่งเป็นสาขาที่ฉันต้องการสำรวจต่อไป