DokumentationStreaming-Antworten
Entwickeln
Streaming-Antworten
Die Antwort erscheint, während das Modell sie schreibt. Die Person liest die ersten Wörter, statt auf das Ende zu warten.
Auf dieser Seite
Streaming aktivieren
Setzen Sie stream auf true. Die Antwort kommt als Server-Sent Events, jedes mit einem Stück des Textes.
stream = client.chat.completions.create(
model="flash",
messages=[
{"role": "user", "content": "Draft a payment reminder."}
],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(
chunk.choices[0].delta.content, end="", flush=True
)
if chunk.usage:
print("\n", chunk.usage.total_tokens, "tokens")const stream = await client.chat.completions.create({
model: "flash",
messages: [
{ role: "user", content: "Draft a payment reminder." },
],
stream: true,
})
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "")
if (chunk.usage)
console.log("\n", chunk.usage.total_tokens, "tokens")
}import type { ChatCompletionChunk } from "openai/resources"
const stream: AsyncIterable<ChatCompletionChunk> =
await client.chat.completions.create({
model: "flash",
messages: [
{ role: "user", content: "Draft a payment reminder." },
],
stream: true,
})
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "")
if (chunk.usage)
console.log("\n", chunk.usage.total_tokens, "tokens")
}curl -N https://api.learnya.ai/v1/chat/completions \
-H "Authorization: Bearer $LEARNYA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "flash",
"stream": true,
"messages": [
{"role": "user", "content": "Draft a payment reminder."}
]
}'Antwort
data: {"id":"chatcmpl-91a7","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
data: {"id":"chatcmpl-91a7","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Dear Ms Keller,"}}]}
data: {"id":"chatcmpl-91a7","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-91a7","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":17,"completion_tokens":96,"total_tokens":113}}
data: [DONE]Das Format
| Event | Was es enthält |
|---|---|
delta.content | Das nächste Textstück |
delta.reasoning_content | Das Schlussfolgern, wenn Sie es angefordert haben |
finish_reason | Der Grund für das Ende, im letzten Textstück |
usage | Die Token-Zählung, in einem letzten Event ohne choices |
[DONE] | Das Ende des Streams |
Die Zeitlimits
- Ein Stream, der zwei Minuten lang still bleibt, wird mit einem letzten Fehler-Event beendet.
- Eine Antwort dauert höchstens zehn Minuten.
- Um eine Antwort abzubrechen, schliessen Sie die Verbindung.