DocumentazioneRisposte in streaming
Sviluppa
Risposte in streaming
La risposta appare mentre il modello la scrive. La persona legge le prime parole invece di aspettare la fine.
In questa pagina
Attivare lo streaming
Imposta stream a true. La risposta arriva come eventi inviati dal server, ognuno con un frammento del testo.
stream = client.chat.completions.create(
model="flash",
messages=[
{"role": "user", "content": "Draft a payment reminder."}
],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(
chunk.choices[0].delta.content, end="", flush=True
)
if chunk.usage:
print("\n", chunk.usage.total_tokens, "tokens")const stream = await client.chat.completions.create({
model: "flash",
messages: [
{ role: "user", content: "Draft a payment reminder." },
],
stream: true,
})
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "")
if (chunk.usage)
console.log("\n", chunk.usage.total_tokens, "tokens")
}import type { ChatCompletionChunk } from "openai/resources"
const stream: AsyncIterable<ChatCompletionChunk> =
await client.chat.completions.create({
model: "flash",
messages: [
{ role: "user", content: "Draft a payment reminder." },
],
stream: true,
})
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "")
if (chunk.usage)
console.log("\n", chunk.usage.total_tokens, "tokens")
}curl -N https://api.learnya.ai/v1/chat/completions \
-H "Authorization: Bearer $LEARNYA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "flash",
"stream": true,
"messages": [
{"role": "user", "content": "Draft a payment reminder."}
]
}'Risposta
data: {"id":"chatcmpl-91a7","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
data: {"id":"chatcmpl-91a7","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Dear Ms Keller,"}}]}
data: {"id":"chatcmpl-91a7","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-91a7","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":17,"completion_tokens":96,"total_tokens":113}}
data: [DONE]Il formato
| Evento | Cosa contiene |
|---|---|
delta.content | Il frammento di testo successivo |
delta.reasoning_content | Il ragionamento, se lo hai richiesto |
finish_reason | Il motivo dell’arresto, sull’ultimo frammento del testo |
usage | Il conteggio dei token, in un ultimo evento senza scelte |
[DONE] | La fine dello stream |
I tempi limite
- Uno stream silenzioso per due minuti viene interrotto, con un ultimo evento di errore.
- Una risposta dura al massimo dieci minuti.
- Per interrompere una risposta, chiudi la connessione.