LLM API Call Beginner's Guide: First Request, OpenAI Protocol & Streaming

This tutorial covers LLM API fundamentals including OpenAI-compatible protocol, making first requests via curl and SDK, multi-turn conversation with message history, key parameters like temperature and top_p, token billing mechanics, streaming responses, and common error handling.

Code Farmer Manor Chronicle
Code Farmer Manor Chronicle
Code Farmer Manor Chronicle
LLM API Call Beginner's Guide: First Request, OpenAI Protocol & Streaming

01 OpenAI-Compatible Protocol: Learn Once, Use Everywhere

The OpenAI /v1/chat/completions request format has become the de facto industry standard. Major model providers (DeepSeek, Qwen, GLM, Kimi) and self-hosted solutions (vLLM, Ollama) all offer OpenAI-compatible endpoints. Switching providers only requires changing the base_url and model parameters; business logic remains unchanged.

Provider Base URLs

DeepSeek: api.deepseek.com Alibaba Qwen: dashscope.aliyuncs.com/compatible-mode/v1 Zhipu GLM: open.bigmodel.cn/api/paas/v4 Local vLLM/Ollama:

localhost:8000/v1

02 First API Call

The raw HTTP POST request using curl:

curl https://api.deepseek.com/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $DEEPSEEK_API_KEY" \
-d '{"model": "deepseek-chat",
"messages": [{"role": "user",
"content": "Explain idempotency in one sentence"}]}'

Using the official OpenAI Python SDK ( pip install openai):

from openai import OpenAI

client = OpenAI(api_key="sk-xxx", base_url="https://api.deepseek.com")

resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user",
    "content": "Explain idempotency in one sentence"}]
)
print(resp.choices[0].message.content)  # model reply
print(resp.usage)                       # token usage

Common pitfall: confusing the api_key (authentication credential) with base_url (gateway address). The key is obtained from each provider's console; the base URL differs per provider — do not blindly copy api.openai.com.

03 Messages: Multi-Turn Conversation Is History Concatenation

LLM services are stateless. "Multi-turn conversation" is implemented by the client appending previous exchanges into the messages array and sending the full history each request.

messages = [{"role": "system", "content": "You are a patient programming tutor"}]

messages.append({"role": "user", "content": "What is JVM GC?"})
messages.append({"role": "assistant",
"content": "GC is garbage collection..."})  # store previous reply verbatim
messages.append({"role": "user",
"content": "What is the difference compared to Python's?"})

resp = client.chat.completions.create(
    model="deepseek-chat", messages=messages)
messages.append({"role": "assistant",
"content": resp.choices[0].message.content})

Message Roles

system — defines persona and rules, highest priority (analogy: job description)

user — user input (analogy: each new task)

assistant — model's historical replies (can also demonstrate desired format) (analogy: work record)

The context window is limited; unlimited history concatenation will eventually exceed it. Production systems use sliding windows and summarization, offloading older content to a vector store for retrieval (see RAG).

04 Core Parameters: temperature and top_p

temperature — lower values make output more deterministic, higher values increase randomness. Practical guidance: 0–0.3 for extraction/code, 0.7 for chat, 1.0+ for creative tasks.

top_p — nucleus sampling: sample from the smallest set of tokens whose cumulative probability exceeds p. Use either temperature or top_p, not both simultaneously.

max_tokens — caps maximum output length; prevents runaway generation and controls cost.

stream — enables streaming (SSE) response; critical for reducing first-token latency (covered in SSE article).

Selection intuition: use temperature=0 for reproducible results (testing, extraction); increase for diversity (copywriting). Adjust only one variable at a time.

05 Token Billing: How Costs Are Calculated

Token is the model's billing unit — not characters or words. Rough conversions:

1 Chinese character ≈ 0.6–1 token

1 English word ≈ 1.3 tokens

1000 Chinese characters ≈ 600–1000 tokens

Three billing rules:

Input and output are priced separately ; output is typically several times more expensive.

Every request charges for the entire message history — longer conversations cost more, motivating history compression.

Responses include exact usage counters : resp.usage.prompt_tokens and resp.usage.completion_tokens. Production systems must log these for cost tracking.

The real cost explosion comes from uncontrolled history concatenation in loops and stuffing entire long documents into prompts. Optimizations include tiered routing, semantic caching, and prompt caching.

06 Streaming Calls: Just One Parameter

Passing stream=True turns the endpoint into an SSE stream that yields incremental chunks:

stream = client.chat.completions.create(
    model="deepseek-chat",
    messages=[{"role": "user",
    "content": "Write a poem about code refactoring"}],
    stream=True
)
for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

First-token latency drops from "wait for full generation" to "first character appears", dramatically improving perceived responsiveness.

07 Common Error Quick Reference

401 Unauthorized — invalid key or key from wrong provider. Fix: verify key matches base_url provider.

400 Bad Request — malformed parameters or context length exceeded. Fix: log request body for debugging; compress history.

429 Too Many Requests — rate limit hit. Fix: exponential backoff retry (1s, 2s, 4s).

Long timeout / no response — proxy buffering or missing timeout setting. Fix: set explicit timeout; disable proxy buffering with X-Accel-Buffering: no.

08 Summary

LLM API call = a single HTTP POST with a messages array; OpenAI-compatible protocol lets one codebase work across providers.

Multi-turn dialogue is an illusion: the client concatenates history and resends it; context window is finite, so production must compress. temperature controls stability, usage controls billing: use 0 for testing/extraction, 0.7 for chat, and always persist token counts.

Streaming is just stream=True; production details covered in the SSE practical guide.

RAG simply means "retrieve results and concatenate into the user message"; Agent function calling just adds a tools field to the request. With this foundation, building higher-level features becomes straightforward.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

streamingerror handlingOpenAI compatibleLLM APItoken billingAPI tutorial
Code Farmer Manor Chronicle
Written by

Code Farmer Manor Chronicle

A heart like drifting clouds, ever at ease; a mind like flowing water, free to roam.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.