Building an AI Chat Backend MVP with FastAPI: Auth, Tools & SSE Streaming

This article walks through building a production-ready AI chat backend using FastAPI, covering user authentication with JWT, per-user chat isolation, local tool calling (time, date diff, calculator), SSE streaming with proper connection handling, and transactional message persistence that leaves no partial records on errors or client disconnects.

Tech Ocean
Tech Ocean
Tech Ocean
Building an AI Chat Backend MVP with FastAPI: Auth, Tools & SSE Streaming

Project Scope and Architecture

The MVP implements ten business endpoints plus a health check: three for authentication (register, login, current user), five for chat management (create, list, detail, rename, delete), one for listing available tools, and one SSE endpoint for sending messages. The project structure is flat, with each file handling a single responsibility. Technologies used include FastAPI, SQLAlchemy 2.0 async with MySQL, PyJWT for tokens, pwdlib (Argon2) for password hashing, LangChain's create_agent for the model-tool loop, and pytest with httpx for testing.

Database Schema

Three tables: users (email, password hash, created_at), chats (id, user_id FK with CASCADE delete, title, system_prompt, tools JSON array, created_at), and messages (id, chat_id FK with CASCADE delete, role user/assistant, content, created_at). Tool invocations are not stored — only final user and assistant messages.

Configuration

All settings load from environment variables via Pydantic BaseSettings. Critical security choices: jwt_secret has no default and requires minimum 32 characters (app fails to start otherwise), preventing weak secrets in production. Model configuration supports any OpenAI-compatible API (default DeepSeek deepseek-flash at https://api.deepseek.com). A llm_fake flag enables a scripted local model for offline testing.

Authentication: Password Hashing and JWT

Registration hashes passwords with Argon2 via pwdlib. Login verifies using constant-time comparison: even for non-existent emails, a dummy hash is verified to prevent timing attacks that reveal registered emails. Both failure cases return identical "email or password wrong" messages. JWT payload contains only sub (user id), iat, and exp (24 hours). Verification hardcodes algorithms=["HS256"] and requires exp and sub claims, blocking algorithm-confusion attacks (e.g., alg: none). The get_current_user dependency opens a short DB session, validates the token, fetches the user, and closes the session immediately — crucial for not holding connections during streaming (see section 6).

Chat Ownership Enforcement

All chat-scoped endpoints use a get_owned_chat dependency that queries by both chat_id and user_id. Missing chats return 404 (not 403) to avoid leaking existence of other users' chats. Example: Bob accessing Alice's chat gets 404. The detail endpoint explicitly eager-loads messages via selectinload because async SQLAlchemy cannot lazy-load relationships. Deletion explicitly removes messages first, ensuring compatibility with SQLite (which doesn't enforce FK cascades by default) and MySQL alike.

Tools and Model Integration

Three local tools are exposed as LangChain @tool functions: current_time, days_between (YYYY-MM-DD format, returns error string on bad input so model can retry), and calculator. The calculator parses expressions into an AST and only allows numeric literals and basic operators (+, -, *, /, //, %, **, unary +/-), rejecting all other syntax. Exponentiation is capped at exponent 100 to prevent resource exhaustion (e.g., 9**9**9). Tools are registered in a dict; chat creation validates tool names upfront (422 on unknown).

The model uses ChatOpenAI from langchain-openai pointing at DeepSeek's OpenAI-compatible endpoint. DeepSeek's reasoning_content field is ignored by ChatOpenAI, so thinking traces never reach the frontend or database. The agent loop runs via LangChain's create_agent: model decides tool calls → executor runs them → results fed back → repeat until final answer. The run_agent function streams two modes simultaneously: messages for token-by-token text (typing effect) and updates for complete tool call/result objects. Tokens come from messages; tool calls and results from updates.

SSE Streaming with FastAPI 0.135+

FastAPI's built-in EventSourceResponse handles SSE: declare response_class=EventSourceResponse and yield ServerSentEvent objects. The framework automatically adds Cache-Control: no-cache, X-Accel-Buffering: no (disables Nginx buffering), and sends : ping comments every 15 seconds of idle to keep connections alive through proxies. Chinese characters are kept readable by using raw_data=json.dumps(data, ensure_ascii=False) instead of the default data parameter which escapes to \uXXXX. Five event types are emitted: token, tool_call, tool_result, error, and done (with message IDs).

Database Connection Management During Streaming

Standard FastAPI yield -based DB dependencies hold the connection for the entire request lifetime. For SSE, that means the connection is occupied for the full model generation time (seconds to minutes), exhausting the pool (default 15 connections). Solution: the send_message endpoint does NOT use the shared get_db dependency. Instead, load_chat_context opens a short session, fetches chat + recent 20 messages, closes the session, and returns a plain dataclass. The get_current_user dependency also uses its own short session for the same reason. This keeps connections free during model streaming.

Transactional Message Persistence: No Partial Records

Messages are only saved after the entire generation succeeds. The flow: collect all tokens into a list (inserting blank lines between tool-separated rounds), then in a single transaction insert both user and assistant messages. Three failure modes handled:

Model/tool error : catch Exception, log full traceback, yield error event with generic message, return without saving.

Client disconnect : catch both asyncio.CancelledError (cancelled while awaiting model) and GeneratorExit (cancelled at yield), log with event count, re-raise to let framework clean up. No save.

DB commit failure : log, yield error, return.

Only on successful commit is the done event emitted with the new message IDs. Frontend treats done as the success signal. HTTP status remains 200 even for errors because SSE headers are sent before streaming starts.

Testing with a Scripted Fake Model

LangChain's built-in fakes don't support bind_tools required by create_agent. A custom ScriptedChatModel implements _astream and bind_tools, driven by a "script" fixture that defines per-turn outputs: text chunks, optional tool calls, and optional fail_after index to simulate mid-generation failures. Tests use a temporary SQLite file (not :memory: because connection pools create isolated in-memory DBs) and set LLM_FAKE=true via environment variables before importing the app. Eleven tests cover: auth flows, timing-attack resistance, token validation, ownership enforcement (404 on others' chats), tool validation, full tool-call flow with event ordering, mid-generation error (no DB records), history propagation to second turn, and chat deletion cascades. Client disconnect is tested manually with curl against a live server because httpx's ASGI transport buffers the full response.

Running the Project

git clone https://github.com/liuy-byte/fastapi-ai-chat
cd fastapi-ai-chat
uv sync
docker compose up -d --wait  # starts MySQL
cp .env.example .env  # fill JWT_SECRET and LLM_API_KEY
uv run uvicorn app.main:app --reload

Generate a secure JWT secret: python -c "import secrets; print(secrets.token_urlsafe(48))". Without an API key, set LLM_FAKE=true to run entirely offline. The /docs endpoint uses HTTPBearer (not OAuth2PasswordBearer) so the Authorize button accepts a raw token. Frontend must use fetch (not EventSource) because the SSE endpoint requires POST with an Authorization header.

Known Limitations (Future Work)

Rate limiting — currently unlimited requests per user.

Refresh tokens — access tokens expire in 24h with no renewal or revocation.

Tool call persistence — only final messages stored; model may re-use historical numbers without re-invoking tools.

Concurrent messages in same chat — race condition on history read/write; needs locking or frontend gating.

Auto-generated titles — user must provide title.

Migrations — tables created from models at startup; needs Alembic for schema changes.

Core design decisions (ownership checks in dependencies, connection-free streaming, commit-then- done) remain stable regardless of added features.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LangChainFastAPIMVPTool CallingSSE StreamingJWT AuthenticationAI Chat BackendSQLAlchemy Async
Tech Ocean
Written by

Tech Ocean

Focused on AI programming, sharing ready-to-use development efficiency solutions.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.