Introduction
Uploading a document should not make a mobile app wait while the backend extracts text, splits pages, creates embeddings, and updates search. Those steps can take time and fail independently. A reliable API accepts the request, saves a job, and lets a separate worker complete it while the app checks progress.
This guide is for backend engineers building document features for mobile or AI products. It assumes Python, FastAPI, and a database. The durable job-store example below uses only Python's standard library and was checked with Python 3.14.7; it is a local reference implementation, not a claim that SQLite is the right queue for a large deployment. The FastAPI route shape follows the official API, but the route and a real document processor must be integration-tested in your environment.
If you are new to APIs, begin with the FastAPI guide. If your product will answer questions over uploaded documents, the RAG guide explains retrieval and generation. This article focuses on the missing operational step: getting documents into that system without losing work.
Separate Acceptance From Completion
The client first uploads a document to controlled storage. The upload endpoint validates size, type, ownership, and a storage reference. A second request creates a processing job and returns HTTP 202 with a job ID. HTTP 202 means the work was accepted, not that the document is searchable.
| Endpoint | Purpose | Typical response |
|---|---|---|
| POST /jobs | Create or reuse a job for an uploaded document | 202 with job ID and status URL |
| GET /jobs/ | Return a job owned by the current user | queued, running, completed, or failed |
| Worker process | Parse, index, and record a result | No public endpoint |
The first two endpoints must enforce document and job ownership. A predictable job ID is not authorization. Do not accept an arbitrary local file path or public URL as a document reference without validation.
FastAPI's BackgroundTasks documentation is useful for small work after a response. It also notes that heavier work may benefit from a separate system such as Celery. An in-process task alone is not a durable job record: if a process disappears, the API needs a way to discover and resume unfinished work.
Record the Job Before Returning 202
Here is a compact SQLite-backed job store for understanding the state transitions. Save it as jobs.py. It demonstrates deduplicating a client retry with a request key, claiming a job under a transaction, expiring a worker lease, and preventing an old worker from marking a reclaimed job complete.
import os
import sqlite3
import time
from contextlib import closing
from pathlib import Path
from uuid import uuid4
DB = Path(os.environ.get("JOB_DB", "jobs.sqlite3"))
MAX_ATTEMPTS = 3
LEASE_SECONDS = 30
def connect():
db = sqlite3.connect(DB, timeout=10)
db.row_factory = sqlite3.Row
return db
def init():
with closing(connect()) as db, db:
db.execute("""
CREATE TABLE IF NOT EXISTS jobs (
id TEXT PRIMARY KEY,
owner_id TEXT NOT NULL,
request_key TEXT NOT NULL,
document_ref TEXT NOT NULL,
state TEXT NOT NULL,
attempts INTEGER NOT NULL DEFAULT 0,
available_at REAL NOT NULL,
lease_until REAL,
lease_token TEXT,
result_ref TEXT,
error TEXT,
UNIQUE(owner_id, request_key)
)
""")
def enqueue(owner_id, request_key, document_ref):
with closing(connect()) as db, db:
db.execute(
"""INSERT INTO jobs
(id, owner_id, request_key, document_ref, state, available_at)
VALUES (?, ?, ?, ?, 'queued', ?)
ON CONFLICT(owner_id, request_key) DO NOTHING""",
(uuid4().hex, owner_id, request_key, document_ref, time.time()),
)
row = db.execute(
"""SELECT id, document_ref FROM jobs
WHERE owner_id = ? AND request_key = ?""",
(owner_id, request_key),
).fetchone()
if row["document_ref"] != document_ref:
raise ValueError("Request key already belongs to another document")
return row["id"]
def claim():
now = time.time()
with closing(connect()) as db, db:
db.execute("BEGIN IMMEDIATE")
db.execute(
"""UPDATE jobs SET state = 'failed', error = 'lease expired'
WHERE state = 'running' AND lease_until <= ?
AND attempts >= ?""",
(now, MAX_ATTEMPTS),
)
row = db.execute(
"""SELECT id, owner_id, document_ref FROM jobs
WHERE attempts < ? AND
((state = 'queued' AND available_at <= ?)
OR (state = 'running' AND lease_until <= ?))
ORDER BY available_at, id LIMIT 1""",
(MAX_ATTEMPTS, now, now),
).fetchone()
if row is None:
return None
token = uuid4().hex
db.execute(
"""UPDATE jobs SET state = 'running',
attempts = attempts + 1, lease_until = ?,
lease_token = ? WHERE id = ?""",
(now + LEASE_SECONDS, token, row["id"]),
)
return {"id": row["id"], "owner_id": row["owner_id"],
"document_ref": row["document_ref"], "token": token}
def complete(job_id, token, result_ref):
with closing(connect()) as db, db:
changed = db.execute(
"""UPDATE jobs SET state = 'completed', result_ref = ?,
lease_until = NULL, lease_token = NULL
WHERE id = ? AND lease_token = ? AND state = 'running'""",
(result_ref, job_id, token),
).rowcount
return changed == 1
def fail(job_id, token, safe_error, retryable):
with closing(connect()) as db, db:
db.execute("BEGIN IMMEDIATE")
row = db.execute(
"""SELECT attempts FROM jobs
WHERE id = ? AND lease_token = ? AND state = 'running'""",
(job_id, token),
).fetchone()
if row is None:
return False
retry = retryable and row["attempts"] < MAX_ATTEMPTS
next_time = time.time() + min(60, 2 ** row["attempts"]) if retry else 0
db.execute(
"""UPDATE jobs SET state = ?, available_at = ?, error = ?,
lease_until = NULL, lease_token = NULL WHERE id = ?""",
("queued" if retry else "failed", next_time, safe_error, job_id),
)
return True
def get_job(owner_id, job_id):
with closing(connect()) as db:
row = db.execute(
"""SELECT id, state, attempts, result_ref, error FROM jobs
WHERE owner_id = ? AND id = ?""",
(owner_id, job_id),
).fetchone()
return dict(row) if row else None
The example uses a short 30-second lease to make failure testing easy. A real job that runs longer needs a heartbeat that renews the lease, or a lease long enough for a bounded task. The completion check stops a stale worker from changing the job row, but the indexing side effect must also be idempotent. For example, index by document ID and document version instead of appending duplicate chunks on each retry. The request key is scoped to the authenticated owner, and reusing it for a different document raises an error.
The fail() transition records a safe error summary, delays retryable failures, and marks a job failed after the final attempt. The sample intentionally does not classify provider, parsing, storage, and validation errors for you. That classification depends on your processor. Do not retry a permanently invalid file forever.
Keep FastAPI Thin
With the job store in place, the API should validate input and expose state. The document reference below is assumed to come from an earlier authenticated upload, not directly from untrusted free text.
from fastapi import Depends, FastAPI, HTTPException, status
from pydantic import BaseModel
from jobs import enqueue, get_job, init
from my_auth import authenticated_user_id # Supply this in your application.
app = FastAPI()
init()
class JobRequest(BaseModel):
request_key: str
document_ref: str
@app.post("/jobs", status_code=status.HTTP_202_ACCEPTED)
def submit_job(
request: JobRequest, user_id: str = Depends(authenticated_user_id)
):
# Verify user_id owns document_ref before calling enqueue().
job_id = enqueue(user_id, request.request_key, request.document_ref)
return {"job_id": job_id, "status_url": f"/jobs/{job_id}"}
@app.get("/jobs/{job_id}")
def job_status(
job_id: str, user_id: str = Depends(authenticated_user_id)
):
job = get_job(user_id, job_id)
if job is None:
raise HTTPException(status_code=404, detail="Job not found")
return job
The authentication import and ownership check are required integration points, not optional security polish. Implement them in your own application, validate the stored upload, and set request size limits before using this with customer documents. Map the repeated-key ValueError to a clear client error. FastAPI's file upload guide covers the upload interface; it does not replace your storage and authorization design.
Run a separate worker process, not a loop inside the FastAPI server. The processor below receives an authenticated owner's document reference from the job store. It must return a stable result reference and avoid duplicate indexing when a job is retried.
from jobs import claim, complete, fail
class InvalidDocument(Exception):
pass
def run_one(process_document):
job = claim()
if job is None:
return False
try:
result_ref = process_document(job["owner_id"], job["document_ref"])
except InvalidDocument:
fail(job["id"], job["token"], "Invalid document", retryable=False)
except (TimeoutError, OSError):
fail(job["id"], job["token"], "Temporary processing error", retryable=True)
else:
complete(job["id"], job["token"], result_ref)
return True
Call run_one() from a worker loop with a short pause when no job is ready. Log unexpected exceptions; the lease will eventually expire, but repeated unexpected failures need an alert. For multi-instance production workloads, use a queue or database approach designed for concurrent workers. PostgreSQL's SKIP LOCKED option is specifically documented for avoiding contention in queue-like tables. The right choice depends on throughput and operational capacity.
Test Durability, Not Just a Happy Path
A local smoke test should prove that a job record survives a new Python process and that a duplicate request key returns the same job ID. Then test the worker and API together:
| Failure case | What to verify |
|---|---|
| API restarts after returning 202 | The stored job is still visible and claimable |
| Client repeats POST after timeout | One logical job exists for that user and request key |
| Worker dies after claiming | An expired lease makes the job eligible again |
| Old worker finishes after a new claim | Its stale token cannot complete the job |
| Indexing succeeds, completion response fails | Retry does not create duplicate document chunks |
| Document is invalid or inaccessible | A safe failure state replaces endless retries |
| User A requests User B's job | The API refuses to reveal its state or result |
Track queue age, running time, attempt count, failure reasons, and the age of expired leases. A single “AI request duration” number hides whether the delay came from upload, parsing, embedding, indexing, or the model. The streaming LLM guide covers delivery of generated text after the knowledge base is ready; it solves a different stage of the product.
Conclusion
A reliable document-processing backend makes an accepted upload visible and recoverable. Persist the job before returning 202, make retries safe, give workers leases, and separate API requests from processing. Then test restarts and duplicate work deliberately.
The SQLite store is a learning-sized core, not a drop-in production service. Add tenant isolation, upload validation, safe failure handling, idempotent indexing, worker heartbeats, and an appropriately scaled queue before handling customer documents.
Planning document AI for a mobile or enterprise product? I help teams build FastAPI backends that can process real workloads and recover cleanly from failures. Book a meeting to review your backend design.
Interested in working together?
Let's discuss your project and explore how I can help bring it to life.
