From Notebook to Production: Deploying a RAG App with Flask and Railway
A recent post refactored my LangChain RAG script so it stopped throwing deprecation warnings at you. It’s a nice script, but it has one big limitation: you have to run it from a terminal. That’s fine for me. It’s not fine for a high school student who wants to ask her phone what the dress code says.
My daughter picked a high school science fair project: a chatbot that answers questions about her school’s student handbook. Her school’s teams are the Red Stickers, so the project is called StickerGPT. I am her mentor, which mostly means I got to watch her take my script and turn it into a real web app that anyone can use. This post walks through what changed between the script and the deployed app.
You can find the code for this post on GitHub.
The Problem With a Script
The command line version does everything, every time. Each run loads the PDF, splits it into chunks, sends every chunk off to be embedded, builds the vectorstore, and only then answers your one question. For a 5 page policy document, nobody notices. For an 81 page student handbook, you’re paying to embed the whole thing again for every single question, and you’re waiting for it too.
A web app flips that around. The expensive part (load, split, embed, index) happens once, when the server starts. The cheap part (retrieve, generate) happens per question. If you remember the table from earlier, it’s the same four stages, just split across two moments in time:
| When | RAG stage | Function |
|---|---|---|
| Server start | Load the document | extract_data() |
| Server start | Chunk the text | split_text() |
| Server start | Embed and store | vectorize_and_store() |
| Each question | Retrieve and generate | answer_question() |
Reusing the Pipeline
The nice thing about breaking the original script into functions is that Flask can just import them. The RAG code lives in rag_faiss_vectorstore.py and still runs from the command line. The web app is a thin layer on top:
import os
from flask import Flask, render_template, request
from rag_faiss_vectorstore import (
DEFAULT_PDF_NAME,
answer_question,
extract_data,
split_text,
vectorize_and_store,
)
PDF_NAME = os.getenv("PDF_NAME", DEFAULT_PDF_NAME)
app = Flask(__name__)
def build_docstorage():
text = extract_data(PDF_NAME)
docs = split_text(text)
return vectorize_and_store(docs, os.getenv("OPENAI_API_KEY"))
# Process the PDF and build the vectorstore once, at server startup.
docstorage = build_docstorage()
@app.route("/", methods=["GET", "POST"])
def index():
answer = None
question = None
if request.method == "POST":
question = request.form.get("question", "").strip()
if question:
answer = answer_question(question, os.getenv("OPENAI_API_KEY"), docstorage)
return render_template("index.html", question=question, answer=answer)
The key line is docstorage = build_docstorage() sitting at module level. Python runs it once when the app is imported, and every request after that reuses the same in-memory FAISS index. No database, no saved index files, nothing to manage. The tradeoff is that the index gets rebuilt every time the server restarts, but embedding the handbook with text-embedding-3-small costs a fraction of a cent, so that’s a trade I’ll take for version 1.
The page itself is a single Jinja template with a text box and an answer card. It uses Bootstrap, because the main audience is students on phones, and Bootstrap makes “looks fine on a phone” nearly free.
One More LangChain Change
If you compare answer_question() in the repo to the version in my last post, you’ll notice it doesn’t use create_retrieval_chain anymore. With LangChain 1.x, the old chain helpers live in the langchain-classic package for backward compatibility. Rather than build a new app on a compatibility package, StickerGPT uses the LangChain Expression Language (LCEL) pipe syntax, which is what LangChain points you toward now:
def format_docs(docs):
return "\n\n".join(doc.page_content for doc in docs)
def answer_question(question, api_key, docstorage):
llm = ChatOpenAI(model=os.getenv("OPENAI_MODEL", "gpt-4o-mini"), api_key=api_key)
qa = (
{"context": docstorage.as_retriever() | format_docs, "question": RunnablePassthrough()}
| PROMPT
| llm
| StrOutputParser()
)
response = qa.invoke(question)
return response
I actually like this better for teaching. You can read the RAG pipeline left to right: the retriever pulls chunks, format_docs glues them into one block of text, the prompt template drops that text into {context}, the LLM answers, and StrOutputParser hands back a plain string instead of a message object. Same retrieval and augmented generation as before, just with nothing hidden.
The prompt keeps the same guardrail as the last version. It tells the model to say it doesn’t know when the handbook doesn’t cover the question. For a tool students might actually trust about school rules, that line matters more than anything else in the file.
Building It With Claude Code
Most of the Flask and deployment work was done with the Claude Code CLI. My daughter described what she wanted, Claude Code proposed changes, and we reviewed them before accepting. A few observations from watching a high schooler work this way:
- It’s a great explainer. Ask it why something is there, and you get a real answer. Several of the comments in the repo, like the ones in
gunicorn.conf.pybelow, came out of those conversations. - It does not replace understanding the RAG pipeline. The pipeline came from my blog post, and she had to understand it well enough to explain it to science fair judges. No tool does that part for you.
- Small, reviewable steps work best. “Make the page look better on a phone” is a good request. “Build me a chatbot” is not.
Why Railway
For hosting, my priorities were simple: ease of deployment first, then price. This is a science fair project, not a startup. I didn’t want to configure a VM, a reverse proxy, and TLS certificates for a chatbot that will mostly get used by a few dozen, or hundred, students and some judges.
Railway fit that well. You connect the GitHub repo, and every push to main triggers a new build and deploy. That means the workflow for my daughter is just: change the code, commit, push, and refresh the site a couple of minutes later. The .python-version file in the repo (3.12) tells the build which Python to use, and requirements.txt handles the rest.
The one secret, OPENAI_API_KEY, goes in Railway’s service variables instead of a .env file. Locally we still use .env, and it’s listed in .gitignore, so the key never ends up on GitHub. (If you take one thing away from this post, make it that one. Bots scrape public repos for API keys within minutes.)
Gunicorn, Not the Flask Dev Server
python app.py runs Flask’s built-in development server, which is great for working on your laptop and not meant for the internet. In production, Railway’s start command is simply:
gunicorn app:app
Gunicorn automatically picks up a gunicorn.conf.py file in the directory it starts from, so all the settings live in the repo instead of a long start command:
import os
# Railway tells the app which port to listen on through the PORT variable
bind = "0.0.0.0:" + os.environ.get("PORT", "8000")
# One worker keeps memory low: every worker loads its own copy of the vector index,
# so 2 workers roughly doubles RAM (and the Railway bill).
# Threads let that single worker handle a few questions at the same time.
workers = 1
threads = 4
# LLM calls can take 10 to 30 seconds. Gunicorn's default 30s timeout
# would kill slow answers mid-response, so give it more room.
timeout = 120
# Send logs to stdout/stderr so they show up in Railway's log viewer
accesslog = "-"
errorlog = "-"
Two settings in there are worth calling out because they’re RAG specific:
workers = 1. The usual Gunicorn advice is several workers per CPU. But each worker is a separate process, and each process runsbuild_docstorage()on its own. More workers means more copies of the FAISS index in memory and more embedding calls on every deploy. One worker with a few threads handles our traffic just fine, since most of a request’s time is spent waiting on OpenAI anyway.timeout = 120. Gunicorn kills any request that runs longer than 30 seconds by default. A slow LLM response can get close to that, and the user just sees an error. Bumping the timeout avoids that.
A Health Check That Means Something
The app also has a tiny /health endpoint:
@app.route("/health")
def health():
return {"status": "ok"}
It looks trivial, but think about when it can answer. Flask can’t serve any route until app.py finishes importing, and app.py can’t finish importing until the vectorstore is built. So if /health returns ok, the PDF loaded, the embeddings came back, and the index is ready. Point Railway’s health check at /health, and a deploy with a broken API key or a missing PDF never replaces the version that’s already working.
Running It Locally
If you want to try it yourself:
$ git clone https://github.com/forrestgumpfan1/stickergpt.git
$ cd stickergpt
$ pip install -r requirements.txt
$ python app.py
Add a .env file with your OPENAI_API_KEY first, then open http://127.0.0.1:5001. The command line version still works, too:
$ python rag_faiss_vectorstore.py "What is the cell phone policy?"
To point it at your own document, pass the PDF as a second argument on the command line, or set the PDF_NAME environment variable for the web app.
What’s Next
Version 1 is deliberately small: one PDF, no user accounts, and an index that lives in memory. That kept the project focused on the part that matters for a science fair, which is understanding how retrieval and generation work together. The obvious next steps are:
- Support multiple policy documents.
- Add user accounts.
- Show which handbook pages each answer came from, so students can check the source.
That last one is my favorite, and it’s a natural extension of the tip from the last post about printing response['context']. If you can see what the model saw, you can decide whether to trust the answer. For students asking about school rules, that’s a pretty good lesson in using AI, science fair or not.