My stack for serving an internal chatbot to 600 employees a day
By Jean-Baptiste de la Broise · 4 August 2026
- tools
- infrastructure
- LLM
- AI
Over the past year we've been running an internal chatbot at MDPI that around 600 employees use daily, for drafting emails, digging through internal knowledge, and answering questions about our own guidelines and training material. None of it runs through an external API: everything is hosted on-premise. Here's the stack behind it, piece by piece.
Model serving
The core of the system is vLLM, exposed through its OpenAI-compatible API, running a custom Qwen3.6-35B-A3B configuration on a range of data center GPUs each with at least 46Gb vRAM. Keeping this on-premise means no data ever leaves our infrastructure, but it also means we own every scaling and hardware decision that a hosted API would otherwise abstract away.
Qwen alone covers about 90% of the use cases we encounter. For the rest (tasks that need very long context or very high complexity) we fall back to GPT models hosted on Azure.
Middleware
Sitting in front of the model, a custom middleware layer handles logging and acts as an internal security boundary it allows prompt filtering and related checks before requests reach the model.
Tools and retrieval
- FastMCP serves our internal tools over MCP, everything from knowledge-base lookups to integrations with other company systems.
- Qdrant is the vector database behind our RAG layer, indexing MDPI's internal knowledge.
- GTE-Qwen2 embeds that knowledge base; SPECTER2 embeds academic articles specifically, since general-purpose embeddings underperform on scholarly text.
- Firecrawl handles web search when the assistant needs to reach outside our own knowledge base.
Monitoring, and the app layer
- Kibana gives us real-time log monitoring.
- LibreChat is the front-end application employees actually talk to.
- Everything runs on Kubernetes, though orchestrating LLM workloads across heterogeneous GPU hardware is one place where Kubernetes shows real limits. That's probably worth its own post.
Usage analytics
A daily Airflow pipeline exports anonymized chat data into Superset, giving us a more detailed picture of usage patterns than raw logs alone.
What people actually use it for
A few examples that come up constantly:
- "Draft an email following the template for use case X."
- "Answer questions about the employee training material."
- "Find how many manuscript we published last year on topic Y"
- "Draft a presentation about this training topic"