From a buzzword to a core part of our workflow: the AI at MDPI
By Jean-Baptiste de la Broise · 5 August 2026
- NLP
- LLMs
- AI in publishing
- Team growth
Five years ago, AI at MDPI meant one GPU workstation running small R&D experiments. Today, nearly every one of our employees is using AI-powered tools daily.
The Inception
In the beginning, I was an intern in a small R&D team that functioned like a startup within the company. My initial project was to develop an internal embedding model, ScilitBERT, that could be used in various applications. This model can turn a scientific text into a vector representation of its meaning, such that similarity can be used to compare different texts.
We first saw the potential applications when we tuned the model on some downstream tasks:
- Journal-Finder that allowed us to recommend journals for authors to publish in based on their submission
- A category classification framework that allows us to provide better filters on Scilit.
The company started to realize the power of AI when we introduced article similarity, allowing us to find semantically similar publications to a target article or query, indexing the whole MDPI portfolio, as well as Scilit in a semantic vector space. We managed to build a map of the published scientific knowledge, and from this simple concept, many applications became possible.
The Expansion
At this stage, the team started to grow, allowing us to develop many applications (a lot of PoCs using Streamlit, without dedicated frontend resources). Most of those are still actively worked on today. We developed tools to find suitable peer reviewers, improve our marketing targeting, recommend articles to read, etc.
The most interesting one at that time was Ethicality, a tool proposing a suite of AI checks on submitted manuscripts, at a time when paper mills were becoming a major concern within academia, with some publishers needing to do fast retraction at scale. Ethicality was our response to these concerns: upload a PDF and get access to different signals: self-citations, retracted references, grey literature detection, out-of-scope citations, AI-generated text detection, etc. It started as a tool for the ethics team, but now it is part of the standard workflow.
Today, not a single manuscript is published without being screened.
This is also a time that was marked by the advent of ChatGPT. As tech enthusiasts, we were fascinated by this technology, but as a publisher, the company and academia as a whole were deeply concerned by it. Indeed, the rise of LLM meant paper milling would become so easy that all our legacy processes were at risk. This is when the idea of Generative AI text detection was first experimented with within the company.
The Maturation
At that stage, we had many promising and useful tools and already many users despite suboptimal UI; the next step was to put some structure in there.
It was decided to invest a lot in the team. In one year, the team size doubled. Until then, I was building end-to-end solutions A-Z; now we have dedicated backend developers, product owners, data team. We collaborate with the frontend team, build integrations with other tools, invest in infrastructure (dedicated production GPU servers). We also made the switch from FAISS indexes to dedicated Qdrant vector database solutions, built many automations and started to deploy using helm and Kubernetes… This was quite a turn and not a simple one to take, but it brought us so much.
This is also when I started to experiment with the use of LLMs within the company. At that time, I deployed our first internal model that made me believe in self-hosted LLMs. This was a Mixtral8x7b, a revolutionary model from 2023 that is totally outdated now. I built a lot of prototype applications with it, and worked on finding the right use cases that match our users' needs for this new technology. We were also very cautious about providing LLM solutions to externals, as the cybersecurity on the technology is not very mature. Going through different PoCs, I found a real audience when I built a RAG application covering our internal policies and guidelines. Knowledge was previously fragmented across countless pages, making it hard to find. This tool reduced search times to seconds. It also lightened the load on our experienced mentors and senior staff, who were previously the 'go-to' experts for information that wasn't easily accessible.
Now we are taking a new step and have started to experiment with agentic capabilities. We provide a secure chat interface with access to many of the tools and knowledge bases built by the team, accessed daily by close to a thousand employees.
We also started to serve Large Language Models and embedding models to other teams at scale, mostly for analytics purposes.
I am still very excited for what the future has to bring to the team. To see how I can keep innovating in the ever-shifting AI and publishing landscape.