LAPSE:2026.0529v1
Published Article

LAPSE:2026.0529v1
Benchmarking generative AI on fermentation knowledge
June 12, 2026
Abstract
With the ongoing advances in generative artificial intelligence (GenAI), the initial skepticism surrounding its tools is gradually diminishing. In fact, tools such as ChatGPT, Copilot and similar, are often used in everyday tasks, both in our personal lives and in educational contexts. Educators may use them for content creation, grading exams, or automating repetitive tasks. Students resort to them to better understand a topic, get feedback on an assignment and brainstorm ideas. Research has shown that, if used correctly, these tools can spur and support both teaching and learning. However, these continuous advancements and the increasing number of available tools also require more research to benchmark all these models and, if possible, provide quantifiable indications of which tool is better to use for which specific subtopic. As such, we created FermBench, a dataset of fermentation knowledge, which can be used to benchmark various large language models (LLMs). The models selected for the experiment are: GPT-5 mini, Claude Sonnet 4.5, Gemini 3 Flash, Mistral Small 3.2, and DeepSeek-V3.2. The performance of the models is scored using the LLM-as-a-Judge approach, with GPT 5.2 and Claude Opus 4.5 as judges. The LLMs agree that the model used in the free tier of ChatGPT is the most factually accurate. DeepSeek and Gemini outperform the other models in readability and helpfulness, while Claude and Gemini provide the most concise answers. Overall, we conclude that the quality of the responses generated is satisfactory for educational contexts. We also provide students with practical guidelines on how to be more critical of responses generated by LLMs and to improve their prompts. All code and curated data are open-source.
With the ongoing advances in generative artificial intelligence (GenAI), the initial skepticism surrounding its tools is gradually diminishing. In fact, tools such as ChatGPT, Copilot and similar, are often used in everyday tasks, both in our personal lives and in educational contexts. Educators may use them for content creation, grading exams, or automating repetitive tasks. Students resort to them to better understand a topic, get feedback on an assignment and brainstorm ideas. Research has shown that, if used correctly, these tools can spur and support both teaching and learning. However, these continuous advancements and the increasing number of available tools also require more research to benchmark all these models and, if possible, provide quantifiable indications of which tool is better to use for which specific subtopic. As such, we created FermBench, a dataset of fermentation knowledge, which can be used to benchmark various large language models (LLMs). The models selected for the experiment are: GPT-5 mini, Claude Sonnet 4.5, Gemini 3 Flash, Mistral Small 3.2, and DeepSeek-V3.2. The performance of the models is scored using the LLM-as-a-Judge approach, with GPT 5.2 and Claude Opus 4.5 as judges. The LLMs agree that the model used in the free tier of ChatGPT is the most factually accurate. DeepSeek and Gemini outperform the other models in readability and helpfulness, while Claude and Gemini provide the most concise answers. Overall, we conclude that the quality of the responses generated is satisfactory for educational contexts. We also provide students with practical guidelines on how to be more critical of responses generated by LLMs and to improve their prompts. All code and curated data are open-source.
Record ID
Keywords
Subject
Suggested Citation
Caccavale F, Krühne U, Gernaey KV, Gargalo CL. Benchmarking generative AI on fermentation knowledge. Systems and Control Transactions 5:2600-2606 (2026) https://doi.org/10.69997/sct.137474
Author Affiliations
Caccavale F: Process and Systems Engineering Center (PROSYS), Department of Chemical and Biochemical Engineering, Technical University of Denmark, Søltofts Plads, Building 228A, 2800 Kgs. Lyngby, Denmark [ORCID]
Krühne U: Process and Systems Engineering Center (PROSYS), Department of Chemical and Biochemical Engineering, Technical University of Denmark, Søltofts Plads, Building 228A, 2800 Kgs. Lyngby, Denmark [ORCID]
Gernaey KV: Process and Systems Engineering Center (PROSYS), Department of Chemical and Biochemical Engineering, Technical University of Denmark, Søltofts Plads, Building 228A, 2800 Kgs. Lyngby, Denmark [ORCID]
Gargalo CL: Process and Systems Engineering Center (PROSYS), Department of Chemical and Biochemical Engineering, Technical University of Denmark, Søltofts Plads, Building 228A, 2800 Kgs. Lyngby, Denmark [ORCID]
[Login] to see author email addresses.
Krühne U: Process and Systems Engineering Center (PROSYS), Department of Chemical and Biochemical Engineering, Technical University of Denmark, Søltofts Plads, Building 228A, 2800 Kgs. Lyngby, Denmark [ORCID]
Gernaey KV: Process and Systems Engineering Center (PROSYS), Department of Chemical and Biochemical Engineering, Technical University of Denmark, Søltofts Plads, Building 228A, 2800 Kgs. Lyngby, Denmark [ORCID]
Gargalo CL: Process and Systems Engineering Center (PROSYS), Department of Chemical and Biochemical Engineering, Technical University of Denmark, Søltofts Plads, Building 228A, 2800 Kgs. Lyngby, Denmark [ORCID]
[Login] to see author email addresses.
Journal Name
Systems and Control Transactions
Volume
5
First Page
2600
Last Page
2606
Year
2026
Publication Date
2026-06-12
Version Comments
Original Submission
Other Meta
PII: 2600-2606-37-SCT-5-2026, Publication Type: Journal Article
Record Map
Published Article

LAPSE:2026.0529v1
This Record
External Link

https://doi.org/10.69997/sct.137474
Publisher Version
Download
Meta
Record Statistics
Record Views
392
Version History
[v1] (Original Submission)
Jun 12, 2026
Verified by curator on
Jun 12, 2026
This Version Number
v1
Citations
Most Recent
This Version
URL Here
https://psecommunity.org/LAPSE:2026.0529v1
Record Owner
PSE Press
Links to Related Works
References Cited
- Kasneci E, Sessler K, Küchemann S, Bannert M, Dementieva D, Fischer F, Gasser U, Groh G, Günnemann S, Hüllermeier E, Krusche S, Kutyniok G, Michaeli T, Nerdel C, Pfeffer J, Poquet O, Sailer M, Schmidt A, Seidel T, Stadler M, Weller J, Kuhn J, Kasneci G. Chatgpt for good? on opportunities and challenges of large language models for education. Learning and Individual Differences 103:102274 (2023) https://doi.org/10.1016/j.lindif.2023.102274
- Rodrigues L, Dwan Pereira F, Cabral L, Gaševi? D, Ramalho G, Ferreira Mello R. Assessing the quality of automatic-generated short answers using GPT-4. Computers and Education: Artificial Intelligence 7:100248 (2024) https://doi.org/10.1016/j.caeai.2024.100248
- Kamalov F, Santandreu Calonge D, Gurrib I. New era of artificial intelligence in education: towards a sustainable multifaceted revolution. Sustainability 15:12451 (2023) https://doi.org/10.3390/su151612451
- Caccavale F, Gargalo CL, Kager J, Larsen S, Gernaey KV, Krühne U. Chatgmp: a case of AI chatbots in chemical engineering education towards the automation of repetitive tasks. Computers and Education: Artificial Intelligence 8:100354 (2025) https://doi.org/10.1016/j.caeai.2024.100354
- Caccavale, F., Gargalo, C. L., Gernaey, K. V., & Krühne, U. (2024). FermentAI: Large language models in chemical engineering education for learning fermentation processes. In Computer aided chemical engineering (Vol. 53, pp. 3493-3498). Elsevier.
- Caccavale F, Aouichaoui ARN, Krühne U, Gernaey KV, Gargalo CL. Fermbench: a new benchmark for measuring the capabilities of llms on fermentation knowledge. Computers and Education: Artificial Intelligence 10:100577 (2026) https://doi.org/10.1016/j.caeai.2026.100577
- Stöhr C, Ou AW, Malmström H. Perceptions and usage of AI chatbots among students in higher education across genders, academic levels and fields of study. Computers and Education: Artificial Intelligence 7:100259 (2024) https://doi.org/10.1016/j.caeai.2024.100259
- Caccavale F, Gargalo CL, Gernaey KV, Krühne U. Towards education 4.0: the role of large language models as virtual tutors in chemical engineering. Education for Chemical Engineers 49:1-11 (2024) https://doi.org/10.1016/j.ece.2024.07.002
- Deschenes A, McMahon M. A survey on student use of generative AI chatbots for academic research. EBLIP 19:2-22 (2024) https://doi.org/10.18438/eblip30512
- Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training.
- Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., ... & Agarwal, S. (2020). Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 1(3), 3.
- Anthropic, Introducting Claude, https://www.anthropic.com/news/introducing-claude, 2023a. [Online; accessed 26-September-2025].
- Google, Gemini, https://blog.google/technology/ai/google-gemini-ai, 2023. [Online; accessed 26-September-2025].
- Google Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al., Gemini: a family of highly capable multimodal models, arXiv preprint arXiv:2312.11805 (2023).
- DeepSeek, DeepSeek-R1 Release, https://api-docs.deepseek.com/news/news250120, 2025. [Online; accessed 26-September-2025].
- DeepSeek, DeepSeek-R1-Lite Release 2024/11/20, https://api-docs.deepseek.com/news/news1120, 2024. [Online; accessed 26-September-2025].
- MistralAI, le Chat, https://mistral.ai/news/all- new- le- chat, 2025a. [Online; accessed 26-September-2025].
- A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017).
- Chiang WL, Gonzalez J, Li D, Li Z, Lin Z, Sheng Y, Stoica I, Wu Z, Xing E, Zhang H, Zheng L, Zhuang S, Zhuang Y. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 :46595-46623 (2023) https://doi.org/10.52202/075280-2020
- Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., ... & Guo, J. (2024). A survey on llm-as-a-judge. The Innovation.
- Wang J, Liang Y, Meng F, Sun Z, Shi H, Li Z, Xu J, Qu J, Zhou J. Is chatgpt a good NLG evaluator? a preliminary study. Proceedings of the 4th New Frontiers in Summarization Workshop :1-11 (2023) https://doi.org/10.18653/v1/2023.newsum-1.1
- Badshah S, Sajjad H. Reference-guided verdict: llms-as-judges in automatic evaluation of free-form QA. Proceedings of the 9th Widening NLP Workshop :251-267 (2025) https://doi.org/10.18653/v1/2025.winlp-main.37
(0.09 seconds)
[0.1 s]

