LLM evaluation service

Kafka-native service that scores LLM and RAG output as events flow through the pipeline.

context
Master's thesis, LUT University, with Telia Finland
updated
stack
  • Python
  • Kafka
  • DeepEval
run
llm-eval-service in the terminal

A developer-facing backend service that makes LLM features measurable. It evaluates model output for groundedness and faithfulness, tracks regressions between versions, and plugs into real systems as a modular, scalable service.

I built it for my master’s thesis in software engineering at LUT University, together with Telia Finland.

Why

LLMs fail in ways classic software rarely does. A change that improves one prompt can degrade another. A model update can change tone, detail or refusal behavior. Without evaluation you end up debugging vibes.

What it checks

How it works

[TODO: architecture, topics and how teams integrate it]

Results

[TODO: thesis results and a link to the thesis or a write-up]


All projects