LLM-based systems fail in specific, repeatable ways starting with wrong output format, ignored rules, missing required fields. In this talk we will cover how to build evals (in Langfuse) to catch these: deterministic checks for structure and format, and LLM-as-a-judge for the cases rules can’t cover. We’ll go through what each eval type looks like in practice and how to read results alongside traces when something breaks. Live demo included
Evals, llm-as-a-judge and building predictable systems
GenAI Cracow #26 - Observability
Wydarzenie i data
GenAI Cracow #26 - Observability · · Software Mind, Kraków, PL